Automatic generation and / or use of text-dependent speaker verification functions

Specific TD-SVs constrained to user-specific words or phrases generate speaker features during interactions, improving authentication accuracy and reducing interaction time and resource use by eliminating the need for re-prompting and additional authentication methods.

JP7771249B2Active Publication Date: 2025-11-17GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024039147
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-10-13
Filing Date
2024-03-13
Publication Date
2025-11-17
Estimated Expiration
2040-12-15

AI Technical Summary

Technical Problem

Existing text-dependent speaker verification (TD-SV) approaches based solely on invocation phrases are insufficient for reliable user verification, leading to prolonged interactions and resource utilization due to the need for re-prompting or additional authentication methods when invocation phrases are not spoken or device sensors are absent.

Method used

Implementing specific TD-SVs constrained to user-specific words or phrases, generating speaker features during typical interactions using neural networks, and comparing speech features to authenticate users without separate enrollment procedures, enabling robust authentication through multiple TD-SVs in a single utterance.

Benefits of technology

Enhances authentication accuracy and reduces interaction time by eliminating the need for re-prompting and additional authentication, conserving client device resources, and supporting sensor-less invocations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007771249000001
    Figure 0007771249000001
  • Figure 0007771249000002
    Figure 0007771249000002
  • Figure 0007771249000003
    Figure 0007771249000003
Patent Text Reader

Abstract

To facilitate interaction between a person and an assistant.SOLUTION: Implementations relate to automatic generation of speaker features of each of one or more particular text-dependent speaker verification (TD-SV) of users. Implementations can generate the speaker features of particular TD-SV using instances of audio data which captures spoken utterance of a corresponding user during normal non-enrollment interactions with an automated assistant via one or more assistant devices. For example, a portion of the instance of audio data can be used in response to: (a) determining a recognized term for the spoken utterance captured by that the portion corresponds to particular TD-SV; and (b) determining that an authentication measure for the user and the spoken utterance satisfies a threshold. Implementations additionally or alternatively relate to utilization of the speaker features for each of one or more particular TD-SV of the user in determining whether to authenticate the spoken utterance of the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Humans can participate in human-computer interactions using interactive software applications referred to herein as “automated assistants” (also referred to as “chatbots,” “interactive personal assistants,” “intelligent personal assistants,” “personal voice assistants,” “conversational agents,” etc.). For example, a human (who may be referred to as a “user” when interacting with an automated assistant) may provide commands / requests to the automated assistant using vocal natural language input (i.e., spoken utterances), which may in some cases be converted to text and then processed, and / or by providing textual (e.g., typed) natural language input. The automated assistant typically responds to the commands or requests by providing responsive user interface output (e.g., audible and / or visual user interface output), controlling a smart device, and / or performing other actions.

[0002] Some user commands or requests can only be fully processed and responded to when the automated assistant authenticates the requesting user. For example, to maintain the security of personal data, user authentication may be required for requests that access the user's personal data when generating a response and / or that incorporate the user's personal data into the response. For example, user authentication may be required to properly respond to requests such as "What's in my calendar tomorrow?" (which requires access to the requesting user's personal calendar data and inclusion of the personal calendar data in the response) and "Message Vivian, I'm going to be late" (which requires access to the requesting user's personal contact data). As another example, to maintain the security of smart devices (e.g., smart thermostats, smart lights, smart locks), user authentication may be required for requests to control one or more of such smart devices.

[0003] Various techniques are utilized for user authentication of automated assistants. For example, when authenticating a user, some automated assistants utilize text-dependent speaker verification (TD-SV) restricted to the assistant's invocation phrases (e.g., "OK, assistant" and / or "Hey, assistant"). With such TD-SV, an enrollment procedure is performed in which the user is explicitly prompted to provide one or more instances of spoken utterances of the invocation phrases to which TD-SV is constrained. Speaker features (e.g., speaker embeddings) for the user can then be generated through processing instances of audio data, each capturing a respective one of the spoken utterances. For example, the speaker features can be generated by processing each instance of audio data with a TD-SV machine learning model to generate a corresponding speaker embedding for each utterance. Speaker features can then be generated as a function of the speaker embeddings and stored (e.g., on the device) for use in TD-SV. For example, the speaker features can be a cumulative speaker embedding that is a function (e.g., average) of the speaker embeddings.

[0004] After the speaker features are generated, the speaker features can be used in verifying that a spoken utterance was spoken by the user. For example, when another spoken utterance is spoken by the user, audio data capturing the spoken utterance can be processed to generate speech features, which can be compared to the speaker features, and a determination can be made as to whether to authenticate the speaker based on the comparison. As one particular example, the audio data can be processed using a speaker recognition model to generate speech embeddings, which can be compared to previously generated speaker embeddings for the user in determining whether to verify that the user is the speaker of the spoken utterance. For example, if a distance metric between the generated speech embedding and the user's speaker embedding meets a threshold, the user can be verified as the user who spoke the spoken utterance. Such verification can be utilized as a measure (e.g., the only measure) for authenticating the user. Summary of the Invention [Problem to be solved by the invention]

[0005] However, existing TD-SV approaches based solely on invocation phrases have various drawbacks. For example, comparing speech features and speaker features is often insufficient to verify a user with sufficient confidence. For example, if the comparison involves generating a distance metric between speech features that are speech embeddings and speaker features that are speaker embeddings, the distance metric may not meet a closeness threshold for verifying the user. As a result, the automated assistant may need to prompt the user to speak the invocation phrase again to retry verification and / or may need to prompt the user for additional authentication data (e.g., prompting the user to provide a PIN or other passcode, prompting the user to interact with a fingerprint sensor, and / or prompting the user to position themselves in front of a camera for facial identification). This lengthens the human-assistant interaction, resulting in longer utilization of client device resources and / or other resources to facilitate the human-assistant interaction. As another example, if a user invokes an automated assistant without using an invoke phrase, squeezing the corresponding device, or looking at the corresponding device (e.g., detected based on an image from the corresponding device's camera) (e.g., instead of invoking by interacting with a hardware button on the corresponding device), TD-SV based solely on the invoke phrase does not work because the invoke phrase was not spoken when invoking the automated assistant. Thus, in such a situation, the automated assistant must prompt the user to speak the invoke phrase and / or prompt the user for additional authentication data. This lengthens the human-assistant interaction and, consequently, the utilization of client device resources and / or other resources to facilitate the human-assistant interaction. As yet another example, TD-SV techniques require the execution of a registration procedure and the resulting utilization of client device resources and / or other resources to perform the registration procedure. [Means for solving the problem]

[0006] Some implementations disclosed herein relate to the automatic generation of speaker features for each of a user's one or more specific text-dependent speaker verifications (TD-SVs). Each user's specific TD-SV is associated with one or more specific words and / or phrases to which it is constrained. For example, a first specific TD-SV can be constrained to a single word such as "door," a second specific TD-SV can be constrained to a single phrase such as "front door," and a third specific TD-SV can be constrained to two or more phonetically similar words such as "light" and "lights." Furthermore, each user's specific TD-SV is associated with speaker features for that specific TD-SV, which are generated according to implementations disclosed herein. The speaker features for a specific TD-SV can be, for example, speaker embeddings. The words and / or phrases to which a specific TD-SV is constrained can be selected based on words and / or phrases associated with a corresponding assistant action for which user verification is at least selectively required.

[0007] Implementations can generate speaker features for a particular TD-SV using instances of audio data capturing a corresponding user's spoken utterance during a typical, non-enrollment interaction with the automated assistant via one or more respective assistant devices. More specifically, a portion of an instance of audio data can be used in response to (a) determining a recognized term for the captured spoken utterance (determined using speech recognition performed on the audio data) where the portion corresponds to a particular TD-SV, and (b) determining that an authentication measure for the user and the spoken utterance meets a threshold indicating sufficient confidence that the spoken utterance was spoken by the user. The authentication measure can be based on one or more factors, such as user fingerprint verification, user face verification, analysis of a verification code entered by the user, a different specific TD-SV for which speaker features have already been generated, and / or a general invocation TD-SV. The general invocation TD-SV is based on processing of the audio data, or preceding audio data preceding the audio data, for the user and one or more general invocation wake words for the automated assistant.

[0008] As one particular example, assume a particular TD-SV of a user that is constrained to the phrase "kitchen thermostat." Further, assume that the user invokes an automated assistant on an assistant device and provides the spoken utterance, "Please set the kitchen thermostat to 72 degrees." The automated assistant can perform speech recognition of the spoken utterance based on the corresponding captured audio data to generate a recognition of "Please set the kitchen thermostat to 72 degrees" for the spoken utterance. The automated assistant can further determine authentication criteria for the utterance, for example, based on facial verification of the user performed based on images captured by a camera on the assistant device before, during, and / or after the spoken utterance. The authentication criteria can be determined to meet a threshold, for example, based on facial verification indicating at least a threshold confidence that the image captures a face corresponding to the user's stored facial embedding.

[0009] In response to determining that the authentication criteria meet a threshold value and in response to determining that the recognition includes the term “kitchen thermostat” (and optionally, that the speech recognition confidence for that term meets a threshold value), the portion of the audio data corresponding to “kitchen thermostat” (e.g., as indicated by the speech recognition) can be used in generating speaker features for a particular TD-SV. For example, the portion of the audio data can be processed using a neural network model to generate embeddings, such as embeddings that include values ​​from a hidden layer of the neural network after processing the audio data. The embeddings can be used in generating the speaker features. For example, the embeddings can be used as speaker features, or the speaker features can be generated as a function of embeddings and other embeddings that are each generated based on respective portions of the audio data that captured “kitchen thermostat” in the spoken utterance when the user's respective verification criteria meet a threshold value. Furthermore, it should be noted that the automated assistant can also perform an assistant action conveyed by the utterance. That is, the automated assistant can perform an assistant action to adjust the setting of the kitchen thermostat to 72 degrees.

[0010] In these and other methods, specific TD-SV speaker features constrained to the phrase "kitchen thermostat" can be generated based on spoken utterances from a typical user interaction with an automated assistant, without any interruption to the human-assistant interaction. This avoids the need for a separate, computationally intensive enrollment procedure to generate specific TD-SV speaker features. Furthermore, by utilizing the portion of audio data that speech recognition has determined corresponds to the phrase "kitchen thermostat" (optionally using a confidence threshold), it is possible to automatically determine the portion of audio data that corresponds to the specific TD-SV phrase. Furthermore, generating speaker features based on that portion only if the user's authentication criteria meet the threshold confirms that the user's specific TD-SV speaker features were actually generated for the user and based on the user's spoken utterances.

[0011] Some implementations disclosed herein additionally or alternatively relate to utilizing speaker features for each of one or more specific TD-SVs of a user when determining whether to authenticate the user's spoken utterance. As an example, speech recognition can be performed based on audio data capturing the spoken utterance to generate a recognition of the spoken utterance. If a recognition term corresponds to a term of a specific TD-SV, the corresponding portion of the audio data is processed to generate speech features, and the speech features are compared to the speaker features of the specific TD-SV when determining whether to authenticate the user. For example, the portion of the audio data can be processed using a neural network model (e.g., a model used in generating speaker features of the TD-SV) to generate embeddings, and the embeddings can be speech features. Furthermore, the embeddings can be compared to the TD-SV and the user's speaker features, which are embeddings. For example, the comparison can include generating a cosine distance measurement or other distance metric. The decision of whether to authenticate the user can be based on the comparison (e.g., based on the generated distance metric). For example, authentication of a user may be contingent on a comparison exhibiting at least a threshold degree of similarity (eg, less than a threshold distance if embeddings are being compared).

[0012] In various implementations, the spoken utterance includes at least a first term corresponding to a first specific TD-SV and a second term corresponding to a second specific TD-SV. For example, the spoken utterance may be "open the garage door," where the first term may be "open" and the second term may be "garage door." In these implementations, a first portion of the audio data corresponding to the first terms may be processed to generate first speech features, and the first speech features are compared with first speaker features of the first specific TD-SV. Furthermore, a second portion of the audio data corresponding to the second terms may be processed to generate second speech features, and the second speech features are compared with second speaker features of the second specific TD-SV. A decision on whether to authenticate the user may be based on both the first comparison of the first speech features with the first speaker features and the second comparison of the second speech features with the second speaker features. For example, authentication of the user may be contingent on both a first comparison indicating at least a threshold similarity and a second comparison indicating at least a threshold similarity. As another example, the first distance metric from the first comparison and the second distance metric from the second comparison may be averaged and / or otherwise combined to generate an overall distance metric, and authentication of the user may be contingent on the overall distance metric indicating at least the threshold similarity. Optionally, the first and second distance metrics may be weighted differently in the averaging or other combination. For example, the first distance metric may be weighted more heavily based on a first speaker feature based on a greater amount of past spoken utterances of the user than a speaker feature of a second speaker feature. As another example, the first distance metric may additionally or alternatively be weighted more heavily based on a speech recognition confidence of a first term indicating greater confidence than a speech recognition confidence of a second term.As another example, additionally or alternatively, the second distance metric may be weighted more heavily based on "garage door" (corresponding to the first TD-SV) containing more characters and / or more syllables than "open" (corresponding to the second TD-SV). In various implementations, the weighting of the distance metric may be a function of two or more features, such as the amount of spoken speech on which the corresponding speaker feature is based, the speech recognition confidence of the corresponding term, and / or the length (e.g., character, syllable, and / or other length) of the corresponding term.

[0013] Implementations that determine whether to authenticate a user by considering multiple different specific TD-SVs of the user and based on a single utterance of the user can improve the accuracy and / or robustness of authentication. This can eliminate the need to prompt for alternative authentication data (e.g., prompting the user to provide a PIN or other passcode, prompting the user to interact with a fingerprint sensor, and / or prompting the user to position themselves in front of a camera for facial identification). This can prevent prolonged human-assistant interactions, reduce the amount of input required by the user during human-assistant interactions, and reduce the amount of client device resources and / or other resources required to facilitate the human-assistant interactions.

[0014] As described above, a particular TD-SV can be limited to words that are not the assistant's invocation wake word and / or words associated with assistant actions that at least selectively require user authentication. Thus, speaker features of a particular TD-SV can be utilized to authenticate a user in situations where the automated assistant is invoked without utilizing an invocation hotword (e.g., instead, invoked in response to a touch gesture, a touch-free gesture, presence detection, and / or the user's gaze). In these and other ways, an automated assistant becomes more robust by enabling user authentication without requiring a prompt for alternative authentication data and / or by enabling robust authentication when the user is interacting with an assistant device that may lack a fingerprint sensor, camera, and / or other sensors for providing alternative authentication data. Furthermore, this may eliminate the need to request alternative authentication data in situations where the user invokes the assistant without utilizing an invocation hotword. This prevents prolonged human-assistant interactions, reduces the amount of input required by the user during human-assistant interactions, and reduces the amount of client device resources and / or other resources required to facilitate the human-assistant interaction.

[0015] The above is provided only as a summary of some implementations. These and / or other implementations are disclosed in more detail herein. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 1 is a block diagram of an example computing environment in which implementations disclosed herein may be implemented. [Figure 2] 1 is a flow diagram illustrating an example method for automatically generating speaker features for each of one or more specific text-dependent speaker verifications, according to various implementations. [Figure 3A]1A-1C illustrate exemplary user spoken utterances, each provided when the user authenticates with an automated assistant. [Figure 3B] FIG. 3B illustrates an example of utilizing audio data for the spoken utterance of FIG. 3A in automatically generating speaker features for a plurality of different specific text-dependent speaker verifications. [Figure 4] 1 is a flowchart illustrating an example method for authenticating a user using one or more specific text-dependent speaker verification methods, according to various implementations. [Figure 5A] FIG. 1 illustrates an example of a spoken utterance and the use of multiple specific text-dependent speaker verification in determining whether to verify the spoken utterance as being spoken by a particular user. [Figure 5B] FIG. 1 illustrates an example of a spoken utterance and the use of multiple specific text-dependent speaker verification in determining whether to verify the spoken utterance as being spoken by a particular user. [Figure 6] FIG. 1 illustrates an exemplary architecture of a computing device. DETAILED DESCRIPTION OF THE INVENTION

[0017] Referring initially to Figure 1, an exemplary environment in which various implementations may be implemented is shown. Figure 1 includes an assistant device 110 (i.e., a client device that runs and / or otherwise has access to an automated assistant client) that runs an instance of an automated assistant client 120. One or more cloud-based automated assistant components 140 can be implemented on one or more computing systems (collectively referred to as "cloud" computing systems) communicatively coupled to assistant device 110 via one or more local and / or wide area networks (e.g., the Internet), generally indicated at 108.

[0018] An instance of automated assistant client 120, optionally through interaction with one or more cloud-based automated assistant components 140, can form what appears from a user's perspective to be a logical instance of an automated assistant with which the user can participate in human-computer interactions. An example of such an automated assistant 100 is shown in FIG. 1.

[0019] Assistant device 110 may be, for example, a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a vehicle computing device (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker, a smart appliance such as a smart television, and / or a wearable device that includes a computing device (e.g., a watch with a computing device, glasses with a computing device, a virtual reality or augmented reality computing device).

[0020] Assistant device 110 may be utilized by one or more users within a home, business, or other environment. Additionally, one or more users may be registered with assistant device 110 and have corresponding user accounts accessible via assistant device 110. The text-dependent speaker verification (TD-SV) described herein may be generated and stored for each enrolled user (e.g., associated with a corresponding user profile) with permission from the associated user. For example, a TD-SV constrained to the term “door” may be stored in association with a first enrolled user and have corresponding speaker features unique to the first enrolled user, and a TD-SV constrained to the term “door” may be stored in association with a second enrolled user and have corresponding speaker features unique to the second enrolled user. TD-SV techniques described herein may be utilized to authenticate an utterance as coming from a particular user (rather than from another enrolled user or a guest user). Optionally, TI-SV techniques, speaker verification, face verification, and / or other verification techniques (e.g., PIN entry) may be utilized additionally or alternatively in authenticating a particular user.

[0021] Additional and / or alternative assistant devices may be provided, and in some of these implementations, a user's specific TD-SV speaker features can be shared among assistant devices for which the user is a registered user. In various implementations, assistant device 110 can optionally operate one or more other applications, such as a messaging client (e.g., SMS, MMS, online chat), a browser, etc., in addition to automated assistant client 120. In some of these various implementations, one or more of the other applications optionally interface with automated assistant 100 (e.g., via an application programming interface) or include their own instance of an automated assistant application (which may also interface with cloud-based automated assistant component 140).

[0022] Automated assistant 100 engages in a human-computer interaction session with a user through user interface input and output devices of client device 110. To protect the user's privacy and / or conserve resources, in many situations, the user must often explicitly invoke automated assistant 100 before the automated assistant can fully process a spoken utterance. Explicit invocation of automated assistant 100 can occur in response to a specific user interface input received at assistant device 110. For example, a user interface input that can invoke automated assistant 100 via assistant device 110 can optionally include activation of a hardware and / or virtual button on assistant device 110. Additionally, the automated assistant client can include one or more local engines, such as an invocation engine, operable to detect the presence of one or more spoken common invocation wake words. The invocation engine can invoke automated assistant 100 in response to detecting one of the spoken invocation wake words. For example, the call engine can call the automated assistant 100 in response to detecting a spoken call wake word, such as "Hey Assistant," "OK Assistant," and / or "Assistant." The call engine can continuously process a stream of audio data frames based on output from one or more microphones of the assistant device 110 (e.g., when not in "inactive" mode) to monitor for the occurrence of a spoken call phrase. While monitoring for the occurrence of a spoken call phrase, the call engine discards (e.g., after temporarily storing in a buffer) any audio data frames that do not contain the spoken call phrase. However, if the call engine detects the occurrence of a spoken call phrase in a processed audio data frame, the call engine can call the automated assistant 100.As used herein, "invoking" automated assistant 100 can include activating one or more previously inactive functions of automated assistant 100. For example, invoking automated assistant 100 can include causing one or more local engines and / or cloud-based automated assistant component 140 to further process the audio data frame and / or one or more subsequent audio data frames based on which invocation phrase was detected (whereas no further processing of the audio data frame occurred prior to the invocation). For example, the local and / or cloud-based component can process the captured audio data using an ASR model in response to the invocation of automated assistant 100.

[0023] 1 is shown as including an automatic speech recognition (ASR) engine 122, a natural language understanding (NLU) engine 124, a text-to-speech (TTS) engine 126, a fulfillment engine 128, and an authentication engine 130. In some implementations, one or more of the illustrated engines may be omitted (e.g., instead implemented solely by the cloud-based automated assistant component 140) and / or additional engines may be provided (e.g., the invocation engine described above).

[0024] The ASR engine 122 can process audio data capturing the spoken utterance to generate a recognition of the spoken utterance. For example, the ASR engine 122 can process the audio data utilizing one or more ASR machine learning models to generate a prediction of recognized text corresponding to the utterance. In some of these implementations, the ASR engine 122 can generate, for each of one or more recognized terms, a corresponding confidence score indicating the confidence that the predicted term corresponds to the spoken utterance.

[0025] The TTS engine 126 can convert text into synthetic speech, which can rely on one or more speech synthesis neural network models. The TTS engine 126 is used, for example, to convert text responses into audio data that includes a synthesized version of the text, which is audibly rendered through the hardware speakers of the assistant device 110.

[0026] The NLU engine 124 determines the semantic meaning of the audio and / or text converted from the audio by the ASR engine and determines an assistant action corresponding to the semantic meaning. In some implementations, the NLU engine 124 determines the assistant action as an intent and / or parameters determined based on the recognition of the ASR engine 122. In some situations, the NLU engine 124 can resolve the intent and / or parameters based on a single utterance of the user, while in other situations, prompts can be generated based on unresolved intents and / or parameters, those prompts rendered to the user, and the user's responses to those prompts utilized by the NLU engine 124 in resolving the intent and / or parameters. In such situations, the NLU engine 124 can optionally work in conjunction with a dialogue manager engine (not shown) that determines unresolved intents and / or parameters and / or generates corresponding prompts. The NLU engine 124 can utilize one or more NLU machine learning models in determining the intent and / or parameters.

[0027] Fulfillment engine 128 can trigger the execution of an assistant action determined by NLU engine 124. For example, if NLU engine 124 determines the assistant action "turning on the kitchen lights," fulfillment engine 128 can send corresponding data (either directly to the lights or to a remote server associated with the light manufacturer) to turn on the "kitchen lights." As another example, if NLU engine 124 determines the assistant action "provide a summary of the user's meetings for today," fulfillment engine 128 can access the user's calendar, summarize the user's meetings for that day, and cause the summary to be visually and / or audibly rendered on assistant device 110.

[0028] 1 includes a comparison module 132, an utterance feature module 134, an other module 136, and a speaker feature module 138. Authentication engine 130 may include additional or alternative modules in other implementations. Authentication engine 130 may determine whether to authenticate a spoken utterance of a particular user registered with assistant device 110.

[0029] The speaker feature module 138 can generate speaker features for each of one or more specific TD-SVs of the user, each using an instance of audio data capturing the corresponding user's spoken utterance during a normal, non-enrollment interaction with the automated assistant via one or more respective assistant devices. For example, the speaker feature module 138 can utilize a portion of the instance of audio data capturing the spoken utterance in generating speaker features for the user's TD-SV in response to (a) determining a recognized term (determined using speech recognition performed on the audio data) for the captured spoken utterance by which the portion corresponds to a specific TD-SV, and (b) determining that authentication criteria for the user and the spoken utterance meet a threshold indicating sufficient confidence that the user spoke the spoken utterance.

[0030] In generating speaker features for TD-SV, the speaker feature module 138 can process portions of the audio data using one of one or more TD-SV models 152A-N corresponding to the TD-SV. In some implementations, one of the TD-SV models 152A-N can be used for the TD-SV and each of multiple additional TD-SVs. For example, the same single TD-SV model can be utilized for all TD-SVs or a subset including multiple TD-SVs. In some other implementations, one of the TD-SV models 152A-N can be used only for the TD-SV. For example, a single TD-SV model can be trained based on multiple users' utterances, where the utterances are limited to those containing TD-SV terms and / or phonetically similar terms. The speaker feature module 138 can store the speaker features generated for the TD-SV in association with the user and in association with the TD-SV. For example, the speaker feature module can store the generated speaker features and associations in the TD-SV feature database 154, which can optionally be local to the assistant device 110. In some implementations, the speaker feature module 138 can perform at least some aspects of at least blocks 268, 270, and 272 of the method 200 of FIG. 2 (described below).

[0031] The speech feature module 134 may generate speech features for each of the spoken utterance and one or more particular TD-SVs using corresponding portions of the audio data that each capture a corresponding term of the TD-SV. For example, in response to determining recognized terms (determined using speech recognition performed on the audio data) for the captured spoken utterance whereby the portions correspond to a particular TD-SV, the speech feature module 134 may utilize the portions of the instance of the audio data that capture the spoken utterance in generating speaker features for the user's TD-SV.

[0032] In generating speech features for the TD-SV, the speech feature module 134 may process the portion of the audio data using one of the one or more TD-SV models 152A-N corresponding to the TD-SV. In some implementations, the speech feature module 134 may perform at least aspects of at least blocks 462, 466 and processing portions of block 464 of the method 400 of FIG. 4 (described below).

[0033] The comparison module 132 may compare each of the one or more speech features generated by the speech feature module 134 for the utterance with their corresponding speaker features for the corresponding TD-SV. For example, three speech features may be generated by the speech feature module 134 for the utterance, each corresponding to a different TD-SV. In such an example, the comparison module 132 may compare each of the three speech features with one of the corresponding speaker features stored in the TD-SV feature database 154 to generate a corresponding distance metric. Each distance metric may indicate how well the speech feature matches the feature of the corresponding speaker. In some implementations, the speaker comparison module 132 may perform at least some aspects of at least the comparison portion of block 464 of the method 400 of FIG. 4 (described below).

[0034] The authentication engine 130 may determine whether to authenticate a user for a spoken utterance based at least in part on the comparison performed by the comparison module 132. In some implementations, the authentication engine 130 may additionally or alternatively utilize metrics generated by other modules 136 for other verifications in determining whether to authenticate a user. For example, the other modules may generate metrics related to text-independent speaker verification, fingerprint verification, face verification, and / or other verifications.

[0035] Cloud-based automated assistant component 140 is optional and can operate in conjunction with and / or be utilized in place of (always or selectively) a corresponding component of assistant client 120. In some implementations, cloud-based component 140 can leverage the virtually limitless resources of the cloud to perform more robust and / or more accurate processing of audio data and / or other data compared to any counterpart of automated assistant client 120. In various implementations, assistant device 110 can provide audio data and / or other data to cloud-based automated assistant component 140 in response to a call engine that detects a spoken call phrase or some other explicit call of automated assistant 100.

[0036] The illustrated cloud-based automated assistant components 140 include a cloud-based ASR engine 142, a cloud-based NLU engine 144, a cloud-based TTS engine 146, a cloud-based fulfillment engine 148, and a cloud-based authentication engine 150. These components may perform similar functions as their automated assistant counterparts (if present). In some implementations, one or more of the illustrated cloud-based engines may be omitted (e.g., instead implemented solely by automated assistant client 120) and / or additional cloud-based engines may be provided.

[0037] 2 is a flow diagram illustrating an example method 200 for automatically generating speaker features for each of one or more particular text-dependent speaker verifications, according to various implementations. For convenience, the operations of the flow chart are described with reference to a system that performs the operations. The system may include various components of various computer systems, such as one or more components of the automated assistant 100. Furthermore, although the operations of the method 200 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0038] In block 252, the system receives audio data capturing the user's spoken utterances via the Assistant device's microphone.

[0039] In block 254, the system performs automatic speech recognition (ASR) on the audio data to generate a recognition of the spoken utterance.

[0040] In block 256, the system determines an assistant action to perform in response to the utterance based on the recognition in block 254. In some implementations, block 256 can include sub-block 257. In sub-block 257, the system performs natural language understanding (NLU) on the recognition to determine intents and parameters. In some of these implementations, the system optionally prompts the first user to clarify the intents and / or parameters (e.g., using visual and / or audible prompts). For example, the user may be prompted to disambiguate between a "play music" intent and a "play video" intent. As another example, the user may be prompted to disambiguate between a "smart device 1" parameter and a "smart device 2" parameter.

[0041] In block 258, the system determines authentication criteria for the user based on the speech and / or other data captured in block 252. For example, the system can perform a specific TD-S based on audio data capturing the speech and generate authentication criteria based on whether and / or the extent to which the specific TD-SV matches the user's corresponding speaker characteristics. As another example, the system can additionally or alternatively perform a general call TD-SV based on audio data preceding the spoken speech and generate authentication criteria based on whether and / or the extent to which the specific TD-SV matches the user's corresponding speaker characteristics. As another example, the system can additionally or alternatively perform text-independent speaker verification (TI-SV) ​​based on audio data capturing the speech and generate authentication criteria based on whether and / or the extent to which the TD-SV matches the user's corresponding speaker characteristics. As another example, the system can additionally or alternatively determine authentication criteria based on face verification (e.g., based on an image from a camera of the assistant device) and / or fingerprint verification (e.g., based on data from a fingerprint sensor of the assistant device).

[0042] In various implementations, the system may determine the authentication criteria based on one or more of how closely the corresponding verifications match the user's corresponding data (i.e., the closer the matching verifications, the more indicative the authentication criteria will be of authentication), how many different verifications have been made (i.e., the more verifications have been made, the more indicative the authentication criteria will be of authentication), and / or which verifications have been made (i.e., some verifications may be weighted more heavily than others).

[0043] In some implementations, block 258 includes sub-block 259, where the system performs further authentication if such further authentication is required by one or more assistant actions determined in block 256. For example, if the assistant action requires more advanced authentication than that indicated by the utterance and / or other currently available data in block 252, the system can prompt the user to provide utterances and / or data that enable further verification. For example, verification may not be possible based on the utterance and / or other currently available data, and the system can prompt the user to speak a general call wake word and / or speak (or otherwise enter) a PIN or other passcode. Also, for example, only general call TD-SV verification may be possible based on the utterance and / or other currently available data, and the action may require additional forms of verification (e.g., face or fingerprint verification), and the system can prompt for such additional forms of verification.

[0044] In block 260, the system performs an assistant action. For example, the assistant action may include generating and rendering audible and / or visual user interface output in response to the utterance, controlling a smart device, making a phone call (e.g., using VoIP), sending a message to another user, and / or performing other actions.

[0045] In block 262, the system determines whether the authentication criteria of block 258 meet a threshold. Note that in some implementations, the threshold in block 262 may be the same as the threshold required by the assistant action in block 259, if either is required in block 259 (e.g., the authentication criteria can still be determined and the threshold can be met even if the assistant action does not require authentication). In some other implementations, the threshold in block 262 may be different from the threshold required by the assistant action in block 259, if required in block 259.

[0046] If, at block 262, the system determines that the authentication criteria do not meet the threshold, the system can proceed to block 264 where the method 200 ends.

[0047] If, in block 262, the system determines that the authentication criteria meet the threshold, the system proceeds to block 266, where the system determines whether the recognition terms correspond to one or more TD-SVs of the user. For example, the utterance may be "turn up the kitchen thermostat," and the system may determine that "turn up" corresponds to the user's first specific TD-SV, "kitchen" corresponds to the user's second specific TD-SV, "thermostat" corresponds to the user's third specific TD-SV, and "kitchen thermostat" corresponds to the user's fourth specific TD-SV. As another example, the utterance may be "what time is it," and the system may determine that there are no terms that correspond to any of the user's specific TD-SVs.

[0048] If, at block 264, the system determines that there are no recognition terms corresponding to one or more TD-SVs of the user, the system proceeds to block 264 and method 200 ends.

[0049] If, in block 264 , the system determines that the recognition term corresponds to one or more TD-SVs of the user, the system proceeds to block 268 .

[0050] In block 268, the system selects a portion of the audio data that contains a particular TD-SV term. For example, the portion may be selected based on the ASR of block 254, indicating that the portion contains a particular TD-SV term.

[0051] In block 270, the system processes the portion of audio data in generating speaker features for the specific TD-SV. For example, the system can process the portion of audio data to generate given speaker features for the portion of audio data directly based on the processing. The speaker features for the specific TD-SV can be generated based on the generated given speaker features and, optionally, based on previously generated speaker features from processing a previous portion of audio data from the user that includes the specific TD-SV terms and meets the authentication threshold. If the speaker features match the given speaker features, they may be the result of, for example, a portion of audio data that is the first portion from the user that includes the specific TD-SV terms and meets the authentication threshold. In block 270, the system further stores the speaker features in association with the specific TD-SV and in association with the user. The speaker features can be stored locally, for example, on the assistant device. In some implementations, speaker features may additionally or alternatively be stored on the remote assistant server, but with limited access only for use in association with the user and / or shared with other assistant devices (such as via local Wi-Fi or other local connection) for which the user is also a registered user.

[0052] In some implementations, block 270 includes sub-block 271A and / or sub-block 271B.

[0053] In sub-block 271A, the system uses a TD-SV model for a particular TD-SV when processing a portion of the audio data. Then, for a portion of the audio data, a given speaker feature can be based on (e.g., corresponds exactly to) values ​​from a set of layer activations of the TD-SV model after processing of the audio data. The TD-SV model for a particular TD-SV can be a neural network model and can be utilized only for the particular TD-SV, or alternatively can be used for multiple additional (e.g., all) particular TD-SVs.

[0054] In sub-block 271B, the system combines the given speaker features generated from processing the portion of the audio data with previously generated speaker features for the user's particular TD-SV. For example, previously generated speaker features for the user's particular TD-SV may be averaged and / or otherwise combined with the given speaker features.

[0055] In block 272, the system determines whether there are more terms identified in block 266 that have not yet been processed. If so, the system performs another iteration of blocks 268 and 270 for another particular TD-SV of the user, using the portion of the audio data that contains those terms.

[0056] Referring now to Figure 3A, five exemplary spoken utterances 300A, 300B, 300C, 300D, and 300E are shown. Each of the exemplary spoken utterances 300A-E is from the same user and was provided when the user was authenticated with an automated assistant. Also shown in Figure 3A are various representations of portions of audio data for the spoken utterances 300A-E, each corresponding to a term that corresponds to a particular TD-SV of the user. More specifically, portions 301A, 301D, and 301E correspond to the term "unlock," portions 303A, 303B, 303C, and 303D correspond to the term "front door," portions 305A, 305B, 305C, 305D, and 305E correspond to the term "door," portions 307A, 307B, 307C, and 307D correspond to the term "front," and portions 309A and 309D correspond to the term "unlock the front door."

[0057] Referring now to FIG. 3B, an example is shown of utilizing the portion of the audio data shown in FIG. 3A in automatically generating speaker features for a number of different specific TD-SVs.

[0058] 3B, portions 305A-E of the audio data corresponding to "door" are utilized to generate speaker features 306 for a user's particular TD-SV, constrained to "door." More specifically, portion 305A is shown as being processed to generate given speaker features 306A corresponding to portion 305A, portion 305B is shown as being processed to generate given speaker features 306B corresponding to portion 305B, portion 305C is shown as being processed to generate given speaker features 306C corresponding to portion 305C, portion 305D is shown as being processed to generate given speaker features 306D corresponding to portion 305D, and portion 305E is shown as being processed to generate given speaker features 306E corresponding to portion 305E. The processing of portions 305A-E can, for example, use a TD-SV model utilized only for a particular TD-SV, or a TD-SV model utilized for a particular TD-SV and multiple additional (e.g., all) particular TD-SVs. Given speaker features 306A-E can be combined to generate speaker feature 306. For example, speaker feature 306 can be an average of given speaker features 306A-E. Optionally, when averaging or otherwise combining given speaker features 306A-E, given speaker features 306A-E can be weighted based on one or more factors. For example, any outliers can be weighted less (e.g., weights are reduced proportionally to the outer ranges) or even not considered at all (i.e., weighted essentially to "0"). As another example, given speaker feature 306A can be weighted based on a corresponding authentication criterion. For example, a given speaker feature 306A may be weighted more heavily than a given speaker feature 306B based on the given speaker feature 306A being based on spoken utterance 300A and the given speaker feature 306B being based on spoken utterance 300B, such that a greater authentication criterion is determined when spoken utterance 300A is provided compared to the authentication criterion determined when spoken utterance 300B is provided.

[0059] In FIG. 3B , portions 307A-D of audio data corresponding to the “front” are utilized to generate speaker features 306 for a user's specific TD-SV and are constrained to the “front.” More specifically, portion 307A is shown as being processed to generate given speaker features 308A corresponding to portion 307A, portion 307B is shown as being processed to generate given speaker features 308B corresponding to portion 307B, portion 307C is shown as being processed to generate given speaker features 308C corresponding to portion 307C, and portion 307D is shown as being processed to generate given speaker features 308D corresponding to portion 307D. The processing of portions 307A-D can, for example, use a TD-SV model utilized only for the specific TD-SV, or a TD-SV model utilized for the specific TD-SV and multiple additional (e.g., all) specific TD-SVs. The given speaker features 308A-D can be combined to generate speaker features 308. For example, the speaker feature 308 may be the average of the given speaker features 308A-D, optionally discarding any outliers of the given speaker features 308A-D.

[0060] In Figure 3B, portions 303A-D of the audio data corresponding to "front door" are utilized to generate speaker feature 304 for a particular TD-SV of a user, constrained to "front door." In particular, speaker feature 304 is for the particular TD-SV constrained to "front door," while speaker feature 306 is for a separate particular TD-SV constrained to the standalone "door," and speaker feature 308 is for an even separate particular TD-SV constrained to the standalone "front." In Figure 3B, portion 303A is shown as being processed to generate given speaker feature 304A corresponding to portion 303A, portion 303B is shown as being processed to generate given speaker feature 304B corresponding to portion 303B, portion 303C is shown as being processed to generate given speaker feature 304C corresponding to portion 303C, and portion 303D is shown as being processed to generate given speaker feature 304D corresponding to portion 303D. The processing of portions 303A-D can, for example, use a TD-SV model that is utilized only for the particular TD-SV, or a TD-SV model that is utilized for the particular TD-SV and multiple additional (e.g., all) particular TD-SVs. The given speaker features 304A-D can be combined to generate speaker feature 304. For example, speaker feature 304 can be an average of the given speaker features 304A-D, optionally discarding any outliers of the given speaker features 304A-D.

[0061] In FIG. 3B , portions 309A and 309B of audio data corresponding to “Please unlock the front door” are utilized to generate speaker features 310 for a user's particular TD-SV, constrained to “Please unlock the front door.” More specifically, portion 309A is shown as being processed to generate given speaker features 310A corresponding to portion 309A, and portion 309B is shown as being processed to generate given speaker features 310B corresponding to portion 309B. The processing of portions 309A and 309B can, for example, use a TD-SV model utilized only for the particular TD-SV, or a TD-SV model utilized for the particular TD-SV and multiple additional (e.g., all) particular TD-SVs. The given speaker features 310A and 310B can be combined to generate speaker feature 310. For example, speaker feature 310 can be an average of given speaker features 310A and 310B.

[0062] 3B , portions 301A, 301D, and 301E of the audio data corresponding to "unlock" are utilized to generate speaker features 302 for a user's particular TD-SV, constrained to "unlock." More specifically, portion 301A is shown as being processed to generate given speaker features 302A corresponding to portion 301A, portion 301D is shown as being processed to generate given speaker features 302D corresponding to portion 301D, and portion 301E is shown as being processed to generate given speaker features 302E corresponding to portion 301E. The processing of portions 301A, 301D, and 301E can, for example, use a TD-SV model utilized only for the particular TD-SV, or can use a TD-SV model utilized for the particular TD-SV and multiple additional (e.g., all) particular TD-SVs. The given speaker features 302A, 302D, and 302E can be combined to generate the speaker feature 302. For example, the speaker feature 302 can be an average of the given speaker features 302A, 302D, and 302E, optionally discarding any outliers.

[0063] In some implementations, each of the five distinct specific TD-SVs can be associated with a corresponding weighting that is utilized when the specific TD-SV is utilized in authenticating a user. Stated another way, some of the specific TD-SVs can be given greater weighting than others of the specific TD-SVs. In some of these implementations, the weighting given to a specific TD-SV can be a function of the amount of audio data on which the corresponding speaker features are based, the number of terms in the corresponding sequence (e.g., whether constrained to only a single term, only sequences of two terms each), and / or other factors. As an example, the TD-SV for "Door" can have a stronger weighting than the TD-SV for "Front Desk," based on speaker features of the TD-SV for "Door" that are based on at least five portions of audio data, while the speaker features of the TD-SV for "Front Desk" are based on only four portions of audio data. As another example, the TD-SV of "front door" may have a stronger weighting than the TD-SV of "front" based at least on the fact that the TD-SV of "front door" is constrained to a sequence of two terms, while the TD-SV of "front" is restricted to only one term.

[0064] 4 is a flowchart illustrating an example method 400 for authenticating a user using one or more particular text-dependent speaker verification methods, according to various implementations. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. The system may include various components of various computer systems, such as one or more components of automated assistant 100. Furthermore, although the operations of method 400 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.

[0065] In block 452, the system receives audio data capturing the user's spoken utterances via the Assistant device's microphone.

[0066] At block 454, the system performs automatic speech recognition (ASR) on the audio data to generate a recognition of the spoken utterance.

[0067] In block 456, the system determines an assistant action to perform in response to the utterance based on the recognition in block 454. In some implementations, block 456 can include sub-block 457. In sub-block 457, the system performs natural language understanding (NLU) on the recognition to determine intent and parameters. In some of these implementations, the system optionally prompts the first user to clarify the intent and / or parameters (e.g., using visual and / or audible prompts).

[0068] In block 458, the system determines whether and / or to what extent to authenticate the user who spoke the utterance based on the assistant action determined in block 456. For example, if the assistant action is limited to intents and / or parameters that do not require any personal data execution (or do not require any particular type of personal data), the system may determine that assistant authentication is not required. On the other hand, for other types of assistant actions, the system may determine that authentication is required and may optionally determine the extent of the required authentication. For example, the system may determine that some or all smart device control assistant actions require authentication. Furthermore, the system may determine that certain smart device control assistant actions (e.g., unlocking a smart lock and / or controlling kitchen appliances) require a higher level of authentication than other smart device control assistant actions (e.g., turning on the lights). The assistant actions that require authentication and / or the extent of authentication may be stored in the assistant device and / or remote assistant server and may be manually configured for a population of users or configured for each user (e.g., based on input by the user indicating the required authentication).

[0069] If, in block 459, the system determines that authentication is not required, the system can optionally proceed to block 460 and perform the assistant action 460 without authentication.

[0070] If, in block 459, the system determines that authentication is required, the system proceeds to block 462 and selects a portion of the audio data that contains the user's particular TD-SV term. For example, the portion can be selected based on the ASR of block 454, indicating that the portion contains the particular TD-SV term.

[0071] In block 464, the system processes the portion of the audio data in generating speech features for the particular TD-SV and compares the generated speech features with previously generated speaker features for the particular TD-SV (e.g., generated based on method 300 of FIG. 3).

[0072] In some implementations, block 464 includes sub-block 465A and / or sub-block 465B. In sub-block 465A, the system uses a TD-SV model for a particular TD-SV when processing a portion of the audio data. The speech features of the portion of the audio data can then be based on (e.g., closely correspond to) values ​​(e.g., embeddings) from a set of layer activations of the TD-SV model after processing of the audio data. The TD-SV model for a particular TD-SV can be a neural network model and can be utilized only for the particular TD-SV, or alternatively can be used for multiple additional (e.g., all) particular TD-SVs.

[0073] In sub-block 465B, the system generates a distance measure when comparing the generated speech features to previously generated speaker features for a particular TD-SV. For example, the speech features may be embeddings and compared to the TD-SV and the user's speaker features, which are also embeddings. The comparison may include generating a cosine distance measure or other distance metric between the two embeddings.

[0074] In block 466, the system determines whether there are more unprocessed terms that correspond to the user's associated particular TD-SV, and if so, the system performs another iteration of blocks 462 and 464 for another particular TD-SV of the user, using the portion of the audio data that contains those terms.

[0075] If, in block 466, the system determines that there are no more unprocessed terms, the system proceeds to block 468. In block 468, the system determines whether to authenticate the user as a function of the comparisons of one or more iterations of block 464. For example, if there have been three iterations of block 464 for three separate TD-SVs of the user, the system may determine whether to authenticate as a function of all three comparisons. For example, the user's authentication may be contingent on each of the comparisons indicating at least a threshold degree of similarity, and / or may be contingent on an overall comparison indicating at least a threshold degree of similarity, the overall comparison being based on an average or other combination of the three individual comparisons. Optionally, if the three individual comparisons are averaged or otherwise combined, they may each be weighted based on any weighting optionally assigned to the user's corresponding TD-SV, as described herein. In block 468, the system may determine whether to authenticate the user based on the comparisons, and optionally based on any additional verification that may be performed based on the available data. For example, the system may also utilize the TI-SV in determining whether to authenticate in block 468.

[0076] If it is determined to authenticate the user in block 468, the system also performs the assistant action. If the system determines not to authenticate the user, the system can prevent the assistant action from being performed and optionally notify the user of the inability to authenticate. Optionally, if the assistant determines not to authenticate the user based on the comparison and / or other available data, the assistant can prompt the user to provide further input for verification (e.g., prompting for a passcode).

[0077] In some implementations, block 468 includes sub-blocks 469A and / or 469B.

[0078] In sub-block 469A, the system determines whether to authenticate based on the comparisons, based on distance metrics optionally generated from the comparisons. For example, a distance metric can be generated for each of the comparisons, and authentication of the user can be contingent on one or more (e.g., all) of the distance metrics indicating a threshold similarity. As another example, distance metrics can be generated for each of the comparisons, and the system can determine an overall distance metric based on averaging and / or otherwise combining the individual distance metrics to generate an overall distance metric, and authentication of the user can additionally or alternatively be contingent on at least the overall distance metric indicating a threshold similarity. Optionally, and as indicated by sub-block 469B in which each of the comparisons is weighted, each of the distance metrics can be weighted differently when averaging and / or otherwise combined. For example, a first distance metric can be weighted more heavily than a second distance metric based on a first speaker feature utilized in generating the first distance metric, which is based on a greater amount of past user spoken utterances than a speaker feature utilized in generating the second distance metric. As another example, additionally or alternatively, the first distance metric can be weighted more heavily than the second distance metric based on the speech recognition confidence for terms corresponding to a particular TD-SV of the first distance metric, indicating a higher confidence than the speech recognition confidence for a second term of a particular TD-SV of the second distance metric. As another example, additionally or alternatively, the first distance metric can be weighted more heavily than the second distance metric based on the first distance metric generated for a TD-SV for a sequence of terms that includes more terms than the TD-SV utilized in generating the second distance metric.

[0079] 5A, an exemplary spoken utterance 500A of a user, "Please unlock the front door," is shown. Also shown in FIG. 5A are portions 501A of audio data that capture the spoken utterance 500A and correspond to "please unlock," portions 503A of audio data that correspond to "front door," portions 505A of audio data that correspond to "door," portions 507A of that audio data that correspond to "front," and portions 509A of that audio data that correspond to "please unlock the front door." The portions can be determined to correspond to corresponding terms based on ASR of the audio data representing them.

[0080] Also shown in Figure 5A is an example of processing portions 501A, 503A, 505A, 507A, and 509A to generate corresponding speech features 502A, 504A, 506A, and 508A. Also shown in Figure 5A is a first comparison 5A1 of speech feature 502A with corresponding speaker features 302 (Figure 3B) for a user's TD-SV constrained to the term "unlock" to generate a first distance metric 512A, a second comparison 5A2 of speech feature 505A with corresponding speaker features 306 (Figure 3B) for a user's TD-SV constrained to the term "door" to generate a second distance metric 516A, and a third comparison 5A3 of speech feature 505A with corresponding speaker features 306 (Figure 3B) for a user's TD-SV constrained to the term "front" to generate a third distance metric 518A. A third comparison 5A3 is made between the speech feature 508A and the corresponding speaker feature 308 (FIG. 3B) for the user's TD-SV to generate a fourth distance metric 514A; a fourth comparison 5A4 is made between the speech feature 504A and the corresponding speaker feature 302 (FIG. 3B) for the user's TD-SV constrained to the term "unlock" to generate a fourth distance metric 514A; and a fifth comparison 5A5 is made between the speech feature 510A and the corresponding speaker feature 310 (FIG. 3B) for the user's TD-SV constrained to the term "front door" to generate a fifth distance metric 520A.

[0081] 5A illustrates generation and utilization of overlapping TD-SV-based utterance features for an utterance 500A. For example, a TD-SV constrained to "front" overlaps with a TD-SV constrained to "front door," which also overlaps with "door," which overlaps with "front door," which constrains "please unlock the front door" overlaps with all other relevant TD-SVs, and portion 507A overlaps with portion 503A, which overlaps with portion 503A, and portion 509A overlaps with all other portions. However, some implementations may prevent or limit overlapping TD-SVs when determining whether to authenticate a user for an utterance. For example, an implementation may utilize only the set of TD-SVs for an utterance that are non-overlapping with each other and that collectively provide the greatest coverage for the entire query (e.g., maximum word-, character-, and / or syllable-based coverage). If multiple sets of TD-SVs are non-overlapping and provide the same amount of coverage for a query, optionally, fewer TD-SVs and / or sets of TD-SVs may be utilized based on a greater amount of past utterance.

[0082] Referring now to FIG. 5B , an example 500B of a user's spoken utterance, "hey assistant, unlock the basement door," is shown. Also shown in FIG. 5B are a portion 501B of audio data (capturing spoken utterance 500B) corresponding to "unlock," a portion 505B of that audio data corresponding to "door," and a portion 520B of that audio data (or previous audio data) corresponding to "hey assistant." The portions of the audio data can be determined to correspond to the corresponding terms based on an ASR indicating such. In some implementations, "hey assistant" is the invoke wake word of an automated assistant. In some of those implementations, portion 520B can be determined to correspond to "hey assistant" based on portion 520B causing a wake word detector to trigger (e.g., based on output generated using a wake word neural network model based on processing portion 520B).

[0083] Also shown in Figure 5B is an example of processing portions 520B, 501B, and 505B to generate corresponding speech features 521B, 502B, and 506B. Also shown in Figure 5B is a first comparison 5B1 of speech feature 502A with corresponding speaker features 530 of the user's TD-SV, constrained to the term "Hey, Assistant" and optionally other wake words (e.g., "OK, Assistant"), to generate a first distance metric 522B. Optionally, corresponding speaker features 530 can be generated based on utterances provided by the user in response to prompts during an explicit enrollment procedure. Also shown in Figure 5B is a second comparison 5B2 of the speech features 502B with the corresponding speaker features 303 (Figure 3B) for the user's TD-SV constrained to the term "unlock" to generate a second distance metric 512B, and a third comparison 5B3 of the speech features 506B with the corresponding speaker features 306 (Figure 3B) for the user's TD-SV constrained to the term "door" to generate a third distance metric 516B.

[0084] 5B is a decision 590B to authenticate a user as a function of the first distance metric 522B, the second distance metric 512B, and the third distance metric 516B. For example, the decision to authenticate a user may be contingent on at least N (e.g., 2) of the metrics meeting a threshold and / or an average (optionally weighted as described herein) of at least N (e.g., 2) of the metrics meeting a threshold.

[0085] 6 is a block diagram of an example computing device 610 that may optionally be utilized to perform one or more aspects of the techniques described herein. In some implementations, one or more of the client computing devices and / or other components may comprise one or more components of the example computing device 610.

[0086] Computing device 610 typically includes at least one processor 614 that communicates with a number of peripheral devices via a bus subsystem 612. These peripheral devices may include, for example, a storage subsystem 624 including a memory subsystem 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices allow a user to interact with computing device 610. The network interface subsystem 616 provides an interface to external networks and is coupled to corresponding interface devices in other computing devices.

[0087] The user interface input devices 622 may include a keyboard, a pointing device such as a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touchscreen integrated into a display, a voice recognition system, an audio input device such as a microphone, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and methods for inputting information onto the computing device 610 or a communications network.

[0088] The user interface output devices 620 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube ("CRT"), a flat panel device such as a liquid crystal display ("LCD"), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 610 to a user, or to another machine or computing device.

[0089] Storage subsystem 624 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 624 may include logic for performing selected aspects of one or more of the methods described herein and / or for implementing various components shown herein.

[0090] These software modules are typically executed by the processor 614 alone or in combination with other processors. The memory 625 used in the storage subsystem 624 can include several memories, including a main random access memory (“RAM”) 630 for storing instructions and data during program execution, and a read-only memory (“ROM”) 632 in which fixed instructions are stored. The file storage subsystem 626 can provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of a particular implementation may be stored by the file storage subsystem 626 in the storage subsystem 624 or in other machines accessible by the processor 614.

[0091] The bus subsystem 612 provides a mechanism that allows the various components and subsystems of the computing device 610 to communicate with each other as intended. Although the bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0092] The computing device 610 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computing device 610 shown in Figure 6 is intended only as a specific example to illustrate some implementations. Many other configurations of the computing device 610 are possible, having more or fewer components than the computing device shown in Figure 6.

[0093] In situations where the systems described herein collect or may utilize personal information about users (or often referred to herein as “participants”), users may be provided with an opportunity to control whether a program or feature collects user information (e.g., information about the user's social networks, social actions or activities, occupation, user preferences, or the user's current geographic location) or whether and / or how the user receives content from content servers that may be more relevant to the user. Also, certain data may be processed in one or more ways before being stored or used so that personally identifiable information is removed. For example, a user's identity may be processed so that personally identifiable information about the user cannot be determined, or the user's geographic location may be generalized to the location from which the geographic location information is obtained (e.g., to the city, zip code, or state level) so that the user's specific geographic location cannot be determined. Thus, users may control how information about them is collected and / or used.

[0094] In some implementations, a processor-implemented method is provided that includes receiving audio data capturing a user's spoken utterance and performing speech recognition on the audio data to generate a recognition of the spoken utterance. The audio data is detected via one or more microphones of the user's assistant device. The method further includes determining an assistant action conveyed by the spoken utterance based on the speech recognition and performing the assistant action in response to receiving the spoken utterance. The method further includes determining that one or more terms of the recognition correspond to a specific text-dependent speaker verification (TD-SV) of the user and determining that authentication criteria for the user and the spoken utterance meet a threshold. The one or more terms are distinct from any common invocation wake words of the assistant device. The method further includes processing portions of the audio data corresponding to the one or more terms in generating speaker features for the user's specific TD-SV in response to determining that the one or more terms of the recognition correspond to the user's specific TD-SV and in response to determining that the authentication criteria meet a threshold.

[0095] These and other implementations disclosed herein can include one or more of the following features.

[0096] In some implementations, the method further includes, after processing the audio data in generating speaker features for the user's specific TD-SV, receiving additional audio data capturing additional spoken utterances of the user, the additional spoken utterances including one or more terms; processing a given portion of the additional audio data corresponding to the one or more terms in generating speech features for the portion of the additional audio data; comparing the speech features with the speaker features for the user's specific TD-SV; and determining, based on the comparison, whether to authenticate the user for the additional spoken utterances. In some versions of these implementations, the method further includes performing speech recognition on the additional audio data to generate additional recognition of the additional spoken utterances; and determining that one or more terms are included in the additional recognition and correspond to the given portion of the additional audio data. In these versions, processing the given portion of the additional audio data and comparing the speech features with the speaker features for the specific TD-SV is responsive to determining that one or more terms are included in the additional recognition and correspond to the given portion of the additional audio data. In some of these versions, the comparing step includes determining a distance criterion between the utterance features and the speaker features, and optionally determining whether to authenticate the user for the additional spoken utterance based on the comparison, the determining step including determining a threshold value that depends on additional assistant interactions conveyed by the additional spoken utterance, and authenticating the user for the additional spoken utterance only in response to determining that the distance criterion satisfies the threshold value.

[0097] In some implementations, the method further includes generating initial speaker features for the user-specific TD-SV based on one or more previous instances of the audio data, each determined to include one or more terms and each determined to be authenticated for the user, prior to receiving the audio data. In some of these implementations, processing the portion of the audio data in generating speaker features for the user-specific TD-SV can include modifying the initial speaker features based on processing the portion of the audio data.

[0098] In some implementations, the method further includes determining an authentication criterion based on a generic call TD-SV based on processing the audio data or preceding audio data preceding the audio data, fingerprint verification of the user, face verification of the user, and / or analysis of a verification code entered by the user. The generic call TD-SV is for the user and for one or more generic call wake words for the assistant device.

[0099] In some implementations, a processor-implemented method is provided that includes receiving audio data capturing spoken speech of a given user and processing a first portion of the audio data to generate first speech features. The audio data is detected via one or more microphones of the given user's assistant device. The method further includes performing a first comparison of the first speech features with first speaker features for a first text-dependent speaker verification (TD-SV) of the user, where the first TD-SV depends on a first set of one or more terms. The method further includes processing a second portion of the audio data to generate second speech features, where the second portion of the audio data is different from the first portion of the audio data. The method further includes performing a second comparison of the second speech features with second speaker features for a second TD-SV of the user, where the second TD-SV depends on a second set of one or more terms that is different from the first set of one or more terms. The method further includes determining to authenticate the user for the spoken utterance based on both the first comparison and the second comparison. The method further includes performing one or more actions based on the spoken utterance in response to determining to authenticate the user for the spoken utterance.

[0100] These and other implementations disclosed herein can include one or more of the following features.

[0101] In some implementations, the first portion of the audio data captures a generic wake word for the assistant device, and the first set of one or more terms on which the first TD-SV depends is composed of the one or more generic wake words for the assistant device. The generic wake word is one of the one or more generic wake words for the assistant device. In some versions of these implementations, before receiving the audio data, the method further includes: performing an enrollment procedure in which a plurality of utterances of the user are collected in response to one or more prompts to speak a generic wake word; and generating first speaker features in response to the plurality of utterances. In some additional or alternative versions of these implementations, the second portion of the audio data does not overlap with the first portion of the audio data, and the second portion of the audio data does not include any of the one or more generic wake words for the assistant device. In some of these additional or alternative versions, the method includes, before receiving the audio data, generating second speaker characteristics based on a plurality of instances of previous audio data, wherein generating the second speaker characteristics based on the plurality of instances of previous audio data is based on determining that the plurality of instances of previous audio data capture at least some of a second set of one or more terms, and are captured when the user is authenticated.

[0102] In some implementations, the second portion of the audio data does not include any common invocation wake word for the assistant device. In some versions of these implementations, the method further includes, before receiving the audio data, generating second speaker features based on multiple instances of the previous audio data. Generating second speaker features based on multiple instances of the previous audio data is based on determining that the multiple instances of the previous audio data capture at least some of the second set of one or more terms, and are captured when the user is authenticated.

[0103] In some implementations, the method further includes performing speech recognition on the audio data to generate a recognition of the spoken utterance and determining that a second set of one or more terms is included in the recognition and corresponds to a second portion of the audio data. In some of these implementations, processing the second portion of the audio data and performing a second comparison of the second speech features with second speaker features of the second TD-SV is responsive to determining that the second set of one or more terms is included in the recognition and corresponds to the second portion of the audio data.

[0104] In some implementations, processing the first portion of the audio data to generate first speech features includes first processing the first portion of the audio data using a neural network model, the first speech features being based on a first set of activations of the neural network model after the first processing. In some versions of those implementations, processing the second portion of the audio data to generate second speech features includes second processing the second portion of the audio data using a neural network model, the second speech features being based on a second set of activations of the neural network model after the second processing. In some alternative versions of those implementations, processing the second portion of the audio data to generate second speech features includes second processing the second portion of the audio data using a second neural network model, the second speech features being based on a second set of activations of the second neural network model after the second processing.

[0105] In some implementations, performing the first comparison includes determining a first distance criterion between the first speech feature and the first speaker feature, and performing the second comparison includes determining a second distance criterion between the second speech feature and the second speaker feature. In some versions of those implementations, determining to authenticate the user for the spoken utterance based on both the first comparison and the second comparison includes determining to authenticate the user based on an overall criterion based on both the first distance criterion and the second distance criterion. In some of those versions, determining to authenticate the user based on an overall criterion based on both the first distance criterion and the second distance criterion includes determining a threshold value that depends on an assistant interaction conveyed by the spoken utterance, and authenticating the user for the additional spoken utterance only in response to determining that the overall distance criterion satisfies the threshold.

[0106] Further, some implementations may include a system that includes one or more user devices, each user device having one or more processors and a memory operatively coupled to the one or more processors, wherein the memory of the one or more user devices stores instructions that, in response to execution of the instructions by the one or more processors of the one or more user devices, cause the one or more processors to perform any of the methods described herein. Some implementations also include at least one non-transitory computer-readable medium that includes instructions that, in response to execution of the instructions by the one or more processors, cause the one or more processors to perform any of the methods described herein. [Explanation of symbols]

[0107] 100 Automated Assistants 110 Assistant Device 120 Automated Assistant Clients 122 Automatic Speech Recognition (ASR) Engine 124 Natural Language Understanding (NLU) Engine 126 Text-to-Speech (TTS) engine 128 Fulfillment Engine 130 Authentication Engine 132 Comparison Module 134 Speech Feature Module 136 other modules 138 Speaker Feature Module 140 Automated Assistant Components 142 Cloud-based ASR engine 144 Cloud-based NLU engine 146 Cloud-based TTS engines 148 Cloud-based Fulfillment Engine 150 Cloud-based Authentication Engines 152A~N TD-SV model 154 TD-SV feature database 200 ways 300A Spoken Utterances 300B Spoken Utterances 300C Spoken Utterances 300D Spoken Utterances 300E Spoken Utterances 301A part 301D part 301E part 302 Speaker Characteristics 302A Speaker Characteristics 302D Speaker Characteristics 302E Speaker Characteristics 303A part 303B part 303C part 303D part 304 Speaker Characteristics 304A Speaker Characteristics 304B Speaker characteristics 304C Speaker Characteristics 304D Speaker characteristics 305A part 305B part 305C part 305D part 305E part 306 Speaker Characteristics 306A Speaker Characteristics 306B Speaker Characteristics 306C Speaker Characteristics 306D Speaker characteristics 306E Speaker Characteristics 307A part 307B part 307C part 307D part 308 Speaker Characteristics 308A Speaker Characteristics 308B Speaker Characteristics 308C Speaker Characteristics 308D Speaker Characteristics 309A part 309B part 309D part 310 Speaker Characteristics 310A Speaker Characteristics 310B Speaker characteristics 400 ways 500A Spoken Utterances 500B Spoken Utterances 501A Audio data part 501A part 502A Speech Features 502B Speech Features 503A Audio data part 503A part 504A Speech Features 505A Audio data section 505A part 505A Speech Features 506A Speech Features 506B Speech Features 507A Audio data part 507A part 508A Speech Features 509A Audio data part 509A part 510A Speech Features 512A First Distance Metric 512B Second Distance Metric 514A Fourth Distance Metric 516A Second Distance Metric 516B Third Distance Metric 518A Third Distance Metric 520A 5th Distance Metric 520B part 521B Speech Features 522B First Distance Metric 530 Speaker Characteristics 610 Computing Devices 612 Bus Subsystem 614 processor 616 Network Interface Subsystem 620 User Interface Output Device 622 User Interface Input Devices 624 Storage Subsystem 625 Memory Subsystem 626 File Storage Subsystem 630 Main Random Access Memory ("RAM") 632 Read Only Memory ("ROM")

Claims

1. 1. A method implemented by one or more processors, comprising: Receiving audio data capturing spoken utterances of a given user, the audio data being detected via one or more microphones of the given user's assistant device; performing speech recognition on the audio data to generate a recognition of the spoken utterance, the performing natural language understanding on the generated recognition to determine intent and parameters, and prompting the given user to clarify the intent and / or the parameters; processing a first portion of the audio data to generate first speech features; performing a first comparison of the first speech features with first speaker features for a first text-dependent speaker verification (TD-SV) of the user, the first TD-SV depending on a first set of one or more terms; processing a second portion of the audio data to generate second speech features, the second portion of the audio data being different from the first portion of the audio data; performing a second comparison of the second speech features with second speaker features for a second TD-SV of the user, the second TD-SV relying on a second set of one or more terms that is different from the first set of one or more terms; determining to authenticate the user for the spoken utterance based on both the first comparison and the second comparison; In response to determining to authenticate the user for the spoken utterance, performing one or more actions based on the spoken utterance; A method comprising:

2. 2. The method of claim 1, wherein the first portion of the audio data captures a generic wake word for the assistant device, the first set of one or more terms on which the first TD-SV depends is comprised of one or more generic wake words for the assistant device, and the generic wake word is one of the one or more generic wake words for the assistant device.

3. before receiving the audio data, performing an enrollment procedure in which a plurality of utterances of the user are collected in response to one or more prompts to speak the general invocation wake word; generating the first speaker feature in response to the plurality of utterances; The method of claim 2 further comprising:

4. 3. The method of claim 2, wherein the second portion of the audio data does not overlap with the first portion of the audio data, and the second portion of the audio data does not include any of the one or more general invoke wake words of the assistant device.

5. before receiving the audio data, generating the second speaker feature based on a plurality of instances of previous audio data, wherein the generating the second speaker feature based on the plurality of instances of previous audio data comprises: capturing at least some of said second set of one or more terms; further comprising a step based on determining that the user is captured when authenticated. The method of claim 4.

6. The method of any one of claims 1 to 5, wherein the second portion of the audio data does not include any general invocation wake word for the assistant device.

7. before receiving the audio data, generating the second speaker feature based on a plurality of instances of previous audio data, wherein the generating the second speaker feature based on the plurality of instances of previous audio data comprises: capturing at least some of said second set of one or more terms; The method of claim 6 , further comprising the step of determining based on the user being captured when authenticated.

8. The method of claim 7, wherein the second set of one or more terms is included in the recognition and corresponds to the second portion of the audio data. further comprising 8. The method of claim 1, wherein processing the second portion of the audio data and performing the second comparison of the second speech features with the second speaker features of the second TD-SV is responsive to determining that the second set of one or more terms is included in the recognition and corresponds to the second portion of the audio data.

9. processing the first portion of the audio data to generate the first speech features includes first processing the first portion of the audio data using a neural network model, the first speech features being based on a first set of activations of the neural network model after the first processing; 9. The method of claim 1, wherein processing the second portion of the audio data to generate the second speech features comprises second processing the second portion of the audio data using the neural network model, and the second speech features are based on a second set of activations of the neural network model after the second processing.

10. processing the first portion of the audio data to generate the first speech features includes first processing the first portion of the audio data using a first neural network model, the first speech features being based on a first set of activations of the first neural network model after the first processing; 9. The method of claim 1, wherein processing the second portion of the audio data to generate the second speech features comprises second processing the second portion of the audio data using a second neural network model, and the second speech features are based on a second set of activations of the second neural network model after the second processing.

11. performing the first comparison includes determining a first distance measure between the first speech feature and the first speaker feature; 11. The method of claim 1, wherein performing the second comparison comprises determining a second distance measure between the second speech feature and the second speaker feature.

12. determining to authenticate the user for the spoken utterance based on both the first comparison and the second comparison; The method of claim 11 , comprising determining to authenticate the user based on an overall criterion based on both the first distance criterion and the second distance criterion.

13. determining to authenticate the user based on the overall criteria based on both the first distance criterion and the second distance criterion; determining a threshold value dependent on the assistant's interaction conveyed by the spoken utterance; authenticating the user for additional spoken utterances only in response to determining that the overall distance criterion satisfies the threshold; and 13. The method of claim 12, comprising:

14. A computer-readable storage medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 13.

15. A client device comprising one or more processors for performing the method of any one of claims 1 to 13.

Citation Information

Patent Citations

  • Voice recognition system

    JP1999312391A

  • Method, apparatus, and system for constructing a user voiceprint model

    JP2018527609A

  • Speaker Verification

    JP2019530888A

  • Text independent speaker recognition

    WO2020117639A2