Hybrid Multilingual Text-Dependent and Text-Independent Speaker Verification

A hybrid multi-language speaker verification system addresses the challenges of language and dialect variations by using both text-dependent and text-independent models, enhancing accuracy and reducing computational load by selective model invocation.

JP7698058B2Active Publication Date: 2025-06-24GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023558522
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-24
Filing Date
2022-03-09
Publication Date
2025-06-24
Estimated Expiration
2042-03-09

AI Technical Summary

Technical Problem

Existing speaker verification systems face challenges in multi-language environments due to variations in languages, dialects, and accents, which affect the accuracy and efficiency of speaker identification.

Method used

A hybrid multi-language text-dependent and text-independent speaker verification system that processes audio data using both TD-SV and TI-SV models. The system receives audio data containing a hotword followed by a query, generates evaluation vectors, and calculates confidence scores to identify the speaker, with the text-independent model being invoked only when text-dependent confidence scores fail to meet a threshold.

Benefits of technology

The system improves speaker verification accuracy across multiple languages and dialects while reducing computational burden by selectively invoking the more intensive text-independent model only when necessary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007698058000001
    Figure 0007698058000001
  • Figure 0007698058000002
    Figure 0007698058000002
  • Figure 0007698058000003
    Figure 0007698058000003
Patent Text Reader

Abstract

The speaker verification method (400) receives audio data (120) corresponding to an utterance (119). A first portion (121) of the audio data characterizing a predefined hotword is processed to generate a text-dependent evaluation vector (214). One or more text-dependent confidence scores (215) are generated. If one of the text-dependent confidence scores meets a threshold, the operations identify the speaker of the utterance as each enrolled user associated with the text-dependent confidence score that meets the threshold. The operations initiate execution of an action without performing speaker verification. If none of the text-dependent confidence scores meet the threshold, the operations process a second portion (122) of the audio data characterizing a query to generate a text-independent evaluation vector (224). One or more text-independent confidence scores (225) are generated. The operations determine whether the identity of the speaker of the utterance comprises any of the enrolled users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to hybrid multi - language text - dependent and text - independent speaker verification.

Background Art

[0002] In speech - enabled environments such as homes and automobiles, users can access information and control various functions by using voice input. The information and / or functions may be personalized for a given user. Therefore, it may be advantageous to identify a given speaker from among a group of speakers associated with the speech - enabled environment.

[0003] Speaker verification (e.g., voice authentication) provides an easy way for a user of a user device to access the user device. In speaker verification, the user can unlock and access the user device by speaking, without the need to manually enter a passcode (e.g., typing) to access the user device.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, since there are multiple different languages, dialects, accents, etc., there are certain challenges in speaker verification.

Means for Solving the Problems

[0006] One aspect of the present disclosure provides a computer-implemented method for speaker verification that, when executed on data processing hardware, causes the data processing to perform operations including receiving audio data corresponding to an utterance captured by a user device. The utterance includes a predetermined hotword followed by a query specifying an action to be performed. The operations also include processing a first portion of the audio data characterizing the predetermined hotword using a text-dependent speaker verification (TD-SV) model to generate a text-dependent evaluation vector representing the acoustic characteristics of the hotword utterance, and generating one or more text-dependent confidence scores. Each text-dependent confidence score indicates the likelihood that the text-dependent evaluation vector matches one of each of one or more text-dependent reference vectors. Each text-dependent reference vector is associated with one of each of one or more different registered users of the user device. The operations further include determining whether any of the one or more text-dependent confidence scores meet a confidence threshold. If any of the one or more text-dependent confidence scores meet the confidence threshold, the operations include identifying the speaker of the utterance as each registered user associated with the text-dependent reference vector corresponding to the text-dependent confidence score that meets the confidence threshold, and initiating execution of the action specified by the query without performing speaker verification on a second portion of the audio data characterizing the query following the hotword. If none of the one or more text-dependent confidence scores meet the confidence threshold, the operations include providing instructions to a text-independent speaker verifier. The instructions, when received by the text-independent speaker verifier, cause the text-independent speaker verifier to generate a text-independent evaluation vector by processing a second portion of the audio data characterizing the query using a text-independent speaker verification (TI-SV) model. The operations further include generating one or more text-independent confidence scores each indicating the likelihood that the text-independent evaluation vector matches one of each of one or more text-independent reference vectors.Each text-independent reference vector is associated with each one of one or more different registered users of the user device. The operation also includes a step of determining, based on one or more text-dependent confidence scores and one or more text-independent confidence scores, whether the identity of the speaker who made the utterance comprises any one of one or more different registered users of the user device.

[0007] Implementations of the present disclosure may include one or more of the following optional features. In some implementations, each of one or more different registered users of the user device has permissions to access different respective sets of personal resources, and execution of an action specified by a query requires access to each respective set of personal resources associated with each registered user identified as the speaker of the utterance. In some examples, the data processing hardware executes a text-dependent speaker verification TD-SV model and is present on the user device. The text-independent speaker verifier executes a text-independent speaker verification TI-SV model and is present on a distributed computing system that communicates with the user device via a network. In these examples, when none of the one or more text-dependent confidence scores meet a confidence threshold, the step of providing instructions to the text-independent speaker verifier comprises the step of transmitting the instructions and the one or more text-dependent confidence scores from the user device to the distributed computing system.

[0008] In some implementations, the data processing hardware resides on top of one of a user device or a distributed computing system that communicates with the user device via a network. Here, the data processing hardware executes both a text-dependent speaker verification (TD-SV) model and a text-independent speaker verification (TI-SV) model. In some embodiments, the text-independent speaker verification (TI-SV) model is more computationally intensive than the text-dependent speaker verification (TD-SV) model. In some embodiments, the operation further comprises detecting a predetermined hotword in the audio data preceding the query by using a hotword detection model. A first portion of the audio data characterizing the predetermined hotword is extracted by the hotword detection model.

[0009] In some examples, text-dependent speaker verification (TD-SV) models and text-independent speaker verification (TI-SV) models are trained on multiple training datasets. Each training dataset is associated with a different respective language or dialect and includes corresponding training utterances spoken in each respective language or dialect by different speakers. Each corresponding training utterance includes a text-dependent portion that characterizes a predetermined hotword and a text-independent (independent) portion that characterizes a query sentence following the predetermined hotword. Here, the TD-SV model is trained on the text-dependent portion of each corresponding training utterance in each training dataset of the multiple training datasets. The TI-SV model is trained on the text-independent portion of each corresponding training utterance in each training dataset of the multiple training datasets. In these examples, the corresponding training utterances spoken in each respective language or dialect associated with at least one of the training datasets may pronounce a different predetermined hotword than the corresponding training utterances of other training datasets. In some additional examples, the TI-SV model is trained on the text-dependent portion of at least one corresponding training utterance in one or more of the multiple training datasets. Further, or alternatively, the query sentence characterized by the text-independent portion of the training utterance includes variable language content.

[0010] In some implementations, when generating text-independent evaluation vectors, the text-independent speaker verifier processes both a first portion of audio data characterizing a predetermined hotword and a second portion of audio data characterizing a query by using a text-independent speaker verification (TI-SV) model. Additionally or alternatively, each of one or more text-dependent reference vectors may be generated by a text-dependent speaker verification (TD-SV) model in response to receiving one or more previous utterances of a predetermined hotword spoken by one of each of one or more different registered users of the user device. Each of one or more text-independent reference vectors may be generated by a text-independent speaker verification (TI-SV) model in response to receiving one or more previous utterances spoken by one of each of one or more different registered users of the user device.

[0011] Another aspect of the present disclosure provides a system for speaker verification. The system includes data processing hardware and memory hardware communicatively coupled to the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving audio data corresponding to an utterance captured by a user device. The utterance includes a predetermined hotword followed by a query that specifies an action to be performed. The operations also include processing a first portion of the audio data characterizing the predetermined hotword using a text-dependent speaker verification (TD-SV) model to generate a text-dependent evaluation vector representing the acoustic features of the hotword utterance, and generating one or more text-dependent confidence scores. Each text-dependent confidence score indicates the likelihood that the text-dependent evaluation vector matches one of each of one or more text-dependent reference vectors. Each text-dependent reference vector is associated with one of each of one or more different registered users of the user device. The operations further include determining whether any of the one or more text-dependent confidence scores meets a confidence threshold. If any of the text-dependent confidence scores meets the confidence threshold, the operations include identifying the speaker of the utterance as each of the registered users associated with the text-dependent reference vector corresponding to the text-dependent confidence score that meets the confidence threshold, and initiating execution of the action specified by the query without performing speaker verification on a second portion of the audio data characterizing the query following the hotword. If none of the one or more text-dependent confidence scores meets the confidence threshold, the operations include providing instructions to a text-independent speaker verifier. The instructions, when received by the text-independent speaker verifier, cause the text-independent speaker verifier to generate a text-independent evaluation vector by processing a second portion of the audio data characterizing the query using a text-independent speaker verification (TI-SV) model.The operation further comprises generating one or more text-independent confidence scores, each indicating the likelihood that a text-independent evaluation vector matches one of each of one or more text-independent reference vectors. Each text-independent reference vector is associated with one of each of one or more different registered users of the user device. The operation also comprises determining, based on the one or more text-dependent confidence scores and the one or more text-independent confidence scores, whether the identity of the speaker who made the utterance comprises any of one or more different registered users of the user device.

[0012] This aspect can comprise one or more of the following optional features. In some implementations, each of one or more different registered users of the user device has permission to access a different respective set of personal resources. Execution of an action specified by a query requires access to each respective set of personal resources associated with each registered user identified as the speaker of the utterance. In some examples, the data processing hardware executes a text-dependent speaker verification TD-SV model and is present on the user device. The text-independent speaker verifier executes a text-independent speaker verification TI-SV model and is present on a distributed computing system that communicates with the user device via a network. In these examples, when none of the one or more text-dependent confidence scores meet a confidence threshold, the step of providing instructions to the text-independent speaker verifier comprises transmitting, from the user device to the distributed computing system, the instructions and the one or more text-dependent confidence scores.

[0013] In some implementations, the data processing hardware resides on either the user device or a distributed computing system that communicates with the user device via a network. Here, the data processing hardware executes both a text-dependent speaker verification (TD-SV) model and a text-independent speaker verification (TI-SV) model. In some embodiments, the text-independent speaker verification (TI-SV) model is more computationally intensive than the text-dependent speaker verification (TD-SV) model. In some embodiments, the operation further comprises detecting a predetermined hotword in the audio data preceding the query by using a hotword detection model. A first portion of the audio data characterizing the predetermined hotword is extracted by the hotword detection model.

[0014] In some examples, a text-dependent speaker verification (TD-SV) model and a text-independent speaker verification (TI-SV) model are trained on multiple training datasets. Each training dataset is associated with a respective different language or dialect and includes corresponding training utterances spoken in that respective language or dialect by different speakers. Each corresponding training utterance includes a text-dependent portion that characterizes a predetermined hotword and a text-independent portion that characterizes a query sentence following the predetermined hotword. Here, the TD-SV model is trained on the text-dependent portion of each corresponding training utterance in each training dataset of the multiple training datasets. The TI-SV model is trained on the text-independent portion of each corresponding training utterance in each training dataset of the multiple training datasets. In these examples, the corresponding training utterances spoken in each respective language or dialect associated with at least one of the training datasets may pronounce a different predetermined hotword than the corresponding training utterances of other training datasets. In some additional examples, the TI-SV model is trained on the text-dependent portion of at least one corresponding training utterance in one or more of the multiple training datasets. Additionally or alternatively, the query sentence characterized by the text-independent portion of the training utterance includes variable language content.

[0015] In some implementations, when generating text-independent evaluation vectors, the text-independent speaker verifier processes both a first portion of audio data characterizing a predetermined hotword and a second portion of audio data characterizing a query, by using a text-independent speaker verification (TI-SV) model. Additionally or alternatively, each of one or more text-dependent reference vectors is generated by a text-dependent speaker verification (TD-SV) model in response to receiving one or more previous utterances of a predetermined hotword spoken by each of one or more different enrolled users of a user device. Each of one or more text-independent reference vectors may be generated by a text-independent speaker verification (TI-SV) model in response to receiving one or more previous utterances spoken by each of one or more different enrolled users of a user device.

[0016] Details of one or more embodiments of the present disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Best Mode for Carrying Out the Invention

[0018] Like reference numerals in the various drawings indicate like elements. In voice-enabled environments such as homes, cars, workplaces, schools, etc., when a user speaks a query or command, a digital assistant can answer the query or execute the command. Such voice-enabled environments can be implemented using a network of connected microphone devices distributed across various rooms or areas of the environment. Through the network of microphones, a user can query (send a query to) the digital assistant verbally without having a computer or other interface in front of them. In some cases, the voice-enabled environment is associated with multiple registered users (e.g., people living in a household). Such examples can apply when a single device such as a smartphone, smart speaker, smart display, tablet device, smart TV, smart home appliance, vehicle infotainment system, etc. is shared by multiple users. Here, the voice-enabled environment may be used by a limited number of users, such as two to six users, in a voice-enabled home, office, or car. Therefore, it is desirable to determine the identity of the specific user who uttered the query. The process of determining the identity of a specific speaker / user may be referred to as speaker verification, speaker recognition, speaker identification, or voice recognition.

[0019] Using speaker verification, in a multi-user environment, a user can issue a query acting on behalf of a specific user or trigger a personalized response. Speaker verification (e.g., voice authentication) provides an easy way for a user of a user device to access the user device. With speaker verification, the user does not need to manually enter a passcode (e.g., typing) to access the user device, so they can unlock the user device and access it by speaking. However, there are certain challenges with speaker verification because there are multiple different languages, dialects, accents, etc.

[0020] In some scenarios, a user makes a query to a digital assistant that requests access to resources related to the user's personal information and / or resources from a set of personal resources related to the user. For example, a particular user (e.g., a user registered with the digital assistant) may ask the digital assistant "When is my meeting with Matt?" or query the digital assistant "Play my music playlist." Here, the user may be one of one or more registered users who each have permission to access their respective personal resource sets (e.g., calendar, music player, email, messaging, contact list, etc.), while being restricted from accessing the personal resources of other registered users. For example, if both John and Meg are registered users of the digital assistant, the digital assistant not only has to determine which of John and Meg made the utterance "When is my meeting with Matt?" and determine when the meeting with Matt is scheduled by accessing the appropriate registered user's calendar, but also needs to respond with details of the scheduled meeting with Matt. Similarly, since John and Meg have their own music playlists, the digital assistant needs to access the music player and ultimately determine which of John and Meg made the utterance "Play my music playlist" in order to output tracks from the appropriate music playlist as audio.

[0021] In a multi-user voice-responsive environment, to determine which user is speaking, a voice-responsive system may include a speaker verification system (e.g., a speaker identification system or a voice authentication system). In the speaker verification system, two types of models can be used to verify the speaker. For the hot word (keyword, wake word, trigger phrase, etc.) part of the utterance, the system can use one or more text-dependent models. On the other hand, for the remaining part of the utterance that generally characterizes the query, the system can use one or more text-independent models. By combining these two types of models, the verification accuracy of speaker verification can be improved, especially at the initial use of the speaker verification system.

[0022] When one or more specific hot words of an utterance (e.g., "Hey, Google" or "Okay, Google") are spoken, a digital assistant running on a user device may be triggered / activated to process (e.g., through automatic speech recognition (ASR)) and execute the query spoken in the utterance following the specific hot word. A hot word detector running on the user device can detect the presence of a specific hot word in the streaming audio captured by the user device and trigger the user device to wake up from a sleep state and start processing (e.g., automatic speech recognition ASR) for subsequent audio data characterizing the query part of the utterance. The hot word detector can extract a first portion of the audio data characterizing the hot word, which can be used as a basis for performing text-dependent speaker verification. The first portion of the audio data may comprise a fixed-length audio segment of about 500 milliseconds (ms) of audio data.

[0023] Generally, a text-dependent model for verifying a speaker's identity (identity) from a first portion of voice data characterizing the hot word of speech is executed on a voice-responsive device. On the other hand, a text-independent model for identifying a speaker from a second portion of voice data characterizing a query following the hot word is executed on a remote server communicating with the voice-responsive device. The text-dependent model can output respective text-dependent speaker vectors. By comparing this text-dependent speaker vector with one or more reference vectors each associated with one or more different registered users of the user device, a first confidence score corresponding to a first likelihood that the speaker who made the speech corresponds to a specific registered user can be determined. The text-independent model can also output respective text-independent speaker vectors. By comparing this text-independent speaker vector with one or more reference vectors each associated with one or more different registered users, a second confidence score corresponding to a second likelihood that the speaker who made the speech corresponds to a specific registered user can be determined. By combining the first and second confidence scores, it can ultimately be determined whether the speaker who made the speech corresponds to a specific registered user.

[0024] Note that in a speaker verification system, there are challenges in training these text-independent models and text-dependent models on a large scale for a wide range of users spanning multiple different languages and dialects. Specifically, it is difficult and time-consuming to obtain training samples of audio data for training models for each language and dialect individually. For low-resource languages, there is a difficult problem because there are few sufficient training samples of audio data. Furthermore, when using text-independent models and text-dependent models separately for each language, a great deal of human and computational effort is required to maintain and update the models in operation, so it is necessary to train new models for new languages that have not been supported so far. For example, in order to train a new text-dependent model and a text-independent model for an added new language, training samples of audio data with speaker labels for the target language must be available.

[0025] To mitigate issues related to the construction and support of multiple speaker verification systems across multiple different languages, the implementation herein is directed to a multilingual speaker verification system having hybrid multilingual text-dependent speaker verification models and text-independent speaker verification models trained in different languages and dialects. By training each of the text-dependent speaker verification model and the text-independent speaker verification model in multiple languages and dialects, the multilingual speaker verification system can not only generalize to unseen languages not used in training, but also maintain speaker verification performance in both high-resource and low-resource languages used in training. As used herein, the multilingual text-dependent speaker verification model and the text-independent speaker verification model each refer to a single respective model that can be used to accurately verify the identification (identity) of speakers speaking different languages or dialects. That is, neither the text-dependent speaker verification model nor the text-independent speaker verification model is dependent on or limited to utterances in a particular single language or dialect. As a result, rather than using different models for different languages, dialects, and / or accents, a single respective model can be trained for each of the text-dependent speaker verification model and the text-independent speaker verification model.

[0026] By using a combination of a text-dependent speaker verification model and a text-independent speaker verification model, the speaker verification performance / accuracy of the speaker verification system is optimized, while adopting a text-independent speaker verification model increases the computational cost. That is, the text-dependent speaker verification model is generally a lightweight model executed on a user device. On the other hand, the text-independent speaker verification model is more computationally intensive than the text-dependent speaker verification model and requires a larger memory footprint. Therefore, the text-independent speaker verification model is suitable for execution on a remote server. In addition to the increased computational cost incurred by executing the text-independent speaker verification model, the waiting time for executing a query also increases in proportion to the time required for the execution of the computation by both the text-dependent speaker verification model and the text-independent speaker verification model. To not only reduce the overall computational burden but also maintain the optimal speaker verification performance / accuracy of the speaker verification system, the implementation herein further targets a speaker verification triage stage that causes the text-independent speaker verification model to perform text-independent speaker verification only when the text-dependent confidence score related to text-dependent speaker verification does not meet the confidence threshold. Otherwise, when the text-dependent confidence score, which indicates the likelihood that the text-dependent evaluation vector generated by the text-dependent speaker verification TD-SV model matches each text-dependent reference vector, meets the confidence threshold, the triage system can permit the speaker verification system to avoid the need for the text-independent speaker verification model to perform text-independent speaker verification.

[0027] Referring to FIG. 1, in some implementations, an exemplary system 100 in an audio-responsive environment includes a user device (user device, user equipment) 102 associated with one or more users 10. The user device 102 communicates with a remote system 111 via a network 104. The user device 102 can correspond to computing devices such as a mobile phone (cellular phone), a computer (laptop or desktop), a tablet, a smart speaker / display, a smart home appliance, smart headphones, a wearable, a vehicle infotainment system, etc., and includes data processing hardware 103 and memory hardware 107. The user device 102 includes or communicates with one or more microphones 106 for capturing utterances from each user 10. The remote system 111 can be a single computer, multiple computers, or a distributed system (e.g., a cloud environment) having scalable / elastic computing resources 113 (e.g., data processing hardware) and / or storage resources 115 (e.g., memory hardware).

[0028] The user device 102 includes a hot word detector 110 (also referred to as a hot word detection model) configured to detect the presence of a hot word in the streaming audio 118 without performing semantic analysis or speech recognition processing on the streaming audio 118. The user device 102 can include an acoustic feature extractor (not shown) implemented as part of the hot word detector 110 or as a separate component for extracting audio data 120 from the utterance 119. For example, the acoustic feature extractor can receive the streaming audio 118 captured by one or more microphones 106 of the user device 102 corresponding to the utterance 119 uttered by the user 10 and extract the audio data 120. The audio data 120 can include acoustic features such as Mel Frequency Cepstral Coefficients (MFCCs) or filter bank energies calculated over a window of the audio signal. In the illustrated example, the utterance 119 uttered by the user 10 includes "Okay, Google, play my music playlist."

[0029] The hotword detector 110 can receive the voice data 120 and determine whether the utterance 119 includes a specific hotword (e.g., okay, Google) uttered by the user 10. That is, the hotword detector 110 can detect the presence of a hotword (e.g., okay, Google) or one or more variations of the hotword (e.g., hey, Google) in the voice data 120, wake up the user device 102 from a sleep state or a standby state, trigger the automatic speech recognition (ASR) system 180, and be trained to perform speech recognition on the hotword and / or one or more other terms following the hotword, e.g., a voice query that specifies an action to be performed following the hotword. In the illustrated example, the query following the hotword of the utterance 119 captured in streaming audio is "Play my music playlist" which specifies an action to provide a response 160 that includes an audio track from John's music playlist for the user device 10 (and / or one or more specified audio output devices) to play for audible output from the speaker, where the digital assistant has access to a music playlist associated with a specific user (e.g., John) 10. Hotwords are useful for "always-on" systems that may pick up sounds other than the sounds directed at the voice-enabled user device 102. For example, the use of a hotword may help the device 102 to identify when a given utterance 119 is directed at the device 102 as opposed to an utterance directed at another individual present in the environment or background chatter. By doing so, the device 102 can avoid triggering computationally expensive processes (such as speech recognition and semantic interpretation) for sounds or utterances that do not include the hotword. In some examples, the hotword detector 110 is a multilingual hotword detector 110 trained in multiple different languages or dialects.

[0030] System 100 includes a multilingual speaker verification system 200 configured to determine the identity of user 10 speaking utterance 119 by processing voice data 120. The multilingual speaker verification system 200 can determine whether the identified user 10 is an authorized user such that a query is executed (e.g., an action specified by the query is performed) only if the user has been identified as an authorized user. Advantageously, the multilingual speaker verification system 200 enables unlocking and accessing user device 102 by speaking an utterance without requiring the user to manually enter (e.g., via typing) or vocalize a passcode to access the user device 102, or to provide some other verification means (e.g., answer a challenge question, provide biometric data, etc.).

[0031] In some examples, the system 100 operates in a multi-user, voice-enabled environment where a plurality of different users 10, 10a - 10n (FIG. 2) are each registered with the user device 102 and have permission to access each set of personal resources (e.g., calendars, music players, email, messaging, contact lists, etc.) associated with that user. The registered user 10 is restricted from accessing personal resources from each set of personal resources associated with other registered users. Each registered user 10 can have each user profile that links to each set of personal resources associated with that user, and other relevant information associated with that user 10 (e.g., user-specified preference settings). Thus, by using the multilingual speaker verification system 200, it is possible to determine which user is speaking the utterance 119 in the multi-user voice-enabled environment 100. For example, in the illustrated example, both John and Meg could be registered users 10 of the user device 102 (or a digital assistant interface running on the user device), and since the digital assistant knows that John and Meg might each have their own music playlists, it needs to access the music player and determine which of John and Meg spoke the utterance 119 "Okay, Google, play my music playlist" in order to ultimately output the track from the appropriate music playlist as audio. Here, the multilingual speaker verification system 200 processes one or more portions 121, 122 of the voice data 120 corresponding to the utterance 119 to identify that John is the speaker of the utterance 119.

[0032] Continuing to refer to FIG. 1, after the hot word detector 110 detects the presence of a hot word (e.g., okay, Google) in the voice data 120, the text-dependent (TD) verifier 210 of the multi-language speaker verification system 200 receives a first portion 121 of the voice data 120 that characterizes the hot word detected by the hot word detector 110. The hot word detector 110 can extract from the voice data 120 an audio segment corresponding to the first portion 121 of the voice data 120. In some examples, the first portion 121 of the voice data 120 comprises a fixed-length audio segment of sufficient length with the voice characteristics of the spoken hot word or other term / phrase that the hot word detector is trained to detect and the text-dependent TD verifier 210 is trained to perform speaker verification. The text-dependent TD verifier 210 processes the first portion 121 of the voice data 120 by using a text-dependent speaker verification (TD-SV) model 212 (FIG. 2), and is configured to output one or more text-dependent (TD) confidence scores 215 each indicating the likelihood that the hot word characterized by the first portion 121 of the voice data 120 was uttered by one of one or more different registered users 10 of the user device 102. It should be noted that a given hot word when spoken in an utterance serves two purposes: determining whether the user 10 is calling the user device 102 to process a subsequent voice query, and determining the identity of the user 10 who made the utterance. The text-dependent (TD) verifier 210 is configured to execute on the user device 102. Further, as will be described in more detail below, the text-dependent speaker verification (TD-SV) model 212 comprises a lightweight model suitable for storage and execution on the user device.

[0033] To improve speaker verification accuracy, the multilingual speaker verification system 200 can also employ a text-independent (TI) verifier 220 to verify the identity of the user 10 who uttered the utterance 119. The text-independent TI verifier 220 processes a second portion 122 of the audio data 120 that characterizes the query following the hotword by using a text-independent speaker verification (TI-SV) model 222 (Figure 2), and the query characterized by the second portion 122 of the audio data 120 may be configured to output one or more text-independent (TI) confidence scores 225 each indicating the likelihood that the query was uttered by one of one or more different registered users 10 of the user device 102. In the illustrated example, the query characterized by the second portion 122 of the audio data 120 comprises "Play my music playlist" (play_my_music_playlist). In some embodiments, the text-independent speaker verification (TI-SV) model 222 additionally processes the first portion 121 of the audio data 120 such that one or more text-independent TI confidence scores 225 are based on both the first and second portions 121, 122 of the audio data 120. In some implementations, the text-independent (TI) verifier 220 receives one or more text-dependent TD confidence scores 215 output from the text-dependent TD verifier 210, and based on the one or more text-dependent TD confidence scores 215 and the one or more text-independent TI confidence scores 225, determines whether the identity of the speaker who uttered the utterance 119 comprises any of one or more different registered users 10 of the user device 102. For example, the text-independent TI verifier 220 can identify the speaker of the utterance 119 as the registered user 10a John.

[0034] The text-independent TI verifier 220 is more computationally intensive than the text-dependent TD verifier 210, and thus, the text-independent TI verifier 220 incurs a higher computational cost for execution than the text-dependent TD verifier 210. Further, the text-independent TI verifier 220 requires a much larger memory footprint than the text-dependent TD verifier 210. Therefore, the text-independent TI verifier 220 is suitable for execution on the remote system 111. However, the text-independent TI verifier 220 may be executed on the user device 102 in other implementations.

[0035] By combining the text-dependent TD verifier 210 and the text-independent TI verifier 220, the accuracy of speaker verification / identification is improved, while there is a trade-off due to the increased computational cost incurred by performing speaker verification with the text-independent TI verifier 220. In addition to the increased computational cost incurred by running the text-independent TI verifier 220, the waiting time for executing a query also increases in proportion to the time required for the text-independent TI verifier 220 to perform additional computations for the long duration of the voice data 120. Without sacrificing the speaker verification performance / accuracy of the multilingual speaker verification system 200 and to reduce the overall computational burden and shorten the waiting time, the multilingual speaker verification system 200 has an intermediate speaker verification triage stage 205 that enables the multilingual speaker verification system 200 to call the text-independent TI verifier 220 only when none of the one or more text-dependent TD confidence scores 215 output from the text-dependent TD verifier 210 meet the confidence threshold. That is, in a scenario where the speaker verification SV triage stage 205 determines that the text-dependent TD confidence score 215 output from the text-dependent TD verifier 210 meets (YES) the confidence threshold, the multilingual speaker verification system 200 bypasses the speaker verification in the text-independent TI verifier 220 and provides the automatic speech recognition ASR system 180 with a speaker verification SV confirmation 208 that identifies the speaker of the utterance 119 as each registered user 10 associated with the text-dependent TD confidence score 215 for the utterance hotword that meets the confidence threshold. When received by the automatic speech recognition ASR system 180, the speaker verification SV confirmation 208 can instruct the automatic speech recognition ASR system 180 to start executing the action specified by the query without the need to perform speaker verification on the second portion 122 of the voice data 120 that characterizes the query following the hotword.In the illustrated example, the automatic speech recognition (ASR) system 180 includes an automatic speech recognition (ASR) model 182 configured to perform speech recognition on a second portion 122 of the speech data 120 (and optionally, in addition to the second portion 122, a first portion 121 of the speech data 120) that characterizes the query.

[0036] The automatic speech recognition (ASR) system 180 also includes a natural language understanding (NLU) module 184 configured to perform an interpretation of the query on the speech recognition results output by the automatic speech recognition (ASR) model 182. Generally, the natural language understanding (NLU) module 184 can perform semantic analysis on the speech recognition results to identify the action to be performed as specified by the query. In the illustrated example, the natural language understanding (NLU) module 184 can determine that execution of the action specified by the query "Play my music playlist" requires access to each set of personal resources associated with each registered user 10 of the user device 102. Thus, the natural language understanding (NLU) module 184 determines that the parameter required to execute the action specified by the query, namely the user ID, is missing. Thus, the natural language understanding (NLU) module 184 uses the SV confirmation 208 to identify a specific registered user (e.g., John) 10a as the speaker of the utterance 119, and thus begins fulfillment of the query by providing an output instruction 185 to execute the action specified by the query. In the illustrated example, the output instruction 185 can instruct the music streaming service to stream music tracks from the music playlist of the registered user John. The digital assistant interface may provide a response 160 to the query that includes the music track for audible output from the user device 102 and / or one or more other devices communicating with the user device 102.

[0037] Since the text-dependent TD confidence score 215 met the confidence threshold, the natural language understanding NLU module 184 was able to rely on the identification (identity) of the registered user determined by the text-dependent TD verifier 210 without waiting for the text-independent TI verifier 220 to perform additional calculations to identify the registered user. As a result, the natural language understanding NLU module 184 was able to expedite the fulfillment of the query.

[0038] In a scenario where the speaker verification SV triage stage 205 determines that none of the one or more text-dependent TD confidence scores 215 output from the text-dependent TD verifier 210 meet the confidence threshold, the speaker verification SV triage stage 205 passes the one or more text-dependent TD confidence scores 215 to the text-independent TI verifier 220 and can instruct the text-independent TI verifier 220 to perform speaker verification on at least a second portion 122 of the voice data 120 that characterizes the query following the hotword during the utterance 119. The text-independent TI verification unit 220 can perform speaker verification by processing the second portion 122 of the voice data 120 to generate one or more text-independent TI confidence scores 225 each indicating the likelihood that the query was uttered by one of each of one or more different registered users 10 of the user device 102. In some implementations, the text-independent TI verifier 220 combines the generated pairs of text-dependent TD confidence scores 215 and text-independent TI confidence scores 225 associated with each registered user 10 to determine a combined confidence score indicating whether the identity of the speaker who made the utterance comprises each registered user 10. For example, if there are four registered users 10a - 10d of the user device, the text-independent TI verifier 220 combines four separate sets of the generated text-dependent TD confidence scores 215 and text-independent TI confidence scores 225 to generate four combined confidence scores each indicating the likelihood that the utterance 119 was uttered by one of each of the four different registered users 10 of the user device. The registered user associated with the highest combined confidence score may be identified as the speaker of the utterance 119.

[0039] In some examples, the text-independent TI verifier 220 combines the text-dependent TD confidence score 215 and the text-independent TI confidence score 225 by averaging the text-dependent TD confidence score 215 and the text-independent TI confidence score 225. In some examples, the text-independent TI verifier 220 calculates a weighted average of the text-dependent TD confidence score 215 and the text-independent TI confidence score 225 to obtain a combined confidence score. For example, the text-dependent TD confidence score 215 may be weighted more heavily than the text-independent TI confidence score 225. In one example, the text-dependent TD confidence score 215 is multiplied by a weight of 0.75, while the text-independent TI confidence score 225 is multiplied by a weight of 0.25. In other examples, the text-independent TI confidence score 225 is weighted more heavily than the text-dependent TD confidence score 215. In some embodiments, the weighting applied to the text-dependent TD confidence score 215 and the text-independent TI confidence score 225 is dynamic such that the weights applied can change over time. That is, the text-dependent TD confidence score 215 may initially be weighted more heavily to reflect that the text-dependent TD verifier 210 may be associated with higher accuracy compared to the text-independent TI verifier 220. However, over time, the text-independent TI verifier 220 is updated based on subsequent utterances of the user and may ultimately become more accurate than the text-dependent TD verifier 210 for performing speaker verification. As a result, the text-independent TI confidence score 225 output by the text-independent TI verifier 220 may ultimately be weighted more heavily than the text-dependent TD confidence score 215 output by the text-dependent TD verifier 210.

[0040] FIG. 2 is a schematic diagram of the multilingual speaker verification system 200 of FIG. 1. The multilingual speaker verification system 200 includes a text-dependent TD verifier 210 having a multilingual text-dependent speaker verification TD-SV model 212, and a text-independent TI verifier 220 having a multilingual text-independent speaker verification TI-SV model 222. In some implementations, each registered user 10 of the user device 102 has access permissions to different respective sets of personal resources, and the execution of a query characterized by the second portion 122 of the voice data 120 requires access to each respective set of personal resources associated with the registered user 10 identified as the speaker of the utterance 119. Here, each registered user 10 of the user device 102 can perform a voice registration process to obtain respective registered user reference vectors 252, 254 from voice samples of a plurality of registered phrases uttered by the registered user 10. For example, the multilingual text-dependent speaker verification TD-SV model 212 can generate one or more text-dependent (TD) reference vectors 252 from predetermined terms (e.g., hot words) within the registered phrases uttered by each registered user 10 that can be combined, e.g., averaged or accumulated in other ways, to form each TD reference vector 252. Further, the multilingual text-independent speaker verification TI-SV model 222 can generate one or more text-independent TI reference vectors 254 from voice samples of the registered phrases uttered by each registered user that are combined, e.g., averaged, or accumulated in other ways, to form each text-independent (TI) reference vector 254.

[0041] One or more registered users 10 can perform voice registration processing using the user device 102. The microphone 106 captures voice samples of these users speaking registered utterances. The multi-language text-dependent speaker verification TD-SV model 212 and the multi-language text-independent speaker verification TI-SV model 222 generate respective text-dependent TD reference vectors 252 and text-independent TI reference vectors 254 therefrom. Further, one or more of the registered users 10 can register with the user device 102 by providing authorization and authentication qualification information to an existing user account of the user device 102. Here, the existing user account can store the text-dependent TD reference vector 252 and the text-independent TI reference vector 254 obtained from previous voice registration processing performed by each user on another device linked to the user account.

[0042] In some embodiments, the text-dependent TD reference vector 252 of the registered user 10 is extracted from one or more voice samples of a predetermined term such as a hot word (e.g., "Okay, Google") that each registered user 10 speaks to wake up the user device from the sleep state. In some implementations, the text-dependent TD reference vector 252 is generated by the multi-lingual text-dependent speaker verification TD-SV model 212 in response to receiving one or more previous utterances of a predetermined hot word spoken by each registered user 10 of the user device 102. For example, the text-dependent speaker verification TD-SV model 212 can be improved / updated / re-trained by using voice data that characterizes a predetermined hot word detected with high confidence by a hot word detector and that results in a text-dependent (TD) evaluation vector 214 associated with a high confidence score that matches the text-dependent TD reference vector 252 stored for a particular registered user. Further, the text-independent TI reference vector 254 of the registered user 10 may be obtained from one or more voice samples in which each registered user 10 speaks different terms / words and phrases of different lengths. For example, the text-independent TI reference vector 254 may be obtained over time from voice samples obtained from a voice conversation of the user 10 with the user device 102 or other devices linked to the same account. In other words, the text-independent TI reference vector 254 may be generated by the multi-lingual text-independent speaker verification TI-SV model 222 in response to receiving one or more previous utterances spoken by the registered user 10 of the user device 102.

[0043] In some examples, the multilingual speaker verification system 200 resolves the identity of the user 10 who uttered the utterance 119 by using a text-dependent (TD) verifier 210. The text-dependent (TD) verifier 210 first identifies the user 10 who uttered the utterance 119 by extracting a text-dependent (TD) evaluation vector 214 that represents the acoustic characteristics of the utterance of the hotword from a first portion 121 of the audio data 120 that characterizes the predetermined hotword uttered by the user. Here, the text-dependent (TD) verifier 210 can execute a multilingual text-dependent speaker verification (TD-SV) model 212 configured to receive the first portion 121 of the audio data 120 as an input and generate the text-dependent (TD) evaluation vector 214 as an output. The multilingual text-dependent speaker verification (TD-SV) model 212 may be a neural network model (e.g., a first neural network 330) trained under machine or human supervision to output the text-dependent (TD) evaluation vector 214.

[0044] When the text-dependent (TD) evaluation vector 214 is output from the multilingual text-dependent speaker verification (TD-SV) model 212, the text-dependent (TD) verifier 210 determines whether the text-dependent (TD) evaluation vector 214 matches any of the text-dependent (TD) reference vectors 252 stored in the user device 102 (e.g., in the memory hardware 107) for the registered users 10, 10a - 10n of the user device 102. As described above, the multilingual text-dependent speaker verification (TD-SV) model 212 may generate the text-dependent (TD) reference vector 252 of the registered user 10 during the voice registration process. Each text-dependent (TD) reference vector 252 can be used as a reference vector corresponding to a voiceprint or unique identifier that represents the acoustic characteristics of each registered user 10 who speaks the predetermined hotword.

[0045] In some implementations, the text-dependent TD verifier 210 uses a text-dependent (TD) scorer 216 that compares the text-dependent TD evaluation vector 214 with each text-dependent TD reference vector 252 associated with each registered user 10a - 10n of the user device 102. Here, the text-dependent TD scorer 216 can generate a score for each comparison that indicates the likelihood that the utterance 119 corresponds to the identity of each registered user 10. Specifically, the text-dependent TD scorer 216 generates a text-dependent (TD) confidence score 215 for each registered user 10 of the user device 102. In some implementations, the text-dependent TD scorer 216 calculates each cosine distance between the text-dependent TD evaluation vector 214 and each text-dependent TD reference vector 252, and generates the text-dependent TD confidence score 215 for each registered user 10.

[0046] When the text-dependent TD scorer 216 generates a text-dependent TD confidence score 215 that indicates the likelihood that the utterance 119 corresponds to each registered user 10, the speaker verification (SV) triage stage 205 determines whether any of the text-dependent TD confidence scores 215 meet a confidence threshold. In some implementations, the speaker verification SV triage stage 205 determines that the text-dependent TD confidence score 215 meets the confidence threshold. In these implementations, the multilingual speaker verification system 200 bypasses the speaker verification in the text-independent TI verifier 220 and, instead, provides the speaker verification SV confirmation 208 to the automatic speech recognition ASR system 108 that identifies the speaker of the utterance 119 as each registered user 10 associated with the text-dependent TD confidence score 215 that meets the confidence threshold.

[0047] Conversely, if the speaker verification SV triage stage 205 determines that none of the text-dependent TD confidence scores 215 meet the confidence threshold, the speaker verification SV triage stage 205 provides the text-dependent TD confidence scores 215 generated by the text-dependent TD verifier 210 and the command 207 to the text-independent TI verifier 220. Here, when the command 207 is received by the text-independent TI verifier 220, the text-independent TI verifier 220 is made to resolve the identity of the user 10 who uttered the utterance 119. The text-independent TI verifier 220 first identifies the user 10 who uttered the utterance 119 by extracting a text-independent (TI) evaluation vector 224 representing the acoustic features of the utterance 119 from a second portion 122 of the acoustic data 120 that characterizes the query following the predetermined hotword. To generate the text-independent TI evaluation vector 224, the text-independent TI verifier 220 may execute a multilingual text-independent speaker verification TI-SV model 222 configured to receive the second portion 122 of the acoustic data 120 as input and generate the text-independent TI evaluation vector 224 as output. In some implementations, the multilingual text-independent speaker verification TI-SV model 222 receives both the first portion 121 and the second portion 122 of the acoustic data 120 and processes both the first portion 121 and the second portion 122 to generate the text-independent TI evaluation vector 224. In some additional implementations, the text-independent speaker verification TI-SV model 222 can process additional acoustic data following the query portion of the utterance 119. For example, the utterance 119 can include a query such as "Send the next message to mom" and can also include additional acoustic corresponding to the content of the message "I'll be home for dinner". The multilingual text-independent speaker verification TI-SV model 222 can be a neural network model (e.g., the second neural network 340) trained under machine or human supervision to output the text-independent TI evaluation vector 224.

[0048] When the text-independent TI verification vector 224 is output from the multi-language text-independent speaker verification TI-SV model 222, the text-independent TI verifier 220 determines whether the text-independent TI evaluation vector 224 matches any of the text-independent TI reference vectors 254 stored in the user device 102 (e.g., in the memory hardware 107) for different registered users 10, 10a to 10n of the user device 102. As described above, the multi-language text-independent speaker verification TI-SV model 222 may generate a text-independent TI reference vector 254 for the registered user 10 during the voice registration process. Each text-independent TI reference vector 254 can be used as a reference vector corresponding to a voiceprint or unique identifier representing the characteristics of the voice of each registered user 10.

[0049] In some implementations, the text-independent TI verifier 220 uses a scorer 226. The scorer 226 compares the text-independent TI evaluation vector 224 with each text-independent TI reference vector 254 associated with each registered user 10a - 10n of the user device 102. Here, the scorer 226 can generate a score for each comparison that indicates the likelihood that the utterance 119 corresponds to the identity of each registered user 10. Specifically, the scorer 226 generates a text-independent (TI) confidence score 225 for each registered user 10 of the user device 102. In some embodiments, the scorer 226 generates the text-independent TI confidence score 225 for each registered user 10 by calculating the cosine distance between the text-independent TI evaluation vector 224 and each text-independent TI reference vector 254. Further, the scorer 226 determines a combined confidence score by combining the pair of the generated text-dependent TD confidence score 215 and the text-independent TI confidence score 225 for each registered user 10. The combined confidence score indicates whether the speaker (identity, identity) who uttered the utterance 119 includes each registered user 10. As described above with respect to FIG. 1, the weights of the text-dependent TD confidence score 215 and the text-independent TI confidence score 225 used to obtain the combined confidence score may be different and / or may change dynamically over time.

[0050] The text-independent TI verifier 220 can identify the user 10 who uttered the utterance 119 as each registered user associated with the highest combined (composite, combined) confidence score. In these embodiments, the text-independent TI verifier 220 provides a speaker verification SV confirmation 208 that identifies the speaker of the utterance 119 as each registered user 10 associated with the highest combined score to the automatic speech recognition ASR system 108. In some examples, the text-independent TI verifier 220 determines whether the highest combined confidence score meets a threshold and identifies the speaker only if the combined confidence score meets the threshold. Otherwise, the text-independent TI verifier 220 can instruct the user device to speak additional verification utterances and / or answer authentication questions.

[0051] Figure 3 shows an example of a multi - language speaker verification training process 300 for training the multi - language speaker verification system 200. The training process 300 can be executed on the remote system 111 of FIG. 1. The training process 300 acquires a plurality of training data sets 310, 310A to 310N stored in the data storage device 301, and trains each of the text - dependent speaker verification TD - SV model 212 and the text - independent speaker verification TI - SV model 222 on the training data set 310. The data storage 301 may exist on the memory hardware 113 of the remote system 111. Each training data set 310 is associated with a different language or dialect, and includes corresponding training utterances 320, 320Aa to 320Nn spoken in each language or dialect by different speakers. For example, the first training data set 310A may be associated with American English and include corresponding training utterances 320Aa to 320An spoken in English by speakers from the United States. That is, the training utterances 320Aa to 320An of the first training data set 310A are all uttered in American - accented English. On the other hand, the second training data set 310B associated with British English includes corresponding training utterances 320Ba to 320Bn spoken in English by speakers from the UK. Therefore, the training utterances 320Ba to 320Bn of the second training data set 310B are spoken in British - accented English and are thus associated with a different dialect (i.e., British accent) from the training utterances 320Aa to 320An associated with the American - accented dialect. In particular, speakers of British - accented English may pronounce some words differently from speakers of another American - accented English. Figure 3 also shows another training data set 310N associated with the Korean language, which includes corresponding training utterances 320Na to 320Nn spoken by Korean speakers.

[0052] In some implementations, the training process 300 trains the multilingual speaker verification system 200 on at least 12 training data sets, each associated with a different language. In additional implementations, the training process 300 trains the multilingual speaker verification system 200 on training utterances 320 that cover 46 different languages and 63 dialects.

[0053] Each corresponding training utterance 320 includes a text-dependent portion 321 and a text-independent portion 322. The text-dependent portion 321 includes an audio segment (audio clip) that characterizes a predetermined hot word (e.g., "Hey, Google") or a variation of a predetermined hot word (e.g., "Okay, Google") uttered in the training utterance 320. The audio segment associated with the text-dependent portion 321 can include a fixed-length audio segment (e.g., 1,175 milliseconds of audio) represented by a sequence of fixed-length frames having audio features (e.g., 40-dimensional log mel filter bank energy features or mel-frequency cepstral coefficients). Here, since the predetermined hot word and its variations are each made detectable by the hot word detector 110 when spoken in the streaming audio 118, one or more terms following the predetermined hot word or its variation can serve as a trigger for the user device to wake up and start voice recognition. In some examples, the fixed-length audio segment associated with the text-dependent portion 321 of the corresponding training utterance 320 that characterizes the predetermined hot word (or its variation) is extracted by the hot word detector 110.

[0054] The same predetermined hotword may be used in multiple different languages. However, since language characteristics such as accents vary depending on the language or dialect, the pronunciation of the same predetermined hotword or its variations will differ depending on the language or dialect. Notably, the hotword detectors 110 placed in some geographical regions may be trained to detect different predetermined hotwords in streaming audio. Therefore, the text-dependent portion 321 of the corresponding training utterances 320 spoken in the languages or dialects associated with these geographical regions may instead characterize different predetermined hotwords. As will become apparent, the trained (learned) multilingual text-dependent speaker verification TD-SV model 212 can distinguish speakers of different languages or dialects based on a predetermined hotword, a variation of the predetermined hotword, or different hotwords specific to a particular language or geographical region. In an additional implementation, the text-dependent portion 321 of some of the training utterances 320 comprises audio segments as follows, in addition to or instead of a predetermined hotword or a variation of the predetermined hotword. That is, the audio segments characterize other terms / phrases such as custom hotwords or commonly used voice commands (e.g., play, pause, volume up / down, call, message, navigate / directions, etc.).

[0055] The text-independent portion 322 of each training utterance 320 comprises an audio segment that characterizes a query sentence spoken in the training utterance 320 following a predetermined hotword characterized by the text-dependent portion 321. For example, the corresponding training utterance 320 may comprise "Okay, Google, what's the weather outside?" (How about it). The text-dependent portion 321 characterizes the hotwords "Okay, Google". The text-independent portion 322 characterizes the query sentence "what's the weather outside?". The text-dependent portion 321 of each training utterance 320 is phonetically constrained by the same predetermined hotword or a variation thereof. However, the vocabulary of the query sentences characterized by each text-independent (independent) portion 322 is not constrained. That is, the duration and phonemes associated with each query sentence are variable. In particular, the language of the spoken query sentences characterized by the text-dependent portion 321 comprises each language associated with the training dataset 310. For example, the query sentence "what's the weather outside?" (What's the weather outside) spoken in English is translated to "Cuales el clima afuera" (What is the climate outside) when spoken in Spanish. In some examples, the audio segment that characterizes the query sentence of each training utterance 320 comprises a variable time in the range from 0.24 seconds to 1.60 seconds.

[0056] Continuing to refer to FIG. 3, the training process 300 trains the first neural network 330 on the text-dependent portions 321 of the training utterances 320, 320Aa - 320Nn spoken in each respective language or dialect associated with each training dataset 310, 310A - 310N. During training, additional information regarding the text-dependent portion 321 may be provided as input to the first neural network 330. For example, a text-dependent (TD) target 323 such as a text-dependent TD target vector corresponding to the ground truth output label for training the text-dependent speaker verification TD-SV model 212 to learn a prediction method may be provided as input to the first neural network 330 during training using the text-dependent TD portion 321. Thus, one or more utterances of a given hotword from each specific speaker may be paired with a specific text-dependent TD target vector 323.

[0057] The first neural network 330 may include a deep neural network formed from a plurality of long short-term memory (LSTM) layers having a projection layer after each LSTM layer. In some examples, the first neural network uses 128 memory cells and the projection size is equal to 64. The multilingual text-dependent speaker verification TD-SV model 212 comprises a trained version of the first neural network 330. The text-dependent TD evaluation vector 214 and the reference vector 252 generated by the text-dependent speaker verification TD-SV model 212 may include d-vectors having an embedding size equal to the projection size of the last projection layer. The training process may use a generalized end-to-end contrastive loss to train the first neural network 330.

[0058] After training, the first neural network 330 generates a multilingual text-dependent speaker verification TD-SV model 212. The trained multilingual text-dependent speaker verification TD-SV model 212 may be pushed to a plurality of user devices 102 associated with users who speak different languages, dialects, or both, and are dispersed across multiple geographical regions. The user device 102 can store and execute the multilingual text-dependent speaker verification TD-SV model 212 to perform text-dependent speaker verification on an audio (voice) segment that characterizes a predetermined hot word detected by the hot word detector 110 in the streaming audio 118. As described above, even if the same hot word is spoken in different languages or locations, users with different languages, dialects, accents, or locations may pronounce the hot word differently. Such pronunciation variations were often inappropriately attributed as speaker identification characteristics to this pronunciation variation due to language or accent in previous speaker verification models trained for only one language. For example, when these prior models interpret the common characteristics of a regional accent as the main characteristic elements of a particular speaker's voice, the rate of false positives in verification increases. However, in reality, those characteristics are common to all users speaking the same or similar accents. The trained multilingual text-dependent speaker verification TD-SV model 212 of the present disclosure can distinguish one user from other users having the same language, dialect, accent, or location.

[0059] Also, the training process 300 trains the second neural network 340 on the text-independent (TI) portion 322 of the training utterances 320, 320Aa to 320Nn spoken in each respective language or dialect associated with each training dataset 310, 310A to 310N. Here, for the training utterance 320Aa, the training process 300 trains the second neural network on the text-independent TI portion 322 that characterizes the query sentence "What's the weather outside?" (What_ is_ the_ weather_ outside) spoken in American English. Optionally, in addition to the text-independent TI portion 322 of the corresponding training utterance 320, the training process can also train the second neural network 340 on the text-dependent TD portion 321 of at least one corresponding training utterance 320 in one or more of the training datasets 310. For example, by using the above training utterance 320Aa, the training process 300 can train the second neural network 340 on the entire utterance "Okay, Google, what's the weather outside?" During training, additional information regarding the text-independent TI portion 322 may be provided as input to the second neural network 340. For example, a text-independent TI target 324 such as a text-independent TI target vector corresponding to the ground truth output label for training the text-independent speaker verification TI-SV model 222 to learn the prediction method may be provided as input to the second neural network 340 during training using the text-independent TI portion 322. Thus, one or more utterances of a query sentence from each specific speaker may be paired with a specific text-independent TI target vector 324.

[0060] The second neural network 340 may include a deep neural network formed from a plurality of LSTM layers each having a projection layer after each LSTM layer. In some examples, the second neural network uses 384 memory cells and the projection size is equal to 128. The multi-lingual text-independent speaker verification TI-SV model 222 comprises a trained version of the second neural network 340. The text-independent TI evaluation vector 224 and the text-independent TI reference vector 254 generated by the text-independent speaker verification TI-SV model 222 may include d-vectors having an embedding size equal to the projection size of the last projection layer. The training process 300 may use a generalized end-to-end contrast loss to train the first neural network 330. In some examples, the trained multi-lingual text-dependent speaker verification TD-SV model 212 is associated with a small memory footprint (e.g., 235 kiloparameters) suitable for execution on the user device 102. However, the trained multi-lingual text-independent speaker verification TI-SV model 222 is more computationally intensive and has a much larger capacity (e.g., 1.3 million parameters) suitable for execution on a remote system.

[0061] Figure 4 comprises a flowchart of an example arrangement of operations of a hybrid multi-lingual text-dependent and text-independent speaker verification method 400. In operation 402, the method 400 comprises receiving audio data 120 corresponding to an utterance 119 captured by the user device 102. The utterance 119 comprises a predetermined hotword followed by a query specifying an action to perform. In operation 404, the method 400 further comprises generating a text-dependent (TD) evaluation vector 214 representing the acoustic features of the hotword utterance 119 by processing a first portion 121 of the audio data 120 characterizing the predetermined hotword using a text-dependent speaker verification (TD-SV) model 212.

[0062] In operation 406, method 400 comprises generating one or more text-dependent (TD) confidence scores 215. The one or more text-dependent (TD) confidence scores 215 each indicate the likelihood that the text-dependent TD evaluation vector 214 matches one of each of the one or more text-dependent (TD) reference vectors 252. Each text-dependent TD reference vector 252 is associated with one of each of one or more different registered users 10 of the user device 102. Method 400 further comprises, in operation 406, determining whether any of the one or more text-dependent TD confidence scores 215 meet a confidence threshold.

[0063] If one of the text-dependent TD confidence scores 215 meets the confidence threshold, method 400, in operation 408, includes identifying the speaker of utterance 119 as each registered user 10 associated with the text-dependent TD reference vector 252 corresponding to the text-dependent TD confidence score 215 that meets the confidence threshold. Method 400 also, in operation 410, includes initiating execution of the action specified by the query without performing speaker verification on the second portion 122 of the voice data 120 characterizing the query following the hot word. If none of the one or more text-dependent TD confidence scores 215 meet the confidence threshold, method 400, in operation 412, includes providing instructions to the text-independent speaker verifier 220 to generate an text-independent (TI) evaluation vector 224 by processing the second portion 122 of the voice data 120 characterizing the query using the text-independent speaker verification (TI-SV) model 222. In operation 414, method 400 also includes generating one or more text-independent (TI) confidence scores 225 each indicating the likelihood that the text-independent TI evaluation vector 224 matches one of each of the one or more text-independent (TI) reference vectors 254. Each text-independent TI reference vector 254 is associated with each of one or more different registered users 10 of the user device 102. In operation 416, method 400 further includes determining whether the identification (identity) of the speaker who uttered utterance 119 comprises any of one or more different registered users 10 of the user device 102 based on the one or more text-dependent TD confidence scores 215 and the one or more text-independent TI confidence scores 225.

[0064] FIG. 5 is a schematic diagram of an exemplary computing device 500 that may be used to implement the systems and methods described herein. Computing device 500 is intended to represent various forms of digital computers, such as a laptop, desktop, workstation, personal digital assistant, server, blade server, mainframe, and other appropriate computers. The components shown here, their connections and relationships, and their functions are intended to be exemplary only as described in this document, and are not intended to limit the implementation of the claimed invention.

[0065] The computing device 500 includes a processor 510, a memory 520, a storage device (storage device) 530, a high-speed interface / controller 540 connected to the memory 520 and the high-speed expansion port 550, and a low-speed interface / controller 560 connected to the low-speed bus 570 and the storage device 530. Each component 510, 520, 530, 540, 550, and 560 is interconnected using various buses and can be implemented on a common motherboard or in other suitable ways. The processor 510 can process instructions for execution within the computing device 500 that include graphical information for a graphical user interface (GUI) to be displayed on an external input / output device such as a display 580 coupled to the high-speed interface 540 and stored in the memory 520 or the storage device 530. In other embodiments, multiple processors and / or multiple buses may be used as appropriate, along with multiple memories and memory types. Also, multiple computing devices 500 may be interconnected with each other such that each device provides some of the necessary operations (e.g., as a server bank, blade server cluster, or multiprocessor system). The processor 510 may be referred to as data processing hardware 510, which includes the data processing hardware 103 of the user device 102 or the data processing hardware 113 of the remote system 111. The memory 720 may be referred to as memory hardware 720, which includes the memory hardware 107 of the user device 102 or the memory hardware 115 of the remote system 111.

[0066] Memory 520 stores information non-temporarily (non-transitorily) within computing device 500. Memory 520 may be a computer-readable medium, a volatile memory unit(s), or a non-volatile memory unit(s). Non-volatile memory 520 may be a physical device used to store programs (e.g., sequences of instructions) or data (e.g., program state information) used by computing device 500, either temporarily or persistently. Examples of non-volatile memory include, but are not limited to, flash memory, read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., commonly used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks and tapes, etc.

[0067] Storage device (storage device 530) can provide large-capacity storage for computing device 500. In some implementations, storage device 530 is a computer-readable medium. In various different implementations, storage device 530 may be a floppy (registered trademark) disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices with a storage area network or other configured devices. In additional embodiments, a computer program product is embodied in an information carrier. The computer program product comprises instructions that, when executed, perform one or more of the methods as described above. The information carrier is a computer machine-readable medium or machine-readable medium such as memory 520, storage device 530, or memory on processor 510, etc.

[0068] The high-speed controller 540 manages the bandwidth-intensive operations of the computing device 500. On the other hand, the low-speed controller 560 manages the low-bandwidth-intensive operations. Such task assignments are merely exemplary. In some implementations, the high-speed controller 540 is coupled to the memory 520, to the display 580 (e.g., via a graphics processor or accelerator), and to the high-speed expansion port 550, and can accept various expansion cards (not shown). In some implementations, the low-speed controller 560 is coupled to the storage device 530 and the low-speed expansion port 590. The low-speed expansion port 590 includes various communication ports (e.g., USB, Bluetooth®, Ethernet®, Wireless Ethernet®), and is coupled to one or more input / output devices such as a keyboard, a pointing device, a scanner, etc., or to a network device such as a switch or a router via a network adapter.

[0069] The computing device 500 may be implemented in a number of different forms as shown. For example, the computing device 500 may be implemented as a standard server 500a, or multiple times within a group of such servers 500a, as a laptop computer 500b, or as part of a rack server system 500c.

[0070] The various implementations of the systems and techniques described in this specification can be realized in digital electronic circuits and / or optical circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can be of special purpose or general purpose and can include one or more computer programs executable and / or interpretable on a programmable system comprising at least one programmable processor coupled to receive data and instructions from, and to transmit data and instructions to, a memory system, at least one input device, and at least one output device.

[0071] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application", an "app", or a "program". Examples of applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and game applications.

[0072] A non-transitory (non-temporary) memory may be a physical device used to temporarily or permanently store a program (e.g., an instruction sequence) or data (e.g., program state information) used by a computing device. The non-transitory memory may be a volatile and / or non-volatile addressable semiconductor memory. Examples of non-volatile memory include, but are not limited to, flash memory, read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (commonly used for firmware such as a boot program). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), and disks and tapes, etc.

[0073] These computer programs (also known as programs, software, software applications or code) comprise machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine language. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, non-transitory computer-readable medium, apparatus and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives the machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0074] The processes and logical flows described in this specification can be implemented by one or more programmable processors, also known as data processing hardware, executing one or more computer programs to operate on input data and generate output. The processes and logical flows can also be implemented by special-purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). Processors suitable for the execution of a computer program include, by way of example, both general and special purpose microprocessors, and any one or more processors of any kind of digital computer. In general, a processor receives instructions and data from a read only memory or a random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer is also operatively coupled to one or more mass storage devices, such as magnetic disks, magneto-optical disks, optical disks, etc., for storing data, receiving data from them, transferring data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, such as, by way of example, semiconductor memory devices, such as EPROM, EEPROM and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CDROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0075] To provide interaction with a user, one or more aspects of the present disclosure can be implemented on a computer having a display device such as, for example, a CRT (cathode ray tube), an LCD (liquid crystal display) monitor, or a touch screen, for displaying information to the user, and optionally, a keyboard and a pointing device such as, for example, a mouse or a trackball, by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback such as visual feedback, auditory feedback, tactile feedback, etc. The input from the user can be received in any form such as acoustic input, voice input, tactile input, etc. Further, the computer can interact with the user by sending documents to or receiving documents from the devices used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.

[0076] Numerous embodiments have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. Accordingly, other embodiments are also within the scope of the following claims.

Claims

1. A computer-implemented method (400) for speaker verification that, when executed on data processing hardware (510), causes the data processing hardware (510) to perform operations, the operations comprising: Receiving audio data (120) corresponding to an utterance (119) captured by a user device (102), the utterance (119) comprising a predetermined hotword followed by a query specifying an action to be performed; receiving the audio data (120); Processing a first portion (121) of the audio data (120) characterizing the predetermined hotword using a text-dependent speaker verification TD-SV model (212) to generate a text-dependent evaluation vector (214) representing the acoustic features of the utterance (119) of the predetermined hotword; Generating one or more text-dependent confidence scores (215) each indicating the likelihood that the text-dependent evaluation vector (214) matches one of each of one or more text-dependent reference vectors (252), each text-dependent reference vector (252) being associated with one of one or more different registered users (10) of the user device (102); generating the text-dependent confidence scores (215); Determining whether any of the one or more text-dependent confidence scores (215) meets a confidence threshold; wherein the operations further comprise: If any of the text-dependent confidence scores (215) meets the confidence threshold, Identifying the speaker of the utterance (119) as each registered user (10) associated with the text-dependent reference vector (252) corresponding to the text-dependent confidence score (212) that meets the confidence threshold, and Starting execution of the action specified by the query without performing speaker verification on a second portion (122) of the audio data (120) characterizing the query following the predetermined hotword; If none of the one or more text-dependent confidence scores (215) meets the confidence threshold, providing an instruction to a text-independent speaker verifier (220); comprising one of: When the command is received by the text-independent speaker verifier (220), the text-independent speaker verifier (220) processing the second portion (122) of the voice data (120) characterizing the query by using a text-independent speaker verification TI-SV model (222) to generate a text-independent evaluation vector (224); generating one or more text-independent confidence scores (225), each indicating a likelihood that the text-independent evaluation vector (224) matches one of each of one or more text-independent reference vectors (254), wherein each text-independent reference vector (254) is associated with one of one or more different registered users (10) of the user device (102), the step of generating the text-independent confidence score (225); determining, based on one or more of the text-dependent confidence scores (215) and one or more of the text-independent confidence scores (225), whether the identification of the speaker who uttered the utterance (119) comprises any of one or more different registered users (10) of the user device (102); A computer-implemented method (400) for causing to execute The text-dependent speaker verification TD-SV model (212) and the text-independent speaker verification TI-SV model (222) are trained on a plurality of training data sets (310), Each of the training data sets (310) is associated with a different respective language or dialect and comprises corresponding training utterances (320) spoken in each respective language or dialect by different speakers, Each corresponding training utterance (320) comprises a text-dependent portion characterizing the predetermined hot word and a text-independent portion characterizing a query sentence following the predetermined hot word, The text-dependent speaker verification TD-SV model (212) is trained on the text-dependent portion of each corresponding training utterance (320) of each of the training data sets (310) of the plurality of training data sets (310), The text-independent speaker verification TI-SV model (222) is trained on the text-independent part of each corresponding training utterance (320) of each of the plurality of training data sets (310). The text-independent speaker verification TI-SV model (222) is trained on the text-dependent part of at least one corresponding training utterance (320) of one or more of the plurality of training data sets (310). A computer-implemented method (400). **Claim 2** Each of one or more different registered users (10) of the user device (102) has access permission to access each different set of personal resources. Execution of the action specified by the query requires access to each set of personal resources associated with each registered user (10) identified as the speaker of the utterance (119). The computer-implemented method (400) according to claim 1. **Claim 3** The data processing hardware (510) executes the text-dependent speaker verification TD-SV model (212) and is present on the user device (102). The text-independent speaker verifier (220) executes the text-independent speaker verification TI-SV model and is present on a distributed computing system (111) that communicates with the user device (102) via a network. The computer-implemented method (400) according to claim 1 or 2. **Claim 4** If none of the one or more text-dependent confidence scores (215) satisfy the confidence threshold. The step of providing the command to the text-independent speaker verifier (220) comprises transmitting the command and one or more of the text-dependent confidence scores (215) from the user device (102) to the distributed computing system (111). The computer-implemented method (400) according to claim 3. **Claim 5** The data processing hardware (510) is present on one of the user device (102) and a distributed computing system (111) that communicates with the user device (102) via a network. The data processing hardware (510) executes both the text-dependent speaker verification TD-SV model (212) and the text-independent speaker verification TI-SV model (222). The computer-implemented method (400) according to claim 1 or 2.

6. The text-independent speaker verification TI-SV model (222) is more computationally intensive than the text-dependent speaker verification TD-SV model (212). The computer-implemented method (400) according to any one of claims 1 to 5.

7. The operation further includes a step of detecting the predetermined hot word in the voice data (120) preceding the query by using a hot word detection model (110), The first portion (121) of the voice data (120) characterizing the predetermined hot word is extracted by the hot word detection model (110). The computer-implemented method (400) according to any one of claims 1 to 6.

8. Each corresponding training utterance (320) spoken in each language or dialect related to at least one of the training data sets (310) pronounces the predetermined hot word different from the corresponding training utterances (320) of other training data sets (310). The computer-implemented method (400) according to any one of claims 1 to 7.

9. The query sentence characterized by the text-independent portion of the training utterance (320) includes variable language content. The computer-implemented method (400) according to any one of claims 1 to 8.

10. When generating the text-independent evaluation vector (224), the text-independent speaker verifier (220) processes both the first portion (121) of the voice data (120) characterizing the predetermined hot word and the second portion (122) of the voice data (120) characterizing the query by using the text-independent speaker verification TI-SV model (222). The computer-implemented method (400) according to any one of claims 1 to 9.

11. Each of the one or more text-dependent reference vectors (252) is generated by the text-dependent speaker verification TD-SV model (212) in response to receiving one or more previous utterances (119) of the predetermined hotword spoken by one of each of one or more different registered users (10) of the user device (102). The computer-implemented method (400) according to any one of claims 1 to 10.

12. Each of the one or more text-independent reference vectors (254) is generated by the text-independent speaker verification TI-SV model (222) in response to receiving one or more previous utterances (119) spoken by one of each of one or more different registered users (10) of the user device (102). The computer-implemented method (400) according to any one of claims 1 to 11.

13. A system (100), wherein the system (100) comprises: data processing hardware (510); memory hardware (720) communicating with the data processing hardware (510), the memory hardware (720) storing a first instruction that, when executed on the data processing hardware (510), causes the data processing hardware (510) to perform an operation; and the operation comprises: receiving audio data (120) corresponding to an utterance (119) captured by a user device (102), the utterance (119) comprising a predetermined hotword followed by a query specifying an action to be performed, the step of receiving the audio data (120); generating a text-dependent evaluation vector (214) representing the voice characteristics of the utterance (119) of the predetermined hotword by processing a first portion (121) of the audio data (120) characterizing the predetermined hotword by using a text-dependent speaker verification TD-SV model; Generating one or more text-dependent confidence scores (215) each indicating the likelihood that the text-dependent evaluation vector (214) matches one of each of one or more text-dependent reference vectors (252), wherein each of the text-dependent reference vectors (252) is associated with one of each of one or more different registered users (10) of the user device (102), the step of generating the text-dependent confidence scores (215); Determining whether any of the one or more text-dependent confidence scores (215) meets a confidence threshold; Comprising, and the operation further comprises When any of the text-dependent confidence scores (215) meets the confidence threshold Identifying the speaker of the utterance (119) as each registered user (10) associated with the text-dependent reference vector (252) corresponding to the text-dependent confidence score that meets the confidence threshold, and Starting execution of the action specified by the query without performing speaker verification on the second portion (122) of the voice data (120) characterizing the query following the predetermined hot word, or When none of the one or more text-dependent confidence scores (215) meets the confidence threshold, comprising providing a second instruction to a text-independent speaker verifier (220); Comprising one of When the second instruction is received by the text-independent speaker verifier (220), the text-independent speaker verifier (220) Processing the second portion (122) of the voice data (120) characterizing the query by using a text-independent speaker verification TI-SV model to generate a text-independent evaluation vector (224); Generating one or more text-independent confidence scores (225) each indicating the likelihood that the text-independent evaluation vector (224) matches one of each of one or more text-independent reference vectors (254), wherein each of the text-independent reference vectors (254) is associated with one of each of one or more different registered users (10) of the user device (102), the step of generating the text-independent confidence scores (225); Based on one or more of the text-dependent confidence scores (215) and one or more of the text-independent confidence scores (225), determining whether the identification of the speaker who uttered the utterance (119) comprises any one of one or more different registered users (10) of the user device (102); A system (100) for causing the execution of; The text-dependent speaker verification TD-SV model (212) and the text-independent speaker verification TI-SV model (222) are trained on a plurality of training data sets (310), Each of the training data sets (310) is associated with a different respective language or dialect and comprises corresponding training utterances (320) spoken in each respective language or dialect by different speakers, Each corresponding training utterance (320) comprises a text-dependent part characterizing the predetermined hot word and a text-independent part characterizing the query sentence following the predetermined hot word, The text-dependent speaker verification TD-SV model (212) is trained on the text-dependent part of each corresponding training utterance (320) of each of the training data sets (310) among the plurality of training data sets (310), The text-independent speaker verification TI-SV model (222) is trained on the text-independent part of each corresponding training utterance (320) of each of the training data sets (310) among the plurality of training data sets (310), The text-independent speaker verification TI-SV model (212) is trained on the text-dependent part of at least one corresponding training utterance (320) of one or more of the training data sets (310) among the plurality of training data sets (310), System (100).

14. Each of one or more different registered users (10) of the user device (102) has access permission to access different respective sets of personal resources, Execution of the action specified by the query requires access to each respective set of personal resources associated with each registered user (10) identified as the speaker of the utterance (119). The system (100) according to claim 13.

15. The data processing hardware (510) executes the text-dependent speaker verification TD-SV model (212) and is present on the user device (102), and The text-independent speaker verifier (220) executes the text-independent speaker verification TI-SV model and is present on a distributed computing system (111) that communicates with the user device (102) via a network. The system (100) according to claim 13 or 14.

16. When none of the one or more text-dependent confidence scores (215) satisfy the confidence threshold, the step of providing the second command to the text-independent speaker verifier (220) includes transmitting the second command and the one or more text-dependent confidence scores (215) from the user device (102) to the distributed computing system (111). The system (100) according to claim 15.

17. The data processing hardware (510) is present on one of the user device (102) and a distributed computing system (111) that communicates with the user device (102) via a network. The data processing hardware (510) executes both the text-dependent speaker verification TD-SV model (212) and the text-independent speaker verification TI-SV model (222). The system (100) according to claim 13 or 14.

18. The text-independent speaker verification TI-SV model (222) is more computationally intensive than the text-dependent speaker verification TD-SV model (212). The system (100) according to any one of claims 13 to 17.

19. The operation further includes detecting the predetermined hotword in the voice data (120) preceding the query by using a hotword detection model (110), and the first portion (121) of the voice data (120) characterizing the predetermined hotword is extracted by the hotword detection model (110). The system (100) according to any one of claims 13 to 18.

20. Each corresponding training utterance (320) spoken in each language or dialect associated with at least one of the training data sets (310) pronounces the predetermined hot word differently from the corresponding training utterances (320) of other training data sets (310). The system (100) according to any one of claims 13 to 19.

21. The query sentence characterized by the text-independent part of the training utterance (320) comprises variable language content. The system (100) according to any one of claims 13 to 20.

22. When generating the text-independent evaluation vector, the text-independent speaker verifier (220) uses the text-independent speaker verification TI-SV model (222) to process both the first part (121) of the audio data (120) characterizing the predetermined hot word and the second part (122) of the audio data (120) characterizing the query. The system (100) according to any one of claims 13 to 21.

23. Each of the one or more text-dependent reference vectors is generated by the text-dependent speaker verification TD-SV model (212) in response to receiving one or more previous utterances (119) of the predetermined hot word spoken by each of one or more different registered users (10) of the user device (102). The system (100) according to any one of claims 13 to 22.

24. Each of the one or more text-independent reference vectors is generated by the text-independent speaker verification TI-SV model (222) in response (160) to receiving one or more previous utterances (119) spoken by each of the one or more different registered users (10) of the user device (102). The system (100) according to any one of claims 13 to 23.

Citation Information

Patent Citations

  • Speaker Verification

    JP2019530888A

  • Speaker Diarization

    JP2020527739A

  • Neural network for speaker verification

    JP2021006913A

  • Text independent speaker recognition

    WO2020117639A2