Speaker verification using collocation information

Incorporating collocation information and multiple speaker models improves speaker verification accuracy by reducing false acceptance rates in noisy environments, addressing the challenge of distinguishing registered users from impostors.

JP7799665B2Active Publication Date: 2026-01-15GOOGLE LLC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023190911
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2014-07-18
Filing Date
2023-11-08
Publication Date
2026-01-15
Estimated Expiration
2035-05-13

AI Technical Summary

Technical Problem

Speaker verification systems face challenges in distinguishing between registered users and impostors, particularly in environments with multiple potential impostors, leading to high false acceptance rates due to limited information available for decision-making.

Method used

Enhance speaker verification systems by incorporating collocation information from publicly available APIs and utilizing multiple speaker models from co-located devices to normalize scores and improve matching decisions.

Benefits of technology

Reduces false acceptance rates by up to 80% through the use of collocation information and normalized scores, enhancing the accuracy of speaker verification in noisy environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007799665000001
    Figure 0007799665000001
  • Figure 0007799665000002
    Figure 0007799665000002
  • Figure 0007799665000003
    Figure 0007799665000003
Patent Text Reader

Abstract

To provide methods, systems, and apparatus, including computer programs encoded on computer storage media, for identifying a user in a multi-user environment.SOLUTION: One of the methods includes the steps of: receiving, by a first user device, an audio signal encoding an utterance; obtaining, by the first user device, a first speaker model for a first user of the first user device; obtaining, by the first user device, for a second user of a second user device that is co-located with the first user device, a second speaker model for the second user or a second score that indicates a respective likelihood that the utterance was spoken by the second user; and determining, by the first user device, that the utterance was spoken by the first user using (i) the first speaker model and the second speaker model or (ii) the first speaker model and the second score.SELECTED DRAWING: Figure 1A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This specification relates to speaker verification. [Background technology]

[0002] In conversational environments such as a home or automobile, a user may use voice input to access information or control various functions. The information and functions may be personalized for a given user. In many user environments, it may be advantageous to identify a given speaker from a group of speakers. Summary of the Invention [Means for solving the problem]

[0003] This specification relates to enhancing speaker verification systems by providing them with more information. For example, some speaker verification systems typically include continuously listening for predefined phrases to wake up a computing device to perform further processing and / or receive more user input, such as voice commands and queries. Such speaker verification systems can distinguish utterances of predefined phrases from a set of registered users at the device and from unknown, unregistered users. In a typical scenario, a particular computing device will detect any utterance of a predefined phrase spoken by someone located relatively close to the device, such as a group of people in a conference room or other diners at a dining table. In some cases, these people may use a speaker verification system compatible with their device. By leveraging collocation information, the speaker verification system associated with each device can detect whether the utterance was spoken by the registered user of the respective device or by another nearby user, such as an impostor. This information can then be used to improve speaker verification decisions.

[0004] In general, one innovative aspect of the subject matter described herein may be embodied in a method including the following operations: receiving, by a first user device, an audio signal encoding an utterance; acquiring, by the first user device, a first speaker model for a first user of the first user device; acquiring, by the first user device, for a second user of a corresponding second user device co-located with the first user device, a second speaker model for the second user or a second score indicative of the likelihood that the utterance was spoken by the second user, respectively; and determining, by the first user device, that the utterance was spoken by the first user using (i) the first speaker model and the second speaker model, or (ii) the first speaker model and the second score. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method. One or more computer systems may be configured to perform particular operations or behaviors by having software, firmware, hardware, or a combination thereof installed on the system that, during operation, causes the system to perform those operations. One or more computer programs may be configured to perform particular operations or behaviors by containing instructions that, when executed by a data processing device, cause the device to perform the operations.

[0005] In general, one innovative aspect of the subject matter described herein may be embodied in a method including the following operations: receiving, by a first user device, an audio signal encoding an utterance; acquiring, by the first user device, a first speaker model for a first user of the first user device; acquiring, by the first user device, a speaker model for each of the users or a score indicative of the likelihood that the utterance was spoken by each of a plurality of other users of other user devices co-located with the first user device; and determining, by the first user device, that the utterance was spoken by the first user using (i) the first speaker model and the plurality of other speaker models, or (ii) the first speaker model and the plurality of scores. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method. One or more computer systems may be configured to perform particular operations or behaviors by having software, firmware, hardware, or a combination thereof installed on the system that, during operation, causes the system to perform those operations. One or more computer programs may be configured to cause a data processing device to perform particular operations or behaviors by containing instructions that, when executed by the device, cause the device to perform the operations.

[0006] In general, one innovative aspect of the subject matter described herein may be embodied in a method including the following operations: receiving, by a first user device, an audio signal encoding an utterance; determining, by the first user device, a first speaker model for a first user of the first user device; determining, by the first user device, one or more second speaker models stored on the first user device for other people who may be co-located with the first user device; and determining, by the first user device, that the utterance was spoken by the first user using the first and second speaker models. Other embodiments of this aspect include corresponding computer systems, apparatuses, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method. One or more computer systems may be configured to perform particular operations or behaviors by having software, firmware, hardware, or a combination thereof installed on the system that causes the system to perform those operations during operation. One or more computer programs may be configured to cause a data processing device to perform particular operations or actions by containing instructions that, when executed by the data processing device, cause the device to perform the actions.

[0007] In general, one innovative aspect of the subject matter described herein may be embodied in a method including the following operations: receiving, by at least one of the computers, an audio signal encoding an utterance; obtaining, by at least one of the computers for each of a plurality of user devices, an identification of a respective speaker model for each user of each user device; and determining, by at least one of the computers, that the utterance was spoken by a particular user of one of the user devices using the identified speaker model. Other embodiments of this aspect include corresponding computer systems, apparatuses, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method. One or more computer systems may be configured to perform particular operations or behaviors by having software, firmware, hardware, or a combination thereof installed on the system that, during operation, causes the system to perform those operations. One or more computer programs may be configured to cause the device to perform particular operations or behaviors by including instructions that, when executed by a data processing device, cause the device to perform those operations.

[0008] In general, one innovative aspect of the subject matter described herein may be embodied by a method including the following operations: receiving, by a first user device, an audio signal encoding an utterance; obtaining, by the first user device, a first score indicative of a likelihood that the utterance was spoken by a first user of the first user device; obtaining, by the first user device, second scores for a second user of a corresponding second user device co-located with the first user device, each indicative of a likelihood that the utterance was spoken by a second user; determining, by the first user device, a combination of the first score and the second score; normalizing, by the first user device, the first score and the second score using the combination of the first score and the second score; and determining, by the first user device, that the utterance was spoken by the first user using the normalized first score and the normalized second score.

[0009] The above-described and other embodiments may each optionally include one or more of the following features, alone or in combination: Obtaining, by the first user device, a second speaker model for the second user or a second score indicative of each likelihood that the utterance was spoken by the second user for a corresponding second user device co-located with the first user device may include obtaining, by the first user device, a second speaker model for the second user or a second score indicative of each likelihood that the utterance was spoken by the second user for the second user device co-located in a physical area near the physical location of the first user device. The method may include performing an action in response to determining that the utterance was spoken by the first user. The method may include analyzing the audio signal to identify a command included in the utterance and performing an action corresponding to the command. The method may include generating, by the first user device, a first score indicating a likelihood that the utterance was spoken by the first user using a portion of the audio signal and a first speaker model. The method may include comparing the first score to a second score to determine a highest score. Determining that the utterance was spoken by the first user may include determining that the first score is the highest score.

[0010] In some implementations, obtaining, by the first user device, a second speaker model for the second user, or a second score indicative of the likelihood that each utterance was spoken by the second user, for a corresponding second user device located in a physical area near the physical location of the first user device, may include obtaining, by the first user device, the second speaker model, and generating, by the first user device, the second score using a portion of the audio signal and the second speaker model.

[0011] In some implementations, obtaining by the first user device a second speaker model for the second user, or a second score indicating the likelihood that each utterance was spoken by the second user, for a second user of a corresponding second user device located in a physical area near the physical location of the first user device may include determining by the first user device that the second user device is located in a physical area near the physical location of the first user device; determining by the first user device that the first user device has a setting that enables the first user device to access the second speaker model; receiving by the first user device the second speaker model; and generating by the first user device a second score using a portion of the audio signal and the second speaker model. Receiving the second speaker model by the first user device may include identifying, by the first user device, one or more third speaker models stored on the first user device and determining, by the first user device, that a subset of the third speaker models may include the second speaker model. The method may include deleting, by the first user device, any third speaker model not included in the subset of third speaker models from the first user device. Receiving the second speaker model by the first user device may include retrieving, by the first user device, the second speaker model from a memory in the first user device. Generating the second score by the first user device may include generating, by the first user device, the second score using the second speaker model stored on the first user device and a portion of the audio signal without requesting the second speaker model from another user device. Receiving the second speaker model by the first user device may include receiving the second speaker model by the first user device from a server. The second user device may include the second speaker model.Receiving the second speaker model by the first user device may include receiving the second speaker model by the first user device from the second user device.

[0012] In some implementations, obtaining, by the first user device, a second speaker model for the second user or a second score indicating the likelihood that each utterance was spoken by the second user for a corresponding second user device located in a physical area near the physical location of the first user device may include determining, by the first user device, that the second user device is located in a physical area near the physical location of the first user device, and receiving, by the first user device, the second score. Receiving, by the first user device, the second score may include receiving, by the first user device, the second score from the second user device. Receiving, by the first user device, the second score may include receiving, by the first user device, the second score from a server. The method may include determining, by the first user device, a device identifier for the second user device; and providing, by the first user device, the device identifier to a server, wherein in response to providing the identifier to the server, the first user device receives a second score from the server.

[0013] In some implementations, the method may include determining, by the first user device, one or more third speaker models stored in the first user device for other people who may be located in a physical area near the physical location of the first user device, and determining, by the first user device, that the utterance was spoken by the first user using (i) the first speaker model, the second speaker model, and the third speaker model, or (ii) the first speaker model, the second score, and the third speaker model. The method may include generating, by the first user device, a first score indicating a likelihood that the utterance was spoken by the first user using a portion of the audio signal and the first speaker model, generating, by the first user device, a respective third score for each of the third speaker models using each third speaker model and the portion of the audio signal, and comparing, by the first user device, the first score, the second score, and the third score to determine a highest score. The method may include determining, for the third user device, by the first user device, a frequency with which the third user device is located in a physical area near the physical location of the first user device, determining, by the first user device, whether the frequency satisfies a threshold frequency, and storing, by the first user device, a third speaker model for the third user of the third user device in a third speaker model in response to determining that the frequency satisfies the threshold frequency. The method may include receiving, by the first user device, input from the first user identifying the third speaker model, and storing, by the first user device, the third speaker model in the third speaker model in response to receiving the input from the user identifying the third speaker model.

[0014] In some implementations, the method may include receiving, by at least one of the computers for each of the user devices, a respective speaker model from each of the user devices, and retrieving, by at least one of the computers for each of the user devices, a respective speaker model from a memory included in at least one of the computers using each identification information.

[0015] In some implementations, the method may include determining, by the first user device, that the normalized first score satisfies a threshold, and determining that the utterance was spoken by the first user is responsive to determining that the normalized first score satisfies the threshold. The method may include determining, by the first user device, that an average of the first score and the second score does not satisfy the threshold, and determining a combination of the first score and the second score is responsive to determining that the average of the first score and the second score does not satisfy the threshold. The method may include determining, by the first user device, that both the first score and the second score do not satisfy the threshold, and determining a combination of the first score and the second score is responsive to determining that both the first score and the second score do not satisfy the threshold. The method may include determining, by the first user device, that the first score does not satisfy a threshold, and determining a combination of the first score and the second score in response to determining that the first score does not satisfy the threshold.

[0016] The subject matter described herein may be implemented in particular embodiments to achieve one or more of the following advantages: In some implementations, the use of an impostor speaker model may reduce action by a user device in response to utterances spoken by persons other than the user of the user device. In some implementations, the system may reduce false positives by 60 to 80 percent when using an impostor speaker model. In some implementations, the system may normalize the final utterance score using a combination of scores for different collocated speakers.

[0017] The details of one or more embodiments of the subject matter herein are set forth in the accompanying drawings and the detailed description below. Other features, aspects, and advantages of the subject matter will become apparent from the detailed description, the drawings, and the claims. [Brief explanation of the drawings]

[0018] [Figure 1A] 1 illustrates an example environment in which one or more user devices AD analyze audio signals encoding vocalizations. [Figure 1B] 1 illustrates an example environment in which one or more user devices AD analyze audio signals encoding vocalizations. [Figure 1C] 1 illustrates an example environment in which one or more user devices AD analyze audio signals encoding vocalizations. [Figure 2] FIG. 1 illustrates an example of a speaker verification system. [Figure 3] FIG. 10 is a flow diagram of a process for determining whether an utterance has been spoken by a user. [Figure 4] FIG. 1 is a block diagram of a computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION

[0019] Like reference numbers and designations in the various drawings indicate like elements.

[0020] A speaker verification system may typically include a process that continuously listens for predefined phrases to wake up the computing device to perform further processing and / or to receive more user input, such as voice commands and queries. Such a speaker verification system may distinguish utterances of hot words from the set of users enrolled on the device and from unknown, unenrolled users.

[0021] Enrollment refers to whether a user provides sample utterances to the system to generate models that can be used to distinguish him or her from other users, known or unknown. The speaker verification process may include comparing the model generated for a given utterance with models generated for a speaker (or speakers) and deciding whether to accept or reject the utterance based on a similarity threshold.

[0022] Speaker verification systems have applicability in a wide range of areas, particularly with regard to recognition quality and anti-fraud effectiveness, and also with a wide range of performance requirements. For example, a speaker verification system used to unlock a device may have higher requirements to provide a low false acceptance of a fraudster than if the system were used on an already unlocked device in a trusted environment. False acceptance can be advantageously mitigated to a lower false rejection (not recognizing the registered user).

[0023] If the verification system only had information provided by the enrolled speaker to make a decision to accept or reject a given utterance, the verification process would be challenging because the set of unknown, possible impostors would be realistically open. As a result, there would be a higher probability that an utterance from an unknown speaker would exceed the similarity threshold for the enrolled speaker, resulting in a false acceptance. This challenge is particularly important in the case of mobile devices, where the availability of possible impostors in the mobile device's vicinity is increasing and constantly changing.

[0024] Speaker verification systems can be improved by providing these systems with more information. In particular, by utilizing collocation information provided by publicly available APIs that may already exist on mobile devices / platforms, the verification system on each device can detect whether a possible fraudster is nearby. Such information can be used to adjust similarity thresholds and share enrolled speaker models to improve matching decisions. In some examples, the system may normalize scores for one or more speaker models using a combination of scores for collocated speakers. For example, a user device may use speaker models stored on the user device and speaker models received from other user devices to generate each score, determine a combination of these scores, and use this combination to normalize each of these scores.

[0025] For example, a user device may generate low scores for an utterance due to background noise. For example, these scores may decrease proportionally to the background noise. In very noisy conditions, such as a moving vehicle or a crowded restaurant, it may be possible that the score of an utterance from a user of a user device does not meet a threshold, e.g., is below the acceptance threshold, and is erroneously rejected. Normalization of the scores may reduce the noise penalty. For example, since the average of many scores each generated using different speaker models does not meet, e.g., is below, the acceptance threshold, normalization would result in improving each of these scores so that the score for the user of the user device should meet, e.g., be greater than, the acceptance threshold.

[0026] Because such matching systems have access to models of possible impostors, they may be able to successfully reject some utterances (e.g., reduce false acceptance rates) in cases where the impostor utterance obtains a similarity score to a registered user higher than an acceptance threshold. For example, if an utterance has an equal or higher score to one of the models in a collection of "impostors," e.g., generated from collocated users, the system may assume that the utterance is likely from an impostor and reject it. Such an approach may be compatible with various types of speaker models, e.g., i-vectors, d-vectors, etc.

[0027] There may be many ways to determine when devices are co-located in a given geographic area. For example, this information may be derived from one or more of Global Positioning System (GPS), Near Field Communication (NFC), Bluetooth, subsonic audio, and / or other sensors and technologies. In some examples, co-located devices may be virtually associated, for example, when the devices are participating in the same call or video conference. In these examples, the devices or servers may determine co-location using calendar entries, email or text messages, or other "soft" concepts.

[0028] Many users may also be co-located in the same area where not all of the users have corresponding user devices, but some of these user devices include speaker models for these users. For example, if five friends are in one of their living rooms and two of these friends have their own mobile devices, to determine which of these friends spoke a particular utterance, a first mobile device may include speaker models for the three friends who do not have mobile devices, and the first and second mobile devices may use these speaker models as well as the speaker model for the friend who owns a device.

[0029] In a typical implementation, a speaker verification system receives an audio signal encoding an utterance and determines whether a score generated using a speaker model satisfies a threshold score value. If the speaker verification system uses only a single speaker model for a particular user of a particular user device, the speaker verification system may generate a score that satisfies the threshold score value for an utterance spoken by another user (e.g., the user's sibling).

[0030] To increase the accuracy of a speaker verification system, the speaker verification system uses many speaker models, such as one for the user and another for the user's sibling. For example, the speaker verification system generates two scores for an audio signal encoding an utterance: one score for the user and another score for the sibling. The speaker verification system compares these scores, both of which may satisfy a threshold score value, to determine which score is best. The speaker verification system is most likely to generate the highest score using a speaker model for a particular person who spoke an utterance compared to when another person spoke the utterance, such that a speaker model for another person is used to generate the highest score.

[0031] If the speaker verification system determines that the user's score, e.g., generated using a speaker model for the user, is the highest, the particular user device may perform an action in response to the utterance. If the speaker verification system determines that the user's sibling's score, e.g., generated using a speaker model for the user's sibling, is the highest, the particular user device does not take an action.

[0032] The speaker verification system uses other speaker models for other users in a physical area near the particular user device, for example, collocated with the particular user device, or scores received from these other user devices, to determine which score is highest and whether the particular user device performs an action in response to the utterance. The speaker verification system may run on the particular device or on another device, such as a server.

[0033] 1A-1C illustrate an example environment 100 in which one or more user devices AD 102a-d analyze audio signals encoding utterances. The user devices AD 102a-d may use one of many different algorithms to determine whether an utterance is likely to have been spoken by a respective user of the user device and whether the user device should perform an action in response to the utterance, or whether the utterance is not likely to have been spoken by a respective user and the user device should not take an action.

[0034] For example, four coworkers may be in a conference room, and a first coworker, e.g., user D, may issue the command, "Okay Google, start the demo." User device A 102a may analyze the audio signal using many speaker models, including, for example, speaker model A 104a for user A and other speaker models for other users of user device A 102a that are occasionally or often in the same physical area as user device A 102a. The other speaker models may be stored in the memory of user device A 102a for a short period of time, such as when user device A 102a recently requested a particular speaker model from another user device B-D 102b-d, or for a longer period of time, such as when there is a high probability that other users are in the same physical area as user device A 102a.

[0035] User device A 102a determines a score for each of the speaker models and determines a highest score from the many scores. User device A 102a may, for example, compare the highest score to a threshold score value to determine whether the highest score satisfies the threshold score value and that the highest score is likely related to user A of user device A 102a. If the highest score does not satisfy the threshold score value, user device A 102a may take no further action and, for example, determine that the utterance was made by a user for whom user device A 102a does not have a speaker model.

[0036] If the user device A 102a determines that the highest score is for user A of the user device A 102a, e.g., the first coworker who issued the command is user A, the user device A 102a performs an action in response to receiving the audio signal. For example, the user device A 102a may launch the requested demo.

[0037] If the highest score is not for user A and user device A 102a determines that the first coworker is not user A, user device A 102a may take no further action with respect to the audio signal. For example, user device A 102a may receive another audio signal with another utterance spoken by the first coworker and take no action in response to the other utterance.

[0038] In some examples, if user devices A-D 102a-d include the same or compatible speaker verification systems, each of user devices A-D 102a-d may share information about each user, such as a speaker model, or information about an analysis of an audio signal encoding an utterance, such as a score. For example, as illustrated in FIG. 1A, a first colleague, such as user D, may utter the utterance 106, "Okay Google, start the demo." A microphone in each of user devices A-D 102a-d may then capture a signal representing this utterance and encode this utterance in an audio signal.

[0039] Each of user devices A-D 102a-d analyzes its respective audio signal using a corresponding speaker model A-D 104a-d to generate a score representing the likelihood that user A-D of each of the user devices uttered the utterance 106, as shown in Figure 1B. In this example, user device A 102a generates a score of 0.76 for user A, user device B 102b generates a score of 0.23 for user B, user device C 102c generates a score of 0.67 for user C, and user device D 102d generates a score of 0.85 for user D.

[0040] Each of the user devices A-D 102a-d shares its score with other user devices. For example, the user devices A-D 102a-d may use one or more sensors, such as GPS, NFC, Bluetooth, subsonic audio, or any other suitable technology, to determine other user devices physically located in an area near each user device. The user devices A-D 102a-d may determine an access setting indicating whether the user device can share its score with another user device and may use this score, for example, to determine whether the other user devices use the same speaker verification system, or both.

[0041] Each of user devices A-D 102a-d compares all of their scores with each other to determine whether the score generated by each user device is the highest score and whether each user device should perform an action in response to the utterance 106. For example, as shown in FIG. 1C , user device D 102d determines that the score generated using speaker model D 104d for user D of user device D 102d is the highest and that the likelihood that the utterance 106 was spoken by user D is greater than the likelihood that the utterance 106 was spoken by another user for the other user devices A-C 102a-c. User device D 102d may perform an action consistent with the utterance 106, such as launching the requested demo 108. User device D 102d may compare the highest score to a threshold score value to ensure that the utterance is likely spoken by user D and not, for example, by another user for whom user device D 102d did not receive a score.

[0042] Similarly, each of the other user devices A-C 102a-c determines that its respective score is not the highest and that it should not take any action. Before determining that its respective score is not the highest, each of the other user devices A-C 102a-c may compare the highest score, for example, to a threshold score value specific to its respective user device, to ensure that there is at least a minimum similarity between the utterance and one of the speaker models and that the utterance is not spoken by another user for whom the other user devices A-C 102a-c do not have their respective speaker model. When the highest score is received from another user device, the other user devices A-C 102a-c may or may not know information about the user, user device, or both corresponding to the highest score. For example, each of the user devices A-D 102a-d may transmit the score to the other user device without any identification information of the user or user device, for example. In some examples, the user device may transmit the score along with an identifier for the user for whom the score was generated.

[0043] 2 is an example of a speaker verification system 200. One or more user devices A-B 202a-b or a server 204 may analyze audio signals encoding an utterance, such as data representing characteristics of the utterance, to determine the user who most likely spoke the utterance. The user devices A-B 202a-b, the server 204, or a combination of multiple of these devices may analyze the audio signals using a speaker model and compare the analysis of the audio signals determined using the speaker model to another analysis of the audio signal to determine whether a particular user spoke the utterance.

[0044] For example, each of user devices A-B 202a-b includes a speaker model A-B 206a-b for the respective user. The speaker models A-B 206a-b may be generated for a particular user using any suitable method, such as having each user speak an enrollment phrase and then extracting Mel-Frequency Cepstral Coefficient (MFCC) features from keyword samples and using these features as a basis for future comparisons, and / or training a neural network using representations of utterances spoken by the particular user.

[0045] Speaker verification module A 208a uses speaker model A 206a for user A of user device A 202a to determine the likelihood that a particular utterance was spoken by user A. For example, speaker verification module A 208a receives an audio signal encoding a particular utterance, such as a representation of the audio signal, and uses speaker model A 206a to generate a score representing the likelihood that the particular utterance was spoken by user A.

[0046] Speaker verification module A 208a may use one or more imposter speaker models 210a stored in user device A 202a to generate, for each of the imposter speaker models 210a, a score that represents the likelihood that a particular utterance was spoken by each user corresponding to the particular imposter speaker model. For example, user device A 202a may receive an audio signal, determine that user device B 202b is located in a physical area near the physical location of user device A 202a, such as in the same room, and request a speaker model for the user of user device B 202b from user device B 202b, such as speaker model B 206b, or from server 204. For example, user device A may send a device identifier for user device B 202b or an identifier for user B to, for example, server 204, as part of the request for speaker model B 206b. User device A 202a stores speaker model B 206b in memory as one of the impostor speaker models 210a, and the speaker verification module 208a generates a score for each of the impostor speaker models 210a.

[0047] The impostor speaker models 210a may include speaker models for other users who may be in a physical area near the physical location of user device A 202a, such as the same room, hallway, or part of a sidewalk or corridor, etc. The impostor speaker models may include speaker models for users who are frequently in the same physical area as user A or user device A 202a, for example, as determined using historical data. For example, user device A 202a may determine that another user device, such as user device C, is in the same physical area as user device A 202a for approximately four hours each workday, this four-hour daily duration being longer than, for example, a three-hour daily threshold duration characteristic of a workday, an average daily duration, etc., and that speaker model C for user C of user device C should be stored in impostor speaker model 210a, e.g., until user A requests removal of speaker model C from impostor speaker model 210a or until the daily duration for user device C no longer satisfies the threshold duration. This frequency may be a particular value, such as four hours each day, or may be a percentage, such as, for example, five percent of the time that user device A 202a detects the particular other user device, or ten percent of the total other user devices detected by user device A 202a are the particular other user device, to name a few.

[0048] In some examples, user A may identify one or more speaker models that user device A 202a should include in impostor speaker model 210a. For example, user device A 202a may receive input to train another speaker model at user device A 202a of a family member or friend of user A. This input may indicate that the other speaker model should be the impostor speaker model, and may be a speaker model for a user other than user A who is not the user of user device A 202a, for example. The other speaker model may be for another user who is frequent in the physical area around user device A 202a, such as a child of user A, in order to reduce or pare down actions performed by user device A 202a in response to utterances spoken by the other user, unless user device A 202a is programmed otherwise.

[0049] For example, if the speaker verification module 208a generates a first score using speaker model A 206a and a respective second score for each of the impostor speaker models 210a, the speaker verification module 208a compares these scores to determine the highest score. If the highest score is generated using speaker model A 206a, the speaker verification module 208a determines that user A spoke a particular utterance and that user device A 202a can take appropriate action, such as, for example, the voice recognition module 212a can analyze the particular utterance to identify a command contained in the particular utterance.

[0050] In one example, one of the impostor speaker models may be for user A's brother, for example, if both brothers have similar voices. The speaker verification module 208a generates a first score for user A and a second score for his brother by analyzing an utterance spoken by one of the brothers using each speaker model. The speaker verification module 208a compares these two scores to determine which score is greater. Each of these is greater than a threshold score, e.g., would trigger an action by user device A 202a solely due to the similarity in the speaker models. If the first score for user A is greater than the second score, user device A 202a performs an action based on the utterance; for example, this action may be determined in part using the voice recognition module 212a. If the second score for user A's brother is greater than the first score, user device A 202a takes no further action, e.g., does not perform an action in response to the particular utterance.

[0051] Some of the impostor speaker models 210a may be used during a particular time of day, a particular day, a particular location, or a combination thereof. For example, if user device A 202a is in user A's family home, user device A 202a may use impostor speaker models for people living in the family home. And, for example, if a co-located user device for one of these people is not detected, these impostor speaker models may not be used.

[0052] In some examples, user devices A-B 202a-b may use settings 214a-b stored in memory to determine whether their respective speaker models, or scores generated using their respective speaker models, can be provided to other user devices using a wireless communication channel 216, such as one generated using near field communication. For example, user device A 202a may receive a particular utterance, determine that user device B 202b is in a physical area near user device A 202a, and request a speaker model from user device B 202b, such as speaker model B 202b, without knowing the particular speaker model being requested. User device B 202b receives this request and analyzes configuration B 214b to determine whether speaker model B 206b can be shared with another device or with the particular user device A 202a, and in response to determining that user device B 202b can share speaker model B 206b, user device B 202b transmits a copy of speaker model B 206b to user device A 202a using the wireless communication channel 216.

[0053] User device A 202a may request a speaker model for user B of user device B 202b, or for all users of user device B 202b, for example, in examples where multiple people may operate a single user device. Speaker model A 206a may include many speaker models in examples where multiple people operate user device A 202a. In these examples, speaker verification module 208a may generate a score for each of the users of user device A 202a, compare these scores with other scores generated using impostor speaker model 210a, and determine the highest score. If the highest score is for one of the users of user device A 202a, user device A 202a may perform an appropriate action, determined at least in part using, for example, speech recognition module 212a.

[0054] The determination of whether to perform an action may be made using a particular type of action, a particular user of user device A 202a, or both. For example, a first user A may have permission to launch any application on user device A 202a, while a second user A may have permission to launch only education applications on user device A 202a.

[0055] In some implementations, one or more of the speaker models are stored on the server 204 instead of or in addition to user device A 202a-b. For example, the server 204 may store speaker models 218 for users A-B of user devices A-B 202a-b. In these examples, user device A 202a or user device B 202b may receive an audio signal encoding an utterance and provide the audio signal, or a portion of the audio signal, such as, for example, a representation of a portion of the audio signal, to the server 204. The server 204 receives an identifier of the user device, speaker model, or user of the user device and, for example, uses the speaker identifier 220 to determine which of the speaker models 218 corresponds to the received identifier.

[0056] In some examples, the server 204, when analyzing a portion of the audio signal, receives identifiers for other speaker models that will be used in addition to the speaker model of the user device. For example, if user device A 202a determines that user device B 202b is physically located in an area near the physical location of user device A 202a, the server 204 may receive the audio signal and identifiers of user devices A-B 202a-b from user device A 202a along with a speaker verification request.

[0057] The server 204 may receive location information from the user device, e.g., along with the audio signal or separately, and may use the location information for the other user devices to determine other user devices that are physically located in an area near the physical location of the user device that provided the audio signal to the server 204. The server 204 may then identify other speaker models 218 for the determined other devices. The server 204 may use the identified other speaker models when generating scores at the server 204 or when providing speaker models to the user devices A-B 202a-b.

[0058] The speaker verification module 222 on the server 204 uses all of the speaker models from the user device that provided the audio signal to the server 204 and the determined other user devices to generate respective scores, each representing the likelihood that each person spoke the particular utterance encoded in the audio signal. The speaker verification module 222 may retrieve the speaker models from memory included in the server 204. The speaker verification module 222 may receive the speaker models from each user device. The server 204 or the speaker verification module 222 determines the highest score and provides a message to each user device indicating that the user of that user device most likely spoke the particular utterance. The server 204 may provide a message to the other user devices indicating that the corresponding other user likely did not speak the utterance.

[0059] In some examples, a particular user device may provide many speaker identifiers to the server 204, such as, for example, for each user of the particular user device, for each imposter speaker model associated with the particular user device, or both. The particular user device may include data indicating the type of model for each of the speaker identifiers, such as user or imposter. The speaker verification module 222 may analyze the audio signal using all of the speaker models 218 corresponding to the received speaker identifiers and determine which speaker model is used to generate the highest score. If the highest score is generated using a model for one of the users of the particular user device, the server 204 provides a message to the particular user device indicating that the user of the particular device most likely spoke the particular utterance. This message may include the speaker identifier for the particular speaker model used to generate the highest score.

[0060] In some implementations, a lower number may represent a higher likelihood that a particular user spoke the utterance compared to a higher number, e.g., a lower number may be a higher score than a higher number.

[0061] In some examples, if a user device has many users, the user device or server 204 may determine a particular speaker model for the current user of the user device. For example, the user device may provide a speaker identifier for the current user to server 204 and indicate that all of the other speaker identifiers for other users of the user device are for imposter speaker models stored on server 204. In some examples, the user device uses the speaker model of the current user to determine whether to perform an action in response to receiving the audio signal, and uses the speaker models for the other users of the user device as the imposter speaker models. The user device may use any suitable method for determining the current user of the user device, such as using a password, a username, or both, to unlock the user device and determine the current user.

[0062] In some implementations, a score is generated for an audio signal using an imposter speaker model or a model received from another user device, and if the score is equal to or greater than a score generated using a speaker model for a user of the particular user device, the particular user device does not take any action in response to receiving the audio signal. In these implementations, if the two scores are the same, the user device does not take any action in response to receiving the audio signal. In other implementations, if two scores for two users of different user devices are the same and are both the highest scores, both user devices corresponding to the two scores may take an action. In implementations where two scores for a model are the same highest score at a single user device, the user device may or may not take an action. For example, if each of the two scores is for a user of a different user device, the user device may take an action. If one of the scores is for a user speaker model and one of the scores is for an imposter speaker model, the user device may not take any action.

[0063] In some implementations, the user device may adjust the threshold depending on the amount of other user devices detected. For example, if no other devices are detected, the threshold may be less restrictive, and if other user devices are detected, e.g., after receiving an audio signal, the threshold may become more restrictive based on the number of other devices detected, e.g., linearly or exponentially, until a maximum threshold is reached. In some examples, one or more scores may be normalized, e.g., using a combination of scores for the same utterance generated using different similarity models. This combination may be an average, sum, or product.

[0064] In some implementations, one or more of user devices A-B 202a-b may periodically detect other user devices in a physical area near each user device. For example, user device B 202b may determine whether another user device is in the same room as user device B 202b every 5 minutes, 10 minutes, or 30 minutes. In some examples, user device B 202b may determine whether another user device is within a predetermined distance from user device B 202b upon determining that user device B 202b has remained in substantially the same area for a predetermined period of time, e.g., user B of user device B 202b is holding user device B 202b but not walking, or user B is remaining in a single room.

[0065] The user devices A-B 202a-b may include personal computers, e.g., smartphones or tablets, and other devices that may be able to send and receive data over the network 224, e.g., mobile communication devices such as wearable devices like watches or thermometers, televisions, and network-connected appliances. The network 224, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof, connects the user devices A-B 202a-b and the server 204.

[0066] 3 is a flow diagram of a process 300 for determining whether an utterance was spoken by a user. For example, the process 300 can be used by the user device A 202a or the server 204 from the speaker verification system 200.

[0067] The process receives an audio signal encoding an utterance 302. For example, a microphone on a user device receives the audio signal and provides the audio signal to a speaker verification module on a first user device or to a server.

[0068] The process obtains a first speaker model for the first user of the first user device (304). For example, the speaker verification module determines that there is a single first user for the first user device and obtains a first speaker model for the first user. In some examples, the speaker verification module determines a current user of the first user device who is currently logged in to the first user device or who most recently logged in to the first user device if the first user device is in a locked state, and obtains a first speaker model for that user.

[0069] In some examples, the speaker verification module determines that there are many users for the first user device and obtains a first speaker model for one of these users. The first user device may then repeat one or more steps in process 300 for other users. For example, the speaker verification module may repeat steps 304 and 306 for each of the users.

[0070] The process uses the portion of the audio signal and the first speaker model to generate a first score indicative of the likelihood that the utterance was spoken by the first user (306). For example, a speaker verification module of the first device uses the entire audio signal and the first speaker model to generate the first score.

[0071] The audio signal may include a variation of the utterance that the speaker verification module may compare against the first speaker model. For example, a microphone may record the utterance and provide the recording of the utterance to a feature extraction module that generates an audio signal that the speaker verification module uses to generate the first score.

[0072] In implementation, if there are many users of the first user device, the speaker verification module compares the scores of each of these many users and selects the highest score. For example, the first user device may have 1 to 5 speaker models, one for each user of the first user device.

[0073] The speaker verification module may compare a score, such as the highest score, to a threshold score value to determine whether the score satisfies the threshold score value. For example, the speaker verification module may determine whether the highest score is higher than the threshold score value if the threshold score value is a minimum required score, or whether the highest score is lower than the threshold score value if the threshold score value is a maximum required score, e.g., the highest score has the smallest numerical value of the scores generated for the user of the first user device.

[0074] If the highest score satisfies the threshold score value, the speaker verification module or another module at the first user device may generate a score for each of the impostor speaker models identified at the first user device, e.g., stored on the first user device or server, and continue process 300 to perform step 308. If the highest score does not satisfy the threshold score value, the user device or server may stop executing process 300. If the first user device or server stops executing process 300, the first user device or server may stop requesting other speaker models or other scores from other user devices.

[0075] The speaker verification module at the first user device, or a similar module at the server, may generate scores for each of the impostor speaker models until a score is generated that is equal to or greater than the highest score for the user of the first user device, at which time the speaker verification module stops execution of process 300. If the speaker verification module determines that there are no more impostor speaker models or that the highest score for the user of the first user device has been compared to the scores for all of the impostor speaker models, including scores for impostor speaker models for other users of other user devices, determined, for example, using steps 308 and 310, the process proceeds to step 312.

[0076] For example, the process determines that one or more second user devices are located in a physical area near the physical location of the first user device (308). The first user device may determine the second user devices using near-field communication. In an example, if the speaker verification module has already determined the first score, the first user device may provide the first score to the other user devices, for example, for use by other speaker verification modules performing similar processing. In some examples, the first user device may provide the first speaker model, other speaker models for other users of the first user device, or a combination of the two, to at least some of the second user devices.

[0077] In some implementations, this process may determine a second user device that is co-located with the first user device but in a different physical location. For example, the first user device may determine that a particular second user device is co-located with the first user device if both the first user device and the second user device are participating in the same call or video conference, or if the second user device is a nearby device participating in the same call or video conference. The devices may be located in the same physical room or in different rooms, each of which includes separate videoconferencing equipment. The first device or the server may determine that the devices are co-located using calendar entries for each user, for example, if the calendar entries for both users are the same and show all of the users participating in the event.

[0078] The process obtains, for each second user of the second user device, a second speaker model for each second user or a second score indicating each likelihood that the utterance was spoken by each second user (310). For example, other speaker verification modules in the second user devices generate respective second scores for each user of the second user device, e.g., using respective second speaker models and other audio signals encoding the same utterance or portions of the same utterance. The first user device receives each of the second scores from the second user device, and may receive many second scores from a single second user device in a single message or many messages if that single second user device has many users.

[0079] In some examples, the server generates some of the second scores and provides these second scores to the first user device. The server may generate a first score for the user of the first user device and provide the first score to the first user device. The server may compare all of these scores and send a message to the device with the highest score. The server may or may not send a message to other devices that do not correspond to the highest score.

[0080] The process determines that the utterance was spoken by a first user (312). For example, the speaker verification module compares the highest score for the first user device with the score of the impostor speaker model stored on the user device, with a second score received from the second user device, or with both. The speaker verification module may stop comparing the highest score of the first user device to the other scores if the speaker verification module determines that one of the other scores is equal to or greater than the highest score of the first user device, and then, for example, stop execution of process 300.

[0081] The process performs an action in response to determining that the utterance was spoken by the first user (314). For example, a speech recognition module analyzes the audio signal and determines a text representation of the utterance encoded in the audio signal. The first user device uses the text representation to determine a command provided by the first user in the utterance and performs an action in response to the command.

[0082] The order of steps in process 300 described above is merely exemplary, and the steps of determining whether an utterance has been spoken by a user may be performed in a different order. For example, a user device may perform, for example, step 302 to determine a second user device located in a physical area near the physical location of the user device before receiving an audio signal.

[0083] In some implementations, process 300 may include some of the steps, which may be divided into additional, fewer, or more steps. For example, a first user device may determine which second user devices have any speaker models, such as imposter speaker models, stored in memory, and request from each second user device only those second speaker models not stored in memory. In these examples, the first user device may delete from memory any imposter speaker models for other users whose respective user devices are no longer in a physical area near the physical location of the first user device, e.g., not currently included in the second user device.

[0084] When deleting an imposter speaker model from memory for a user device that is no longer in a physical area near the physical location of the first user device, the first user device may retain any imposter speaker models for other users that are flagged as not for deletion. For example, one of the imposter speaker models may be for a friend who is often in a physical area near the physical location of the first user device. The first user device may retain one of the imposter speaker models for the friend even if the first user device does not detect another user device operated by the friend.

[0085] Embodiments of the subject matter and functional operations described herein may be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed herein and their structural equivalents, or one or more combinations of these. Embodiments of the subject matter described herein may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by or to control the operation of a data processing apparatus. Alternatively, or in addition, the program instructions may be encoded in an artificially generated propagated signal, such as, for example, a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a suitable receiving apparatus for execution by the data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or one or more combinations of these.

[0086] The term "data processing apparatus" refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including, by way of example, a programmable processor, computer, or multiprocessor or computer. The apparatus may also, or in addition, include special purpose logic circuitry, such as, for example, an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus may optionally include, in addition to hardware, code that creates an execution environment for a computer program, such as code that builds processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.

[0087] A computer program, which may also be referred to or described as a program, software, software application, module, software module, script, or code, may be written in any type of programming language, including compiled or interpreted languages, or declarative or procedural languages, and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored in a single file dedicated to the program in question, or in many coordinated files, e.g., files storing one or more modules, subprograms, or portions of code, or in a file that holds other programs or data, such as one or more scripts stored in a markup language document. A computer program may be deployed to be executed on one computer or on many computers located at one site or distributed across many sites and interconnected by a communications network.

[0088] The processes and logic flows described herein may be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and an apparatus may be implemented as, special purpose logic circuitry such as, for example, an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0089] A computer suitable for executing a computer program typically includes a general-purpose or special-purpose microprocessor or both, or any other type of central processing unit. Typically, the central processing unit will receive instructions and data from a read-only memory or a random-access memory or both. The essential elements of a computer are a central processing unit for executing or carrying out instructions and one or more memory devices for storing instructions and data. Typically, a computer also includes one or more mass storage devices for storing data, such as, for example, magnetic disks, magneto-optical disks, or optical disks, or may be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices. However, a computer need not have such devices. Furthermore, a computer may be incorporated into another device, such as, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as, for example, a universal serial bus (USB) flash device, to name a few.

[0090] Computer-readable media suitable for storing computer program instructions and data include, by way of example, all types of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0091] To provide for interaction with a user, embodiments of the subject matter described herein may be implemented in a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, a pointing device, such as a mouse or trackball, for providing input to the computer, and a keyboard. Other types of devices may be used to provide for interaction with a user as well. For example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user may be received in any form, including acoustic, speech, or tactile input. Additionally, a computer may interact with a user by sending documents to and receiving documents from a device used by the user, for example, by sending web pages to a web browser on the user's device in response to a request received from the web browser.

[0092] Embodiments of the subject matter described herein may be implemented in a computing system that includes back-end components, such as, for example, a data server, or that includes middleware components, such as, for example, an application server, or that includes front-end components, such as, for example, a client computer having a graphical user interface or web browser through which a user can interact with implementations of the subject matter described herein, or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, such as, for example, a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0093] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, such as HTML pages, to a user device, e.g., to display the data to and receive user input from a user interacting with the user device acting as a client. Data generated at the user device, e.g., as a result of user interaction, may be received from the user device at the server.

[0094] 4 is a block diagram of computing devices 400, 450 that may be used as a client, a server, or multiple servers to implement the systems and methods described herein. Computing device 400 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 450 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, smart watches, head-worn devices, and other similar computing devices. The components, their connections and relationships, and their functions illustrated herein are merely exemplary and are not meant to limit the practice of the invention(s) described and / or claimed herein.

[0095] Computing device 400 includes a processor 402, memory 404, a storage device 406, a high-speed interface 408 connecting to memory 404 and a high-speed expansion port 410, a low-speed bus 414, and a low-speed interface 412 connecting to storage device 406. Each of components 402, 404, 406, 408, 410, and 412 are interconnected using various buses and may be mounted on a common motherboard or otherwise suitably mounted. Processor 402 may process instructions for execution within computing device 400, including instructions stored in memory 404 or storage device 406, for displaying graphical information for a GUI on an external input / output device, such as a display 416 coupled to high-speed interface 408. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and types of memory, as appropriate. Also, multiple computing devices 400 may be connected, each providing a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).

[0096] The memory 404 stores information within the computing device 400. In one implementation, the memory 404 is a computer-readable medium. In one implementation, the memory 404 is a volatile memory unit. In another implementation, the memory 404 is a non-volatile memory unit.

[0097] The storage device 406 can provide mass storage for the computing device 400. In one implementation, the storage device 406 is a computer-readable medium. In various different implementations, the storage device 406 can be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, flash memory, or other similar solid-state memory device, or devices in a storage area network or other configuration. In one implementation, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer- or machine-readable medium, such as memory 404, the storage device 406, or memory in the processor 402.

[0098] High-speed controller 408 manages high-bandwidth-intensive operations for computing device 400, while low-speed controller 412 manages low-bandwidth-intensive operations. Such duty allocation is merely exemplary. In one implementation, high-speed controller 408 is coupled to memory 404, to display 416 (e.g., via a graphics processor or accelerator), and to high-speed expansion port 410, which may receive various expansion cards (not shown). In an implementation, low-speed controller 412 is coupled to storage device 406 and low-speed expansion port 414. The low-speed expansion port, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled to one or more input / output devices, such as a keyboard, pointing device, scanner, or networking device, such as a switch or router, e.g., via a network adapter.

[0099] As illustrated in the figure, computing device 400 may be implemented in many different forms. For example, it may be implemented as a standard server 420, or any number of such servers within a group. It may also be implemented as part of a rack server system 424. Additionally, it may be implemented in a personal computer, such as a laptop computer 422. Alternatively, components from computing device 400 may be combined with other components in a mobile device (not shown), such as device 450. Each such device may include one or more of computing devices 400, 450, and the entire system may consist of many computing devices 400, 450 communicating with each other.

[0100] Computing device 450 includes, among other components, a processor 452, memory 464, input / output devices such as a display 454, a communications interface 466, and a transceiver 468. Device 450 may also be provided with a storage device such as a microdrive or other device to provide additional storage. Each of components 450, 452, 464, 454, 466, and 468 are interconnected using various buses, and some of the components may be mounted on a common motherboard or otherwise suitable.

[0101] The processor 452 may process instructions for execution within the computing device 450, including instructions stored in the memory 464. The processor may also include separate analog and digital processors. The processor may provide for coordination of other components of the device 450, such as the user interface, applications executed by the device 450, and control of wireless communications by the device 450.

[0102] The processor 452 may communicate with a user via a control interface 458 and via a display interface 456 coupled to a display 454. The display 454 may be, for example, a TFT LCD display or an OLED display, or other suitable display technology. The display interface 456 may comprise appropriate circuitry for driving the display 454 to provide graphics and other information to the user. The control interface 458 may receive commands from the user and convert them for issuance to the processor 452. Additionally, an external interface 462 may be provided for communication with the processor 452 to enable near-area communication of the device 450 with other devices. The external interface 462 may comprise, for example, for wired communication (e.g., via a docking procedure) or for wireless communication (e.g., via Bluetooth or other such technology).

[0103] Memory 464 stores information within computing device 450. In one implementation, memory 464 is a computer-readable medium. In one implementation, memory 464 is a volatile memory unit. In another implementation, memory 464 is a non-volatile memory unit. Expansion memory 474 may also be provided and connected to device 450 via expansion interface 472, which may include, for example, a SIMM card interface. Such expansion memory 474 may provide additional storage space for device 450 or may also store applications or other information for device 450. Specifically, expansion memory 474 may include instructions for performing or supplementing the processes described above and may also include secure information. Thus, for example, expansion memory 474 may be provided as a security module for device 450 and may be programmed with instructions that permit secure use of device 450. Additionally, secure applications may be provided via a SIMM card along with additional information to place identifying information on the SIMM card in an unhackable manner.

[0104] The memory may include, for example, flash memory and / or MRAM memory as described above. In one implementation, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods as described above. The information carrier is a computer- or machine-readable medium, such as memory 464, expansion memory 474, or memory on processor 452.

[0105] Device 450 may communicate wirelessly via communication interface 466, which may include digital signal processing circuitry, if necessary. Communication interface 466 may provide communication under various modes or protocols, such as GSM voice calls, SMS, EMS, or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS, among others. Such communication may occur, for example, via radio frequency transceiver 468. Additionally, short-range communication may occur, for example, using Bluetooth, Wi-Fi, or other such transceivers (not shown). Additionally, GPS receiver module 470 may provide additional wireless data to device 450, which may be used appropriately by applications running on device 450.

[0106] Device 450 may also communicate audibly using audio codec 460. Audio codec 460 may receive spoken information from a user and convert it into usable digital information. Audio codec 460 may also generate audible sounds for the user, for example, via a speaker in the handset of device 450. Such sounds may include sounds from voice telephone calls, recorded sounds (e.g., voice messages, music files, etc.), and may also include sounds generated by applications running on device 450.

[0107] As illustrated in the figure, the computing device 450 may be implemented in many different forms. For example, it may be implemented as a cellular telephone 480. Furthermore, it may be implemented as part of a smartphone 482, a personal digital assistant, or other similar mobile device.

[0108] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of what is to be claimed, but rather as details of features that may be specific to particular embodiments. Some features that are described herein in the context of a separate embodiment may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable subcombination. Furthermore, while features may be described above as operating in several combinations and as originally claimed, in some cases, one or more features from a claimed combination may be removed from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0109] Similarly, although operations are illustrated in a particular order in the figures, this should not be understood as requiring such operations to be performed in the particular order illustrated, or in sequential order, or that all of the illustrated operations be performed to achieve desired results. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packed into many software products.

[0110] In situations where the systems discussed herein may collect or utilize personal information about a user, the user may be provided with the opportunity to control whether and / or how a program or feature collects user information, such as, for example, a speaker model, user preferences, or the user's current location, or receives content from a content server. Additionally, some data may be handled in one or more ways before being stored or used, such that personally identifiable information is removed. For example, the user's identity may be handled so that personally identifiable information is not determined for the user, or the user's geographic location from which location information is obtained may be generalized, for example, to the city, zip code, or state level, such that the user's specific location is not determined. Thus, the user may have control over how information about the user is collected and used by the content server.

[0111] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims may be performed in a different order and still achieve desirable results. As one example, the processes illustrated in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desired results. In some cases, multitasking and parallel processing may be advantageous. For example, a module that performs similarity score calculations, such as part of a speaker verification module, may be implemented in hardware, such as directly in a digital signal processing (DSP) unit. [Explanation of symbols]

[0112] 100 Environment 102a User Device A 102b User Device B 102c User Device C 102d User Device D 104a Speaker Model A 104b Speaker Model B 104c Speaker Model C 104d Speaker Model D 106 User D: Okay Google, start the demo. 108 Launch Demo 200 Speaker Verification System 202a User Device A 202b User Device B 204 Server 206a Speaker Model A 206b Speaker Model B 208a Speaker Verification Module 208b Speaker Verification Module 210a Impostor Speaker Model 210b Impostor Speaker Model 212a Speech Recognition Module 212b Voice Recognition Module 214a Setting A 214b Setting B 216 wireless communication channels 224 Network 400 computing devices 402 processor 404 Memory 406 Storage Devices 408 high-speed interface 410 High-Speed ​​Expansion Port 412 slow interface 414 Low-Speed ​​Expansion Port 416 Display 420 Server 422 laptop computer 424 Rack Server System 450 computing devices 452 processor 454 Display 456 Display Interface 458 Control Interface 460 Audio Codec 462 External Interface 464 memory 466 Communication Interface 468 Transceiver 470 GPS Receiver Module 472 Extended Interface 474 Extended Memory 480 Cellular Phone 482 smartphones

Claims

1. When executed on data processing hardware, the data processing hardware: acts of receiving audio data corresponding to utterances of voice commands captured by a user device, the user device having a plurality of different users; determining a particular user of the plurality of different users of the user device as the speaker of the utterance based on a comparison of the audio data with corresponding speaker verification data stored in memory hardware for each user of the plurality of different users of the user device, for each user of the plurality of different users of the user device, obtaining the corresponding speaker verification data and corresponding imposter verification data; operations of generating a corresponding speaker verification score using the corresponding speaker verification data and the audio data, generating a corresponding impostor verification score using the corresponding impostor verification data and the audio data, and comparing the corresponding impostor verification score to the corresponding speaker verification score, wherein the corresponding speaker verification score indicates a likelihood that the utterance of the voice command was spoken by a corresponding user of the plurality of different users of the user device, and the corresponding impostor verification score indicates a likelihood that the utterance of the voice command was spoken by an impostor associated with the corresponding user; in response to determining that the corresponding speaker match score is highest, identifying the speaker of the utterance of the voice command as the particular user of the plurality of different users of the user device associated with the highest corresponding speaker match score; and based on determining the particular user of the plurality of different users of the user device as the speaker of the utterance, providing a message to the user device including a speaker identifier associated with the particular user, and outputting the message on the user device; A computer-implemented method for causing a computer to perform operations including:

2. The operation is 10. The computer-implemented method of claim 1, further comprising, prior to the act of identifying the speaker of the utterance of the voice command, an act of determining that the highest corresponding speaker match score satisfies an acceptance threshold.

3. 10. The computer-implemented method of claim 1, wherein obtaining the corresponding speaker verification data comprises obtaining a corresponding speaker verification model for each of the plurality of different users of the user device.

4. The computer-implemented method of claim 3 , wherein at least one of the corresponding speaker verification models comprises an i-vector speaker verification model.

5. 4. The computer-implemented method of claim 3, wherein at least one of the corresponding speaker verification models comprises a d-vector speaker verification model.

6. 2. The computer-implemented method of claim 1, wherein receiving the audio data corresponding to the utterance of the voice command comprises receiving the audio data corresponding to the utterance of the voice command preceded by a specific predefined hotword captured by the user device while in a locked state.

7. 7. The computer-implemented method of claim 6, wherein the user device is configured to respond to a voice command upon receiving the specific predefined hotword while in the locked state.

8. The computer-implemented method of claim 1 , wherein the data processing hardware resides on a server in communication with the user device.

9. data processing hardware; memory hardware in communication with the data processing hardware and storing instructions; Equipped with The instructions, when executed by the data processing hardware, cause the data processing hardware to: acts of receiving audio data corresponding to utterances of voice commands captured by a user device, the user device having a plurality of different users; determining a particular user of the plurality of different users of the user device as the speaker of the utterance based on a comparison of the audio data with corresponding speaker verification data stored in the memory hardware for each user of the plurality of different users of the user device, for each user of the plurality of different users of the user device, obtaining the corresponding speaker verification data and corresponding imposter verification data; operations of generating a corresponding speaker verification score using the corresponding speaker verification data and the audio data, generating a corresponding impostor verification score using the corresponding impostor verification data and the audio data, and comparing the corresponding impostor verification score to the corresponding speaker verification score, wherein the corresponding speaker verification score indicates a likelihood that the utterance of the voice command was spoken by a corresponding user of the plurality of different users of the user device, and the corresponding impostor verification score indicates a likelihood that the utterance of the voice command was spoken by an impostor associated with the corresponding user; in response to determining that the corresponding speaker match score is highest, identifying the speaker of the utterance of the voice command as the particular user of the plurality of different users of the user device associated with the highest corresponding speaker match score; and based on determining the particular user of the plurality of different users of the user device as the speaker of the utterance, providing a message to the user device including a speaker identifier associated with the particular user, and outputting the message on the user device; A system that causes an operation including

10. The operation is 10. The system of claim 9, further comprising, prior to the act of identifying the speaker of the utterance of the voice command, the act of determining that the highest corresponding speaker match score satisfies an acceptance threshold.

11. 10. The system of claim 9, wherein the act of obtaining the corresponding speaker verification data comprises an act of obtaining a corresponding speaker verification model for each of the plurality of different users of the user device.

12. The system of claim 11 , wherein at least one of the corresponding speaker verification models comprises an i-vector speaker verification model.

13. The system of claim 11 , wherein at least one of the corresponding speaker verification models includes a d-vector speaker verification model.

14. 10. The system of claim 9, wherein receiving the audio data corresponding to the utterance of the voice command comprises receiving the audio data corresponding to the utterance of the voice command preceded by a specific predefined hotword captured by the user device while in a locked state.

15. 15. The system of claim 14, wherein the user device is configured to respond to a voice command upon receiving the particular predefined hotword while in the locked state.

16. The system of claim 9 , wherein the data processing hardware resides on a server in communication with the user device.

Citation Information

Patent Citations

  • Method of recognizing talker and device therefor

    JP1998247092A

  • Performance environment setting device utilizing voice recognition, and method

    JP2000099076A

  • Information processor and musical instrument provided with the information processor

    JP2002007014A

  • Service center and order receiving method

    JP2002279245A

  • Device access using voice authentication

    JP2014517366A