Passive and Continuous Multi-Speaker Voice Biometrics

A passive and continuous voice biometric system addresses enrollment and profile management challenges by using machine learning for unsupervised speaker recognition and adaptive thresholding, ensuring accurate and dynamic speaker identification in diverse environments.

JP7817946B2Active Publication Date: 2026-02-19PINDROP SECURITY INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022561448
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-04-15
Filing Date
2021-04-15
Publication Date
2026-02-19
Estimated Expiration
2041-04-15

AI Technical Summary

Technical Problem

Existing voice biometric systems face challenges in efficiently enrolling new speakers and maintaining accurate speaker profiles, especially in dynamic environments with multiple speakers, leading to issues like outdated models and false authentication.

Method used

A passive and continuous voice biometric authentication system that allows for flexible enrollment and profile management, using machine learning to identify and update speaker profiles without explicit user interaction, enabling unsupervised speaker recognition and adaptive thresholding for improved accuracy.

Benefits of technology

Enables seamless and accurate speaker identification and profile management, reducing false acceptance/rejection rates and adapting to changing speaker characteristics in real-time, suitable for IoT devices and over-the-top services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007817946000001
    Figure 0007817946000001
  • Figure 0007817946000002
    Figure 0007817946000002
  • Figure 0007817946000003
    Figure 0007817946000003
Patent Text Reader

Abstract

The embodiments described herein provide a voice biometric authentication system that implements a machine learning architecture capable of passive, active, continuous, or static computation, or a combination thereof. The system passively and / or continuously enrolls speakers, in some cases actively and / or statically as the speaker speaks into or around an edge device (e.g., car, TV, radio, phone). The system identifies users on the fly without requiring new speakers to mirror prompted utterances to reconfigure the computation. The system manages speaker profiles as speakers provide utterances to the system. The machine learning architecture implements a passive and continuous voice biometric authentication system, potentially without knowledge of speaker identity. The system may create identities in an unsupervised manner and passively enroll and recognize known or unknown speakers. This system offers personalization and security across a wide range of applications, including over-the-top services and media content for IoT devices (e.g., personal assistants, vehicles), and call centers.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 010,504, filed April 15, 2020, the entire contents of which are incorporated by reference.

[0002] This application is generally related to U.S. Patent Application No. 15 / 262,748, filed September 12, 2016, entitled "End-To-End Speaker Recognition Using Deep Neural Network," which issued as U.S. Patent No. 9,824,692, and is hereby incorporated by reference in its entirety.

[0003] This application is generally related to U.S. patent application Ser. No. 15 / 890,967, filed Feb. 7, 2018, entitled "Age Compensation in Biometric Systems Using Time-Interval, Gender and Age," which issued as U.S. Patent No. 10,672,403, and is hereby incorporated by reference in its entirety.

[0004] This application is generally related to U.S. patent application Ser. No. 15 / 910,387, filed Mar. 2, 2018, entitled "Method and Apparatus for Detecting Spoofing Conditions," which issued as U.S. Patent No. 10,692,502, and is hereby incorporated by reference in its entirety.

[0005] This application is generally related to U.S. Patent Application No. 17 / 155,851, filed January 22, 2021, entitled "Robust Spoofing Detection System Using Deep Residual Neural Networks," the entire contents of which are incorporated herein by reference.

[0006] This application is generally related to U.S. patent application Ser. No. 17 / 192,464, filed March 4, 2021, entitled "Systems and Methods of Speaker-Independent Embedding for Identification and Verification from Audio," which is hereby incorporated by reference in its entirety.

[0007] This application relates generally to systems and methods for training and deploying audio processing neural networks. [Background technology]

[0008] The emergence of Internet of Things (IoT) devices has led to newer channels for machines to interact with voice commands. In many cases, interactions with devices involve performing operations on private and sensitive data. Many new mobile apps and home personal assistants enable financial transactions using voice-based interactions with devices. Call centers, especially interactions with human agents at call centers, are no longer the only instance of voice-based interactions for institutions that manage sensitive personal information. It is essential to reliably verify the identity of callers / speakers who access and manage user accounts by operating various edge or IoT devices or by contacting call centers, according to a uniform level of accuracy and security.

[0009] Automatic speech recognition (ASR) and automatic speaker verification (ASV) systems are often used for security and authentication functions and other voice-based operations. Most implementations of voice biometrics use active and static enrollment, typically assuming a known link between voice utterances and speaker identity. Active enrollment is when a user is prompted with an enrollment phase in which the user must repeat a passphrase or, typically, speak freely until criteria defined by the voice biometric system are met. Active enrollment is often combined with static enrollment when a user first sets up their respective device or begins using an over-the-top service. Active enrollment can be time-consuming, and voice biometric deployments can be hindered by the possibility that users may opt out of enrollment. Furthermore, static enrollment can result in voice models that become outdated or produce inaccurate matches as more people want to use the voice biometric system and as these people's voices change.

[0010] Additionally, over-the-top (OTT) services may differ from other services that use automatic voice verification, such as banking, because OTT services may require the identification of individual speakers from multiple speakers at once, as opposed to solely determining whether a speaker meets predefined criteria associated with a speaker profile (e.g., determining whether a speaker's voice matches a voice corresponding to a speaker profile, regardless of the profiles of any other speakers). Maintaining a system that can actively differentiate between speakers can be difficult, especially when multiple speakers are speaking to each other simultaneously or intermittently. Furthermore, providing content for individuals or configuring edge devices can be difficult when the system identifies speech from multiple individuals at once, such as when multiple people gather to watch videos, listen to music together, or get into a car together.

[0011] Therefore, what is needed is an improved approach to enrolling new speakers as the service operates, and to providing content for speakers with pre-established and / or non-established speaker profiles or configuring edge devices so that the system can distinguish the profiles from each other for speech matching. Summary of the Invention

[0012] Disclosed herein are systems and methods that can address the above-mentioned shortcomings and may also provide any number of additional or alternative benefits and advantages. The embodiments described herein provide a flexible voice biometric authentication system that allows for passive, active, continuous, or static operation, or some hybrid combination thereof. In particular, the systems and methods described herein provide a method for passive and / or continuous speaker enrollment, in some cases actively and / or statically enrolling speakers as they speak into or around an edge device (e.g., a car, television, radio, or phone). By implementing such a system and method, a device can identify new users on the fly without requiring them to mirror their utterances, which would prompt the device to actively reconfigure each time a new speaker desires to set up a profile. The systems and methods further provide a method for organizing and reorganizing a speaker's profile as the speaker provides utterances to the system to maintain an up-to-date speaker profile and avoid false authentication acceptance and / or rejection. The systems and methods provide a passive and continuous voice biometric authentication system, in some cases, without knowledge of the speaker's identity. Systems and methods may create identities in an unsupervised manner, passively enrolling and recognizing individual speakers, in some cases when the system identifies speakers who do not meet the criteria for any stored user profile. Such systems and methods may be used for personalization and security purposes across a wide range of applications, including IoT (e.g., identifying a car driver and configuring the car's settings based on settings associated with the driver's speaker profile), for over-the-top services (e.g., identifying television viewers to provide relevant content), and / or call center use cases.

[0013] In one embodiment, a computer-implemented method includes: extracting, by a computer, an inbound embedding of an inbound speaker by applying a machine learning model to an inbound audio signal; generating, by the computer, a similarity score based on a distance between the inbound embedding and a voiceprint stored in a speaker profile in a speaker profile database; and, in response to the computer determining that the similarity score of the inbound embedding does not satisfy a similarity threshold, generating, by the computer, a new speaker profile for the inbound speaker in the speaker profile database that includes the inbound embedding, wherein the new speaker profile is a database record that stores the inbound embedding as a new voiceprint.

[0014] In another embodiment, a system comprises: a speaker database comprising a non-transitory machine-readable storage medium configured to store data records comprising a speaker profile; and a computer comprising a processor, the processor configured to: extract an inbound embedding of the inbound speaker by applying a machine learning model to the inbound speech signal; generate a similarity score based on a distance between the inbound embedding and a voiceprint stored in a speaker profile in the speaker profile database; and, in response to the computer determining that the similarity score of the inbound embedding does not satisfy a similarity threshold, generate a new speaker profile for the inbound speaker in the speaker profile database comprising the inbound embedding, the new speaker profile being a database record that stores the inbound embedding as a new voiceprint.

[0015] In another embodiment, a computer-implemented method includes receiving, by a computer, an inbound speech signal comprising a plurality of utterances of a plurality of inbound speakers; applying, by the computer, a machine learning architecture to the inbound speech signal to extract a plurality of inbound embeddings corresponding to the plurality of inbound speakers; and, for each inbound speaker of the plurality of inbound speakers, generating, by the computer, one or more similarity scores based on the inbound embedding of the inbound speaker, wherein each similarity score for the inbound speaker is calculated based on the inbound embedding and a speaker profile. and one or more voiceprints stored in a speaker profile database, indicating a distance between the inbound speaker and one or more voiceprints stored in the speaker profile database; identifying, by the computer, a closest voiceprint of the inbound speaker from the one or more voiceprints, the closest voiceprint corresponding to a maximum similarity score of the one or more similarity scores generated for the inbound speaker; and for each maximum similarity score that satisfies the one or more similarity score thresholds, updating, by the computer, the speaker profile database to include an inbound embedding of the inbound speaker having the maximum similarity score that satisfies the one or more similarity score thresholds.

[0016] In another embodiment, a system comprises: a speaker database comprising a non-transitory machine-readable storage medium configured to store data records comprising speaker profiles; and a computer comprising a processor, the processor receiving an inbound speech signal comprising a plurality of utterances of a plurality of inbound speakers; applying a machine learning architecture to the inbound speech signal to extract a plurality of inbound embeddings corresponding to the plurality of inbound speakers; and generating, for each inbound speaker of the plurality of inbound speakers, one or more similarity scores based on the inbound speaker's inbound embedding, wherein the inbound utterances are correlated to the one or more similarity scores. The system is configured to: generate a speaker profile database containing a plurality of inbound speaker similarity scores, each of which indicates a distance between the inbound embedding and one or more voiceprints stored in the speaker profile database; identify a closest voiceprint for the inbound speaker from the one or more voiceprints, the closest voiceprint corresponding to a maximum similarity score of the one or more similarity scores generated for the inbound speaker; and, for each maximum similarity score that satisfies one or more similarity score thresholds, update the speaker profile database to include an inbound embedding of the inbound speaker having a maximum similarity score that satisfies the one or more similarity score thresholds.

[0017] In another embodiment, a computer-implemented method includes receiving, by a computer, an inbound speech signal of an inbound speaker from an end user device via a content server; applying, by the computer, a machine learning model to the inbound speech signal to extract an inbound embedding of the inbound speaker; generating, by the computer, a similarity score for the inbound embedding based on a distance between the inbound embedding and a voiceprint stored in a speaker profile in a speaker database, the similarity score satisfying one or more similarity score thresholds; identifying, by the computer, one or more speaker features in the speaker profile that correspond to one or more content features of the content server; and transmitting, by the computer, the one or more speaker features associated with the inbound speaker to a media content server.

[0018] In another embodiment, a system comprises a speaker database comprising a non-transitory machine-readable storage medium configured to store a plurality of speaker profiles; and a server comprising a processor, the processor configured to: receive an inbound speech signal of an inbound speaker from an end user device via a content server; apply a machine learning model to the inbound speech signal to extract an inbound embedding of the inbound speaker; generate a similarity score for the inbound embedding based on a distance between the inbound embedding and a voiceprint stored in a speaker profile in the speaker database, the similarity score satisfying one or more similarity score thresholds; identify one or more speaker features in the speaker profile that correspond to one or more content features of the content server; and transmit the one or more speaker features associated with the inbound speaker to a media content server.

[0019] In another embodiment, a method includes obtaining, by a computer, a speaker profile associated with a speaker that includes one or more embeddings of the speaker; determining, by the computer, a maturity level of the speaker's voiceprint based on a false acceptance rate of the one or more embeddings and one or more maturity factors; and updating, by the computer, one or more similarity thresholds of the speaker profile according to the maturity level and the one or more maturity factors.

[0020] In another embodiment, a system comprises a speaker profile database comprising a non-transitory machine-readable medium configured to store a plurality of speaker profiles; and a computer comprising a processor configured to: obtain a speaker profile associated with a speaker that includes one or more embeddings of the speaker; determine a maturity level of the speaker's voiceprint based on a false acceptance rate of the one or more embeddings and one or more maturity factors; and update one or more similarity thresholds of the speaker profile according to the maturity level and the one or more maturity factors.

[0021] In another embodiment, a device-implemented method includes receiving, by the device, an inbound audio signal comprising an utterance of an inbound speaker; applying, by the device, an embedding extraction model to the inbound audio signal to extract an inbound embedding of the inbound speaker; generating, by the device, one or more similarity scores for the inbound embeddings based on a relative distance between the inbound embeddings and one or more voiceprints stored in a non-transitory machine-readable medium; identifying, by a computer, a speaker identifier associated with the inbound speaker's voiceprint in response to determining that the similarity score generated using the voiceprint satisfies a similarity threshold; and transmitting, by the device, the speaker identifier to a content server.

[0022] In another embodiment, a system includes a speaker database comprising a non-transitory machine-readable storage medium configured to store data records comprising speaker profiles; and a device comprising a processor configured to receive an inbound speech signal comprising an utterance of an inbound speaker; apply an embedding extraction model to the inbound speech signal to extract an inbound embedding of the inbound speaker; generate one or more similarity scores for the inbound embeddings based on a relative distance between the inbound embeddings and one or more voiceprints stored in the speaker database; and, in response to determining that the similarity score generated using the voiceprint satisfies a similarity threshold, identify, by the computer, a speaker identifier associated with the inbound speaker's voiceprint and send the speaker identifier to a content server.

[0023] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are intended to provide further explanation of the invention as claimed. [Brief explanation of the drawings]

[0024] The present disclosure may be better understood by reference to the following drawings, in which components are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the present disclosure, in which reference characters refer to corresponding parts throughout the various views.

[0025] [Figure 1] 1 illustrates components of a system that uses speech processing machine learning operations. [Figure 2] The machine learning models and other machine learning architectures illustrate components of a system that uses audio processing machine learning operations implemented on a local device. [Figure 3A] 1 illustrates the operational steps of a method for actively registering and identifying (authenticating) a user. [Figure 3B]1 illustrates the operational steps of a method for actively registering and identifying (authenticating) a user. [Figure 4] 1 illustrates the operational steps of a method for adaptive thresholding in an audio processing system. [Figure 5] 1 illustrates the execution steps of a method for identifying and evaluating strong and weak speech in speech processing; [Figure 6] 1 illustrates the operational steps of a method for clustering speakers during speech processing. [Figure 7A] 1 illustrates the operational steps of a method for correcting label identifiers (eg, speaker identifiers, subscriber identifiers) of one or more voiceprints according to current and / or historical information. [Figure 7B] 1 illustrates an example of label correction using a particular speaker's cluster and other putative speaker's clusters. [Figure 8] 1 illustrates the operational steps of a method for speech processing using a passive and continuous registration arrangement. [Figure 9] 1 illustrates the operational steps of a method for audio processing of an audio signal using a mixed active-passive and continuous registration arrangement. [Figure 10] 1 illustrates the operational steps of a method for speech processing of an audio signal using an active and continuous registration arrangement. [Figure 11A] 1 illustrates components of a system using speech processing machine learning operations, where the machine learning model is implemented by a vehicle. [Figure 11B] 1 illustrates components of a system using speech processing machine learning operations, where the machine learning model is implemented by a vehicle. [Figure 12] 10 illustrates an example table of a similarity threshold scheduler based on a maturity factor and a false acceptance rate. DETAILED DESCRIPTION OF THE INVENTION

[0026] Reference will now be made to the exemplary embodiments illustrated in the drawings, and specific language will be used herein to describe the same. It will nevertheless be understood that no limitation of the scope of the invention is thereby intended. Alterations and further modifications of the features of the invention illustrated herein, and further applications of the principles of the invention as illustrated herein, which will occur to those skilled in the art and in possession of this disclosure, are to be considered within the scope of the invention.

[0027] Voice biometrics for speaker recognition and other operations (e.g., authentication) typically rely on speaker samples and models or feature vectors (sometimes called "embeddings") generated from a universe of samples for a particular speaker. As an example, during a training phase (or retraining phase), a server or other computing device runs a speech recognition engine (e.g., artificial intelligence and / or machine learning program software) that has been trained to recognize and distinguish instances of speech using multiple training speech signals. The machine learning architecture outputs specific results according to corresponding inputs and evaluates the results according to a loss function by comparing expected outputs with observed outputs. The training operation then adapts weights or hyperparameters (of the neural network in the machine learning architecture) and reapplies the machine learning architecture to the inputs until the expected and observed outputs converge. The server then fixes the hyperparameters and, in some cases, disables one or more layers of the neural network architecture used for training.

[0028] After training the machine learning architecture, the server can further refine and develop the machine learning architecture to recognize a particular speaker during the enrollment operation for that speaker. The speech recognition engine can generate an enrollee model embedding (sometimes called a "voiceprint") using embeddings extracted from the enrollee speech signal comprising the speaker's utterance. During a subsequent inbound speech signal, the server references the voiceprint stored in the speaker profile to determine whether the subsequent speech signal involves a known speaker based on matching the inbound embedding extracted from the subsequent inbound speech signal against the enrollee's voiceprint.

[0029] These approaches have generally been successful and are appropriate for detecting subscribers in the context of evaluating inbound calls to a call center. A more flexible and less transparent approach to registration and deployment operations may be desirable in other contexts, where users would prefer a more fluid or less structured experience, such as when they are watching television or operating certain IoT or voice-enabled devices (e.g., vehicles, smart appliances, personal assistance).

[0030] Machine learning models for voiceprint and content services The embodiments described herein disclose systems and methods for biometric authentication, including voice recognition, for voice-based interfaces, management, authorization, and content personalization. A computing device executes software programming that implements various types of machine learning architecture layers or operations, including Gaussian matrix models (GMMs) and neural networks for processing audio signals. The machine learning architecture generally includes any number of machine learning models that, for example, generate feature vectors to extract embeddings that represent or model aspects of an input audio signal; perform classification according to the embeddings; generate model embeddings (e.g., voiceprints) for a particular speaker based on one or more embeddings extracted from that speaker's utterances; and cluster similar-sounding speech or audio signal characteristics based on similarities or differences between the extracted features or between the embeddings compared to stored / expected features. Once the voiceprint is created, the system can detect utterances from an individual by comparing embedding extractions from the utterance to the voiceprint and generate a similarity score indicating the likelihood that the utterance should be associated with the voiceprint (e.g., spoken by the speaker represented by the voiceprint). If the similarity score satisfies one or more pre-configured similarity scores, the system may associate the utterance embedding with the speaker and / or include the utterance embedding in the voiceprint to update the voiceprint. Although the description of particular embodiments refers to training operations, the embodiments disclosed herein generally assume that the training phase is complete and begin with the enrollment phase.

[0031] Passive Registration The voice biometric authentication systems described herein generally include a machine learning architecture trained on a server or other computing device. Accordingly, embodiments proceed with voice processing in an enrollment phase and a deployment phase, where the server performs enrollment operations in an active or passive enrollment configuration. In an active enrollment configuration, an audio user interface or a visual user interface (e.g., telephone, television screen) presents a user with prompts instructing the user to speak various phrases. In response to the prompts, the user may, for example, repeat a passphrase or speak freely until one or more criteria (e.g., number of utterances, time between inputs) are met. Administrative user input defines the criteria, or one or more machine learning models automatically establish or adjust the criteria.

[0032] Passive enrollment does not require user prompting, although some implementations use a mixed approach (e.g., active and passive). Rather, the server passively applies various machine learning layers to the audio signal to extract features, perform classification, and other operations without requiring user recognition. Beneficially, downstream operations using the speaker modeling operations described herein (e.g., biometric authentication, content personalization, speaker diarization) occur in a seamless and frictionless manner, thereby requiring the user to change or interrupt their dialogue and saving time for enrollment or other operations.

[0033] Continuous Registration The server may further implement static or continuous voice processing operations during the enrollment and deployment phases. In static operations, the voice biometric authentication system enrolls a speaker once or at a fixed time, and the server does not perform enrollment operations from new incoming voice signals or update the machine learning architecture or model (e.g., parameters, voiceprints) after the initial enrollment.

[0034] By implementing continuous voice processing operations, the voice biometric authentication system uses and benefits from new incoming voice signals. A machine learning architecture may initially implement static voice processing operations to actively enroll a speaker at a fixed time, and the server may further incorporate new utterances and develop this machine learning model to enroll and detect new speakers at a later time. By implementing continuous voice processing operations, the continuous voice processing operations may detect new utterances and compare extracted embeddings of the new utterances with predefined criteria (e.g., voiceprints, similarity scores, authentication data, user information, device information) to identify and enroll new speakers over time. The criteria may include, for example, various types of features extracted from biometric information, speaker / user information, device information, and metadata received with data input from an end-user device. A voice biometric authentication system performing continuous enrollment operations may passively capture and analyze enrollment input (e.g., enrollment voice signals) containing various features and other types of information. The system then automatically detects new speakers based on the features of the input speech signal and generates new voiceprints (e.g., model embeddings) that the system references to identify enrolled speakers or distinguish unknown new speakers. Such systems may also perform continuous enrollment and periodic updates of speaker profiles and speaker voiceprints to avoid staleness, which can result in increased false rejection or false acceptance rates.

[0035] Condition-dependent adaptive thresholding Some embodiments of the speaker recognition system use fixed thresholding. For single-speaker verification using fixed thresholding, the server compares the speaker embedding extracted from the inbound speech signal with the enrolled embedding (voiceprint) and calculates a similarity score or predicted score. If the similarity score meets a predefined, speaker-independent, fixed threshold, the computing device matches the inbound speech sample. Otherwise, the computing device rejects the inbound speech signal or reports a failed predicted score. For multi-speaker verification or open-set identification using fixed thresholding with a number of speakers (N), the computing device extracts N inbound speaker embeddings from the inbound speech signal and compares the N inbound embeddings with the number of voiceprints (V) to calculate a similarity score of N / V, whereby the computing device compares each inbound embedding with each voiceprint. The computing device then outputs N different similarity scores. The computing device considers only the maximum similarity score for each speaker, where the maximum similarity score represents the closest match between a particular speaker embedding and a particular voiceprint. If the maximum similarity score for a particular speaker embedding satisfies a predefined, speaker-independent, fixed threshold, the computing device matches or identifies the corresponding speaker in the multi-speaker audio sample.

[0036] In some cases, using a fixed thresholding operation allows some voiceprints to mature faster than other models. Voiceprints based on poor quality metrics may cause a computing device to incorrectly accept a speaker at a certain non-acceptance rate over some maturation factor. Maturation factors include, for example, the number of enrolled utterances. For example, a speaker voiceprint enrolled with 50 utterances will be much more mature than one enrolled with only one utterance. Another example of a maturation factor is the overall duration of the net utterance. For example, a speaker voiceprint model enrolled with one utterance 30 seconds long will be more mature than another model enrolled with one utterance that is only 2 seconds long. Yet another example is speech quality. For example, a speaker model enrolled with one utterance collected in clear conditions with relatively low noise (high SNR, low T60) will be more mature than a model enrolled with one utterance collected in noisy and relatively reverberant conditions (low SNR, high T60).

[0037] Some embodiments of the speaker recognition system use condition-dependent adaptive thresholding. By implementing condition-dependent adaptive thresholding, the system accounts for maturity deficiencies and increases the accuracy rate to meet the desired false acceptance rate (or desired false identification rate) of the machine learning architecture trained to recognize or authenticate the speaker. The system continuously adjusts the similarity threshold for matching individual speaker voiceprints based on the voiceprint's maturity and maturity threshold. In some cases, the server may determine different similarity thresholds for individual speaker profiles based on maturity factors or combinations of such factors associated with a particular speaker profile embedding and voiceprint. The server generates and updates a set of similarity score thresholds for a given speaker according to target false acceptance rates and maturity factors configured according to a management configuration received from a management device. For model embeddings (voiceprints), the system determines or updates the similarity threshold according to different acceptable or target false acceptance rates and / or maturity factor thresholds, such as the number of utterances added to the voiceprint. As one example, as the system adds utterances to the speaker embedding model, the system may increase the similarity threshold, resulting in the system having a better representation of the speaker (e.g., an increase in utterances associated with the speaker). As another example, the similarity threshold for a given voiceprint may decrease as the configured false positive rate increases according to user configuration input to the server.

[0038] A system implementing condition-dependent adaptive thresholding may use the received utterance to update a speaker embedding model. For example, in some embodiments, the server uses double thresholding, in which the server generates a similarity score for an inbound embedding extracted from an inbound utterance by comparing the inbound embedding to a voiceprint, and then evaluates the similarity score against a higher and lower threshold for the particular speaker. If the similarity score exceeds the higher threshold, the server matches or authenticates the speaker. The server then adds the inbound embedding to the voiceprint and adds the inbound utterance to the speaker profile as a new utterance associated with the speaker. If the similarity score exceeds the lower threshold, the server matches or authenticates the inbound speaker. In situations where the prediction score meets the lower threshold but not the higher threshold, the server stores the inbound embedding and inbound utterance in a list of weak embeddings, which is a memory location that serves as a buffer or quarantine for embeddings that were close enough to the voiceprint to match the speaker but not similar enough to update the voiceprint, possibly due to poor sound quality or background noise. When the server updates the voiceprint (another aspect of the machine learning architecture), the system may calculate a new similarity score for the stored weak embeddings and utterance to the voiceprint to determine whether the stored weak embeddings and utterance become sufficiently similar to the updated voiceprint to exceed the higher threshold, and may therefore add the utterance to the model as a new utterance. In some embodiments, the server also includes one or more lists of strong embeddings used by the server to generate the voiceprint.

[0039] Unsupervised Clustering The voice biometric authentication system may be capable of identifying multiple speakers at once using unsupervised clustering methods. The server generates clusters by running any number of clustering algorithms or operations to calculate similarity scores and may reference any number of features or types of data, including voiceprints. Clusters are associated with multiple speakers, up to a threshold number of speakers, and identify speakers in real time based on the utterances most similar to the speakers' respective clusters. For example, a media content server of a media service issues subscriber identifiers to households or power users of the households and then assigns a predetermined number of users in a media database. The speaker profile database generates one or more speaker profiles according to the number of users associated with the subscriber identifier. The server performing the clustering operation references the speaker profiles or media database to determine the number of users assigned to the subscriber identifier and uses the assigned number of users as the threshold number of speakers. Based on a clustering operation, such as comparing multiple embeddings extracted for multiple speakers in the inbound speech signal, the server generates a similarity score, identifies the closest matching voiceprint, and compares the similarity score of the closest voiceprint with a similarity threshold for the respective voiceprint or with a default similarity threshold.

[0040] The voice biometric authentication system may use incremental clustering (e.g., continuous clustering) and / or systematic clustering (e.g., hierarchical clustering) techniques to build clusters of individual speakers to ensure an efficient and accurate clustering method that can be used for passive and continuous enrollment and authentication. The system may use incremental clustering operations unless certain criteria are met (e.g., a scheduled time and the scheduled time interval has elapsed, the system has processed a predetermined number of utterances since the system previously used hierarchical clustering, the system has identified more than a threshold number of speakers, etc.), in which case the system may perform systematic clustering operations.

[0041] To use incremental clustering, for example, a voice biometric authentication system may determine a similarity score of a new utterance relative to a group of existing clusters. The system may identify the highest similarity score and determine whether the similarity score exceeds a predetermined threshold. If the similarity score exceeds the threshold, the system may add the utterance to the cluster associated with the similarity score. If not, the system may create a new cluster with the utterance as the first utterance. The system may implement incremental clustering for each new utterance that the system ingests to maintain an up-to-date speaker embedding model for each individual speaker, while minimizing the processing resources required to do so.

[0042] To use hierarchical clustering, for example, a voice biometric authentication system may access each of the stored utterances in the system and shuffle the utterances between clusters. The system may compare each of the utterances in a cluster to each other and cluster together the utterances with the highest similarity. Additionally or alternatively, the system may compare voiceprints to each other and combine those voiceprints with the highest similarity score that also meets a voiceprint similarity threshold.

[0043] Because each clustering methodology has its own advantages and disadvantages (e.g., incremental clustering may be faster but less accurate, while systematic clustering may be more accurate but require a large amount of computer resources), using a combination of the two methodologies over time may cover the deficiencies of both methods and allow the system to create mature, accurate speaker embedding models. The system may intermittently perform incremental clustering operations along with systematic clustering operations to improve the accuracy rate of the speaker embedding model while avoiding overly frequent use of systematic clustering to conserve processing resources. This combination ensures efficient and accurate clustering that is suitable for passive and continuous enrollment and authentication.

[0044] Label Correction Using a reorganization clustering operation may require the voice biometric authentication system to implement a set of label correction operations. For example, to accurately transfer labels to anonymous clusters (e.g., newly generated clusters, unassigned clusters) created through the reorganization reclustering operation, the system may calculate pairwise similarities between clusters from the disorganized clusters and clusters organized using the reorganization operation. The system may create a similarity matrix by calculating pairwise similarities between each of the old and new clusters and identify clusters that are most similar to each other as matching clusters. The system may transfer labels from the old clusters to the new matching clusters. Because the system may store associations between labels and information about the labels (e.g., content preferences), the system may create and / or maintain associations between the new clusters and any information that was associated with the previous clusters through the transferred labels.

[0045] Content Personalization and Control Some over-the-top services may create, or use third-party services to create, profiles of individuals when they use the respective services to provide them with content (e.g., image content, video content, audio content, etc.) or content recommendations. These services may do so from active input by the individual (e.g., a user may enter preferences indicating the types of content they like or provide information about themselves, such as their age), or the service may maintain profiles about individuals in a database and identify various types of content that users view while using the over-the-top service. By implementing the systems and methods described herein, the system may use voice data received from the over-the-top service (e.g., via an edge device) to identify individuals viewing content using the service and provide the individual's identifier to services that provide relevant content to the individual.

[0046] For example, a voice biometric authentication system may receive an utterance from a speaker and use machine learning techniques (e.g., clustering and / or neural network architectures) to identify a speaker profile for the speaker from a speaker profile database (sometimes referred to as an “analytics database”) maintained by the system. The system retrieves one or more identifiers for the speaker (or the speaker's household, a group associated with the speaker, etc.) from the speaker profile and indicates the speaker to a media service. The media service may use the identifier to identify a profile (e.g., a consumer profile, a user profile) associated with the identifier from a media content database maintained by the service. The service may provide content and / or content suggestions to the speaker or the speaker's edge device based on the profile associated with the identifier. In some cases, the system may identify content to provide to the speaker or edge device based on the speaker's identifier itself. Because the system may determine the identifier in an unsupervised system as an anonymized identifier (e.g., a hashed version of the identifier), the identifier may maintain the anonymity of the speaker, such that neither the system nor the overtop service can derive personally identifiable information from the identifier regarding a particular individual to whom the system or service provides content.

[0047] In some cases, a voice biometric authentication system may use speaker profiles for age-related parental control. To do so, the system may store an association between a speaker profile and a flag indicating a speaker's age-related characteristic (e.g., whether a person is over a certain age or the speaker's age). The system may obtain the age characteristic associated with the profile via user input or automatically based on the utterances the system uses to build the speaker's profile. In some cases, the system may receive the age characteristic from a third-party service. The system may provide the age-related characteristic to an over-the-top service used to select content to provide to the speaker or otherwise stop the speaker from viewing content if the individual does not meet the age-related speaker characteristic.

[0048] In some cases, a voice biometric authentication system may use the systems and methods described herein to stop a speaker from impersonating another speaker and viewing content associated with the impersonated speaker's speaker profile (e.g., in a replay attack in which an individual plays a recording of another speaker). For example, a child may play a recording of their parents speaking to overcome age-related restrictions on an over-the-top service. The system may detect that the child is playing the recording and, instead of identifying the speaker profile of the child's parent, may generate a warning and / or transmit a signal to the over-the-top service indicating that the child is impersonating his or her parents to stop the service from providing age-restricted content to the child. Thus, the system may determine whether an individual is impersonating another individual and generate a warning to stop the service from providing content to unauthorized users.

[0049] To prevent replay attacks, the server may evaluate additional authentication data received from the speaker and compare the authentication data with expected authentication data, such as additional passwords or passphrases or other required information.

[0050] In some embodiments, the server prevents replay attacks by evaluating various types of data and features for spoofing conditions. The server runs a machine learning architecture model trained to evaluate inbound audio signals for artifacts indicative of spoofing conditions, as described in U.S. Patent Application No. 15 / 910,387, issued as U.S. Patent No. 10,692,502, which is incorporated herein by reference. Replay attacks involving recordings of parental speech may include certain low-level features found in the played-back recording that are not typically present in actual live speech. For example, recorded audio samples may consistently introduce audio artifacts related to frequency, frequency range, dynamic power range, reverberation, noise levels in certain frequency ranges, etc., and at least some of these artifacts may be imperceptible without the use of specialized speech processing techniques and / or equipment, such as those disclosed herein. The machine learning architecture includes a spoof detection classifier trained to identify such spoofing conditions (e.g., quality, artifacts) and authentic inbound speech. The server references the parent's voiceprint to identify potential spoofing conditions. For example, a parent may consistently provide utterances using only a particular device that produces voice signals with a particular low-level voice quality, thus resulting in an embedding based on those qualities and a parent voiceprint based on particular consistent low-level features. The speaker profile may capture and store particular types of low-level features for later use in distinguishing between spoofed and genuine access. When a child inputs an inbound replay utterance, the server determines a similarity score and / or a spoof score based on comparing the inbound replay embedding with the parent's voiceprint score and comparing the inbound replay embedding with the artifact features used to generate the parent's voiceprint score. Additional examples of spoof detection that may be implemented by the server to prevent replay attacks can be found in U.S. Patent Application No. 17 / 192,464, which is incorporated herein by reference.

[0051] Device Personalization and Control Some devices may use speaker profiles, such as those described herein, to configure and / or customize edge devices. For example, the system may associate each voice profile with a vehicle configuration and store the voice profiles locally in the vehicle (or in the cloud with an identifier associated with the vehicle). An individual associated with one speaker profile may be associated with one or more of temperature, radio volume, window settings, etc., while another individual may be associated with different settings. Because the system may store profiles for multiple individuals, when the system captures speech and identifies a speaker profile based on the speech, the system may communicate with other applications on the device to automatically adjust the device's configuration based on the settings in the speaker profile.

[0052] Exemplary System Components Audio processing for media content systems For ease of explanation and understanding, the embodiments described herein describe a computing system that uses voice processing and user data analysis in the context of a content distribution system. However, the embodiments are not limited to such implementations, and the processes described herein may be used in any number of systems that may benefit from passive or active speaker identification, or continuous or static speaker identification, to process individual speaker or multi-speaker voice biometric authentication. Nearly any system that receives, processes, and identifies speakers in audio input may implement the systems and processes for multi-speaker identification or recognition of multi-speaker voice biometric authentication described herein. Non-limiting examples of systems that may use the voice processing and data analysis herein include IoT devices (sometimes referred to as edge devices) (e.g., smart appliances, vehicles), call centers and similar help desks or service centers, secure authentication systems or services (e.g., office or home security), and surveillance or intelligence systems, among others.

[0053] Additionally, embodiments herein use voice processing operations to identify a speaker as a particular known or unknown user. However, embodiments are not limited to voice biometrics and may incorporate and process any number of additional types of biometrics to identify a speaker as a particular user. Non-limiting examples of additional types of biometrics that embodiments may incorporate and process include eye scans (e.g., retina or iris recognition), faces (e.g., facial recognition), fingerprints or handprints (e.g., fingerprint recognition), user behavior (e.g., "behavior prints") when accessing a monitoring system (e.g., key presses, menu access, content selection, speed of input or selection), or any combination of biometric information.

[0054] FIG. 1 illustrates components of a system 100 that uses audio processing machine learning operations. The system 100 includes an analysis system 101, a content system 110, and an end-user device 114. The analysis system 101 includes an analysis server 102, an analysis database 104, and an administrator device 103. The content system 110 includes a content server 111 and a content database 112. Embodiments may include additional or alternative components, or omit certain components from those illustrated in FIG. 1, and still fall within the scope of the present disclosure. For example, it may be common to include multiple content systems 110, or for the analysis system 101 to have multiple analysis servers 102. Embodiments may include any number of devices capable of performing the various features and tasks described herein, or may otherwise implement them. For example, FIG. 1 illustrates the analysis server 102 as a computing device separate from the analysis database 104. In some embodiments, the analysis database 104 is integrated into the analysis server 102.

[0055] The various components of system 100 are interconnected with one another through hardware and software components of one or more public or private networks. Non-limiting examples of such networks may include a local area network (LAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a wide area network (WAN), and the Internet. Communications over the networks may be performed according to various communication protocols, such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communication protocols.

[0056] Overview and Infrastructure The analytics service provides customers, such as media content services or enterprise call centers, with speech processing and computing services for analyzing data received from end users. Non-limiting examples of analytics services include user identification, speaker recognition (e.g., speaker diarization), user authentication, and end-user related data analysis. The analytics service operates an analytics system 101 comprising various hardware, software, and networking components configured to host and provide analytics services to one or more content systems 110 of one or more media content services (e.g., Netflix®, TiVO®).

[0057] The content media service operates a content system 110 comprising hardware, software, and networking components configured to host cloud-based media services, such as over-the-top (OTT) or digital streaming services, that provide media content to end-user devices 114. The content system 110 identifies and provides personalized content for end-users, such as content recommendations, content restrictions (e.g., parental controls), and / or advertisements.

[0058] In operation, the provider server 111 (of the content system 110) receives various types of input data from the end-user device 114 and forwards the input data to the analytics server 102 (of the analysis system 101). The analytics server 102 uses the input data forwarded from the provider server 111 to perform the various analysis processes described herein and then transmits various resulting outputs from the analysis processes to the provider server 111. The provider server 111 uses the output received from the analytics server 102 to identify and generate personalized content based on, for example, user actions or behavior (e.g., viewing habits), interactions between the end-user device 114 and the content system 110, user characteristics (e.g., age), and user identity, among other types of information for content personalization.

[0059] While the content system 110 or the analysis system 101 may typically identify a user based on, for example, subscription information (e.g., subscriber identifier) ​​or user credentials (e.g., username, password), the analysis system 101 described herein additionally or alternatively identifies a user (instead of the content system 110) based on user input and verbal utterances captured by the end-user device 114.

[0060] In some cases, the end user device 114 actively captures user input data, where the end user actively interacts with the end user device 114 (e.g., speaks a "wake" word, presses a button, or makes a gesture). In some cases, the end user device 114 passively captures user input data, where the end user passively interacts with the end user device 114 (e.g., speaks to another user, and the end user device 114 automatically captures the speech without any affirmative action by the user). Various types of input, such as sound or voice data captured by a microphone of the end user device 114 or user input entered through a user interface presented by the end user device 114, represent ways in which a user interacts with the end user device 114. The captured sound includes background noise (e.g., ambient noise) and / or speech of one or more speaking users. Additionally or alternatively, user input may include video (or images) of the user (e.g., facial expressions, gestures) captured by or uploaded to the end user device 114. User input to a user interface may include interface input to a physical or graphical user interface, such as touch input swiping across the device, using the device with gestures, pressing buttons on the device (e.g., dual-tone multi-frequency (DTMF) tones on a keypad), entering text, capturing biometric information such as a fingerprint, etc.

[0061] The content server 111 receives user input from the end-user device 114 as user input data. The content server 111 performs various processing operations on the user input data, such as identifying or extracting various forms of metadata, performing one or more authentication operations using the user input data, anonymizing or obfuscating certain types of data (e.g., generating a hash of one or more identifiers), among other possible operations. The content server 111 may transform, modify, and / or enrich the user input data before transmitting it to the analytics server 102. Non-limiting examples of types of data in the user input data include, among others, audio signals, user interface input (e.g., user requests, user commands), authentication input (e.g., user credentials, biometrics), various identifiers (e.g., subscriber identifiers, user identifiers), and various types of communication metadata.

[0062] The analysis system 101 receives audio signals from the provider server 111 as data files or data streams in any number of machine-readable data formats (e.g., WAV, MP3, MP4, MPEG, JPG, TIF, PNG, MWV). The audio signals may include speech in addition to background noise. In some configurations, before the analysis system 101 receives the audio signals, the content system 110 may tag the audio signals with information about the interaction, such as metadata or speaker-independent characteristics. Non-limiting examples of such information may include the time of the interaction, the date of the interaction, the type of end-user device 114, the microphone type, the location of the interaction (e.g., bedroom, living room, restaurant), the particular end-user device 114 associated with the interaction (e.g., a particular smart TV 114a, virtual assistant 114d), a subscriber identifier associated with the interaction, a unique identifier associated with the interaction (e.g., an automatic number identifier), etc.

[0063] Components of the analysis system 101, such as the analysis server 102, generate voiceprints, update voiceprints, predict similarity scores, identify (or authenticate) speakers in audio signals, and relabel voiceprints to provide user identification (or authentication) services. The analysis system 101 may provide similarity scores (or labels) associated with speakers in audio signals to customers of the analysis system 101. For example, the analysis system 101 may transmit speaker identifiers to the content system 110 based on audio signals associated with a dialogue.

[0064] Content System The content system 110 transmits or streams media data to subscriber end-user devices 114. A subscriber represents a customer of the content system 110, but may also represent a collection of one or more users. For example, a subscriber may represent a household, as well as family and / or guest members, who access the services of the content system 110 using an end-user device 114. A family member may access media content if at least one member of the family registers as a subscriber with the content system 110. The provider server 111 transmits media content to the end-user device 114 for users (including subscribers and the subscriber's guests) based on various user interactions with the end-user device 114.

[0065] Components of the content system 110 capture various types of input data (e.g., metadata) received from end-user devices 114, which the content system 110 adds to the input data that is forwarded over one or more networks to the analysis server 102. A computing device of the content system 110, such as the content server 111, assigns a subscriber identifier to the inbound input data (including audio signals) before forwarding the input data to the analysis system 101. A subscriber identifier is a data value, tag, or other form of data that indicates a subscriber or customer of the content system 110. The subscriber identifier associates various types of data, including audio signals, received by the content system 110 along with a particular input from a particular subscriber with a particular subscriber.

[0066] A subscriber defines or otherwise associates a collection of users (e.g., households) who access the services of the content system 110 using a subscriber identifier common to the collection of users. In such cases, the content database 112 includes a data record for each of the users (speakers), sometimes referred to as a "speaker profile." In operation, the analysis system 101 receives the subscriber identifier and speech signal as input data, but for privacy purposes, the analysis system 101 need not include personally identifiable information. For example, the content server 111 may generate and send an anonymized version of the subscriber identifier or speaker identifier to the analysis system 101, thereby protecting the private information of the subscriber or particular user by preventing the subscriber's private information from being directly transmitted (or otherwise ascertainable) to the analysis system 101.

[0067] The content system 110 stores media and subscriber information in a content database 112, allowing the content system 110 to identify subscribing users (e.g., subscriber accounts) associated with subscriber identifiers. For example, the content database 112 stores subscriber data records in a lookup table that include the subscriber identifier, the audio signal (including the speaker identifier, user characteristics, speaker-independent characteristics, and other metadata received from the analysis system 101), and the subscriber account.

[0068] The content system 110 forwards audio signals (using a subscriber identifier) ​​to the analysis system 101 according to preconfigured trigger conditions. For example, the content system 110 receives an audio signal from an end-user device 114 and forwards the audio signal to the analysis system 101. The audio signal may include, for example, one or more utterances from a user requesting media content (or other services) from the content system 110. The content system 110 may forward the audio signal associated with the request to the analysis system 101 to identify the user in real time for a set of users. The analysis system 101 may transmit speaker information (e.g., speaker characteristics, speaker identifiers, speaker-independent characteristics, metadata) to the content system 110. The content system 110 may use the user information to respond to the speaker request with personalized content. In some cases, the content server 111 may forward the audio signal to the analysis system 101 in response to an instruction or query received from another device in the system 100, such as the analysis server 102 or the administrator device 103.

[0069] Content Server In some embodiments, content server 111 may host and execute software processes and services for identifying speech in audio signals, converting audio signals from one format to a different format (e.g., converting a media file from a WAV file format to an MP3 file format), preprocessing audio signals, anonymizing audio signals (e.g., associating a hash identifier with audio data), extracting biometric features associated with speakers in audio signals, etc. For example, content server 111 is configured to detect audio events in audio signals. Content server 111 may also be configured to perform automatic speech recognition (ASR) on audio signals to capture the content of the audio signals (e.g., user requests to consume content).

[0070] The content server 111 provides content to users who actively interact with the end user device 114. For example, a user may speak into the end user device 114. The content server 111 may also provide content to users who passively interact with the end user device 114 (e.g., to users who speak within a predetermined proximity to the end user device 114). The content server 111 may transmit user interface data or content (e.g., computer files, data streams), including television program or advertising recommendations, to a speaker based on the speaker's identity as indicated by the analytics server 102.

[0071] Content Database The content database 112 of the content system 110 stores various types of data records, including subscriber data, speaker profiles, and media content for streaming to the end-user devices 114. For example, the content database 112 may store a library of content. The content database 112 may also store audio signals, speaker identifiers, speaker characteristics, speaker-independent characteristics, and other metadata associated with interactions received from the end-user devices 114 or the analytics server 102.

[0072] The content database 112 may also store subscriber information such as the account owner or household, the authorized number of users for the account, the authorized number of devices associated with the account, the current number of users associated with the account, the current number of devices associated with the account, the authorized geographic area of ​​operation (e.g., an account may be prohibited in some countries and authorized in others), purchasing options (e.g., requiring a password before any purchase), billing information (e.g., credit card information, billing address, shipping address), identifiers associated with the subscriber (e.g., subscriber identifier, household identifier), speaker identifiers associated with the subscriber identifier, anonymization of information (e.g., hash functions, cryptographic keys), etc.

[0073] A subscriber or user profile in the content database 112 may store a particular speaker's viewing history, speaker information (e.g., name, age, birthday, gender, religion), security credentials (e.g., login credentials, biometrics), and preferences. Non-limiting examples of preferences may include content that the speaker has previously viewed, liked, bookmarked, or in which the user has otherwise expressed interest. In operation, the content server 111 identifies a particular speaker based on login credentials or metadata or as determined by the analytics server 102, and determines particular media for the speaker based on preferences. A subscriber profile includes or is associated with one or more speaker profiles. In some cases, a speaker profile in the content database 112 corresponds to a speaker profile stored in the analytics database 104 (sometimes referred to as a speaker database).

[0074] In some implementations, a subscriber or speaker profile includes content restrictions or controls that instruct the content server 111 to prohibit delivery of certain media to certain speakers (e.g., parental controls for underage users). Content restrictions (in a subscriber or user profile) correspond to appropriate age ratings (e.g., R, PG-13, TV-MA) or content characteristics stored in media data records in the content database 112. Content characteristics are data values ​​that indicate parental / discretionary advisories or extreme or objectionable types of content, such as tobacco / drug use, flashing, nudity, and violence, among others. Content characteristics correspond to user or speaker characteristics stored in the content database 112 or analytics database 104 as user or speaker profile information.

[0075] At least one user profile of the subscriber profile is designated a power user profile (e.g., a parent speaker profile) that has privileges to configure content restrictions for the entire subscriber profile (e.g., household) or for specific users (e.g., child speakers). The power user operates a user interface on the end user device 114 to input various content restriction configurations into the configuration interface. The content restriction configurations indicate and configure content restrictions according to a specific user profile identifier (e.g., username), user age, or specific content / speaker characteristics. The content server 111 can typically determine that a particular user has logged into a content service by referencing the user's user identifier. The content server 111 receives the user identifier in login credentials (which contain the user identifier) ​​or uses a speaker identifier returned from the analytics server 102 to identify the user identifier in the content database 112. The content server 111 uses the user identifier to query the content database 112 to determine the level of privileges assigned to the particular user and any corresponding content restrictions.

[0076] In some embodiments, the configuration restrictions further indicate means for overriding or updating the content restrictions, such as an additional challenge input (e.g., PIN, password) or biometric input (e.g., fingerprint, voiceprint). In operation, after the analytics server 102 identifies a particular user (e.g., speaker) from the input data (e.g., audio signal), the content server 111 consults the content database 112 to identify the content restrictions assigned to the user and applies the content restrictions to content requested, queried, or presented to the user according to appropriate age ratings or content characteristics stored in the media data record.

[0077] End-User Devices The end-user device 114 may be any device that allows a user to control any audio or visual interface or otherwise operate a content service. Non-limiting examples of the end-user device 114 may include IoT devices such as a smart TV 114a, a remote control 114b, a set-top box 114c, or a virtual assistant 114d or other edge device or mobile computing device. The smart TV 114a may be a TV configured to connect to a network, such as the Internet network, and configured with a microphone and / or camera. The TV remote control 114b may be a controller configured to control content displayed on the TV. The set-top box 114c (e.g., a cable box, Slingbox, Apple TV, Roku Streaming Stick, TiVo Stream, Amazon Fire) includes any media streaming and / or storage device with a processor and non-transitory storage media configured to perform the various processes described herein. The set-top box 114c may communicate with the media content system 110 via one or more networks to upload and download various types of content information, user information, and device information. An IoT device may be a telecommunications-oriented device (e.g., a mobile phone) or a computing device configured to implement voice-over-IP (VoIP) telecommunications or other network communications (e.g., cellular, Internet). An IoT device comprises hardware and software components configured to stream data over a TCP / IP network or other computing network channel. A personal assistant 114d may be a virtual assistant device (e.g., Alexa®, Google Home®), a smart appliance, an automobile, or other smart device capable of running software applications and / or performing voice interface operations.

[0078] The end-user device 114 may include a processor and / or software capable of utilizing the communication capabilities of a paired or otherwise networked device. The end-user device 114 may be configured with a microphone, accelerometer, gyroscope, camera, fingerprint scanner, interaction buttons (directional buttons, numeric buttons, etc.), joystick, or any combination. The end-user device 114 may include hardware (e.g., microphone) and / or software (e.g., codec) for detecting and converting sound (e.g., oral speech, ambient noise) into an electrical audio signal. While the content system 110 may collect and store audio signals in the content database 112, the analysis system 101 typically avoids storing audio signals or purges any audio signals stored in memory.

[0079] The content server 111 receives input data from the end-user device 114 and may perform various pre-processing operations, such as converting the audio signal of the input data, identifying and associating a subscriber identifier with the input data, or storing various characteristics of the audio signal in the content database 112. The content server 111 then forwards the audio signal (and subscriber identifier) ​​to the analysis server 102 (of the analysis system 101) over one or more networks.

[0080] The content server 111 temporarily stores the audio signal in a non-transitory machine-readable storage medium, such as a buffer or cache memory, for a predetermined period of time. Additionally or alternatively, the content server 111 stores the audio signal in a content database 112.

[0081] The content server 111 forwards the input data (e.g., audio signals, metadata) to a computing device (e.g., the analysis server 102) of the analysis system 101. In some configurations, the content server 111 transmits the input data to the analysis server 102 according to a preconfigured trigger condition, such as a predetermined interval, or in response to the content server 111 receiving an input audio signal from an end-user device 114.

[0082] In some embodiments, the content server 111 or the end user device 114 continuously captures and stores audio recordings, even before the end user device 114 detects active user input (e.g., a wake word). For example, when a group of people discuss what to watch, the end user device 114 captures audio. After the group decides on a particular program to watch, the group typically becomes silent and one person utters the wake word, announcing the program the group has decided to watch. Because the group is silent, the end user device 114 and content server 111 benefit from capturing and storing the audio of the group discussion for a period of time before the wake word. In this way, the content server 111 and the analytics server 102 identify audio signals that begin some time before the active input. In such embodiments, when a user operates the end user device 114 to actively capture audio, the content system 110 retrieves the stored audio signal captured by the end user device 114 some time before the user actively operated the end user device 114. The content system forwards both the audio signal associated with the trigger condition and the retrieved audio signal to the analysis server 102 .

[0083] The analysis server 102 uses the received audio signal and, in some embodiments, additional types of data (e.g., subscriber identifier, user credentials, metadata) to determine a speaker identifier. The analysis server 102 transmits the speaker identifier (and speaker characteristics, speaker-independent characteristics, and other metadata) to the content server 111. The content server 111 maps the speaker identifier to a subscriber identifier and / or to information about a particular speaker to determine the content that the content server 111 should deliver to the speaker.

[0084] Analysis Server The analysis server 102 may be any computing device equipped with one or more processors and software and capable of performing the various processes and tasks described herein. The analysis server 102 may host or communicate with databases 112 and 104 and may receive audio signals, speaker-independent characteristics, and subscriber identifiers from the content system 110. While FIG. 1 shows a single analysis server 102, the analysis server 102 may include any number of computing devices. In some configurations, the analysis server 102 may comprise any number of computing devices operating in a cloud computing or virtual machine configuration. In some embodiments, a computing device of the content system 110 (e.g., the content server 111) partially or completely performs the functions of the analysis server 102.

[0085] The analytics server 102 executes various software-based processes that take input audio signals (e.g., audio recordings of speaker utterances, subscriber identifiers, user identifiers, metadata), for example, from the content server 111, query the analytics database 104, and apply various machine learning operations to the audio data. The machine learning algorithms implement any number of techniques or algorithms (e.g., Gaussian matrix models (GMMs), neural networks) to perform various operations described herein, such as detecting audio events, extracting embeddings, generating or updating enrolled voiceprints, and identifying / authenticating one or more users whose utterances are in the audio signal.

[0086] The analysis server 102 queries speaker profiles stored in the analysis database 104 to identify known or new speakers in the audio signal, generate new or temporary speaker profiles, and / or update speaker profiles in the analysis database 104. Using the subscriber identifier or other metadata received with the audio data, the analysis server 102 identifies a voiceprint (e.g., a suspect voiceprint) associated with the received subscriber identifier, creates a similarity matrix between pairs of embeddings, clusters similar embeddings based on the distance of each of the embeddings, creates a similarity matrix of similarity scores, determines a maximum similarity score using various thresholds, determines strong and weak embeddings, stores the weak embeddings in the analysis database 104, updates the voiceprints using the strong embeddings, and identifies the speaker by evaluating the voiceprint associated with the maximum similarity score.

[0087] Embedding Extraction The analytic server 102 executes machine-executable software for implementing one or more machine learning architectures, each including any number of layers configured to perform specific operations, such as audio data ingestion, preprocessing operations, data augmentation operations, embedding extraction, loss function operations, and classification operations, among others. To perform various operations, the one or more machine learning architectures include any number of models or layers, such as an input layer, an embedding extractor layer, a fully connected layer, a loss layer, and a classifier layer, among others. The analytic server 102 executes audio processing software including one or more machine learning models and layers. For ease of explanation, the analytic server 102 is described as executing a single machine learning architecture with an embedding extractor, although in some embodiments, multiple machine learning architectures (including neural network architectures) may be used.

[0088] The analysis server 102 receives the audio signal from the content server 111 and extracts various types of features from the audio signal. The analysis server 102 performs audio event detection or other voice activity detection to distinguish between background noise, silence, and speakers in the audio signal. For example, the analysis server 102 may preprocess the audio data (e.g., filter the audio signal to reduce noise, parse the audio signal into frames or subframes, and perform various normalization or scaling operations), run voice activity detection (VAD) software or VAD machine learning, and / or extract features (e.g., one or more spectrotemporal features) from portions (e.g., frames, segments) or from substantially all of the audio signal. Features extracted from the audio signal may include Mel Frequency Cepstral Coefficients (MFCCs), Mel filter banks, linear filter banks, bottleneck features, etc.

[0089] In some embodiments, the content server 111 applies a VAD or ASR engine to the input audio signal that the content server 111 receives from the end user device 114. The content server 111 transmits the speech portion of the input audio signal to the analysis server 102, which applies a machine learning model for speaker recognition (e.g., an embedding extractor) to the input audio signal.

[0090] The analysis server 102 extracts embeddings from the audio signal using a neural network architecture (e.g., a deep neural network (DNN), a convolutional neural network (CNN)), a Gaussian mixture model (GMM), or other machine learning methods. The analysis server 102 may represent the embeddings using an x-vector, a CNN-vector, an i-vector, etc.

[0091] As an example, the analysis server 102 may train a machine learning architecture to perform a VAD operation that parses a set of speech portions and a set of non-speech portions from an audio signal. When the VAD is applied to features extracted from the audio signal, the VAD may output a binary result (e.g., speech detected, no speech detected) or a continuous value (e.g., the probability that speech occurs) for each frame (or subframe) of the audio signal. The speech portions of the audio signal may be referred to as speech. The audio signal may include speech from multiple speakers. The audio signal may also include overlapping sounds (e.g., speech and ambient background noise). The analysis server 102 may use speaker detection or other conventional speaker segmentation solutions to determine the start and end of speech.

[0092] In some embodiments, the content server 111 executes a VAD or ASR machine learning model to identify speech portions in user input that the content server 111 receives from the end-user device 114. In such embodiments, the content server 111 transmits speech portions of an audio signal containing speech from one or more speakers to the analysis server 102. The analysis server 102 does not need to perform VAD or ASR operations before extracting the embeddings.

[0093] Replay attack calculation The analysis server 102 may determine whether the audio signal captured by the end user device 114 is an authentic audio signal or a replay attack (e.g., an audio signal captured by a microphone in a physical and reverberant space and presented to the microphone of the end user device 114 using a replay device). In determining whether the speech portion of the audio signal is a replay attack, the analysis server 102 may incorporate, determine, or query the content system 110 for speaker-independent characteristics such as the end user device 114 type and microphone type. Additionally, the analysis server 102 may apply one or more trained machine learning models to the inbound audio signal to identify spoofing conditions based on various artifacts in the inbound audio signal and corresponding artifact features in the parent's voiceprint.

[0094] Active and Static Registration The analysis server 102 can use a machine learning architecture to recognize a particular speaker during the enrollment phase of a particular enrollee speaker. The machine learning architecture can use the enrollee speech signal having speech segments (or utterances) with the enrollee to generate an enrollee voice feature vector (sometimes called a "voiceprint"). During subsequent active or passive interaction with the end-user device 114, the analysis server 102 extracts embeddings from the captured speech signal and compares the embeddings to the voiceprint to determine whether the subsequent captured speech signal involves the enrollee.

[0095] The analytic server 102 registers users during a registration phase (e.g., a predetermined registration time). For example, the analytic server 102 may actively register users during initialization of a new end-user device 114. Additionally or alternatively, the analytic server 102 may actively register users annually (e.g., to update existing voiceprints and generate new voiceprints for new users).

[0096] During the enrollment phase, the analysis server 102 may prompt the user for an enrollment phrase that the user repeats until the analysis server 102 receives enough speech (e.g., the speech portion of the audio signal) to recognize a particular speaker using the speaker's voiceprint.

[0097] The analysis server 102 creates a voiceprint based on the enrollment embeddings during the enrollment phase. During the enrollment phase, the analysis server 102 extracts embeddings from one or more utterances so that the analysis server 102 can mathematically identify the user of a particular signal. The analysis server 102 may determine that sufficient embeddings have been extracted during the enrollment phase when the analysis server 102 receives a net utterance of duration exceeding a threshold. Additionally or alternatively, the analysis server 102 may determine that sufficient embeddings have been extracted during the enrollment phase when the analysis server 102 receives a predetermined number of enrollment signals (e.g., two fingerprint scans and two utterances, five different utterances).

[0098] Continuous and passive registration Additionally or alternatively, the analysis server 102 may continuously enroll users. Instead of enrolling users for a predetermined (or specified) period, as in static enrollment, the analysis server 102 may enroll users at any time by creating voiceprints associated with received utterances. Because the utterances received during continuous passive enrollment can be of any duration or quality, the maturity of the voiceprints created during continuous passive enrollment may vary.

[0099] For example, the analysis server 102 may receive an audio signal (and subscriber identifier) ​​forwarded from the content server 111. The analysis server 102 may use VAD software, an embedding extractor model, or other machine learning models to extract features and embeddings from the audio signal. The analysis server 102 applies a machine learning architecture to the audio signal to extract embeddings for a particular speaker.

[0100] Unsupervised Clustering In some embodiments, the analytics server 102 may receive a subscriber identifier forwarded from the content server 111. The analytics server 102 queries the analytics database 104 to retrieve a voiceprint associated with the received subscriber identifier. The voiceprint is an embedded estimate of the voiceprint based on the association with the subscriber identifier. Each voiceprint and associated unique speaker identifier is linked to at least one subscriber identifier.

[0101] In some embodiments, the analysis server 102 may not receive the subscriber identifier forwarded from the content server 111. Instead of evaluating the similarity of the embeddings and voiceprints associated with a putative speaker cluster (a speaker cluster is a putative speaker cluster based on an association with a subscriber identifier), the analysis server 102 will evaluate the similarity of the embeddings with a set of voiceprints. The set of voiceprints may include voiceprints recently (e.g., within a predetermined time) transmitted to the analysis database 104, voiceprints associated with particular speaker characteristics, and voiceprints associated with particular speaker-independent characteristics.

[0102] The analytic server 102 may cluster (or otherwise associate similar embeddings) using a sequential clustering algorithm (e.g., k-means clustering). The analytic server 102 may cluster voiceprints by creating a similarity matrix and determining the similarity of each of the clusters (voiceprints) to other voiceprints. In some configurations, the analytic server 102 may evaluate the similarity of a cluster to a voiceprint by evaluating the distance from each of the embeddings in the cluster to the voiceprint's centroid. If the cluster similarity meets one or more thresholds, the analytic server 102 may merge the two voiceprints into a single voiceprint (e.g., by averaging the voiceprints).

[0103] In some implementations, the analytic server 102 clusters inbound and / or stored embeddings, for example, by randomly generating centroids and associating embeddings with the centroids. The analytic server 102 clusters the embeddings based on the relative distance between the embeddings and the centroids. The analytic server 102 moves the centroids to new relative positions based on minimizing the average distance of each of the embeddings associated with the centroid. Each time a centroid is moved, the analytic server 102 recalculates the distance between the embeddings and the centroids. The analytic server 102 repeats the clustering process until a stopping criterion is met (e.g., the embeddings do not change clusters, the sum of distances is minimized, or a maximum number of iterations is reached). In some configurations, the analytic server 102 measures the distance between the embeddings and the centroids using Euclidean distance. In some configurations, the analytic server 102 measures the distance between the embeddings and the centroids based on the correlation of features in the embeddings. The distance between the embeddings and the centroids is indicated using a similarity score. The more similar the embedding is to the centroid, the higher the similarity score. The analysis server 102 tracks the similarity scores between each of the embeddings and each of the centroids of the similarity matrix.

[0104] Additionally or alternatively, the analytic server 102 may treat each embedding as a centroid. The analytic server 102 clusters the embeddings based on the distance from the centroid embedding to other embeddings. Distance measures may include, for example, the smallest maximum distance to other embeddings, the smallest average distance to other embeddings, and the smallest sum of squared distances to other embeddings.

[0105] The analysis server 102 may also cluster the embeddings with voiceprints using a sequential clustering algorithm. A cluster represents a collection of utterances similar to a particular speaker (e.g., a speaker cluster), where the voiceprint represents the centroid of the speaker cluster. The analysis server 102 uses a speaker identifier to identify the speaker cluster. The speaker identifier protects the speaker's private information by anonymizing the speaker and distinguishing one speaker cluster from another.

[0106] In some configurations, the analysis server 102 associates metadata (or speaker characteristics, speaker identifiers) with the voiceprint. The metadata associated with the voiceprint can include the quality of the audio signal. For example, the audio signal may contain speech in clear conditions (e.g., high signal-to-noise ratio (SNR), low reverberation time (T60)). Additionally or alternatively, the audio signal may contain speech in noisy conditions (e.g., low SNR, high T60). The metadata can also include the overall duration of the net utterance. The analysis server 102 can determine the duration of the utterance by summing the duration of each of the utterances in a speaker cluster. The metadata can also include the total number of utterances in the speaker cluster, speaker characteristics, speaker identifiers, and / or speaker-independent characteristics.

[0107] The analysis server 102 may also determine speaker characteristics based on information input by the speaker, information received from a content server, or information identified by running various machine learning models. Speaker characteristics may include, for example, the speaker's age, the speaker's gender, the speaker's emotional state, the speaker's dialect, the speaker's accent, and the speaker's tone, among others. In some embodiments, for example, the analysis server 102 may identify specific speaker age characteristics by applying machine learning models, such as those described in Sadjadi et al., “Speaker Age Estimation on Conversational Telephone Speech Using Senone Posterior Based I-Vectors,” IEEE ICASSP, 2016, and Han et al., “Age Estimation from Face Images: Human vs. Machine Performance,” ICB 2013. In some embodiments, the analysis server 102 may identify specific speaker gender characteristics by applying machine learning models, such as those described in Buyukyilmaz et al., “Voice Gender Recognition Using Deep Learning,” Advances in Computer Science, 2016. Each of the references listed above in this paragraph is incorporated herein by reference.

[0108] The analysis server 102 may create a voiceprint associated with a speaker based on one or more embeddings. In some configurations, the analysis server 102 creates the voiceprint using mature embedding clusters (e.g., clusters based on enrollment embeddings). The mature voiceprint (or mature embedding clusters) may contain sufficient biometric information for the analysis server 102 to identify the speaker using a speaker identifier. In some configurations, the analysis server 102 may create the voiceprint by averaging the enrollment embeddings. In some configurations, the analysis server 102 may associate metadata (or user characteristics, speaker identifiers) with the voiceprint.

[0109] Metadata associated with a voiceprint may include information about the utterances (represented by embeddings) in the voiceprint. The metadata may include the quality of the audio data associated with a particular embedding (or voiceprint). For example, a microphone may capture audio data in clear conditions (e.g., high signal-to-noise ratio (SNR), low reverberation time (T60)). The microphone may also capture audio data in noisy conditions (e.g., low SNR, high T60). The metadata may also include the overall duration of the net utterance. The analysis server 102 may determine the duration of an utterance by summing the duration of each utterance in the voiceprint. The metadata may also include the total number of utterances in a cluster. For example, the analysis server 102 may receive one 10-second utterance. Additionally or alternatively, the analysis server 102 may receive five 1-second utterances. The metadata may also include user characteristics, speaker identifiers, and / or speaker-independent characteristics.

[0110] The analysis server 102 generates a similarity matrix of similarity scores. The analysis server 102 determines the similarity score for each particular speaker by evaluating the relative distance between the particular speaker's extracted embedding and the voiceprints of other putative speakers stored in the database. Additionally or alternatively, the analysis server 102 may evaluate the similarity scores using cosine similarity or probabilistic linear discriminant analysis (PLDA). The analysis server 102 determines the embedding that is most similar to the speaker cluster by identifying the maximum similarity score between each of the embeddings and the speaker cluster.

[0111] Strong and weak speech The analytic server 102 may evaluate the maximum similarity score using various thresholds. For example, the analytic server 102 may compare the maximum similarity score to both a lower threshold and an upper threshold. Additionally or alternatively, the analytic server 202 may use one or more algorithms to combine the upper and lower thresholds to estimate an optimal threshold for a particular voiceprint.

[0112] If the analysis server 102 determines that the maximum similarity score for a particular embedding does not reach the low similarity threshold, the analysis server 102 determines that the speaker is likely a new, unknown user. The analysis server 102 will use the particular embedding to generate a new speaker profile. The speaker profile includes a voiceprint, a speaker cluster (e.g., an embedding associated with the voiceprint), a speaker identifier, and metadata.

[0113] If the analysis server 102 determines that the maximum similarity score satisfies the low similarity threshold, the analysis server 102 may identify (or authenticate) the speaker. Additionally, the analysis server 102 determines, based on the weak utterance, that the embedding contributing to the similarity score is a weak embedding. A weak embedding lacks sufficient similarity to the corresponding voiceprint to readily characterize the weak embedding as part of a particular speaker cluster. The analysis server 102 may store and / or update the set of weak embeddings in the analysis database 104.

[0114] The analytic server 102 may evaluate the set of weak embeddings based on trigger criteria. Non-limiting examples of trigger criteria include periodic weak embedding evaluation and a threshold number of stored weak embeddings. In response to the analytic server 102 identifying the trigger criteria, the analytic server 102 may recalculate the similarity score of each embedding in the set of weak embeddings to the voiceprint associated with the set of weak embeddings. In response to the similarity score exceeding the threshold, the analytic server 102 may update the voiceprint with the weak embedding. Additionally or alternatively, the analytic server 102 may remove one or more weak embeddings from the set of weak embeddings.

[0115] If the analysis server 102 determines that the maximum similarity score satisfies the high similarity threshold in addition to the low similarity threshold, the analysis server may determine that the embedding contributing to the similarity score is a strong embedding instead of a weak embedding. A strong embedding is an embedding that is very similar to the voiceprint (e.g., close in terms of relative distance). The analysis server 102 updates the speaker cluster to include this embedding. The analysis server 102 may recalculate the voiceprint based on this new embedding. The analysis server 102 may weight embeddings identified as strong embeddings differently from embeddings that are not identified as strong embeddings. For example, the analysis server 102 may update the voiceprint by taking a weighted average of the embeddings. Additionally or alternatively, the analysis server 102 may update the list of strong embeddings associated with known speakers.

[0116] The analytic server 102 may query the analytic database 104 for the set of weak embeddings based on a trigger condition. For example, the analytic server 102 may query the analytic database 104 periodically (e.g., weekly) or when a predetermined number of sets of weak embeddings is reached. The analytic server 102 may recalculate the similarity score of each of the embeddings in the set of weak embeddings to the voiceprint associated with the set of weak embeddings. Based on the maximum similarity score exceeding various thresholds (e.g., lower and / or higher thresholds), the analytic server 102 may update the voiceprint with a weak embedding. A weak embedding may become a strong embedding as the voiceprint evolves over time (e.g., becomes older, more accurate, and more mature). Additionally or alternatively, the analytic server 102 may remove one or more weak embeddings in the set of embeddings associated with the voiceprint.

[0117] Condition-dependent adaptive thresholding The analysis server 102 may determine the upper and lower thresholds used in adaptively determining strong / weak embeddings. Using condition-dependent adaptive thresholding, the analysis server 102 determines the upper and lower thresholds based on the maturity of the voiceprint associated with the maximum similarity score.

[0118] If the analysis server 102 determines that the voiceprint is mature (e.g., meets the maturity threshold), the analysis server 102 may not use the high and low similarity thresholds when evaluating the maximum similarity score. For example, the analysis server 202 may use one or more algorithms to combine the high and low similarity thresholds to estimate an optimal threshold for a particular voiceprint. The analysis server 102 uses the optimal threshold in evaluating the maximum similarity score for the embedding and voiceprint.

[0119] The analysis server 102 may use one or more maturity factors to determine whether a voiceprint is mature. Non-limiting examples of maturity factors include the number of enrollment utterances, the overall duration of the net speech throughout each utterance, and the quality of the audio from the audio signal associated with the voiceprint. The analysis server 102 may use any number of algorithms to determine whether a voiceprint is mature. For example, the server may compare a maturity factor (e.g., the number of utterances) to a pre-configured maturity threshold corresponding to the maturity factor (e.g., a threshold number of utterances). As another example, the server may statistically or algorithmically combine maturity factors and compare the combined maturity factor to a pre-configured maturity threshold corresponding to the combined maturity factor.

[0120] If the analysis server 102 determines that the voiceprint is not mature (e.g., does not meet the maturity threshold), the analysis server 102 uses condition-dependent adaptive thresholding to determine high and low similarity thresholds for the particular voiceprint.

[0121] The analysis server 102 minimizes the false acceptance rate (FAR) by utilizing a condition-dependent adaptive threshold. The false acceptance rate is the rate at which the analysis server 102 falsely authenticates and / or falsely identifies a speaker. An administrator (e.g., using the administrator device 103) may determine (or pre-configure) the FAR (e.g., 0.5%, 1%, 2%, 3%, 4%, or 5%). Additionally or alternatively, the content system 110 may request a specific FAR associated with their speaker identification / authentication, or a machine learning model may algorithmically determine the FAR.

[0122] Because the FAR varies based on the maturity of the voiceprint, the analytics server 102 applies different thresholds (e.g., high and low similarity thresholds) to different voiceprints when evaluating the similarity of embeddings with the voiceprint. In some configurations, the analytics database 104 may store a table of thresholds (e.g., a threshold scheduler) at various FARs for one or more particular conditions.

[0123] output The analysis server 102 may transmit one or more speaker identifiers based on the audio signals to the content server 111. Additionally or alternatively, the analysis server 102 may transmit a similarity score associated with the speaker identifier to the content server 111. The content server 111 may map the received speaker identifier to a human speaker (and to a subscriber identifier if the speaker identifier was not previously associated with the subscriber identifier). For example, the analysis server may use a lookup table to map the speaker identifier to a particular human speaker. The analysis server 102 may also transmit a speaker profile (including speaker-independent characteristics, speaker characteristics, and metadata) associated with the speaker identifier.

[0124] The content system 110 may store user preferences associated with the speaker identifier. Non-limiting examples of preferences may include content the user has previously viewed, liked, bookmarked, or otherwise expressed interest in. The content system 110 may stream personalized content to the speaker based on the received speaker identifier. If no preferences associated with the speaker identifier are stored (e.g., a new user), the content system 110 may stream generic content to the user.

[0125] The analytics server 102 may identify an environmental setting that describes the speaker or the speaker's current situation and environment. The machine learning models executed by the analytics server 102 may include audio event classification models and / or environmental classification models, such as background noise or specific sounds (e.g., dishwasher, truck) that are classifiable, or the overwhelming amount of energy for the inbound signal. The analytics server 102 may transmit the speaker identifier to the content server 111 along with an indicator of the environmental setting associated with a particular content characteristic. For example, a speaker interacting with a smart TV 114a at a restaurant or party with only adult speakers may cause the content server 111 to generate different suggested content for the end-user device 114 than a different situation where the speaker interacts with a smart TV 114a in a living room with child speakers.

[0126] In some configurations, the content server 111 references the output of the analytic server 102 to restrict access to particular subscriptions and the number of authorized users. If the analytic server 102 receives a subscriber identifier from the content server 111, the analytic server 102 may use the speaker profile (and associated speaker identifiers, speaker characteristics, speaker-independent characteristics, and metadata) to determine whether a speaker is authorized to access a subscriber account based on the authorization rules and restrictions associated with the particular subscriber identifier. The analytic server 102 may transmit an indication to the content system 110 of whether the speaker identifier is authorized for the particular subscriber identifier. Additionally or alternatively, the content server 111 may use the speaker profile information to authenticate a speaker for particular restricted content. For example, the age of a speaker identified in a speaker profile may authorize one or more speakers to consume age-restricted content (e.g., based on parental control or a particular appropriate age rating).

[0127] In some embodiments, the analysis server 102 enables or instructs the content server 111 to enable age-restricted content based on input data (e.g., inbound audio signals, authentication data, metadata, end-user device data 114) that the analysis server 102 receives from the content server 111.

[0128] Label Correction In some configurations, the analysis system 101 corrects label identifiers (e.g., speaker identifiers, subscriber identifiers). The analysis system 101 corrects labels by clustering voiceprints (e.g., hierarchical clustering). Correcting label identifiers minimizes the likelihood of small cumulative identification / recognition errors and increases the purity of speaker clusters.

[0129] The analysis server 102 may correct the label identifiers of the voiceprints created during the deployment phase. The analysis server 102 corrects the label identifiers before they are transmitted to the content server 111. The analysis server 102 may also correct the label identifiers of the voiceprints stored in databases (e.g., content database 112, analysis database 104).

[0130] The analytics server 102 corrects the label identifiers in response to identifying criteria that trigger label correction. Non-limiting examples of trigger criteria include, among others, periodic time intervals or a preconfigured label correction schedule, performing a clustering or reclustering operation, identifying or otherwise receiving a certain number of new speaker identifiers, or generating a certain number of voiceprints.

[0131] In some configurations, even if the analysis server 102 identifies the trigger criteria, the analysis server 102 may determine not to correct the label identifier. If the voiceprint retrieved by the analysis server 102 is associated with a high confidence, the analysis server 102 may determine not to correct the label identifier. The analysis server 102 may associate the confidence based on whether the analysis server 102 created the voiceprint during an active enrollment phase or a passive enrollment phase. The analysis server may determine that a voiceprint created during an active enrollment phase has a higher confidence than a voiceprint created during a passive enrollment phase. A voiceprint created based on active enrollment embedding may be considered pure and mature. Additionally or alternatively, the analysis server 102 may determine not to correct the label identifier if the analysis system 101 is running slowly.

[0132] The analytics server 102 may query a database (e.g., the content database 112 or the analytics database 104) and retrieve a set of speaker profiles (including voiceprints, speaker clusters including embeddings, subscriber identifiers, or speaker identifiers). The set of speaker profiles may be recently accessed and / or modified speaker profiles (e.g., speaker profiles retrieved by the analytics server in the past two days), speaker profiles associated with particular speaker characteristics, speaker profiles associated with particular speaker-independent characteristics, and / or speaker profiles associated with other metadata.

[0133] In some configurations, the analysis server 102 may calculate pairwise similarities between the retrieved voiceprints and the embeddings extracted from the audio signal. The retrieved voiceprints are considered the old labeled set, and the embeddings from the audio signal are considered the new anonymous set. Additionally or alternatively, the analysis server 102 may calculate pairwise similarities between the retrieved voiceprints.

[0134] The analysis server 102 migrates label identifiers associated with the old labeled set to the new anonymous set based on the similarity between the voiceprints in the new anonymous set and the voiceprints in the old labeled set. The analysis server 102 determines the similarity of the voiceprints by evaluating close voiceprints (e.g., by a Euclidean distance measure, a correlation-based measure). If the analysis server 102 determines that the voiceprints are close (e.g., the relative distance satisfies a threshold), the voiceprints and associated label identifiers may be merged. The analysis server 102 may migrate label identifiers associated with the old labeled set to the new anonymous set such that the speaker identifiers and / or subscriber identifiers in the new anonymous set replace the speaker identifiers and / or subscriber identifiers in the old labeled set. In addition to migrating the label identifiers, the analysis server 102 may determine a new centroid of the merged voiceprint by averaging the centroid of the old labeled set and the centroid of the new anonymous set. In some configurations, the analytic server 102 may compare user characteristics before migrating the labels of the old labeled set to the new anonymous set. In some configurations, the analytic server 102 updates the analytic database and / or the content database with the migrated labels.

[0135] User authentication and parental control As discussed herein, the analytics server 102 may determine the identity of a user interacting with the end-user device 114 by comparing the similarity of the extracted embeddings with the embeddings / voiceprints stored in the analytics database 104. Once the analytics server 102 identifies one or more users, it may transmit user identifiers, user characteristics (e.g., age, gender, emotion, dialect, accent, etc.), user-independent characteristics, and / or metadata to the content server 111. In some configurations, the analytics server 102 (or the content server 111 using the information transmitted from the analytics server 102) may use the transmitted information to authenticate the identified users. For example, the user's age may authorize the user to view content that is above a certain age limit.

[0136] In some configurations, the analysis server 102 (or the content server 111) may determine whether a user is authorized to view content based on the identified user. For example, the analysis server 102 may identify an 8-year-old boy watching television. The analysis server 102 may identify another speaker with elevated privileges based on an analysis of the audio signal, where the analysis server 102 identifies speaker profiles of two speakers whose voiceprints match the extracted embeddings of the two speakers. For example, the analysis server 102 may identify the parent of a child in the same audio signal as the 8-year-old boy. The presence of an adult male in proximity to the 8-year-old boy may result in the analysis server 102 (or the content server 111) approving the 8-year-old boy to view certain content.

[0137] analytical database The analytics database 104 may store the FAR of a particular content system 110, speaker identifiers (and voiceprints) associated with subscriber identifiers (e.g., lookup tables), extracted embeddings (e.g., weak embeddings), user characteristics, trained machine learning models (e.g., for extracting embeddings, for performing VAD operations), etc.

[0138] The analysis database 104 may store the cluster embedding as a voiceprint if the embedding within the cluster satisfies one or more thresholds (e.g., the analysis server 102 determines that the cluster is a mature enrollment cluster). The analysis server 102 may determine that an enrollment cluster is mature if the duration of the utterances within the cluster (represented by the embedding within the cluster) satisfies a threshold, the number of utterances within the cluster satisfies a threshold, some combination, etc.

[0139] Additionally or alternatively, the analytics database 104 may store the clustered embeddings as voiceprints if the clustered embeddings are not associated with a speaker identifier and / or a subscriber identifier. The analytics database 104 may store voiceprints even if the voiceprints are not mature.

[0140] In some configurations, the analytical database 104 may purge (delete, erase) stored voiceprints. For example, the analytical database 104 may retrieve instructions from the content server 111 to delete a speaker identifier associated with a subscriber identifier (e.g., a subscriber may decide to unsubscribe from the content system 110's services). Additionally or alternatively, the analytical database 104 may delete a stored voiceprint after a predetermined amount of time has passed. Additionally or alternatively, the analytical database 104 may delete a stored voiceprint if the analytical database 104 (or analytical server 102) determines that the voiceprint meets certain criteria. For example, the embedding in the voiceprint cluster is based on synthetic speech.

[0141] During operation, the analytical database 104 may receive audio data and a subscriber identifier from the content server 111. The analytical database 104 may look up a speaker identifier (and associated voiceprint) associated with the subscriber identifier. The analytical database 104 may also look up upper / lower thresholds associated with the voiceprint and any user characteristics. The analytical database 104 may also look up FARs associated with the content system 110. The analytical database 104 may update the voiceprint entry, upper / lower thresholds, FARs, and user characteristics in the lookup table. The analytical database 104 may also add the voiceprint entry, subscriber identifier, upper / lower thresholds, FARs, and user characteristics to the lookup table.

[0142] Administrator Device The administrator device 103 of the analysis system 101 is a computing device that enables personnel of the analysis system 101 to perform various administrative tasks or user-performed identification, security, or authentication operations. The administrator device 103 may be any computing device equipped with a processor and software and capable of performing the various tasks and processes described herein. Non-limiting examples of the administrator device 103 may include a server, a personal computer, a laptop computer, a tablet computer, etc. In operation, a user may use the administrator device 103 to configure the operation of various components within the system 100, such as the analysis server 102, and may further enable the user to issue queries and commands to various components of the system 100. For example, the administrator device 103 may be used to determine the FAR associated with the content system 110.

[0143] Client-side audio processing for media content systems For ease of explanation and understanding, embodiments described herein refer to using such techniques in the context of a content distribution system that operates in part according to speaker utterances and voice input. However, embodiments are not so limited and may be used in any number of systems or products that may benefit from passive (or active) enrollment, continuous (or static) enrollment, or continuous identification / recognition of multi-speaker voice biometrics. For example, the multi-speaker voice biometric systems and identification / recognition operations described herein may be implemented in any system that receives and identifies audio input (e.g., a car or smart appliance, or an edge device / IoT device such as a call center).

[0144] Additionally, embodiments herein use voice processing operations to identify a speaker as a particular known or unknown user. However, embodiments are not limited to voice biometrics and may incorporate and process any number of additional types of biometrics to identify a speaker as a particular user. Non-limiting examples of additional types of biometrics that embodiments may incorporate and process include eye scans (e.g., retina or iris recognition), faces (e.g., facial recognition), fingerprints or handprints (e.g., fingerprint recognition), user behavior (e.g., "behavior prints") when accessing a monitoring system (e.g., key presses, menu access, content selection, speed of input or selection), or any combination of biometric information.

[0145] FIG. 2 illustrates components of a system 200 that uses audio processing machine learning operations, with machine learning models and other machine learning architectures implemented on a local device. The system 200 includes an analysis system 201, a media content system 210, and an end-user device 214. The analysis system 201 includes an analysis server 202, an analysis database 204, and an administrator device 203. The content system 210 includes a content server 211 and a media content database 212. Embodiments may include additional or alternative components, or omit certain components from those illustrated in FIG. 2, and still fall within the scope of the present disclosure. For example, it may be common to include multiple content systems 210, or for the analysis system 201 to have multiple analysis servers 202. Additionally or alternatively, the analysis system 201, or portions of the analysis system 201, may be embedded in the end-user device 214. Embodiments may include or otherwise implement any number of devices capable of performing the various features and tasks described herein. 2 depicts the analytic server 202 as a separate computing device from the analytic database 204. In some embodiments, the analytic database 204 is integrated into the analytic server 202.

[0146] System 200 of Figure 2 is similar in operation to system 100 of Figure 1, except that in Figure 2, end user device 214 performs various audio processing and data analysis operations. For example, in response to a user's active interaction with end user device 214, end user device 214 captures user input data (e.g., an audio signal). Instead of end user device 214 forwarding the captured audio signal to content system 210 (as in Figure 1), one or more machine learning architectures are applied to the audio signal before forwarding it to content system 210 and / or analysis system 201.

[0147] In some configurations, the end-user device 214 may host and execute software processes and services to identify or extract various forms of metadata, modify, convert, or enrich the audio signal, identify speech in the audio signal, extract biometric features associated with speakers in the audio signal, and identify / authenticate speakers in the audio signal. Additionally or alternatively, the end-user device 214 may filter the audio signal (de-noise the audio signal), convert the format of the audio signal, parse (or segment) the audio signal, run VAD software (or VAD machine learning), perform ASR, and scale the audio signal.

[0148] In some configurations, the end user device 214 stores audio signals from a database or other source in a non-transitory, machine-readable storage medium, such as a buffer or cache memory, for a predetermined period of time. In response to a trigger condition, the end user device 214 may retrieve the stored audio signal and forward the stored audio signal and the audio signal associated with the trigger condition to the analysis system 201. For example, if a user actively operates the end user device 214 to capture sound, the end user device 214 may retrieve the stored audio signal captured by the end user device 214 for the time immediately prior to the user actively operating the end user device 214. In some embodiments, the end user device 214 forwards the audio signal to the analysis system 201 to process the audio signal and identify speakers in the audio signal. Additionally or alternatively, the end user device 214 may process the audio signal and identify speakers in the audio signal. The end user device 214 may process the audio signal by extracting embeddings from the audio signal.

[0149] During operation, the end user device 214 may apply various machine learning operations to the audio signal. For example, the end user device 214 may extract embeddings by running VAD software or other machine learning architectures configured to extract features from the audio signal. In some embodiments, the end user device 214 may forward the extracted embeddings to the analysis system 201, as described in FIG. 1 , so that the analysis system 201 may cluster the embeddings with the voiceprints retrieved from the analysis database 204. The analysis server 202 may identify speakers based on the similarity of utterances in the audio signal to stored voiceprints. The analysis server 202 may forward the speaker identifiers and other metadata to the content system 210.

[0150] Additionally or alternatively, if the voiceprints are stored on the end user device 214, the end user device 214 may cluster the embeddings with the stored voiceprints and generate a similarity matrix describing the similarity score of each extracted embedding compared to one or more voiceprints stored on the end user device 214. Additionally or alternatively, the end user device 214 may query the analysis database 204 in the analysis system 201, as described in FIG. 1, to retrieve suspect or probable voiceprints of the extracted embeddings based on association with subscriber identifiers or speaker characteristics.

[0151] In some configurations, the end-user device 214 forwards the similarity matrix to the analysis system 201 so that the analysis system 201 can compare the maximum similarity score to high and low similarity thresholds using condition-dependent adaptive thresholding, as described in FIG. 1 . The analysis server 201 can determine weak and strong speech and update the voiceprint. Additionally or alternatively, the end-user device 214 can query the analysis database 204 in the analysis system 201 to retrieve the FAR, maturity threshold, and similarity threshold, as described in FIG. 1 . The end-user device 214 can update the voiceprint based on comparing the maximum similarity score in the similarity matrix to the similarity thresholds. The end-user device 214 can store the updated voiceprint and transmit the updated voiceprint to the analysis database 204.

[0152] In some configurations, in response to trigger criteria, the analysis server 202 may perform label correction to correct the label identifiers, as described in Figure 1. Additionally or alternatively, the end-user device 214 may perform label correction and forward the updated label identifiers to the analysis system 201 and / or the content system 210.

[0153] If the end user device 214 uses a speaker identifier to identify a speaker from the audio signal, the end user device 214 may transmit the speaker identifier and any metadata to the content system 210 so that the content system 210 can determine personalized content to stream to the identified speaker. The end user device 214 may also transmit a speaker profile (including the speaker identifier, extracted embeddings, metadata, and voiceprint) to the analysis system 201 (e.g., the analysis database 204).

[0154] Example Operations Active and static enrollment in voice processing authentication systems FIG. 3 describes the phases a server goes through to identify (or authenticate) a user. FIG. 3A illustrates operational steps of a method 300a for actively registering a user during a registration phase, according to one embodiment. FIG. 3B illustrates operational steps of a method 300b for identifying (or authenticating) a user during a deployment phase. A server (e.g., an analysis server) of the analysis system executes machine-readable software code that performs the methods 300a, 300b described below, although one or more processors of any number of computing devices may perform the various operations of methods 300a, 300b. Some embodiments may include operations in addition to, fewer than, or different from those described in methods 300a, 300b.

[0155] Referring to FIG. 3A, in step 302, the server prompts the user for an enrollment signal. In some configurations, the server may prompt the user for an enrollment signal once or at regular time intervals (e.g., every six months). In some configurations, the server may prompt the user for an enrollment signal, for example, when the device is powered up for the first time at a predetermined interval or when the user accesses a configuration interface for registering / enrolling a new speaker user. In some configurations, the server may prompt the user for an enrollment signal based on instructions (from the user or other administrator) to perform the enrollment phase. Prompts may include prompting the user to place their finger on a fingerprint sensor, prompting the user to speak a particular phrase, prompting the user to speak naturally, prompting the user to appear within a digital bounding box on the display, etc.

[0156] In step 304, the server receives an enrollment signal. The enrollment signal is a signal received during a designated enrollment phase (e.g., active enrollment of a user). The enrollment signal may be distinct from other types of signals because the end-user device and the server receive the enrollment signal during the enrollment phase and in response to a prompt (as in step 302). For example, the server receives an enrollment signal that the end-user device transmits to the server in response to the user pressing a button on the end-user device and speaking a specific utterance in response to an audio or visual prompt. The server receives the audio signal directly or through an intermediary server (e.g., a content server). In some configurations, the server receives the enrollment utterance as well as any number of enrollment signals (e.g., multiple speech portions (utterances) of the signal) and biometric data (e.g., one or more fingerprint angles) for enrolling the biometric data.

[0157] In some configurations, the server receives the audio signal, which is a data file or data stream containing audio data in a machine-readable format. The audio signal includes audio recordings containing any number of speaker utterances from any number of speakers. The audio data may also include data or metadata received with the audio signal. For example, the audio data may include speaker-related information (e.g., user / speaker identifier, subscriber / household identifier, user biometrics) or metadata related to the communication protocol or medium (e.g., TCP / IP header data, phone number / ANI).

[0158] In step 306, the server extracts embeddings from the enrollment signal by applying a machine learning architecture including various machine learning models (e.g., embedding extractors). An embedding is a mathematical representation of the biometric information (or biometric features) in the enrollment signal. The server may extract features from the utterance including Mel Frequency Cepstral Coefficients (MFCCs), Mel filter banks, linear filter banks, bottleneck features, etc. The server extracts features from the input speech signal using machine learning models configured to extract features and generate speaker embeddings.

[0159] This type of enrollment signal may define how the server extracts the embedding. For example, a user may provide a fingerprint as the enrollment signal. In other examples, the server may extract features associated with a digital fingerprint image, including ridges, valleys, and minutiae, and / or the server may extract the features using a machine learning model configured to extract features from image data. The server uses the machine learning model configured to extract features to extract various types of features and, in conjunction with the speaker embedding, generate a corresponding embedding for a particular type of biometric used for user recognition.

[0160] At decision step 308, the server determines whether the enrollment is mature. The server may determine that the enrollment is mature when the server extracts enough embeddings (or biometric information) to satisfy a threshold number of embeddings or other information. For example, the server may determine that the enrollment is mature when the server receives a predetermined duration of net speech. Additionally or alternatively, the server may determine that the enrollment is mature when the server receives a predetermined number of different enrollment signals (e.g., two fingerprint scans and two speeches, five different utterances). The mature voiceprint (or mature embedding cluster) may contain enough biometric information for the analysis server 102 to identify the speaker using a speaker identifier.

[0161] If enrollment is not mature, the server prompts the user with an enrollment signal (e.g., step 302). The server prompts the user with additional enrollment signals (comprising enrollment utterances) until enrollment is mature. As an example, if the server receives an enrollment utterance with a first type of content (e.g., a username), the server prompts the user with a second utterance (e.g., the user's birth date). As another example, if the server receives biometric information (e.g., a fingerprint), the server prompts the user with an audio signal containing the enrollment utterance.

[0162] If the enrollment is mature, the server proceeds to step 310. In step 310, the server creates a voiceprint of the user (sometimes referred to as an enrollee voiceprint). The server statistically or algorithmically combines the enrollment embeddings to extract the voiceprint of the enrolled speaker user. In some implementations, a cluster of enrollment embeddings extracted from the enrollment signal represents a collection of utterances similar to a particular speaker (e.g., a speaker cluster) whose voiceprint represents the centroid of the speaker cluster.

[0163] In step 312, the server updates the speaker profile. A speaker profile is a data record associated with each user (or speaker) who enrolls during the enrollment phase. The speaker profile includes a voiceprint, speaker clusters (e.g., embeddings associated with the voiceprint), and metadata. As described in FIG. 1, the metadata may include information about the utterance, the quality of the enrollment signal, the overall duration of the net utterance, the total number of utterances in each cluster, speaker characteristics, and speaker-independent characteristics. A speaker identifier 313 is associated with each speaker profile to distinguish and identify specific speakers. As described in FIG. 1, the speaker identifier may not include or represent any personal identification information. The server may also generate a new speaker identifier or request a new speaker identifier for the new speaker profile from the content server.

[0164] In step 314, the server may transmit the speaker identifiers to a third-party system (e.g., a content server, a call center). The third-party system may map the received speaker identifiers to human speakers. For example, the third-party system may use a lookup table to map the speaker identifiers to specific human speakers. Additionally or alternatively, the third-party system may associate a subscriber identifier or other household and / or group identifier with each speaker identifier. The server may also transmit a speaker profile (including speaker-independent characteristics, speaker characteristics, and metadata) associated with the speaker identifier.

[0165] 3B, in step 322, the server may receive an inbound signal containing biometric information. The inbound signal may be an image of a user, an utterance in an audio signal, a fingerprint, etc. In some configurations, the server receives the inbound signal in response to the end device actively capturing user input data. For example, a user may actively interact with the end user device (e.g., speak a "wake" word, press a button, perform a gesture). Additionally or alternatively, the end user device 114 passively captures user input data, and the end user passively interacts with the end user device 114 (e.g., speaks to another user and the end user device 114 automatically captures the utterance without any affirmative action by the user).

[0166] In step 324, the server extracts an embedding from the inbound signal. In some configurations, the server may perform preprocessing on the inbound signal (e.g., segment the inbound signal, scale the inbound signal, denoise the inbound signal). Additionally or alternatively, the server may identify events in the inbound signal. For example, the server may detect audio events by executing VAD software. The VAD software may distinguish silence from speech. Additionally or alternatively, the server may perform object detection (or recognition, identification) in the inbound signal. For example, the server may be configured to recognize faces and distinguish them from hands. The server may be configured to extract an embedding using an identified biometric portion of the inbound signal (e.g., speech, fingerprint).

[0167] The inbound signal may be an audio signal including an audio recording containing any number of speaker utterances from any number of speakers. The audio data may also include data or metadata received along with the audio signal. For example, the audio data may include speaker-related information (e.g., user / speaker identifier, subscriber / household identifier, user biometrics) or metadata related to a communication protocol or medium (e.g., TCP / IP header data, phone number / ANI).

[0168] The server may extract the embeddings using machine learning models configured to extract features of the audio signal. The features include low-level spectrotemporal features from various speaker utterances. The features may also include various data associated with a user or speaker (e.g., subscriber identifier, speaker identifier, biometrics), or metadata values ​​extracted from protocol information or data packets (e.g., IP addresses). The machine learning architecture, including various machine learning models (e.g., embedding extractor models), extracts estimated speaker embeddings based on the features extracted from the input audio signal.

[0169] In step 326, the machine learning architecture generates one or more similarity scores by comparing the extracted features or embeddings for a particular speaker with corresponding features or embeddings of other putative speakers and / or with corresponding features or embeddings of speaker clusters stored in the database. A cluster represents a collection of utterances similar to a particular speaker, where the speaker voiceprint represents the centroid of the speaker cluster. The server applies the machine learning architecture to the audio signal to extract an embedding for the particular speaker. The server compares the embedding or features with the speaker clusters stored in the speaker profile database and then determines a similarity score for the speaker. For each particular speaker, the server generates a set of similarity scores based on the relative distance between the speaker embedding (extracted from the input audio signal) and the voiceprint stored in the speaker profile database.

[0170] Additionally or alternatively, the server performs a clustering operation according to certain features extracted from the input audio signals and determines one or more clustering similarity scores for each user based on the features.

[0171] The server identifies the speaker cluster and speaker pair with the highest similarity score 327. For each speaker, the server outputs the highest similarity score 327 calculated for that particular speaker, which represents the most likely match between the speakers.

[0172] In decision step 328, the server determines whether the maximum similarity score for each voiceprint satisfies one or more thresholds. The server determines whether the inbound speech signal contains utterances of a new or known speaker by evaluating the similarity scores of the corresponding embeddings or features against known voiceprints or expected features.

[0173] If the maximum similarity score meets one or more thresholds, the server determines that the speaker is likely associated with a known enrolled user 329. For example, the known enrolled user 329 may be a user who enrolled during the enrollment phase of FIG. 3A. On the other hand, if the server determines that the maximum similarity score does not reach one or more thresholds, the server determines that the speaker is likely a new, unknown user 331. The server may use the particular voiceprint to generate a new speaker profile and speaker identifier. The speaker profile includes the voiceprint, a speaker cluster (e.g., an embedding associated with the voiceprint), a speaker identifier, and metadata.

[0174] In step 330, the server outputs the speaker identifier and speaker profile (of the known enrolled user 329 and / or the new unknown user 331) to one or more downstream applications. In some configurations, the server authenticates the user according to the speaker profile of the known enrolled user 329. Additionally or alternatively, downstream operations may use the speaker identifier to identify, authenticate, and / or authorize a particular speaker. The downstream application may perform different functions depending on whether the speaker is a known enrolled user or a new unknown user. For example, the downstream application may execute functionality, unlock, or perform operations based on the identified and / or authenticated enrolled user. If the user is a new unknown user 331, the downstream application may restrict the user's access to the application's software, functionality, or information.

[0175] Adaptive thresholding for speech processing 4 illustrates operational steps of a method 400 for adaptive thresholding in a speech processing system. While method 400, described below, is performed by a server of an analysis system (e.g., an analysis server) executing machine-readable software code, any number of processors of any number of computing devices may perform the various operations of method 400. Embodiments may include operations in addition to, fewer than, or different from those described in method 400. The server applies a machine learning architecture, including any number of machine learning models and algorithms, to the input speech signal to perform the various operations. While the server and machine learning architecture perform method 400 during an enrollment phase, the server and machine learning architecture may perform the various operations of method 400 during the enrollment phase, during a deployment phase, or as a successive combination of such phases.

[0176] In step 402, the server receives an input audio signal. The server receives the input audio signal directly from an end-user device (as in FIG. 2) or via a computing device of a third-party system (e.g., a content system, a call center system) (as in FIG. 1). The input audio signal includes the speech of a speaker user, in which case the input audio signal is a data file (e.g., a WAV file, an MP3 file) or a data stream. The server performs various pre-processing operations on the input audio signal, such as parsing the audio signal into speech segments or frames or performing one or more transform operations (e.g., a fast Fourier transform), among other potential operations.

[0177] The input audio signal may be an enrollment audio signal or an inbound audio signal, in which the server receives the input audio during an enrollment phase or a deployment phase. The audio signal may also include data or metadata received with the audio signal. For example, the audio signal may include speaker-related information (e.g., user / speaker identifier, subscriber / household identifier, user biometrics) or metadata related to the communication protocol or medium (e.g., TCP / IP header data, phone number / ANI).

[0178] The server receives the speech signal and extracts various types of features from the speech signal. The features include low-level spectrotemporal features from various speaker utterances. The features may also include various data related to the user or speaker (e.g., subscriber identifier, speaker identifier, biometrics), or metadata values ​​extracted from protocol information or data packets (e.g., IP addresses). A machine learning architecture, including various machine learning models (e.g., embedding extractor models), extracts estimated speaker embeddings based on the features extracted from the input speech signal.

[0179] In some embodiments, the server or machine learning architecture applies a voice activity detection (VAD) model to the audio signal. A VAD is a machine learning model trained to detect instances of speech in the audio signal and extract or otherwise identify segments of the audio signal that contain the detected speech. In some cases, the VAD generates an abbreviated, speech-only audio signal. The server may store the speech segments and / or the abbreviated audio signal in a database (e.g., a speaker profile database, a voiceprint database) or some other non-transitory machine-readable storage medium.

[0180] In step 404, for each particular user, the server stores the embedding or voiceprint in a database record (representing a speaker profile) in the analysis database. The server also stores various types of user information associated with the particular speaker, including, for example, a user-specific similarity threshold generated for the user. If the speaker profile is new or otherwise lacks a user-specific similarity threshold, the server stores a pre-configured default similarity threshold in the speaker profile.

[0181] In decision step 406, the server determines whether the voiceprint satisfies one or more maturity thresholds, which indicate whether the voiceprint is mature or stable. The server identifies one or more maturity factors associated with the speaker profile and the utterances. Non-limiting examples of maturity factors include the number of enrollment utterances, the overall duration of the net speech throughout each utterance, and the quality of the audio from the speech signal.

[0182] The server may use any number of algorithms to determine whether a voiceprint is mature. For example, the server compares a maturity factor (e.g., number of utterances) to a preconfigured maturity threshold corresponding to the maturity factor (e.g., a threshold number of utterances). As another example, the server statistically or algorithmically combines maturity factors and compares the combined maturity factor to a preconfigured maturity threshold corresponding to the combined maturity factor. A preconfigured false acceptance rate defines the preconfigured maturity threshold. This false acceptance rate is manually entered by an administrative user or algorithmically determined by various machine learning models. As false acceptance increases, the maturity threshold increases, thereby increasing the likelihood that the server will determine that the maturity factor does not meet the maturity threshold and the voiceprint is not sufficiently mature.

[0183] In some embodiments, the server uses tiered maturity thresholds corresponding to tiered false acceptance rates. For example, the server may store a table of maturity thresholds (e.g., a threshold schedule) with various false acceptance rates. Referring to FIG. 12, example 1200 illustrates a threshold scheduler based on a single maturity factor (number of enrollment embeddings) and various false acceptance rates. In some implementations, the threshold schedule may associate similarity scores with several maturity factors (e.g., number of enrollments 1202). Column 1202 illustrates a maturity threshold based on a single maturity factor (e.g., number of enrollment utterances 1202). Columns 1204 and 1206 each illustrate a different predetermined false acceptance rate (FAR). Column 1204 illustrates a similarity threshold based on a predetermined FAR of 0.5%. Column 1206 illustrates a similarity threshold based on a predetermined FAR of 5%. The similarity threshold illustrated in column 1204 is considered a high similarity threshold. The similarity threshold in column 1206 is considered a low similarity threshold.

[0184] Referring again to FIG. 4 , in step 410, the server adjusts the speaker's similarity threshold in response to the server determining (in step 408) that the voiceprint is not mature because it does not meet the maturity threshold. The server adjusts the similarity threshold according to the false acceptance rate, whereby the server increases or decreases the similarity threshold to meet a desired level of accuracy, as represented by the false acceptance rate. For example, the server increases the similarity threshold for a particular speaker if the maturity factor (e.g., number of utterances) does not meet a given maturity factor threshold. The server updates the similarity score so that the server evaluates the voiceprint according to the false acceptance rate of future inbound speech signals. The server repeatedly prompts the speaker for additional utterances or enrollment embeddings until the maturity threshold is met.

[0185] Referring again to FIG. 12 , in one example, the maturity threshold may be set at 10 enrollment utterances. If the server determines that the voiceprint does not meet the maturity threshold (e.g., the input speech signal received in step 402 was of the seventh enrollment utterance), then, depending on the FAR (of the particular subscriber identifier or third-party system), the high similarity score associated with the speaker is 4.29 (e.g., score threshold 1208) and the low similarity score is 0.1 (e.g., score threshold 1210). The server stores the similarity scores associated with the speaker in a speaker database. If the speaker speaks again, the thresholds used in evaluating the similarity of future inbound speech signals will be 4.29 (e.g., score threshold 1208) for the high similarity score associated with the speaker and 0.1 (e.g., score threshold 1210) for the low similarity score.

[0186] Referring again to Figure 4, in step 412, the server stores the voiceprint and similarity threshold in a speaker profile and applies the voiceprint and similarity threshold to future inbound speech signals if the server determines (in step 408) that the voiceprint is mature. The server applies the voiceprint and similarity threshold to future inbound speech signals that are deemed to contain speech from a particular speaker. In some implementations, the server stores the voiceprint and similarity threshold in a speaker profile database. In some implementations, the server stores the voiceprint and similarity threshold on an end user device or other device in communication with the server.

[0187] 12, in the example described above, the maturity threshold is set at 10 enrollment utterances. If the server determines that the voiceprint satisfies the maturity threshold (e.g., the input speech signal received in step 402 was of the 10th enrollment utterance), then, depending on the FAR (of the particular subscriber identifier or third-party system), the high similarity score associated with the speaker is 4.37 (e.g., score threshold 1212) and the low similarity score is 0.15 (e.g., score threshold 1214). The server stores the similarity scores associated with the speaker in a speaker database. If the speaker speaks again, the thresholds used in evaluating the similarity of future inbound speech signals will be that the high similarity score associated with the speaker is 4.29 (e.g., score threshold 1208) and the low similarity score is 0.1 (e.g., score threshold 1210).

[0188] Unsupervised Clustering 5 illustrates the execution steps of a method 500 for identifying and evaluating strong and weak speech in audio processing. Method 500, described below, is performed by a server of an analysis system (e.g., an analysis server) executing machine-readable software code, although any number of processors of any number of computing devices may perform the various operations of method 500. Embodiments may include operations in addition to, fewer than, or different from those described in method 500. The server applies a machine learning architecture, including any number of machine learning models and algorithms, to perform the various operations.

[0189] In step 502, the server receives an input audio signal from an end-user device and extracts various types of features from the input audio signal. The input audio signal includes a data file or data stream containing audio data in a machine-readable format. The audio data includes audio recordings containing any number of speaker utterances from any number of speakers. The input audio signal includes the speech of a speaker user, in which case the input audio signal is a data file (e.g., a WAV file, an MP3 file) or a data stream. The input audio signal may be an enrollment audio signal or an inbound audio signal, where the server receives the input audio during an enrollment phase or a deployment phase. The server extracts various types of features, such as spectrotemporal features or metadata, from the input audio signal. Additionally or alternatively, the server performs various pre-processing operations on the input audio signal, such as parsing the audio signal into speech segments or frames or performing one or more transform operations (e.g., a fast Fourier transform), among other potential operations.

[0190] In step 504, the server compares the extracted embeddings or features for the speaker to speaker clusters stored in a speaker profile database and then determines a similarity score for the speaker. A cluster represents a collection of utterances similar to a particular speaker, where the speaker voiceprint represents the centroid of the speaker cluster. The server applies a machine learning architecture with one or more machine learning models to the speech signal to extract the embeddings for the particular speaker. For each particular speaker, the server generates a set of similarity scores based on the relative distance between the speaker embedding (extracted from the input speech signal) and the voiceprints stored in the speaker profile database.

[0191] Additionally or alternatively, the server performs a clustering operation according to certain features extracted from the input audio signals and determines one or more clustering similarity scores for each user based on the features.

[0192] In step 506, the server identifies each pair of speaker and cluster with the highest similarity score. For each speaker, the server outputs the highest similarity score calculated for that particular speaker, representing the most likely match between the speakers.

[0193] In decision step 508, the server determines whether the similarity score of a particular speaker's embedding (or features) satisfies one or more similarity thresholds. The server determines whether the input speech signal contains an utterance of a new or known speaker by evaluating the similarity score of the corresponding embedding or feature against known voiceprints or expected features. If the server determines that the similarity score of a particular embedding satisfies the similarity threshold, the server similarly determines that the embedding is likely associated with a known, registered user. On the other hand, if the server determines that the similarity score of a particular embedding does not reach the similarity threshold, the server determines that the embedding is likely a new user.

[0194] In some embodiments (such as in FIG. 6 ), the server compares the output similarity score for a particular speaker with a low and a high similarity threshold, providing a level of granularity that can control poor quality embeddings resulting from poor quality utterances. In such embodiments, if the server determines that the similarity score for a particular embedding satisfies the high similarity threshold, the server similarly determines that the embedding is likely associated with a known, registered user. On the other hand, if the server determines that the similarity score for a particular embedding does not reach the low similarity threshold, the server determines that the embedding is likely a new user. As discussed below, if the similarity score is between the low and high thresholds, the server stores the associated speech data and similarity score in a buffer memory, speaker profile, or other isolated memory location.

[0195] In step 510, the server generates a new voiceprint and a new speaker profile in response to the server determining (in step 508) that the embedding for a particular speaker does not meet the similarity threshold. The server generates the new speaker profile in an analysis database or another memory location configured to store temporary or guest speaker profiles. The server stores various types of data that the server identifies or extracts from the input audio signal (e.g., speaker identifier, embedding, voiceprint, features, metadata, device information) in the new speaker profile.

[0196] In step 512, the server updates the stored voiceprint and the existing enrollment (or known) speaker profile in the database in response to the server determining (in step 508) that the embedding of a particular known speaker satisfies the similarity threshold. The server updates the known speaker profile of the known speaker profile in the analysis database or in a temporary speaker profile or guest speaker profile. The server stores various types of data that the server identifies or extracts from the input audio signal (e.g., speaker identifier, embedding, voiceprint, features, metadata, device information) in the known speaker profile.

[0197] In optional step 514, the server performs one or more reclustering operations to reevaluate the speaker clusters and update the speaker profiles in the speaker profile database. The server performs the reclustering operations in response to a particular trigger condition. Non-limiting examples of trigger conditions may include, among others, a preconfigured periodic time interval or when the server receives a threshold number of utterances associated with a subscriber (e.g., a household) or user. The server extracts speaker features or embeddings to generate a new voiceprint or update an existing voiceprint in the speaker profile. The server recalculates the speaker's similarity score based on the relative distance between the extracted features or embeddings for each particular utterance and each particular voiceprint or other type of cluster centroid. The server stores the new or updated voiceprint in a new or updated speaker profile, along with various types of data associated with the speaker.

[0198] In some embodiments, the reclustering operation performed by the server is a hierarchical clustering operation (as in FIG. 7). Hierarchical clustering minimizes the likelihood of small cumulative identification / recognition errors and increases the purity of speaker clusters. The server determines relative distances between voiceprints, or other comparative differences or clustering algorithms (e.g., PLDA, cosine distance), to identify existing or new voiceprints that satisfy a similarity score threshold.

[0199] Discrimination of strong and weak speech for speech processing FIG. 6 illustrates operational steps of a method 600 for clustering speakers during speech processing. While method 600, described below, is performed by a server of an analysis system (e.g., an analysis server) executing machine-readable software code, any number of processors of any number of computing devices may perform the various operations of method 600. Embodiments may include operations in addition to, fewer than, or different from those described in method 600. The server applies a machine learning architecture, including any number of machine learning models and algorithms, to the input speech signal to perform the various operations. While the server and machine learning architecture perform method 600 during an enrollment phase, the server and machine learning architecture may perform the various operations of method 600 during the enrollment phase, during a deployment phase, or as a successive combination of such phases.

[0200] The server executes method 600 during enrollment operations (active or passive), deployment operations, and / or reclustering database update operations. For clustering operations, the server applies the operations of a trained machine learning architecture to current or past audio signals, where the audio signals may include enrollment audio signals, inbound audio signals (received during the deployment phase), or stored audio signals. The machine learning architecture includes any number of machine learning models and various other operations that the server applies to particular audio signals, including preprocessing (e.g., feature extraction) and clustering operations. The clustering operation uses speaker utterances in the audio signals to facilitate new or known speaker recognition. The server extracts features or feature vectors (e.g., embeddings) from the audio signals and then clusters the extracted information (e.g., features, embeddings) into clusters corresponding to speakers present in the audio signals. While method 600 includes unsupervised clustering operations, in some embodiments, the server may perform supervised clustering operations.

[0201] In step 602, the server receives an input audio signal from an end-user device and extracts various types of features from the audio signal. The audio signal may be a data file or a data stream containing audio data in a machine-readable format. The audio data includes audio recordings containing any number of speaker utterances from any number of speakers. The audio data may also include data or metadata received with the audio signal. For example, the audio data may include speaker-related information (e.g., user / speaker identifier, subscriber / household identifier, user biometrics) or metadata related to the communication protocol or medium (e.g., TCP / IP header data, phone number / ANI).

[0202] The server receives the audio signal and extracts various types of features from the audio data. The features include low-level spectrotemporal features from various speaker utterances. The features may also include various data related to the user or speaker (e.g., subscriber identifier, speaker identifier, biometrics), or metadata values ​​extracted from protocol information or data packets (e.g., IP addresses). A machine learning architecture, including various machine learning models (e.g., embedding extraction models), extracts estimated speaker embeddings based on the features extracted from the input audio signal.

[0203] In step 604, the machine learning architecture generates one or more similarity scores by comparing the extracted features or embeddings for a particular speaker with corresponding features or embeddings of other putative speakers and / or with corresponding features or embeddings of speaker clusters stored in a database. A cluster represents a collection of utterances similar to a particular speaker, where the speaker voiceprint represents the centroid of the speaker cluster. The server applies the machine learning architecture to the audio signal to extract an embedding for the particular speaker. The server compares the embedding or features with the speaker clusters stored in a speaker profile database and then determines a similarity score for the speaker. For each particular speaker, the server generates a set of similarity scores based on the relative distance between the speaker embedding (extracted from the input audio signal) and the voiceprints stored in the speaker profile database.

[0204] Additionally or alternatively, the server performs a clustering operation according to certain features extracted from the input audio signals and determines one or more clustering similarity scores for each user based on the extracted features.

[0205] The server identifies the speaker cluster and speaker pair with the highest similarity score, and for each speaker, the server outputs the highest similarity score calculated for that particular speaker, which represents the most likely match between the speakers.

[0206] In step 606, the server determines whether the similarity score for a particular speaker satisfies a low or high similarity threshold. The server determines whether the input speech signal contains an utterance of a new or known speaker by evaluating the similarity score of the corresponding embedding or feature against known voiceprints or expected features. If the server determines that the similarity score for a particular speaker embedding satisfies the high similarity threshold, the server determines that the speaker is likely associated with a known, enrolled user. On the other hand, if the server determines that the similarity score for a particular embedding does not reach the low similarity threshold, the server determines that the speaker is likely a new, unknown user. If the server determines that the similarity score meets the low threshold but does not reach the high threshold, the server determines that the speaker is likely a new, unknown user.

[0207] In step 608, the server generates a new voiceprint and a new speaker profile in response to the server determining (in step 606) that the embedding for a particular speaker does not meet the low similarity threshold. The server generates the new speaker profile in an analysis database or another memory location configured to store temporary or guest speaker profiles. The server stores various types of data that the server identifies or extracts from the input audio signal (e.g., speaker identifier, embedding, voiceprint, features, metadata, device information) in the new speaker profile.

[0208] In step 610, the server generates or updates a list of weak embeddings for the known user in response to the server determining (in step 606) that an embedding meets the lower threshold but falls short of the high threshold. The list of weak embeddings acts as a buffer or isolated storage location associated with a particular speaker, but the embedding (and utterance) lacks sufficient similarity to the corresponding voiceprint to immediately characterize the weak embedding as part of a particular speaker cluster. The weak embeddings are stored with the audio data and various types of data (e.g., voice recordings, utterances, metadata, embeddings) potentially originating from the known user.

[0209] In step 612, the server updates the stored voiceprints and existing enrollment (or known) speaker profiles in the database in response to the server determining (in step 606) that the embeddings of a particular known speaker satisfy the similarity threshold. The server updates a list of strong embeddings associated with the known user. The list of strong embeddings includes embeddings that the server uses to generate the known user's voiceprint or cluster. The server updates the known speaker's speaker profile in the analysis database or in a temporary speaker profile or guest speaker profile. The server stores in the known speaker profile various types of data that the server receives along with the strong embeddings that identify or extract from the input audio signal (e.g., speaker identifiers, embeddings, voiceprints, features, metadata, device information).

[0210] In optional step 614, the server performs a reclustering operation to update the clusters in the database and update the database accordingly. The server performs one or more reclustering operations to reevaluate the speaker clusters and update the speaker profiles in the speaker profile database. The server performs the reclustering operation in response to a specific trigger condition. Non-limiting examples of trigger conditions may include, among others, a preconfigured periodic time interval or when the server receives a threshold number of utterances associated with a subscriber (e.g., a household) or user. The server extracts speaker features or embeddings to generate a new voiceprint or update an existing voiceprint in the speaker profile. The server recalculates the speaker's similarity score based on the relative distance between the extracted features or embeddings for each particular utterance and each particular voiceprint or other type of cluster centroid. The server stores the new or updated voiceprint in a new or updated speaker profile along with various types of data associated with the speaker.

[0211] In some cases, the server re-evaluates each list of weak embeddings to determine whether the weak embeddings are sufficiently similar to a particular known speaker, or any other speaker. Re-clustering may update one or more voiceprints, clusters, or thresholds. As a result, one or more weak embeddings may better match a particular voiceprint or cluster according to the server's recalculated similarity scores. If the server determines that a particular weak embedding satisfies the similarity threshold for a particular voiceprint, the server adds the weak embeddings and associated audio data to the speaker profile corresponding to the particular voiceprint and updates the voiceprint and speaker profile according to the weak embeddings.

[0212] In some cases, the server reevaluates each list of strong embeddings to determine whether the strong embeddings remain sufficiently similar to the particular known speaker or any other speaker. As a result of the reclustering operation, one or more strong embeddings may no longer match the particular speaker voiceprint or cluster well, or may better match another voiceprint or cluster, according to the server's recalculated similarity scores. If the server determines that a particular strong embedding no longer satisfies the similarity threshold for the particular voiceprint, the server removes the strong embedding and associated audio data from the speaker profile corresponding to the particular voiceprint and updates the voiceprint and speaker profile according to the remaining strong embeddings. If the server determines that a particular strong embedding satisfies the similarity threshold for the particular voiceprint, the server adds the strong embedding and associated audio data to the speaker profile corresponding to the particular voiceprint and updates the voiceprint and speaker profile according to the strong embeddings.

[0213] Label Correction FIG. 7A illustrates operational steps of a method 700a for correcting one or more voiceprint label identifiers (e.g., speaker identifiers, subscriber identifiers) according to current and / or historical information. While method 700a, described below, is performed by a server of an analysis system (e.g., an analysis server) executing machine-readable software code, any number of processors of any number of computing devices may perform the various operations of method 700a. Embodiments may include additional, fewer, or different operations than those described in method 700a. The server applies a machine learning architecture, including any number of machine learning models and algorithms, to the input voice signal to perform the various operations. While the server and machine learning architecture perform method 700a during an enrollment phase, the server and machine learning architecture may perform the various operations of method 700a during the enrollment phase, during a deployment phase, or as a successive combination of such phases.

[0214] The server receives an audio signal from an end-user device and extracts various types of features from the audio signal. The audio data includes audio recordings containing any number of speaker utterances from any number of speakers. The server extracts various types of features from the audio signal. The features include low-level spectrotemporal features from the various speaker utterances. The features may also include various data related to a user or speaker (e.g., subscriber identifier, speaker identifier, biometrics), or metadata values ​​extracted from protocol information or data packets (e.g., IP addresses). A machine learning architecture, including various machine learning models (e.g., an embedding extractor model), extracts embeddings of the inbound speaker based on the features extracted from the audio signal.

[0215] The server generates one or more similarity scores by comparing the extracted features or embeddings for a particular speaker with corresponding features or embeddings of other putative speakers and / or with corresponding features or embeddings of voiceprints stored in a database. The server determines the similarity score by assessing the relative distance between the embedding and the voiceprint. The server may determine the relative distance according to a distance measure, such as a Euclidean distance measure and / or a correlation-based measure. Additionally or alternatively, the server may assess the similarity of the embeddings and voiceprints by, for example, using a cosine similarity approach or probabilistic linear discriminant analysis (PLDA) to determine the similarity score. If the maximum similarity score associated with each embedding does not satisfy one or more thresholds, the server may use the embeddings to create a new speaker profile (e.g., identify the new speaker with a new cluster, a new voiceprint, and a new speaker identifier).

[0216] 7B illustrates an example 700b of label correction using a particular speaker's cluster and other putative speaker's clusters. As described above, in one example, the server creates a new speaker profile if the maximum similarity score associated with an embedding does not satisfy one or more thresholds. Clusters 721, 723, 725, 727, and 729 (collectively referred to as "clusters 720") represent embedding clusters extracted from the speech signal. The server may not associate any of clusters 720 with a speaker identifier. Additionally or alternatively, the server may decide to associate a new speaker identifier with cluster 720.

[0217] Referring again to FIG. 7A , in step 702, the server obtains label identifiers (e.g., subscriber identifiers, speaker identifiers) and voiceprints. The server obtains label identifiers from a database or generates new label identifiers for the unknown speaker. The server obtains one or more new label identifiers and voiceprints from the newly created speaker profile. Additionally or alternatively, the server obtains previous label identifiers and voiceprints by querying a speaker database and retrieving data for specific speaker profiles. The server retrieves, for example, speaker profiles associated with the subscriber identifier, speaker profiles recently accessed and / or modified by the server (e.g., speaker profiles retrieved by the analysis server in the past two days), speaker profiles associated with specific speaker characteristics, speaker profiles associated with specific speaker-independent characteristics, and / or speaker profiles associated with other metadata. The retrieved speaker profiles may include specific label identifiers (e.g., subscriber identifiers, speaker identifiers) and voiceprints. The server may also fetch all speaker profiles from one or more databases.

[0218] 7B, as an example, when a server performs a reclustering operation on database records (e.g., speaker profiles) associated with a particular subscriber identifier, the server retrieves the previous label identifier and voiceprint associated with the particular subscriber identifier. In response to querying the database for the subscriber identifier associated with cluster 720 extracted from the audio signal, the server receives clusters known 731, known 733, known 735, and known 737 (collectively referred to as "known clusters 730").

[0219] Referring again to FIG. 7A, in step 704, the server generates voiceprint pair similarity scores by calculating pairwise similarities between various voiceprints. The server compares each voiceprint with each of the other voiceprints to calculate the voiceprint pair similarity score. The server calculates each particular voiceprint pair similarity score by, for example, evaluating the relative distance between the two voiceprints of the pair, the cosine similarity of the two voiceprints, or the PLDA of the two voiceprints. The server identifies the best-matching voiceprint pair by identifying the respective maximum voiceprint pair similarity score for each of the voiceprints.

[0220] 7B, voiceprint similarity score matrix 738 identifies a similarity score for each of clusters 720 compared to each of known clusters 730. Each cell in similarity score matrix 738 represents a similarity score comparing a known cluster of known clusters 730 to a cluster of clusters 720. The maximum cluster-pair similarity score for each of clusters 720 and known clusters 730 is identified at 736.

[0221] 7A, in step 706, the server identifies each particular maximum voiceprint pair similarity score. Optionally, the server may determine whether each particular maximum voiceprint pair similarity score satisfies a preconfigured relabeling threshold (sometimes referred to as a transition threshold). The relabeling threshold may be the same as the similarity threshold described herein used to match inbound embeddings to particular voiceprints.

[0222] In some implementations, voiceprints may be associated with different relabeling thresholds depending on the maturity level of a given voiceprint in a voiceprint pair. For example, a particular voiceprint may be associated with a low similarity threshold and a high similarity threshold in a speaker profile. In one example, the server compares the maximum voiceprint pair similarity score to, for example, the low similarity threshold. The server may also statistically or algorithmically combine thresholds (e.g., thresholds associated with each of the voiceprints).

[0223] In step 708, the server relabels or transfers the label identifiers. The server may relabel or transfer the label identifier associated with either voiceprint to a new or updated speaker profile associated with the cluster. In addition to relabeling the label identifiers, the server may decide to merge the voiceprints by averaging each of the voiceprints or by applying a machine learning architecture to the voiceprints to algorithmically combine the voiceprints.

[0224] 7B, the server updates or corrects the label identifiers of cluster 720 by transferring the label identifiers of known cluster 730 and replacing the label identifiers of cluster 720. For example, cluster 721 is most similar to known cluster 733, as indicated by a maximum similarity score of 54.2, which is greater than the similarity scores of cluster 721 and the other clusters in known cluster 730 (e.g., 0.3, 3.1, and −5.9, respectively).

[0225] Similarly, the server determined that cluster 723 is most similar to known cluster 735, as indicated by a maximum similarity score of 39.8, which is greater than the similarity scores of cluster 723 and the other clusters in known cluster 730 (e.g., −0.5, 2.5, and 5.4, respectively). Although the maximum similarity score of 39.8 associated with cluster 723 and known cluster 735 is less than the maximum similarity score of 54.2 associated with cluster 721 and known cluster 733, the server still migrated the known cluster 735 label identifier to cluster 723. The server's relabeling or migration of the label identifiers associated with clusters 721, 723, 725, and 727 indicates that clusters in cluster 721, 723, 725, and 727 were previously identified by the server (e.g., known cluster 731, known cluster 733, known cluster 735, and known cluster 737). Because cluster 729 is not sufficiently similar to any of the estimated speaker profiles, the server created a new speaker profile for cluster 729. The server generated a new label identifier, known 740, for cluster 729 to indicate that the server created a new speaker profile.

[0226] Fully passive and continuous registration FIG. 8 illustrates operational steps of a method 800 for audio processing using a passive and continuous enrollment arrangement. Method 800, described below, is performed by a server of an analysis system (e.g., an analysis server) executing machine-readable software code, although any number of processors of any number of computing devices may perform the various operations of method 800. Embodiments may include operations in addition to, fewer than, or different from those described in method 800. The server applies a machine learning architecture, including any number of machine learning models and algorithms, to the input audio signal to perform the various operations. While the server and machine learning architecture perform method 800 during an enrollment phase, the server and machine learning architecture may perform the various operations of method 800 during the enrollment phase, during a deployment phase, or as a continuing combination of such phases.

[0227] In step 802, the server receives an input audio signal containing one or more utterances of one or more speakers. The server receives the input audio signal from an end user device and extracts various types of features from the input audio signal, where the server receives the input audio signal directly from the end user device or via an intermediary device (e.g., a third-party server). The input audio signal includes a data file or data stream containing audio data in a machine-readable format. The audio data includes audio recordings containing any number of speaker utterances from any number of speakers. The input audio signal includes the speech of a speaker user, where the input audio signal is a data file (e.g., a WAV file, an MP3 file) or a data stream. The input audio signal can be an enrollment audio signal or an inbound audio signal, where the server receives the input audio during an enrollment phase or a deployment phase. The server extracts various types of features, such as spectrotemporal features or metadata, from the input audio signal. Additionally or alternatively, the server may perform various pre-processing operations on the input audio signal, such as parsing the audio signal into speech segments or frames, or performing one or more transform operations (e.g., a fast Fourier transform), among other potential operations.

[0228] In step 804, the server applies a machine learning architecture, including any number of machine learning models, to the features extracted from the input speech signal. An embedding extraction model of the machine learning architecture uses the features extracted from the input speech signal to extract an inbound embedding for the inbound speaker.

[0229] In decision step 806, the server determines whether a database (e.g., a voiceprint database, a speaker profile database) is empty. The database contains data records for particular households, subscribers, or other collections of individuals who are customers of a third-party content service or data analytics service. The data records include speaker profiles for particular speakers associated with speaker identifiers. In some implementations, the server determines whether a portion of the database is empty. For example, the server determines whether the database contains any speaker profiles (e.g., household speaker profiles) associated with a particular subscriber identifier.

[0230] In step 808, if the server determines (in step 806) that the database is not empty, the server generates similarity scores for the speaker embeddings generated for the utterance based on the relative distance between each embedding and each voiceprint stored in the database. For each particular speaker, the server outputs the maximum similarity score 809, which represents the closest match (by similarity score) of a particular inbound speaker to a particular voiceprint.

[0231] For each particular inbound speaker embedding, the server determines whether the corresponding maximum similarity score 809 satisfies a low-similarity threshold, at decision step 812. In some cases, the low-similarity threshold is a pre-configured default value or an adaptive threshold tailored to the particular voiceprint and putative enrolled speaker (as in FIG. 4).

[0232] If the server determines (in step 812) that a particular maximum similarity score 809 satisfies the low threshold, then the server determines whether the maximum similarity score 809 satisfies a high similarity threshold, in decision step 814. In some cases, the high similarity threshold is a preconfigured default value or an adaptive threshold tailored to a particular voiceprint and putative enrolled speaker (as in FIG. 4).

[0233] In step 816, if the server determines (in step 814) that the maximum similarity score 809 satisfies the high threshold, the server updates the list of strong embeddings. The server uses strong embeddings to generate voiceprints. For example, the server updates a particular voiceprint using a particular inbound strong embedding. The server may also update the corresponding speaker profile to include the updated voiceprint. Because the maximum similarity score 809 of a particular embedding satisfies the high threshold, the server determines that the particular inbound embedding is likely to be the putative enrolled speaker.

[0234] In step 822, after updating the list of strong embeddings, the server updates the database containing speaker profiles to include the strong inbound embeddings. The server adds the inbound embeddings to the particular speaker profile whose voiceprint most closely matches the inbound embedding. The server updates the speaker profile with the speaker identifier associated with the voiceprint that most closely matches the inbound embedding.

[0235] In step 820, the server creates a new speaker profile in the database if the server determines (in step 806) that the database is empty or (in step 812) that the maximum similarity score 809 for a particular inbound embedding does not meet the low similarity threshold. The server assigns a new speaker identifier to the new speaker profile if the server receives the speaker identifier from a third-party server. In some implementations, the server receives a hashed (or otherwise obfuscated) version of the corresponding speaker identifier used by the third-party server, thereby maintaining speaker privacy by preventing the server from receiving any personally identifiable information about a particular speaker.

[0236] In some cases, the new speaker profile is a temporary or guest speaker profile with a limited, predetermined life cycle. In such cases, the server or database purges the temporary profile's data from the database after a preconfigured time for maintaining the temporary profile. The server or database also restarts this life cycle clock for maintaining the temporary profile for each instance in which the server identifies another inbound embedding that meets the high or low threshold for matching the threshold temporary voiceprint (as in step 812 or step 814).

[0237] If the server determines that the voiceprint of the temporary profile is mature, the server converts the temporary profile to a permanent speaker profile in the database, thereby updating the speaker profile in the current step 820. For example, the server determines that a temporary profile is received as mature when it receives a threshold number of embeddings from a particular speaker that satisfy a high threshold for the temporary voiceprint (as in step 814).

[0238] In step 818, the server updates the list of weak embeddings if the maximum similarity score 809 of a particular embedding does not meet the higher threshold for the closest matching voiceprint in the database (in step 814), but the inbound embedding already meets the lower threshold (in step 812). The list of weak embeddings effectively acts as a buffer or quarantine containing embeddings that potentially match the corresponding closest voiceprint. The server can reference these weak embeddings in later operations, such as a reclustering operation, to determine whether to include the weak embedding in the speaker profile of the closest voiceprint.

[0239] In step 822, the server updates the database to include the embeddings and speaker information. For a particular speaker embedding, the database receives one or more updates, such as an updated list of strong embeddings (from step 816), an updated list of weak embeddings (from step 818), or a new data record (from step 820). The database stores the updates of step 822 along with a speaker identifier 823 associated with the embedding, speaker profile information, or list of embeddings. The speaker identifier 823 is an anonymized value representing a user identifier of the content system so that the content server does not reveal any personal information about the speaker to the analytics server.

[0240] In step 824, the server outputs the speaker identifier 823 and any associated information about the speaker profile required by the content server for downstream operations. The server transmits the speaker identifier 823 to, for example, a computing device in the media content system, an end user device, or any other device that performs a particular downstream operation. In some implementations, the server also transmits additional speaker profile information from the speaker profile stored in the database record.

[0241] Active / Passive Mix and Continuous Registration FIG. 9 illustrates operational steps of a method 900 for speech processing a speech signal using a mixed active-passive and continuous enrollment configuration. A server or other computing device performs method 900; in some embodiments, the server may perform method 900 during an enrollment phase. The server performs method 900 during an active enrollment phase and an expansion phase. During the active enrollment phase, the server receives enrollment speech signals from enrolled speakers responding to one or more audio and / or visual prompts. The prompts request the enrolled speakers to audibly respond, and the server receives the responses as enrollment speech signals to generate an enrolled voiceprint, enroll the particular speaker, and generate corresponding speaker profile data in the database. During the expansion phase, the server receives inbound speech signals that the server evaluates to identify the particular inbound speaker. The server can also continuously passively enroll unknown speakers. If the server cannot identify the inbound speaker, the server passively enrolls the unknown speaker and generates a new data record for the new speaker in the database.

[0242] In step 902, the server receives an enrollment speech signal from an enrolled user that includes the speaker user's utterance. The server receives the enrollment speech signal from an end-user device via an intermediate content server. The server extracts features from the enrollment speech signal. The features may be spectrotemporal features, and in some cases, may be various types of data or metadata. The server may also perform various pre-processing operations on the enrollment speech signal, such as various data augmentation operations. The server may extract features from the utterance including Mel Frequency Cepstral Coefficients (MFCCs), Mel filter banks, linear filter banks, bottleneck features, etc. In some cases, the server extracts features from the enrollment speech signal using a machine learning model configured to extract features and generate a speaker embedding.

[0243] In step 904, the server applies a machine learning architecture, including any number of machine learning models, to the features extracted from the input speech signal. An embedding extraction model of the machine learning architecture uses the features extracted from the input speech signal to extract an inbound embedding for the inbound speaker. The server extracts the embedding from the enrollment signal by applying the machine learning architecture, including various machine learning models (e.g., an embedding extractor).

[0244] This type of enrollment signal may define how the server extracts the embedding. For example, a user may provide a fingerprint as the enrollment signal. In other examples, the server may extract features associated with a digital fingerprint image, including ridges, valleys, and minutiae, and / or the server may extract the features using a machine learning model configured to extract features from image data. The server uses the machine learning model configured to extract features to extract various types of features and, in conjunction with the speaker embedding, generate a corresponding embedding for a particular type of biometric used for user recognition.

[0245] In decision step 906, the server determines whether the enrollment voiceprint is mature. The server may determine that enrollment is mature when it extracts enough embeddings (or biometric information) to meet a threshold number of embeddings or other information such that the server can mathematically identify the user of a particular signal. For example, the server may determine that enrollment is mature when it receives a threshold duration of net speech from one or more audio signals containing speech from the enrollee. Additionally or alternatively, the server may determine that enrollment is mature when it receives a predetermined number of different enrollment signals (e.g., two fingerprint scans and two utterances, five different utterances).

[0246] If enrollment is incomplete (e.g., not mature), the server prompts the user with an enrollment signal (as in step 902). The server prompts the user with additional enrollment signals (having enrollment utterances) until enrollment is mature. As an example, if the server receives an enrollment utterance with a first type of content (e.g., a username), the server prompts the user with a second utterance (e.g., the user's date of birth). As another example, if the server receives biometric information (e.g., a fingerprint), the server prompts the user with an audio signal containing the enrollment utterance.

[0247] If the enrollment is mature, the server proceeds to step 908. In step 908, the server creates a new enrollment voiceprint for the enrolled speaker user. The server statistically or algorithmically combines the enrollment embeddings to extract the enrolled speaker user's voiceprint. The server stores the enrollment voiceprint in a speaker profile, along with any various other non-identifying information about the enrollee. The server may also generate a new speaker identifier or request a new speaker identifier from the content server.

[0248] In step 910, the server updates the database, for example by storing the new registered voiceprint and other information (e.g., speaker identifier, subscriber identifier) ​​in the new speaker profile, thereby registering / enrolling the new speaker user.

[0249] In step 914, the server outputs the new speaker identifier and any associated information about the new speaker profile requested by the content server for downstream operations. The server transmits the new speaker identifier to, for example, a computing device in the media content system, an end user device, or any other device that performs certain downstream operations. In some implementations, the server also transmits additional new speaker profile information from the new speaker profile stored in the new database record.

[0250] Following the active enrollment phase, the server transitions the machine learning architecture into a deployment phase, where the server evaluates inbound speech signals of enrolled speakers, but also continues to passively enroll new, unrecognized speakers.

[0251] In step 916, an inbound speech signal containing one or more utterances of one or more inbound speakers is received from the end user device via the intermediate content server. The server extracts features from the inbound speech signal. The features may be spectrotemporal features, and in some cases may be various types of data or metadata. The server may also perform various preprocessing operations on the inbound speech signal, such as various data augmentation operations. The server may extract features from the utterances including Mel Frequency Cepstral Coefficients (MFCCs), Mel filter banks, linear filter banks, bottleneck features, etc. In some cases, the server extracts features from the inbound speech signal using a machine learning model configured to extract features and generate speaker embeddings. In some configurations, the server may perform preprocessing operations on the inbound signal (e.g., splitting the inbound signal, scaling the inbound signal, denoising the inbound signal).

[0252] In step 918, the server applies a machine learning architecture to the inbound speech signal to extract an embedding for each of the inbound speakers. The server extracts each inbound embedding from the inbound speech signal based on features extracted from the inbound speech signal.

[0253] In step 920, the server generates a similarity score for each inbound speaker embedding generated for the utterance based on the relative distance between the particular inbound embedding and each voiceprint stored in the database. For each particular speaker, the server outputs the maximum similarity score that represents the particular inbound speaker's closest match (by similarity score) with the particular voiceprint.

[0254] For each particular inbound speaker embedding, the server determines whether the corresponding maximum similarity score satisfies a low-similarity threshold, at decision step 922. In some cases, the low-similarity threshold is a pre-configured default value or an adaptive threshold tailored to a particular voiceprint and putative known enrolled speaker (as in FIG. 4).

[0255] In step 924, the server creates a new speaker profile in the database if the server determines (in step 922) that the maximum similarity score for a particular inbound embedding does not meet the low similarity threshold. The server assigns a new speaker identifier to the new speaker profile if the server receives the speaker identifier from a third-party server. In some implementations, the server receives a hashed (or otherwise obfuscated) version of the corresponding speaker identifier used by the third-party server, thereby maintaining speaker privacy by preventing the server from receiving any personally identifiable information about a particular speaker.

[0256] In decision step 926, the server determines whether the maximum similarity score satisfies a high similarity threshold if the server determines (in step 922) that the maximum similarity score satisfies the low threshold. In some cases, the high similarity threshold is a preconfigured default value or an adaptive threshold tailored to a particular voiceprint and putative enrolled speaker (as in FIG. 4).

[0257] In step 928, if the maximum similarity score for a particular embedding does not meet the higher threshold for the closest matching voiceprint in the database (in step 926), but the inbound embedding already meets the lower threshold (in step 922), the server updates the list of weak embeddings. The list of weak embeddings effectively acts as a buffer or quarantine containing embeddings that potentially match the corresponding closest voiceprint. The server can reference these weak embeddings in later operations, such as a reclustering operation, to determine whether to include the weak embedding in the speaker profile for the closest voiceprint.

[0258] In step 930, if the server determines (in step 926) that the maximum similarity score satisfies the high threshold, the server updates the list of strong embeddings. The server uses strong embeddings to generate voiceprints. For example, the server updates a particular voiceprint using a particular inbound strong embedding. The server may also update the corresponding speaker profile to include the updated voiceprint. Because the maximum similarity score for a particular embedding satisfies the high threshold, the server determines that the particular inbound embedding is likely to be the putative enrolled speaker.

[0259] In step 910, the server updates the database to include the embeddings and speaker information. For a particular speaker embedding, the database receives one or more updates, such as an updated list of strong embeddings (from step 930), an updated list of weak embeddings (from step 928), or a new data record (from step 924). The database stores the updates of step 910 along with the speaker identifier associated with the embedding, speaker profile information, or list of embeddings. The speaker identifier is an anonymized value representing a user identifier of the content system so that the content server does not reveal any personal information about the speaker to the analytics server.

[0260] In step 914, the server outputs the speaker identifier, new or known identifier, and any associated information about the speaker profile requested by the content server for downstream operations. The server transmits the speaker identifier, for example, to a computing device in the media content system, an end user device, or any other device that performs the particular downstream operation. In some implementations, the server also transmits additional speaker profile information from the speaker profile stored in the database record.

[0261] Active and continuous enrollment FIG. 10 illustrates operational steps of a method 1000 for speech processing of speech signals using an active and continuous enrollment configuration. A server or other computing device performs the method 900; in some embodiments, the server may perform the method 900 during the enrollment phase. The server performs the method 900 during the active enrollment phase and the deployment phase. During the active enrollment phase, the server receives enrollment speech signals from enrolled speakers responding to one or more audio and / or visual prompts. The prompts request the enrolled speakers to audibly respond, and the server receives the responses as enrollment speech signals to generate enrolled voiceprints, enroll the particular speakers, and generate corresponding speaker profile data in the database. During the deployment phase, the server receives inbound speech signals, which the server evaluates to identify particular inbound speakers. Unlike the embodiment of FIG. 9, the active enrollment phase is mandatory to enroll all users. While the server passively evaluates speech signals continuously, the server does not passively enroll new, unrecognized speakers. The clustering operation is semi-supervised, but the clustering may be more constrained or simpler than in other embodiments because the server is pre-configured with a known number of clusters, for example, the database assigns a pre-configured number of enrolled speakers who have completed enrollment.

[0262] In step 1002, the server receives an enrollment speech signal from an enrolled user that includes the speaker user's utterance. The server receives the enrollment speech signal from an end-user device via an intermediate content server. The server extracts features from the enrollment speech signal. The features may be spectrotemporal features, and in some cases, may be various types of data or metadata. The server may also perform various pre-processing operations on the enrollment speech signal, such as various data augmentation operations. The server may extract features from the utterance including Mel Frequency Cepstral Coefficients (MFCCs), Mel filter banks, linear filter banks, bottleneck features, etc. In some cases, the server extracts features from the enrollment speech signal using a machine learning model configured to extract features and generate a speaker embedding.

[0263] In step 1004, the server applies a machine learning architecture, including any number of machine learning models, to the features extracted from the input speech signal. An embedding extraction model of the machine learning architecture uses the features extracted from the input speech signal to extract an inbound embedding for the inbound speaker. The server extracts the embedding from the enrollment signal by applying the machine learning architecture, including various machine learning models (e.g., an embedding extractor).

[0264] At decision step 1006, the server determines whether the enrollment voiceprint is mature. The server may determine that enrollment is mature when it extracts enough embeddings (or biometric information) to meet a threshold number of embeddings or other information such that the server can mathematically identify the user of a particular signal. For example, the server may determine that enrollment is mature when it receives a threshold duration of net speech from one or more audio signals containing speech from the enrollee. Additionally or alternatively, the server may determine that enrollment is mature when it receives a predetermined number of different enrollment signals (e.g., two fingerprint scans and two utterances, five different utterances).

[0265] If enrollment is incomplete (e.g., not mature), the server prompts the user with an enrollment signal (as in step 1006). The server prompts the user with additional enrollment signals (having enrollment utterances) until enrollment is mature. As an example, if the server receives an enrollment utterance with a first type of content (e.g., a username), the server prompts the user with a second utterance (e.g., the user's birth date). As another example, if the server receives biometric information (e.g., a fingerprint), the server prompts the user with an audio signal containing the enrollment utterance.

[0266] If the enrollment is mature, the server proceeds to step 1008. In step 1008, the server creates a new enrollment voiceprint for the enrolled speaker user. The server statistically or algorithmically combines the enrollment embeddings to extract the enrolled speaker user's voiceprint. The server stores the enrollment voiceprint in a speaker profile, along with any various other non-identifying information about the enrollee. The server may also generate a new speaker identifier or request a new speaker identifier from the content server.

[0267] In step 1010, the server updates the database, for example by storing the new registered voiceprint and other information (e.g., speaker identifier, subscriber identifier) ​​in the new speaker profile, thereby registering / enrolling the new speaker user.

[0268] In step 1014, the server outputs the new speaker identifier and any associated information about the new speaker profile requested by the content server for downstream operations. The server transmits the new speaker identifier to, for example, a computing device in the media content system, an end user device, or any other device that performs certain downstream operations. In some implementations, the server also transmits additional new speaker profile information from the new speaker profile stored in the new database record.

[0269] Following the active enrollment phase, the server transitions the machine learning architecture into a deployment phase, where the server evaluates inbound speech signals of enrolled speakers, but also continues to passively enroll new, unrecognized speakers.

[0270] In step 1016, an inbound speech signal containing one or more utterances of one or more inbound speakers is received from the end user device via the intermediate content server. The server extracts features from the inbound speech signal. The features may be spectrotemporal features, and in some cases may be various types of data or metadata. The server may also perform various pre-processing operations on the inbound speech signal, such as various data augmentation operations. The server may extract features from the utterances including Mel Frequency Cepstral Coefficients (MFCCs), Mel filter banks, linear filter banks, bottleneck features, etc. In some cases, the server extracts features from the inbound speech signal using a machine learning model configured to extract features and generate speaker embeddings. In some configurations, the server may perform pre-processing operations on the inbound signal (e.g., splitting the inbound signal, scaling the inbound signal, denoising the inbound signal).

[0271] In step 1018, the server applies a machine learning architecture to the inbound speech signal to extract an embedding for each of the inbound speakers. The server extracts each inbound embedding from the inbound speech signal based on features extracted from the inbound speech signal.

[0272] In step 1020, the server generates a similarity score for each inbound speaker embedding generated for the utterance based on the relative distance between the particular inbound embedding and each voiceprint stored in the database. For each particular speaker, the server outputs the maximum similarity score that represents the particular inbound speaker's closest match (by similarity score) with the particular voiceprint.

[0273] For each particular inbound speaker embedding, the server determines whether the corresponding maximum similarity score satisfies a low-similarity threshold, at decision step 1022. In some cases, the low-similarity threshold is a pre-configured default value or an adaptive threshold tailored to a particular voiceprint and putative known enrollment speaker (as in FIG. 4).

[0274] In decision step 1026, the server determines whether the maximum similarity score satisfies a high similarity threshold if the server determines (in step 1022) that the maximum similarity score satisfies the low threshold. In some cases, the high similarity threshold is a preconfigured default value or an adaptive threshold tailored to a particular voiceprint and putative enrolled speaker (as in FIG. 4).

[0275] In step 1028, if the maximum similarity score for a particular embedding does not meet the higher threshold for the closest matching voiceprint in the database (in step 1026), but the inbound embedding already meets the lower threshold (in step 1022), the server updates the list of weak embeddings. The list of weak embeddings effectively acts as a buffer or quarantine containing embeddings that potentially match the corresponding closest voiceprint. The server can reference these weak embeddings in later operations, such as a reclustering operation, to determine whether to include the weak embedding in the speaker profile for the closest voiceprint.

[0276] In step 1026, if the server determines (in step 1026) that the maximum similarity score satisfies the high threshold, the server updates the list of strong embeddings. The server uses strong embeddings to generate voiceprints. For example, the server updates the particular voiceprint using the particular inbound strong embedding. The server may also update the corresponding speaker profile to include the updated voiceprint. Because the maximum similarity score for the particular embedding satisfies the high threshold, the server determines that the particular inbound embedding is likely to be the putative enrolled speaker.

[0277] In step 1010, the server updates the database to include the embeddings and speaker information. For a particular speaker embedding, the database receives one or more updates, such as an updated list of strong embeddings (from step 1030) or an updated list of weak embeddings (from step 1028). The database stores the updates of step 1010 along with the speaker identifier associated with the embeddings, speaker profile information, or list of embeddings. The speaker identifier is an anonymized value representing a user identifier of the content system so that the content server does not reveal any personal information about the speaker to the analytics server.

[0278] In step 1024, if the server determines (in step 1022) that a particular inbound speaker embedding does not meet the low threshold, the server generates a warning or other instruction for the intermediate content server or end user device indicating that the speaker is not recognized. The server may also be pre-configured to provide additional information to the content server.

[0279] In step 1014, the server outputs various types of information to the content server. In some cases, the server transmits the speaker identifier and any associated information about the speaker profile required by the content server for downstream operations. In such cases, the server transmits the speaker identifier to, for example, a computing device in the media content system, an end-user device, or any other device that performs a particular downstream operation. In some implementations, the server also transmits additional speaker profile information from a speaker profile stored in a database record. Alternatively, the server transmits a warning or other instruction to the content server indicating that a particular speaker was not recognized by the server.

[0280] Additional Exemplary Embodiments 11A-11B show components of a system 1100 that uses audio processing machine learning operations, where the machine learning models are implemented by a vehicle or other edge device (e.g., a car, a home assistant device, a smart appliance).

[0281] The vehicle includes a microphone 1108 configured to capture audio waves 1110 containing speech and convert the audio waves 1110 into audio signals for audio processing operations. The vehicle includes computing hardware and software components (shown as an analysis computer 1102 and a speaker database 1104) configured to perform various audio processing operations described herein. The components and operations described in system 1100 are similar to those of FIGS. 1-2. While system 100 of FIG. 1 located many of the machine learning audio processing operations on analysis server 102, content system 110 may perform certain operations in some embodiments. While system 200 of FIG. 2 located many of the machine learning audio processing operations on end user device 214, end user device 214 may still rely on analysis system 201 or content system 210 for various operations and database information. However, the vehicle-based system 1100 of FIG. 11 attempts to encapsulate much of the voice processing computation and data within the vehicle-based system 1100, with relatively little reliance on external system infrastructure devices.

[0282] The analysis computer 1102 receives input data signals from a microphone 1108 and performs various pre-processing operations, such as VAD and ASR, to identify speech. The analysis computer 1102 and any number of machine learning models are applied to extract features, extract embeddings, and compare the embeddings to voiceprints stored in a speaker database 1104. The analysis computer 1102 is coupled to various electronic components of the vehicle, such as the infotainment system, engine, door locks, and other components of the vehicle. The analysis computer 1102 receives voice instructions from the driver or passenger to activate or adjust various options in the vehicle.

[0283] In some embodiments, analysis computer 1102 employs parental control operations or other restrictions on vehicle functions. Speaker profiles stored in speaker database 1104 include specific function restrictions that prohibit analysis computer 1102 from performing certain operations. For example, analysis computer 1102 may detect a registered speaker embedding for a child speaker profile by running an embedding extraction model, or alternatively, by running a known machine learning model for determining age using voice. Analysis computer 1102 then inhibits activation of the starter or ignition, thereby preventing the engine from starting, until analysis computer 1102 detects an embedding that matches the voiceprint of an authorized user in speaker database 1104. This feature is adapted not only for parental control but also for theft deterrence, whereby analysis computer 1102 prevents the engine from starting until analysis computer 1102 positively detects an inbound audio signal containing an inbound embedding from a registered speaker user.

[0284] Analysis computer 1102 actively or passively enrolls the driver (e.g., a first parent), secondary driver (e.g., a second driver), and passengers (e.g., children). As an example, when a driver first purchases a vehicle, a GUI displayed via the infotainment device presents a prompt requesting the driver to speak a specific phrase, thereby submitting an enrollment utterance that is captured by microphone 1108. Analysis computer 1102 then performs various processes described herein to generate a voiceprint of the driver, which analysis computer 1102 stores in the driver's speaker profile in speaker database 1104. The driver also inputs information about the driver via the GUI, such as name and specific preferences related to seating position, radio station, child lock activation, security preferences, headlight delay, etc. This speaker information is stored in the driver's speaker profile. The driver may also input, via the GUI, the number of speakers (e.g., in the car) expected to operate system 1100. The analysis computer 1102 generates clusters and voiceprints according to the number of expected speakers. The analysis computer 1102 may generate speaker profiles according to the number of expected speakers, or the analysis computer 1102 generates speaker profiles by performing active enrollment operations, passive enrollment operations, and / or continuous enrollment operations described herein. For example, the analysis computer 1102 performs continuous, passive enrollment operations to generate voiceprints for a child's speaker profile. The driver or child may input various types of speaker information about the child via the infotainment system's GUI. This speaker information may include, for example, driver permissions, such as parental controls, as referred to herein.

[0285] In some embodiments, analysis computer 1102 uses a static enrollment configuration, whereby analysis computer 1102 does not accept unknown speaker embeddings as new enrollments. In addition, analysis computer 1102 performs an authentication function that rejects authentication of unrecognized voiceprints and does not allow speakers to access certain vehicle functions. For example, analysis computer 1102 may be used in specialized vehicles (e.g., police cars, delivery trucks) to restrict unauthorized access to the vehicle and vehicle operations.

[0286] Analysis computer 1102 performs continuous enrollment operations to generate speaker profiles and various types of data representing speaker information, which analysis computer 1102 stores in analysis database 1104 according to speaker voiceprints. Non-limiting examples of speaker information may include operational permissions for various computer-based features of the vehicle, such as turning on or operating the ignition, opening doors (e.g., child locks), among others. As an example, a driver (e.g., a parent) indicates, via the infotainment system's GUI, various permissions for a child speaker profile and the child's name associated with the speaker profile's speaker identifier and the child's voiceprint. When analysis computer 1102 identifies a new voiceprint for a new speaker (child), the infotainment system generates a GUI prompt indicating to the driver (or other known enrolled speaker) (e.g., a parent) that analysis computer 1102 has identified and generated a new voiceprint for the child. The parent may enter one or more inputs confirming which particular speaker profile (or speaker identifier) ​​is associated with the child's new voiceprint. Analysis computer 1102 then stores the new voiceprint in the child's speaker profile in analysis database 1104. The child may enter various types of speaker information into the speaker profile via the GUI, such as preference configurations for seating preferences or climate preferences. After enrolling the child, analysis computer 1102 passively identifies when the child is present in the vehicle according to speech received by microphone 1108. Analysis computer 1102 then commands the infotainment system or associated control system to function (e.g., change seating position) according to the preference data in the child's speaker profile.

[0287] In one embodiment, a computer-implemented method includes receiving, by a computer, an inbound speech signal comprising a plurality of utterances of a plurality of inbound speakers; applying, by the computer, a machine learning architecture to the inbound speech signal to extract a plurality of inbound embeddings corresponding to the plurality of inbound speakers; and, for each inbound speaker of the plurality of inbound speakers, generating, by the computer, one or more similarity scores based on the inbound embedding of the inbound speaker, wherein each similarity score for the inbound speaker is based on the inbound embedding and a speaker profile. and one or more voiceprints stored in the database, indicating a distance between the voiceprint and the inbound speaker; identifying, by the computer, a closest voiceprint of the inbound speaker from the one or more voiceprints, the closest voiceprint corresponding to a maximum similarity score of the one or more similarity scores generated for the inbound speaker; and, for each maximum similarity score that satisfies the one or more similarity score thresholds, updating, by the computer, the speaker profile database to include an inbound embedding of the inbound speaker having the maximum similarity score that satisfies the one or more similarity score thresholds.

[0288] The method may further include identifying, by the computer, one or more voiceprints stored in a speaker profile database based on the subscriber identifier received with the inbound voice signal.

[0289] The method may further include determining, by the computer, that the maximum similarity score of the inbound speaker embedding satisfies the one or more similarity scores, and identifying, by the computer, a speaker profile in the speaker database that contains a closest voiceprint of the inbound speaker, the speaker profile including a speaker identifier.

[0290] The method may further include determining, by the computer, that the maximum similarity score for the inbound speaker satisfies a first similarity threshold and is less than a second similarity threshold, the first similarity threshold being relatively lower than the second similarity threshold, and updating, by the computer, a list of weak embeddings stored in the speaker database to include an inbound embedding for the inbound speaker.

[0291] The method may further include: performing, by the computer, a reclustering operation on one or more speaker profiles in a speaker database associated with the subscriber identifier; updating, by the computer, a maximum similarity score of inbound embeddings in the list of weak embeddings based on the reclustering operation; and in response to determining that the maximum similarity score of the inbound embeddings in the list of weak embeddings satisfies one or more similarity score thresholds, updating, by the computer, the speaker profile database to include an inbound embedding of the inbound speaker having a maximum similarity score that satisfies the one or more similarity score thresholds.

[0292] The method may further include detecting, by the computer, a trigger condition for performing a reclustering operation on the subscriber identifier, and performing, by the computer, a hierarchical clustering operation on the plurality of voiceprints of one or more speaker profiles associated with the subscriber identifier.

[0293] The method may further include updating, by the computer, a speaker identifier associated with the new voiceprint cluster from an existing speaker profile, the new voiceprint cluster being generated by applying a hierarchical clustering operation to one or more speaker profiles associated with the subscriber identifier.

[0294] The method may further include receiving, by the computer, inbound voice signals from the end user device via the intermediary server, and transmitting, by the computer, a speaker identifier associated with each inbound speaker having a maximum similarity score that satisfies one or more similarity score thresholds.

[0295] The method may further include receiving, by the computer, an inbound speech signal comprising a plurality of utterances from a plurality of inbound speakers.

[0296] The method may further include applying, by the computer, a machine learning architecture to the one or more features to identify one or more audio events to detect environmental settings, and transmitting, by the computer, an indicator of the environmental settings and each speaker identifier associated with each inbound speaker having a maximum similarity score that satisfies one or more similarity score thresholds to a mediation server.

[0297] The method may further include, for each maximum similarity score that is below one or more similarity score thresholds, generating, by the computer, a new speaker profile associated with the subscriber identifier that includes the inbound embedding, and updating, by the computer, the speaker profile database to include the new speaker profile and inbound embedding for the inbound speaker.

[0298] In one embodiment, a system comprises: a speaker database comprising a non-transitory machine-readable storage medium configured to store data records comprising speaker profiles; and a computer comprising a processor, the processor receiving an inbound speech signal comprising a plurality of utterances of a plurality of inbound speakers; applying a machine learning architecture to the inbound speech signal to extract a plurality of inbound embeddings corresponding to the plurality of inbound speakers; and generating, for each inbound speaker of the plurality of inbound speakers, one or more similarity scores based on the inbound speaker's inbound embedding, wherein the inbound speaker wherein each similarity score indicates a distance between the inbound embedding and one or more voiceprints stored in a speaker profile database; identifying a closest voiceprint for the inbound speaker from the one or more voiceprints, the closest voiceprint corresponding to a maximum similarity score of the one or more similarity scores generated for the inbound speaker; and, for each maximum similarity score that satisfies one or more similarity score thresholds, updating the speaker profile database to include an inbound embedding of the inbound speaker having a maximum similarity score that satisfies the one or more similarity score thresholds.

[0299] The computer may be configured to identify one or more voiceprints stored in a speaker profile database based on a subscriber identifier received with the inbound voice signal.

[0300] The computer may be configured to determine that the maximum similarity score of the inbound speaker embedding satisfies one or more similarity scores and to identify a speaker profile in the speaker database that contains the closest voiceprint of the inbound speaker, the speaker profile including a speaker identifier.

[0301] The computer may be configured to determine that the maximum similarity score for the inbound speaker satisfies a first similarity threshold and is less than a second similarity threshold, the first similarity threshold being relatively lower than the second similarity threshold, and to update a list of weak embeddings stored in a speaker database to include an inbound embedding for the inbound speaker.

[0302] The computer may be configured to: perform a reclustering operation on one or more speaker profiles in a speaker database associated with the subscriber identifier; update a maximum similarity score of an inbound embedding in the list of weak embeddings based on the reclustering operation; and, in response to determining that the maximum similarity score of an inbound embedding in the list of weak embeddings satisfies one or more similarity score thresholds, update the speaker profile database to include an inbound embedding of an inbound speaker having a maximum similarity score that satisfies the one or more similarity score thresholds.

[0303] The computer may be configured to detect a trigger condition for performing a reclustering operation on the subscriber identifier and to perform a hierarchical clustering operation on a plurality of voiceprints of one or more speaker profiles associated with the subscriber identifier.

[0304] The computer may be configured to update a speaker identifier associated with a new voiceprint cluster from an existing speaker profile, the new voiceprint cluster being generated by applying a hierarchical clustering operation to one or more speaker profiles associated with the subscriber identifier.

[0305] The computer may be configured to receive inbound voice signals from an end user device via an intermediary server and transmit a speaker identifier associated with each inbound speaker having a maximum similarity score that satisfies one or more similarity score thresholds.

[0306] The computer may be configured to receive an inbound speech signal that includes a plurality of utterances from a plurality of inbound speakers.

[0307] The computer may be configured to apply a machine learning architecture to the one or more features to identify one or more audio events to detect environmental settings, and transmit to a mediation server an indicator of the environmental settings and each speaker identifier associated with each inbound speaker having a maximum similarity score that satisfies one or more similarity score thresholds.

[0308] The computer may be configured to, for each maximum similarity score that falls short of one or more similarity score thresholds, generate a new speaker profile associated with the subscriber identifier that includes the inbound embedding, and update the speaker profile database to include the new speaker profile and inbound embedding for the inbound speaker.

[0309] In one embodiment, a computer-implemented method includes: receiving, by a computer, an inbound speech signal of an inbound speaker from an end user device via a content server; applying, by the computer, a machine learning model to the inbound speech signal to extract an inbound embedding of the inbound speaker; generating, by the computer, a similarity score for the inbound embedding based on a distance between the inbound embedding and a voiceprint stored in a speaker profile in a speaker database, the similarity score satisfying one or more similarity score thresholds; identifying, by the computer, one or more speaker features in the speaker profile that correspond to one or more content features of the content server; and transmitting, by the computer, the one or more speaker features associated with the inbound speaker to a media content server.

[0310] The method may further include identifying, by the computer, a speaker profile based on the subscriber identifier received from the content server.

[0311] The subscriber identifier received by the computer may be a first anonymized identifier, and the subscriber identifier may be associated with one or more speaker identifiers for one or more speakers, and each speaker identifier may be a second anonymized identifier corresponding to a respective speaker profile.

[0312] The one or more speaker characteristics may include at least one of an age characteristic and a gender characteristic.

[0313] The method may further include determining, by the computer, at least one of an age characteristic and a gender characteristic of the inbound speaker by applying a second machine learning model to the inbound speech signal.

[0314] The method may further include determining, by the computer, the age of the inbound speaker based on age characteristics stored in a speaker profile of the inbound speaker.

[0315] The method may further include receiving, by the computer, speaker information from the content server indicating at least one speaker characteristic of the inbound speaker; storing, by the computer, the at least one speaker characteristic in a speaker profile of the inbound speaker; and identifying, by the computer, the speaker profile based in part on the at least one speaker characteristic of the speaker profile.

[0316] The method may further include receiving, by the computer, inbound authentication data from the content server along with the inbound audio signal, and authenticating, by the computer, the inbound speaker based on the similarity score satisfying the similarity threshold and the inbound authentication data satisfying expected authentication data stored in the speaker profile.

[0317] The authentication data may include at least one of end-user device information, metadata associated with the end-user device, speaker information, and biometric information.

[0318] The method may further include extracting, by the computer, one or more features from the inbound voice signal; and calculating, by the computer, a spoof score indicating a likelihood that the inbound voice signal contains a spoof condition based on the one or more features by applying a second machine learning model.

[0319] The method may further include identifying, by the computer, one or more media content files in a media content database that have one or more content characteristics corresponding to the one or more speaker characteristics.

[0320] The method may further include enabling, by the computer, age-restricted content in the content database based on age characteristics of one or more speaker characteristics of a speaker that satisfies a corresponding age-restriction characteristic in the one or more content characteristics of the age-restricted content.

[0321] In one embodiment, a system includes a speaker database comprising a non-transitory machine-readable storage medium configured to store a plurality of speaker profiles; and a server comprising a processor configured to: receive an inbound speech signal of an inbound speaker from an end user device via a content server; apply a machine learning model to the inbound speech signal to extract an inbound embedding of the inbound speaker; generate a similarity score for the inbound embedding based on a distance between the inbound embedding and a voiceprint stored in a speaker profile in the speaker database, the similarity score satisfying one or more similarity score thresholds; identify one or more speaker features in the speaker profile that correspond to one or more content features of the content server; and transmit the one or more speaker features associated with the inbound speaker to a media content server.

[0322] The computer may be further configured to identify a speaker profile based on the subscriber identifier received from the content server.

[0323] The subscriber identifier received by the computer may be a first anonymized identifier, and the subscriber identifier may be associated with one or more speaker identifiers for one or more speakers, and each speaker identifier may be a second anonymized identifier corresponding to a respective speaker profile.

[0324] The one or more speaker characteristics may include at least one of an age characteristic and a gender characteristic.

[0325] The computer may be further configured to determine at least one of an age characteristic and a gender characteristic of the inbound speaker by applying a second machine learning model to the inbound speech signal.

[0326] The computer may be further configured to determine the age of the inbound speaker based on age characteristics stored in a speaker profile of the inbound speaker.

[0327] The computer may be further configured to receive speaker information from the content server indicating at least one speaker characteristic of the inbound speaker, store the at least one speaker characteristic in a speaker profile of the inbound speaker, and identify the speaker profile based in part on the at least one speaker characteristic in the speaker profile.

[0328] The computer may be further configured to receive inbound authentication data from the content server along with the inbound speech signal, and authenticate the inbound speaker based on the similarity score satisfying the similarity threshold and the inbound authentication data satisfying expected authentication data stored in the speaker profile.

[0329] 21. The system of claim 20, wherein the authentication data includes at least one of end-user device information, metadata associated with the end-user device, speaker information, and biometric information.

[0330] The computer may be further configured to extract one or more features from the inbound speech signal and calculate a spoof score indicative of a likelihood that the inbound speech signal contains a spoof condition based on the one or more acoustic features by applying a second machine learning model.

[0331] The computer may be further configured to identify one or more media content files in the media content database having one or more content characteristics corresponding to the one or more speaker characteristics.

[0332] The computer may be further configured to enable age-restricted content in the content database based on an age characteristic of one or more speaker characteristics of a speaker that satisfies a corresponding age-restriction characteristic in the one or more content characteristics of the age-restricted content.

[0333] In one embodiment, a computer-implemented method includes obtaining, by a computer, a speaker profile associated with a speaker, the speaker profile including one or more embeddings of the speaker; determining, by the computer, a maturity level of the speaker's voiceprint based on a false acceptance rate of the one or more embeddings and one or more maturity factors; and updating, by the computer, one or more similarity thresholds of the speaker profile according to the maturity level and the one or more maturity factors.

[0334] The method may further include generating, by the computer, a speaker profile in response to determining that the speaker is a new user of the media content system, the speaker profile being obtained.

[0335] The method may further include: extracting, by the computer, an inbound embedding of the speaker by applying a machine learning architecture to the inbound speech signal; generating, by the computer, a similarity score for the inbound embedding based on a relative distance between the inbound embedding and the voiceprint; and identifying, by the computer, the speaker profile stored in a speaker database in response to determining that the similarity score for the inbound embedding satisfies at least one similarity threshold of the one or more similarity thresholds for the speaker profile.

[0336] At least one threshold may include a low similarity threshold, and one or more thresholds may include a low similarity threshold and a high similarity threshold.

[0337] The method may further include, in response to determining, by the computer, that the similarity score of the inbound embedding satisfies a high similarity threshold of the one or more similarity thresholds, updating the speaker's voiceprint according to the inbound embedding, wherein the computer determines a level of maturity of the voiceprint after updating the voiceprint.

[0338] The method may further include updating, by the computer, one or more maturity factors of the voiceprint based on the inbound embeddings used to update the voiceprint, and the computer determines a level of maturity of the voiceprint after updating the one or more maturity factors.

[0339] The method may further include, in response to the computer determining that the maturity level of the voiceprint does not satisfy a maturity threshold, generating, by the computer, a prompt requesting additional inbound embeddings associated with the speaker; for each additional inbound embedding that satisfies one or more similarity thresholds, updating, by the computer, the voiceprint according to the additional inbound embedding; and updating, by the computer, one or more maturity factors of the voiceprint based on each additional inbound embedding.

[0340] The method may further include updating one or more similarity thresholds and, in response to determining, by the computer, that the maturity level of the voiceprint satisfies the maturity threshold, increasing one or more thresholds of the speaker profile.

[0341] The one or more maturity factors may be at least one of the number of embeddings of the speaker, the duration of the net utterances occurring in the one or more embeddings, and the quality of the speech from the one or more embeddings.

[0342] The method may further include determining, by the computer, a false acceptance rate based on the configuration input received from the administrative computer.

[0343] In one embodiment, a system includes a speaker profile database comprising a non-transitory machine-readable medium configured to store a plurality of speaker profiles; and a computer comprising a processor configured to: obtain a speaker profile associated with a speaker, the speaker profile including one or more embeddings of the speaker; determine a maturity level of the speaker's voiceprint based on a false acceptance rate of the one or more embeddings and one or more maturity factors; and update one or more similarity thresholds of the speaker profile according to the maturity level and the one or more maturity factors.

[0344] The computer may be further configured to generate a speaker profile in response to determining that the speaker is a new user of the media content system to obtain the speaker profile.

[0345] The computer may be further configured to extract an inbound embedding of the speaker by applying a machine learning architecture to the inbound speech signal to obtain a speaker profile; generate a similarity score for the inbound embedding based on a relative distance between the inbound embedding and the voiceprint; and identify a speaker profile stored in a speaker database in response to determining that the similarity score for the inbound embedding satisfies at least one similarity threshold of the one or more similarity thresholds for the speaker profile.

[0346] At least one threshold may include a low similarity threshold, and one or more thresholds may include a low similarity threshold and a high similarity threshold.

[0347] The computer may be further configured to update the speaker's voiceprint according to the inbound embedding in response to determining that the similarity score of the inbound embedding satisfies a high similarity threshold of the one or more similarity thresholds, and the computer determines a level of maturity of the voiceprint after updating the voiceprint.

[0348] The computer may be further configured to update one or more maturity factors of the voiceprint based on the inbound embeddings used to update the voiceprint, and the computer determines a level of maturity of the voiceprint after updating the one or more maturity factors.

[0349] The computer may be further configured to, in response to determining that the maturity level of the voiceprint does not satisfy a maturity threshold, generate a prompt requesting additional inbound embeddings associated with the speaker, and for each additional inbound embedding that satisfies one or more similarity thresholds, update the voiceprint according to the additional inbound embedding, and update one or more maturity factors of the voiceprint based on each additional inbound embedding.

[0350] The computer may be further configured to increase one or more thresholds of the speaker profile in response to determining that the maturity level of the voiceprint satisfies the maturity threshold to update one or more similarity thresholds of the device.

[0351] The one or more maturity factors may include at least one of the number of embeddings of the speaker, the duration of net utterances occurring in the one or more embeddings, and the quality of the speech from the one or more embeddings.

[0352] The computer may be further configured to determine a false acceptance rate based on configuration input received from the administrative computer.

[0353] In one embodiment, a device-implemented method includes: receiving, by the device, an inbound audio signal comprising an utterance of an inbound speaker; applying, by the device, an embedding extraction model to the inbound audio signal to extract an inbound embedding of the inbound speaker; generating, by the device, one or more similarity scores for the inbound embeddings based on a relative distance between the inbound embeddings and one or more voiceprints stored in a non-transitory machine-readable medium; identifying, by a computer, a speaker identifier associated with the inbound speaker's voiceprint in response to determining that the similarity score generated using the voiceprint satisfies a similarity threshold; and transmitting, by the device, the speaker identifier to a content server.

[0354] The device may include at least one of a smart TV, a set-top box, an edge device, a remote control, and a mobile communication device.

[0355] The method may further include receiving, by the device from the content server, media content for display via a display device coupled to the device.

[0356] The device may transmit a request for media content to a content server along with the speaker identifier.

[0357] The method may further include extracting, by the device, one or more features from the inbound audio signal, the features comprising one or more types of biometric features including at least one audio feature.

[0358] The method may further include receiving, by a microphone of the device, audio waves comprising a plurality of utterances from a plurality of inbound speakers, and converting, by the device, the audio waves into an inbound audio signal comprising the plurality of utterances.

[0359] The inbound speech signal may include multiple utterances of multiple inbound speakers, and the device may apply an embedding extraction model to the inbound speech signal to generate multiple similarity scores for the multiple inbound embeddings corresponding to the multiple inbound speakers, where the multiple similarity scores may be generated according to multiple voiceprints stored in a non-transitory machine-readable storage device of the device.

[0360] The method may further include, in response to the device determining that a second similarity score among the plurality of similarity scores for a second inbound embedding among the plurality of embeddings falls short of the at least one similarity threshold, generating, by the device, a second speaker profile for the second inbound speaker that encompasses the second inbound embedding and a second speaker identifier in non-transitory machine-readable storage.

[0361] The method may further include generating, by the device, one or more enrollment prompts for display to the inbound speaker requesting one or more enrollment voice signals of the inbound speaker; generating, by the device, a voiceprint of the inbound speaker based on one or more enrollment embeddings extracted from the one or more enrollment voice signals; in response to the device determining that the maturity level of the voiceprint does not satisfy the maturity threshold, generating, by the device, a second enrollment prompt requesting a second enrollment voice signal; and updating, by the device, the voiceprint of the inbound speaker based on the second enrollment embedding extracted from the second enrollment voice signal.

[0362] The method may further include extracting, by the device, one or more acoustic features from the inbound voice signal; and calculating, by the device, a spoof score indicative of a likelihood that the inbound voice signal contains a spoof condition based on the one or more acoustic features by applying a second machine learning model.

[0363] In one embodiment, a system includes a speaker database comprising a non-transitory machine-readable storage medium configured to store data records including a speaker profile; and a device comprising a processor configured to receive an inbound speech signal including an utterance of an inbound speaker; apply an embedding extraction model to the inbound speech signal to extract an inbound embedding of the inbound speaker; generate one or more similarity scores for the inbound embeddings based on a relative distance between the inbound embeddings and one or more voiceprints stored in the speaker database; and, in response to determining that the similarity score generated using the voiceprint satisfies a similarity threshold, identify, by the computer, a speaker identifier associated with the inbound speaker's voiceprint; and transmit the speaker identifier to a content server.

[0364] The device may be at least one of a smart TV, a set-top box, an edge device, a remote control, and a mobile communication device.

[0365] The device may be further configured to receive media content from a content server for display via a display device coupled to the device.

[0366] The device may be further configured to transmit a request for the media content to the content server along with the speaker identifier.

[0367] The device may be further configured to extract one or more features from the inbound voice signal, the features comprising one or more types of biometric features including at least one voice feature.

[0368] The system or device may further comprise a microphone configured to receive audio waves comprising multiple utterances from multiple inbound speakers, and the device may be further configured to convert the audio waves into an inbound audio signal comprising the multiple utterances.

[0369] The inbound speech signal may comprise multiple utterances of multiple inbound speakers. The device may be configured to apply an embedding extraction model to the inbound speech signal to generate multiple similarity scores for the multiple inbound embeddings corresponding to the multiple inbound speakers, where the device generates the multiple similarity scores according to multiple voiceprints stored in a speaker database.

[0370] The device may be further configured to determine that a second similarity score among the plurality of similarity scores for a second inbound embedding among the plurality of embeddings falls short of satisfying at least one similarity threshold, and generate a second speaker profile for the second inbound speaker that encompasses the second inbound embedding and a second speaker identifier in non-transitory machine-readable storage.

[0371] The device may be further configured to generate one or more enrollment prompts for display to the inbound speaker requesting one or more enrollment voice signals of the inbound speaker, generate a voiceprint of the inbound speaker based on one or more enrollment embeddings extracted from the one or more enrollment voice signals, and, in response to the device determining that the maturity level of the voiceprint does not satisfy the maturity threshold, generate a second enrollment prompt requesting a second enrollment voice signal, and update the voiceprint of the inbound speaker based on the second enrollment embedding extracted from the second enrollment voice signal.

[0372] The device may be further configured to extract one or more acoustic features from the inbound speech signal and calculate a spoof score indicative of a likelihood that the inbound speech signal contains a spoof condition based on the one or more acoustic features by applying a second machine learning model.

[0373] The various illustrative logic blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.

[0374] Computer software-implemented embodiments may be implemented in software, firmware, middleware, microcode, hardware description languages, or any combination thereof. A code segment or machine-executable instruction may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0375] The actual software code or specialized control hardware used to implement these systems and methods is not a limitation of the present invention. Accordingly, the operation and behavior of the systems and methods have been described without reference to specific software code, with the understanding that software and control hardware may be designed to implement the systems and methods based on the description herein.

[0376] If implemented in software, the functions may be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of a method or algorithm disclosed herein may be embodied in a processor-executable software module, which may reside on a computer-readable or processor-readable storage medium. Non-transitory computer-readable or processor-readable media includes both computer storage media and tangible storage media that facilitate transfer of a computer program from one place to another. Non-transitory processor-readable storage media may be any available medium that can be accessed by a computer. By way of example, and not limitation, such non-transitory processor-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other tangible storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer or processor. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically while discs reproduce data optically with a laser. Combinations of the above should also be included within the scope of computer-readable media. Furthermore, the operations of a method or algorithm may reside as one or any combination or set of code and / or instructions on a non-transitory processor-readable medium and / or computer-readable medium, which may be incorporated into a computer program product.

[0377] The previous description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.

[0378] While various aspects and embodiments have been disclosed, other aspects and embodiments are contemplated. The various aspects and embodiments disclosed are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.

Claims

1. 1. A computer-implemented method comprising: extracting, by a computer, an inbound embedding of the inbound speaker by applying a machine learning model to the inbound speech signal; generating, by the computer, a similarity score based on a distance between the inbound embedding and a voiceprint stored in a speaker profile in a speaker profile database; In response to the computer determining that the similarity score of the inbound embedding does not satisfy a similarity threshold, generating, by the computer, in the speaker profile database a new speaker profile for the inbound speaker that includes the inbound embedding, the new speaker profile being a database record that stores the inbound embedding as a new voiceprint for the inbound speaker, the new voiceprint satisfying a maturity threshold corresponding to one or more maturity factors for the new voiceprint, the one or more maturity factors including at least one of a number of utterances associated with the new voiceprint, an overall net speech duration of the entire utterance associated with the new voiceprint, or a phonetic quality of the utterance associated with the new voiceprint.

2. receiving, by the computer, the inbound voice signal from an end user device via an intermediate server; The method of claim 1 , further comprising transmitting, by the computer, a new speaker identifier associated with the new speaker profile to the intermediate server.

3. 10. The method of claim 1, further comprising extracting, by the computer, one or more features from the inbound speech signal, wherein the computer generates the inbound embedding by applying the machine learning model to the one or more features extracted from the inbound speech signal.

4. extracting, by the computer, a second inbound embedding from a second inbound speech signal by applying the machine learning model to the second inbound speech signal; generating, by the computer, a second similarity score based on the distance between the second inbound embedding and the new voiceprint stored in the new speaker profile; In response to the computer determining that the second similarity score of the second inbound embedding satisfies the similarity threshold, 2. The method of claim 1, further comprising: updating, by the computer, the new voiceprint of the inbound speaker based on the second inbound speech signal.

5. receiving, by the computer, a subscriber identifier associated with the inbound voice signal; 10. The method of claim 1, further comprising: identifying, by the computer, one or more speaker profiles associated with the subscriber identifier stored in the speaker profile database, wherein the computer generates one or more similarity scores for the inbound embedding based on one or more voiceprints stored in the one or more speaker profiles associated with the subscriber identifier.

6. The computer generates the new voiceprint based on one or more inbound embeddings, and the method further comprises: identifying, by the computer, the one or more maturity factors of the new voiceprint based on the one or more inbound embeddings; The method of claim 1 , further comprising determining, by the computer, a level of maturity of the new voiceprint based on the one or more maturity factors.

7. 7. The method of claim 6, further comprising updating, by the computer, a new similarity threshold for the new speaker profile in response to the computer determining that the maturity level of the new voiceprint satisfies the maturity threshold.

8. In response to the computer determining that the maturity level is below the maturity threshold, generating, by the computer, an active registration prompt, the active registration prompt including a user interface configured to display a request for additional inbound voice signals; extracting, by the computer, an additional embedding from the additional inbound speech signal; 7. The method of claim 6, further comprising: updating, by the computer, the new voiceprint according to additional embeddings extracted from the additional inbound voice signal.

9. 7. The method of claim 6, further comprising, in response to the computer determining that the level of maturity satisfies the maturity threshold, updating, by the computer, the new speaker profile from a temporary profile to a permanent profile.

10. 1. A system comprising: a speaker profile database comprising a non-transitory machine-readable storage medium configured to store data records containing speaker profiles; a computer including a processor, the processor including: extracting an inbound embedding of the inbound speaker by applying a machine learning model to the inbound speech signal; generating a similarity score based on a distance between the inbound embedding and a voiceprint stored in a speaker profile in the speaker profile database; In response to the computer determining that the similarity score of the inbound embedding does not satisfy a similarity threshold, generating a new speaker profile for the inbound speaker in the speaker profile database that includes the inbound embedding, the new speaker profile being a database record that stores the inbound embedding as a new voiceprint for the inbound speaker, the new voiceprint satisfying a maturity threshold corresponding to one or more maturity factors for the new voiceprint, the one or more maturity factors including at least one of a number of utterances associated with the new voiceprint, an overall net speech duration of an entire utterance associated with the new voiceprint, or a phonetic quality of the utterances associated with the new voiceprint.

11. The computer receiving the inbound voice signal from an end user device via an intermediate server; The system of claim 10 , further configured to: transmit a new speaker identifier associated with the new speaker profile to the intermediate server.

12. 12. The system of claim 11, wherein the end user device is at least one of a smart television, a media device coupled to a television, and an edge device.

13. 11. The system of claim 10, wherein the computer is further configured to extract one or more features from the inbound speech signal, and wherein the computer generates the inbound embedding by applying the machine learning model to the one or more features extracted by the computer from the inbound speech signal.

14. The computer extracting a second inbound embedding from a second inbound speech signal by applying the machine learning model to the second inbound speech signal; generating a second similarity score based on the distance between the second inbound embedding and the new voiceprint stored in the new speaker profile; In response to the computer determining that the second similarity score of the second inbound embedding satisfies a similarity threshold, 11. The system of claim 10, further configured to: update the new voiceprint of the inbound speaker based on the second inbound voice signal.

15. The computer receiving a subscriber identifier associated with the inbound voice signal; 11. The system of claim 10, further configured to: identify one or more speaker profiles associated with the subscriber identifier stored in the speaker profile database, wherein the computer generates one or more similarity scores for the inbound embedding based on one or more voiceprints stored in the one or more speaker profiles associated with the subscriber identifier.

16. 16. The system of claim 15, wherein the subscriber identifier is associated with one or more speaker identifiers, and each speaker profile is associated with a corresponding speaker identifier.

17. 17. The system of claim 16, wherein at least one of the subscriber identifier and each speaker identifier is an anonymized identifier.

18. The computer generates the new voiceprint based on one or more inbound embeddings, the computer: identifying the one or more maturity factors of the new voiceprint based on the one or more inbound embeddings; The system of claim 10 , further configured to: determine a level of maturity of the new voiceprint based on the one or more maturity factors.

19. The computer in response to determining that the maturity level is below the maturity threshold; generating an active registration prompt, the active registration prompt including a user interface configured to display a request for an additional inbound voice signal; extracting an additional embedding from the additional inbound speech signal; 20. The system of claim 18, further configured to: update the new voiceprint according to additional embeddings extracted from the additional inbound voice signal.

20. 20. The system of claim 18, wherein the computer is further configured to update the new speaker profile from a temporary profile to a permanent profile in response to the computer determining that the maturity level satisfies the maturity threshold.

Citation Information

Patent Citations

  • Speaker recognition method

    JP2007010995A

  • Speaker identifying device and speech recognition device, and speaker identifying program and program for speech recognition

    JP2008250089A

  • Voiceprint authentication processing method and device

    JP2018508799A

  • Device and method for privacy-preserving vocal interaction

    JP2019109503A

  • Speech recognition method and device

    JP2019528476A