Seamless customization of machine learning models

By using machine learning models to process voice data and update the user verification model, the problem of model training in the prior art requires a large number of manual operations and computing resources, seamless model update and continuous learning are achieved, and the accuracy and efficiency of the model are improved.

CN119948561APending Publication Date: 2025-05-06QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380066653.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-23
Filing Date
2023-07-21
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art requires a large amount of manual operation and computing resources when training machine learning models to achieve user verification, and it is difficult to achieve seamless updates and continuous learning of models.

Method used

By receiving voice data from the first user, processing voice data using a machine learning model to generate user verification scores and evaluate data quality, if certain criteria are met, the second user verification model is updated and the voice data is stored as a training example.

Benefits of technology

Seamless model updates and continuous learning are achieved, reducing manual operation and computational costs, and improving model accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948561A_ABST
    Figure CN119948561A_ABST
Patent Text Reader

Abstract

Certain aspects of the present disclosure provide techniques and apparatus for improved machine learning. Voice data from a first user is received. In response to determining that the speech data includes an utterance of the defined keyword, a user authentication score is generated by processing the speech data using a first user authentication machine learning (ML) model, and a quality of the speech data is determined. In response to determining that the user authentication score and the determined quality meet one or more defined criteria, a second user authentication ML model is updated based on the voice data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. patent application serial number 17 / 934,833 filed on September 23, 2022, which is hereby incorporated by reference into this application. Background Art

[0003] Aspects of the present disclosure relate to machine learning.

[0004] Machine learning architectures have been used to provide solutions to a wide variety of computational problems. Training machine learning models to perform accurately and reliably typically requires large amounts of data (and large amounts of computing resources) that are typically not available in common deployment systems (e.g., on an end user's smartphone). In addition, in many solutions, some level of user customization (e.g., training the model using data specific to the end user so that each user has a corresponding personalized model) is desirable for improved model performance. However, such customization requires personalized data for each user, and conventional approaches typically require the user to provide a large amount of manual effort and incur a large amount of computing overhead to implement model customization. Summary of the invention

[0005] Certain aspects provide a processor-implemented method for training a machine learning model for user verification, the method comprising: receiving voice data from a first user; in response to determining that the voice data includes an utterance of a defined keyword: generating a user verification score by processing the voice data using a first user verification machine learning (ML) model; and determining a quality of the voice data; and in response to determining that the user verification score and the determined quality meet one or more defined criteria, updating a second user verification ML model based on the voice data.

[0006] Certain aspects provide a processor-implemented method for performing user verification using machine learning, the method comprising: receiving voice data from a first user; in response to determining that the voice data includes an utterance of a defined keyword: generating a user verification score by processing the voice data using a first user verification machine learning (ML) model; and determining a quality of the voice data; and in response to determining that the user verification score and the determined quality meet one or more defined criteria, storing the voice data as a training example.

[0007] Other aspects provide: a processing system configured to perform the aforementioned methods and those described herein; a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of the processing system, cause the processing system to perform the aforementioned methods and those described herein; a computer program product embodied on a computer-readable storage medium comprising code for performing the aforementioned methods and those further described herein; and a processing system comprising components for performing the aforementioned methods and those further described herein.

[0008] The following description and associated drawings set forth in detail certain illustrative features of the one or more aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The drawings depict certain aspects of the one or more aspects and therefore are not to be considered limiting of the scope of the present disclosure.

[0010] Figure 1 Depicted are example workflows for seamlessly updating machine learning models and / or training and refining new models without manual user re-enrollment.

[0011] Figure 2 Depicts an example workflow for continuous learning of a machine learning model.

[0012] Figure 3 Depicts an example workflow for improved federated learning of machine learning models.

[0013] Figure 4 is a flowchart depicting an example method for seamlessly updating a machine learning model.

[0014] Figure 5 is a flow chart depicting an example method for continuous learning of a machine learning model.

[0015] Figure 6 is a flow chart depicting an example method for improved federated learning for machine learning models.

[0016] Figure 7 is a flow chart depicting an example method for updating a user authentication model.

[0017] Figure 8 is a flow chart depicting an example method for evaluating speech data for improved machine learning.

[0018] Fig. 9 An example processing system configured to perform various aspects of the present disclosure is depicted.

[0019] To facilitate understanding, identical reference numerals have been used, where possible, to designate identical elements that are common to the figures. It is contemplated that elements and features of one aspect may be beneficially incorporated in other aspects without further recitation. DETAILED DESCRIPTION

[0020] Aspects of the present disclosure provide apparatus, methods, processing systems, and non-transitory computer-readable media for improved machine learning model customization.

[0021] In some aspects, user data may be collected and automatically evaluated and validated to provide seamless model updates and continuous learning in a manner that improves model performance (e.g., resulting in improved customization and, therefore, improved model accuracy), reduced manual work (e.g., because the user need not perform data collection, validation, or labeling), reduced computational overhead (e.g., because only verified and high-quality data is stored and the remaining data may be immediately and automatically discarded), and generally improved functionality of the computing system.

[0022] In some examples discussed herein, user voice verification is treated as a technical problem that can be solved using a customized machine learning model trained and refined using aspects of the present disclosure. However, aspects of the present disclosure can be easily applied to a wide variety of model training and refinement scenarios, including applications to other user customizations (e.g., facial recognition for a specific user) and general non-customized training and refinement (e.g., general collection and verification of data for various machine learning purposes).

[0023] In some aspects, machine learning models are used to verify or validate voice, gesture, or other sensory-based commands to detect user input or otherwise initiate various actions. For example, the trained model can be used to authenticate a user based on their voice before allowing the user to make further requests or commands (e.g., unlock a smartphone or modify data).

[0024] In many cases, this verification and / or keyword detection is performed using one or more global general models (e.g., trained to recognize keywords using voice data from multiple users), where the global model can be further refined or fine-tuned for a specific user. While this can provide sufficient data to train a model that may not be usable if limited to data from a specific user, these global models often respond incorrectly due to various challenges, including the user's dialect, pitch, etc. For example, a global model may activate or verify voice data that does not actually include a defined keyword or an utterance that is not actually spoken by an authorized user. Similarly, a global model may fail to identify or verify voice data that has an indicated keyword and / or is spoken by an authorized user.

[0025] Therefore, user-based customization can significantly improve the accuracy of these models. In conventional methods, users are usually required to explicitly provide the necessary user data to implement such customization. For example, the user may be prompted to repeat a keyword or phrase several times (e.g., as a model or user registration step, where the model is refined or fine-tuned for a specific user), thereby allowing the system to refine the global model for the speech data of a specific user. However, such existing customization is cumbersome and time-consuming. In addition, since it requires manual operation from the user, conventional methods cannot practically deploy frequent model or architecture updates because the user will need to repeatedly perform such registration. This reduces the performance of the model and the system because model changes cannot be implemented arbitrarily or on a large scale.

[0026] Aspects of the present disclosure may provide improved and automated data verification and model training to achieve such seamless training and customization. In some aspects, these processes are referred to as "seamless" to indicate that these processes (e.g., new models may be deployed) may be performed without manual operation on the part of the user (and in some aspects, the user may not even know that a new model is being deployed). In some aspects, currently deployed models (e.g., validation models currently used to verify voice data) are used to achieve automatic collection of new examples, which may then be used to refine or fine-tune new models for the user. In some aspects, the user device may continue to extract user utterances after registration and store them in a local storage device on the device. This may be used as labeled training data, and user privacy may be maintained by preventing data from being sent to any other device and / or retaining data locally (e.g., it may be used locally without leaving the registered device).

[0027] In some aspects, after collecting a sufficient amount of good quality samples (e.g., greater than 200), back propagation or model training can be performed so that the new model can be updated and accuracy can be improved. Using aspects of the present disclosure, during this continuous learning, no re-registration is required, and training can be performed on the user device (if sufficient computing resources are available) or shared with a local or remote host, as described in more detail below.

[0028] Example workflow for seamless updates of machine learning models

[0029] Figure 1An example workflow 100 is depicted for seamlessly updating a machine learning model and / or training and refining a model without the need for manual user re-enrollment. As used herein, the term "training" is generally used interchangeably with other terms such as refinement, fine-tuning, updating, and the like. In the illustrated example, the workflow 100 can be used to train or enroll using voice data for user verification. However, as described above, aspects of the present disclosure can be readily applied to a wide variety of machine learning tasks. In some aspects, the workflow 100 is performed by an edge device (also referred to as an end device or user device), such as a smartphone, smart watch, laptop, smart speaker, or digital assistant. In at least one aspect, the method 100 is performed by a computing system such as the one described below with reference to Fig. 9 The processing system 900 executes.

[0030] In the illustrated workflow 100, voice data 105 is received and evaluated by keyword component 110A. Voice data 105 generally corresponds to audio data, which may be received from a microphone (and corresponding audio processing system or component) or other components or sources. In various aspects, voice data 105 may include audio data formatted in a variety of formats (e.g., in pulse code modulation (PCM) format, in Mel frequency cepstral coefficient (MFCC) format, etc.). Although labeled as voice data 105, in some aspects, the input data may not actually include the user's voice. For example, voice data 105 may be an audio recording that is evaluated to determine whether it contains the user's voice completely, whether the voice contains one or more keywords or phrases, whether the speech is made by an authorized user, etc.

[0031] In one aspect, the keyword component 110A can be used to provide a first stage or initial processing of the voice data 105. In some aspects, the keyword component 110A can correspond to or use a lightweight algorithm or machine learning model that has a relatively small memory footprint and / or low computational burden compared to more robust models. For example, the keyword component 110A can be used to perform an initial evaluation of the voice data 105 as it is collected (e.g., continuously as the user device records the voice data) to determine whether the voice data 105 includes utterances of one or more defined keywords or phrases.

[0032] In one aspect, the keyword component 110A can generate a binary output indicating whether the speech data 105 includes an utterance of a keyword or phrase. As used herein, a "keyword" can include a single word, multiple words, a phrase, etc. In some aspects, the keyword component 110A generates a continuous value (e.g., between zero and one) indicating the probability or likelihood that the speech data 105 includes an utterance of a keyword or phrase. In one such aspect, the output score can be compared to one or more thresholds to determine whether to initiate downstream processing (e.g., where if the score is less than a threshold such as 0.7, the speech data 105 can be discarded).

[0033] In some aspects, because the keyword component 110A is lightweight, it can be used to efficiently process this large amount of data in order to determine whether additional downstream processing should be performed (e.g., by the keyword component 110B). That is, the keyword component 110A can generally require fewer computing resources than the keyword component 110B or other components. Therefore, the keyword component 110A can be effectively used to reduce the overall computing requirements of the system because the (more robust) downstream processing is only selectively performed, rather than being performed on all incoming voice data 105.

[0034] In some aspects, the keyword component 110A uses the machine learning model to identify defined keywords that are shared among all users. That is, the keyword component 110A can be used to detect a set of one or more specific keywords defined during training and used by any user and device using the model. In other aspects, the keyword component 110A can additionally or alternatively use the machine learning model to identify or detect custom keywords or phrases (e.g., selected by each individual user).

[0035] In one aspect, if the keyword component 110A does not detect an utterance of a keyword or phrase in the speech data 105, the data is discarded. In the illustrated workflow 100, if the keyword component 110A does detect an utterance of a keyword or phrase, the speech data 105 is passed to a set of downstream components, including a quality component 115, a keyword component 110B, and a validation component 120.

[0036] The keyword component 110B may generally use a machine learning model to perform similar analysis and functionality as the keyword component 110A (e.g., keyword detection / identification in the speech data 105). However, in one aspect, the keyword component 110B may use a more robust model (e.g., a model with more trainable parameters, layers, etc.) or an algorithm with a relatively higher computational cost than the keyword component 110A. Since the keyword component 110B is only used to evaluate the speech data 105, once it is verified by the keyword component 110A, the additional computational cost imposes a lower burden on the system than if it were used for all speech data 105. The keyword component 110B is generally more accurate than the keyword component 110A, and therefore produces fewer false positives and / or false negatives.

[0037] The keyword component 110B is generally used to verify or validate the output of the keyword component 110A. For example, the keyword component 110B may similarly use a machine learning model to attempt to identify or detect the utterance of one or more keywords or phrases (which may be static or user-defined). As described above, the keyword component 110B may output a binary indication and / or a continuous score indicating whether the speech data 105 includes a keyword utterance.

[0038] In the illustrated workflow 100, the quality component 115 is used to evaluate the quality of the speech data 105. For example, the quality component 115 can evaluate the speech data 105 to determine whether it contains clipping (e.g., where portions of the audio are clipped, such as due to the amount or magnitude of the waveform exceeding a threshold), what the signal-to-noise ratio (SNR) of the speech data 105 is (e.g., indicating the level of background noise in the audio data), the duration of the speech data 105, whether the speech data 105 contains distortion, the keyword ratio in the speech data 105 (e.g., the percentage of speech data 105 corresponding to keywords compared to non-keyword audio), etc.

[0039] In one aspect, the quality component 115 (or the evaluation component 125, as discussed in more detail below) can evaluate these metrics to determine whether the speech data 105 is of sufficient quality. For example, the system can determine the number of samples or portions of the speech data 105 that contain clipping, and determine whether the number or percentage is less than a threshold. Similarly, the system can determine whether the SNR meets or exceeds a minimum threshold (e.g., 12 dB), whether the duration meets or exceeds a minimum time length, whether the distortion is below a defined threshold, whether the keyword ratio meets or exceeds a threshold, etc.

[0040] In some aspects, the quality component 115 and / or the evaluation component 125 performs the analysis and generates an overall quality score (e.g., a continuous value indicating the quality of the speech data 105) and / or a binary value indicating whether it is of sufficiently high quality. For example, the quality component 115 can determine that the quality is sufficient only if none of the individual thresholds or criteria are violated, only if fewer than a defined number of thresholds or criteria are violated (e.g., if the duration is short, but all other quality criteria are met), etc.

[0041] In one aspect, the verification component 120 can be used to provide user verification based on the voice data 105. For example, the verification component 120 can use a user verification or voice machine learning model to determine whether the voice data 105 includes the voice of an authorized user (e.g., the owner or user of the device on which the system operates) rather than the voice of an unauthorized user. In some aspects, as described above, the verification component 120 uses a customized or personalized model (e.g., a global model that has been fine-tuned or refined using voice data from an authorized user). In some aspects, the verification component 120 can generate a binary value indicating whether the voice data 105 includes speech from an authorized user, a continuous value indicating the confidence that the voice data 105 includes speech from an authorized user, etc.

[0042] In the illustrated workflow 100, scores or other data generated by the quality component 115, the keyword component 110B, and the verification component 120 are provided to the evaluation component 125. Although depicted as separate components for conceptual clarity, in some aspects, the operations of the evaluation component 125 can be implemented by one or more other components, such as within the quality component 115, the keyword component 110B, and / or the verification component 120.

[0043] In one aspect, the evaluation component 125 can compare the generated scores or other data to confirm whether defined criteria (e.g., minimum and / or maximum thresholds) are met. For example, as described above, the evaluation component 125 can confirm that the quality score (or the value of the underlying quality metric) meets the defined criteria, that the confidence or probability generated by the keyword component 110B and / or the verification component 120 meets or exceeds a defined threshold, etc.

[0044] In some aspects, the system can use relatively higher thresholds for some scores compared to typical use. For example, during typical (e.g., non-training) use, the system can use default minimum values ​​for keyword scores (generated by keyword component 110B) and / or user verification scores (generated by verification component 120). If the voice data 105 meets these default thresholds, the system can proceed normally (e.g., to unlock the user device, retrieve requested data, etc.).

[0045] However, in some aspects, for the purpose of collecting and verifying data for automatic training or fine-tuning, the system may use relatively high thresholds. For example, during ordinary use, the system may accept a keyword score of 0.8 for ordinary purposes, but refuse to use the data as a training example unless it has a keyword score of 0.9 or greater. In various aspects, the specific thresholds used for each score may vary depending on the specific implementation. If higher thresholds are used, the system can ensure that the model is trained (or retrained) only on data with a higher probability that the (automatically generated) labels are accurate.

[0046] In some aspects, if the voice data 105 fails to meet either criterion (e.g., because the quality score is below a threshold, or because the user authentication score is below a threshold), then the evaluation component 125 can determine that the voice data 105 is insufficient for training (even if it is otherwise sufficient for general purposes, such as unlocking the device). Therefore, the voice data 105 can be discarded.

[0047] As shown, in the event that the evaluation component 125 determines that the speech data 105 meets the criteria, the speech data 105 is stored in a repository of training data 130. The repository may be located within a local user device, remotely located on another user device, and / or remotely located on a shared device (such as in the cloud). In some aspects, the training data 130 includes the speech data 105 itself (e.g., PCM or MFCC data). In some aspects, the training data 130 includes extracted features of the speech data 105 (e.g., extracted by a feature extractor, not depicted in the illustrated example), as described below with reference to Figure 2 Discuss in more detail.

[0048] In one aspect, once the training data 130 meets the defined criteria, the training component 135 can use it to generate an updated validation model 140. For example, once the training data 130 includes a minimum number of examples (e.g., five), the training component 135 can refine, fine-tune, or otherwise train or update the user validation machine learning model using stored speech data 105 that is known to belong to the user with a sufficiently high confidence, and to narrate the keyword with a sufficiently high confidence and to be of a sufficiently high quality.

[0049] In some aspects, to refine the model, the training component 135 processes the voice data 105 (or features extracted therefrom) to generate an output score (e.g., a probability or confidence indicating that the voice data 105 belongs to an authorized user). This score can then be compared to a known truth value (e.g., the fact that the voice data 105 actually came from the user), and the difference between the generated score and the truth value can be used to refine or update one or more parameters of the user verification model.

[0050] In one aspect, the updated validation model 140 may include an updated or refined version of a model currently used by the validation component 120, an updated architecture, an entirely new model or architecture, etc. For example, the updated validation model 140 may be the same architecture as the currently used model, but with refined parameters based on fine-tuning and / or based on additional global training. Similarly, the updated validation model 140 may have the same architecture, but with updated hyperparameters. For another example, the updated validation model 140 may be an entirely new architecture.

[0051] In the illustrated example, the updated validation model 140 can then be deployed / instantiated by the validation component 120 to process new speech data 105. In this way, the system can seamlessly and without downtime or manual re-enrollment bring the new model online, thereby improving the accuracy and efficiency of the system.

[0052] Using workflow 100, during ordinary use, the system can collect, verify, and store speech data 105 that is acceptable for training (e.g., for a customized global model) for a particular user. This can allow the system to seamlessly introduce updated or completely new model architectures and instances (e.g., a model with the same architecture but with updated parameters) without laborious re-enrollment. In addition, the system does not need to store or otherwise maintain training examples between such retraining, thereby reducing the system's storage footprint. That is, the system does not need to store the original enrollment examples (or any other speech data) except when a new model is introduced. This reduces long-term storage requirements. Moreover, because the examples can be deleted (or otherwise archived to a secure storage location) after retraining, the memory or storage requirements are relatively short-term and limited, and the user's privacy is enhanced.

[0053] Example workflow for continuous learning of machine learning models

[0054] Figure 2 An example workflow 200 for continuous learning of a machine learning model is depicted. As used herein, continuous learning refers to an ongoing learning process (e.g., updating model parameters) based on input data during runtime, allowing the system to learn using ever-increasing amounts of data and continuously updated data that may more accurately reflect the current environment or exogenous factors than the original training data.

[0055] In the illustrated example, workflow 200 can be used to continuously learn or refine a user verification model using voice data. However, as described above, aspects of the present disclosure can be easily applied to a variety of machine learning tasks. In some aspects, workflow 200 is performed by an edge device (also referred to as a terminal device or user device), such as a smart phone, smart watch, laptop, smart speaker, or digital assistant.

[0056] In some aspects, workflow 200 is similar to Figure 1 As discussed above and in more detail below, workflow 100 can be used to generate positive examples (e.g., voice data from an authorized user) for an enrollment process of a validation model for a particular user. In one aspect, workflow 200 can be used to generate both positive and negative examples to enable more extensive continuous learning of various models.

[0057] In the illustrated example, the various parts of the workflow 200 are related to Figure 1 The workflow 100 partially or completely overlaps. For example, as described above with reference to Figure 1 As discussed, in workflow 200, speech data 105 may be received and processed using keyword component 110A and then passed to quality component 115, keyword component 110B, and verification component 120. The general operations and functionality of keyword component 110A, quality component 115, keyword component 110B, and verification component 120 in workflow 200 may generally correspond to, reflect, or otherwise include the operations and functionality discussed above with reference to workflow 100.

[0058] For example, keyword component 110A may perform an initial detection or search for keywords or phrases in voice data 105 and selectively provide voice data 105 to quality component 115, keyword component 110B, and verification component 120. Quality component 115 may generally be used to assess or determine the quality (e.g., duration, SNR, etc.) of voice data 105, as described above. Keyword component 110B may generally be used to perform more accurate keyword or phrase identification / detection, as described above. Verification component 120 may generally be used to verify that voice data 105 was spoken or uttered by an authorized user, as described above.

[0059] In the illustrated workflow 200, the evaluation component 225 evaluates the quality score and the keyword score, as described above. For example, the evaluation component 225 can confirm that the keyword score meets or exceeds a threshold (e.g., a minimum confidence) to ensure that the speech data 105 is well suited for training a machine learning model. Similarly, the evaluation component 225 can evaluate the quality score or data (e.g., SNR, duration, etc.) to confirm that the speech data 105 is of sufficient quality for accurate training. In one aspect, if any criterion is not met, the evaluation component 225 can discard the speech data 105, as described above.

[0060] In the illustrated example, the evaluation component 225 does not consider the verification score. That is, the system can determine to use the voice data 105 for training even if it does not include the voice of an authorized user. Figure 1Compared to workflow 100 (which can be used to perform initial enrollment or training of a custom model), workflow 200 can be used to provide continuous training or refinement of a validation model. In one such aspect, the system can use both positive examples (e.g., examples in which voice data 105 includes the voice of an authorized user) and negative examples (e.g., examples in which voice data 105 includes the voice of another (unauthorized) user) to train or refine the model. In some aspects, negative examples can be limited to voice data spoken by an unauthorized user, but including keywords and of sufficiently high quality. In some aspects, negative examples can similarly include other data, such as data that does not include keywords (depending on the specific model being trained or refined).

[0061] In at least one aspect, the output of the verification component 120 can be used to label the speech data 105 (eg, as a positive example or a negative example), rather than to determine whether to use the speech data 105 for training.

[0062] In the illustrated aspect, if the speech data 105 meets the criteria, the labeling component 227 can be used to assign, attach, or otherwise associate a label to the speech data 105. For example, based on the output of the verification component 120, the labeling component 227 can determine whether the speech data 105 is a positive example or a negative example (e.g., whether it includes the recorded speech of an authorized user, or the recorded speech of an unauthorized user or other user), and generate a corresponding label.

[0063] In the illustrated example, these labeled examples are then stored in training data 230. The repository may be located within a local user device, remotely on another user device, and / or remotely on a shared device (such as in the cloud). In some aspects, training data 230 includes speech data 105 itself (e.g., PCM or MFCC data). In some aspects, training data 230 includes extracted features of speech data 105. For example, labeling component 227 (or another component) may process speech data 105 to extract or generate one or more features of the speech data (e.g., using a feature extraction model), rather than storing speech data 105.

[0064] Typically, the features may have a smaller memory footprint than the original speech data 105. This may allow them to be stored (e.g., in the training data 230) with a significantly reduced memory or storage footprint. In this way, the system may collect and store a large number of training examples (e.g., 200 or more) without significant overhead.

[0065] In one aspect, once the training data 230 meets the defined criteria, the training component 235 can use it to generate an updated validation model 240. For example, the criteria can include a minimum total number of examples, a minimum number of positive examples, a minimum number of negative examples, and / or a specified ratio of positive examples to negative examples (e.g., so that there are a considerable number of positive examples and negative examples, but not an overwhelming number of positive examples). In some deployments, positive examples may be relatively more common than negative examples. For example, during ordinary use, if voice data 105 of sufficient quality and containing keywords or phrases is recorded, it is substantially more likely to be spoken by an authorized user (e.g., the owner of the device) than by an unauthorized user. In other deployments, the opposite may be true.

[0066] Thus, in some aspects, the system (or remote system) may intelligently inject or provide examples as needed to balance the training data 230 and prevent overfitting. For example, if the training data 230 is unbalanced and includes significantly more positive examples than negative examples, the updated verification model 240 may have reduced accuracy and precision (e.g., fail to reliably identify unauthorized users).

[0067] In some aspects, the system (or a remote server, such as the system that initially trained the global validation model) can thus selectively provide examples as needed. For example, if the training data 230 includes a sufficient number of positive examples but insufficient negative examples, the system can inject a plurality of validated negative examples into the training data 230 to ensure that an accurate model is trained or refined.

[0068] When the training criteria are met, in the illustrated workflow 200, the training component 235 can refine, fine-tune, or otherwise train, retrain, or update the user-verified machine learning model using the stored speech data 105 or extracted features (known to include keywords with sufficiently high confidence and be of sufficiently high quality).

[0069] In some aspects, to refine the model, the training component 235 processes the speech data 105 (or features extracted therefrom) to generate an output score (e.g., a probability or confidence indicating that the speech data 105 belongs to an authorized user). The score can then be compared to a known true value (e.g., a label assigned by the labeling component 227), and the difference between the generated score and the true value can be used to refine or update one or more parameters of the model.

[0070] In one aspect, the updated validation model 240 may include an updated or refined version of a model currently used by the validation component 120, an updated architecture, an entirely new model or architecture, etc. For example, the updated validation model 240 may be the same architecture as the currently used model, but with refined parameters based on fine-tuning and / or based on additional global training. Similarly, the updated validation model 240 may have the same architecture, but with updated hyperparameters. For another example, the updated validation model 240 may be an entirely new architecture.

[0071] In the illustrated example, the updated validation model 240 can then be deployed / instantiated by the validation component 120 to process new speech data 105. In this way, the system can provide continuous learning and refinement of the validation model, ensuring continued accuracy without downtime or manual re-enrollment, thereby improving the accuracy and efficiency of the system.

[0072] Using workflow 200, during normal use, the system can collect, verify, and store speech data 105 (or corresponding features) that are acceptable for training (e.g., for customizing a global model and / or for fine-tuning a custom model) for a particular user or user group. This can allow the system to seamlessly provide updated models without laborious re-enrollment or training.

[0073] Example workflow for improved federated learning of machine learning models

[0074] Figure 3 An example workflow 300 for improved federated learning (referred to in some aspects as proxy federated learning) of a machine learning model is depicted. In the illustrated example, the workflow 300 can be used for initial training and / or continuous learning or refinement of a model using voice data from multiple devices and / or users. However, as described above, aspects of the present disclosure can be readily applied to a wide variety of machine learning tasks.

[0075] In some aspects, proxy federated learning involves offloading some or all of the training process to a separate device or system, whereas conventional federated learning involves each participating system retaining its respective training data locally. For example, in some aspects, a user device may collect and / or label data, and send the labeled data (or features) to a remote (trusted) system to train or refine a model. In some aspects, a user device may perform training locally (using local data), and send model updates to a centralized system that aggregates updates from multiple user devices to generate a new model.

[0076] In some aspects, workflow 300 is similar to Figure 1 Workflow 100 and Figure 2As discussed above and in more detail below, workflow 300 can be used to implement federated learning (e.g., training or retraining of a model on one or more other systems or devices), thereby reducing the computational load on edge devices.

[0077] In the illustrated example, the various parts of the workflow 300 are related to Figure 1 Workflow 100 and / or Figure 2 The workflow 200 partially or completely overlaps. For example, as described above with reference to Figure 1 and Figure 2 As discussed, in the workflow 300, the voice data 105 may be received and processed to detect or identify keywords, check the quality of the data, verify that the voice data corresponds to an authorized user, etc. (e.g., by the ML components 310A and 310N). In the illustrated example, one portion 305A of the workflow 300 is executed within a first environment (e.g., by a first edge device or user device), while a second portion 305N is executed in another environment (e.g., by a second edge device or user device).

[0078] Although two parts 305A and 305N (collectively referred to as parts 305) are depicted for conceptual clarity, in various aspects, any number of devices or environments may participate in workflow 300. In addition, in some aspects, each environment (e.g., each part 305) may be executed by various devices of a single user. For example, part 305A may be executed by a user's smart phone, while part 305N is executed by the user's smart speaker / digital assistant. In some aspects, some or all of the environments (e.g., each part 305) may be executed by various devices of different users. For example, part 305A may be executed by a first user's device, while part 305N is executed by a second user's device, where each user is considered an authorized user for their environment / device.

[0079] In the illustrated example, in portion 305A, speech data 105A is processed by ML component 310A. Similarly, in portion 305N, speech data 105N is processed by ML component 310N. Operation and functionality of ML components 310A and 310N (collectively, ML components 310) may generally correspond to, mirror, or otherwise include the operation and functionality of keyword component 110A, quality component 115, keyword component 110B, verification component 120, and / or evaluation components 125 and 225, as described above with reference to Figure 1 and Figure 2 discussed.

[0080] For example, something like Figure 1 and Figure 2The keyword component 110A of the ML component 310 can perform an initial search for keywords or phrases in the voice data 105 and selectively provide the voice data 105 to downstream components. Figure 1 and Figure 2 As discussed above with respect to the quality component 115 of the ML component 310, the ML component 310 may also be used to evaluate or determine the quality (e.g., duration, SNR, etc.) of the speech data 105, as described above. Figure 1 and Figure 2 In a similar manner to the keyword component 110B, the ML component 310 may be used to perform more accurate keyword or phrase identification / detection, as described above. Figure 1 and Figure 2 The verification component 120, ML component 310 can also be used to verify whether the voice data 105 is spoken or uttered by an authorized user of the environment or device, as described above.

[0081] In the illustrated workflow 300, similar to Figure 1 and Figure 2 The ML component 310 may also evaluate the quality score, keyword score, and / or validation score, as described above, using the evaluation component 125 and / or 225. For example, the ML component 310 may confirm that the keyword score meets or exceeds a threshold (e.g., a minimum confidence level) to ensure that the speech data 105 is well suited for training a machine learning model. Similarly, the ML component 310 may evaluate the quality score or data (e.g., SNR, duration, etc.) to confirm that the speech data 105 is of sufficient quality for accurate training. In one aspect, if any of the criteria are not met, the ML component 310 may discard the speech data 105, as described above.

[0082] In some aspects, as described above with reference to Figure 1 As discussed above, the ML component 310 may also evaluate the verification score to ensure that it meets the defined criteria (e.g., if the workflow 300 is being used to perform enrollment of a single user). Figure 2 As discussed, the ML component 310 may not evaluate a validation score (eg, if the workflow 300 is being used to provide continuous learning, or if the system otherwise requires or expects both positive and negative examples).

[0083] In the illustrated aspect, if the ML component 310 determines that the corresponding speech data 105 meets the defined criteria, as described above, then the tag component 327A or 327N (collectively, the tag component 327) can be used to assign, attach, or otherwise associate a tag to the speech data 105. Specifically, the tag component 327A can be used to tag the data in the portion 305A, and the tag component 327N can be used to tag the data in the portion 305N.

[0084] In some aspects, the specific labels assigned may vary depending on the underlying training objectives. For example, based on the output of the user verification portion of the ML component 310, the labeling component 327 can determine whether the speech data 105 is a positive example or a negative example (e.g., whether it includes the recorded voice of an authorized user, or the recorded voice of an unauthorized user or other user), and generate a corresponding label. For another example, based on the output of the keyword identification portion of the ML component 310, the labeling component 327 can determine whether the speech data 105 is a positive example or a negative example (e.g., whether it includes a defined keyword or phrase), and generate a corresponding label.

[0085] In the illustrated example, these labeled examples are then sent to another device or system, as represented by portion 329. Specifically, the training examples from each discrete portion 305 are sent and stored in shared training data 330. This repository may be located on another user device (e.g., a device that also participates in federated learning and performs data collection and labeling), a non-participating user device (e.g., a desktop computer or server managed by the user), and / or on a remote system (such as in the cloud).

[0086] In some aspects, the training data 330 includes the speech data 105 itself (e.g., PCM or MFCC data). That is, each user device may send the original speech data 105. In some aspects, each user device (e.g., the labeling component 327) may alternatively perform feature extraction, thereby allowing the training data 330 to include the extracted features of the speech data 105. As described above, the features may typically have less memory usage compared to the original speech data 105. This can allow them to be sent and stored (e.g., in the training data 330) with significantly reduced memory or storage usage. In this way, the system can reduce the bandwidth and network burden of the workflow 300 while also collecting and storing a large number of training examples (e.g., 200 or more) without significant burden.

[0087] In one aspect, once training data 330 meets defined criteria, training component 335 can use it to generate updated model 340. For example, as described above, the criteria can include a minimum total number of examples (e.g., in the case of a user registering a new model), a minimum number of positive examples, a minimum number of negative examples, and / or a specified ratio of positive examples to negative examples (e.g., so that there are a fair number of positive and negative examples, but not an overwhelming number of positive or negative examples).

[0088] In some deployments, positive examples may be relatively more common than negative examples. For example, during ordinary use, if voice data 105 of sufficient quality is recorded and contains a keyword or phrase, it is substantially more likely to be spoken by an authorized user (e.g., the owner of the device) than by an unauthorized user. In other deployments, the opposite may be true. Similarly, during ordinary use, if voice data 105 of sufficient quality is collected, it is substantially more likely to contain a defined keyword or phrase than not include the keyword.

[0089] Thus, in some aspects, the system (or remote system) may intelligently inject or provide examples as needed to balance the training data 330 and prevent overfitting. For example, if the training data 330 is unbalanced and includes significantly more positive examples than negative examples, the updated model 340 may have reduced accuracy and precision (e.g., fail to reliably identify unauthorized users).

[0090] In some aspects, the system corresponding to portion 329 can thus selectively provide examples as needed. For example, if the training data 330 includes a sufficient number of positive examples but insufficient negative examples, the system can inject a plurality of verified negative examples into the training data 330 to ensure that an accurate model is obtained for training or refinement.

[0091] When the training criteria are met, in the illustrated workflow 300, the training component 335 can refine, fine-tune, or otherwise train, retrain, or update one or more machine learning models (e.g., a user verification model, a keyword detection model, etc.) using the stored speech data 105 or the extracted features (each of which is known to be of sufficiently high quality, and corresponding labels).

[0092] In some aspects, to refine the model, the training component 335 processes the speech data 105 (or features extracted therefrom) to generate an output score (e.g., a probability or confidence indicating that the speech data 105 belongs to an authorized user and / or a probability or confidence indicating that the speech data 105 includes an utterance of a defined keyword or phrase). The score can then be compared to a known true value (e.g., a label assigned by the labeling component 327), and the difference between the generated score and the true value can be used to refine or update one or more parameters of the model.

[0093] In one aspect, the updated model 340 may include an updated or refined version of a model currently used by the ML component 310, an updated architecture, an entirely new model or architecture, etc. For example, the updated model 340 may be the same architecture as the currently used model, but with refined parameters based on fine-tuning and / or based on additional global training. Similarly, the updated model 340 may have the same architecture, but with new parameters generated using updated hyperparameters. For another example, the updated model 340 may be an entirely new architecture.

[0094] In the illustrated example, the updated model 340 may then be deployed / instantiated to each ML component 310 participating in the federated learning workflow 300 to process new speech data 105. In this manner, the system may provide for continuous learning and refinement of models (e.g., keyword detection models and / or user verification models), ensuring continued accuracy without downtime or manual re-enrollment, thereby improving the accuracy and efficiency of the system.

[0095] Using workflow 300, during normal use, the system can collect, validate, and store speech data 105 (or corresponding features) that are acceptable for training (e.g., for customizing a global model and / or for fine-tuning a custom model) using a federated learning environment (e.g., a collection of edge devices and one or more centralized systems). This can allow the system to seamlessly provide updated models without laborious re-enrollment or training.

[0096] Example method for seamless updates of machine learning models

[0097] Figure 4 4 is a flow chart of an example method 400 for seamlessly updating a machine learning model. In the illustrated example, the method 400 can be used to train or enroll using voice data for user authentication. However, as described above, aspects of the present disclosure can be readily applied to a wide variety of machine learning tasks. In some aspects, the method 400 is performed by an edge device (also referred to as an end device or user device), such as a smart phone, smart watch, laptop, smart speaker, or digital assistant. In at least one aspect, the method 400 is Figure 1 The workflow 100 provides additional details.

[0098] At block 405, the user device receives voice data (eg, Figures 1 to 3 As described above, the voice data may generally correspond to recorded audio (e.g., by one or more microphones of the user device). For example, the voice data may correspond to PCM data, MFCC data, etc. As described above, although referred to as voice data for conceptual clarity, the voice data may or may not actually include voice. That is, the voice data may include a recording of the user's voice, or may be background noise or other data.

[0099] At block 410, the user device evaluates the speech data to determine whether it contains one or more defined keywords or phrases. For example, the user device may use a lightweight initial machine learning model (e.g., Figure 1 and Figure 2 The keyword component 110A of the embodiment of the present invention determines whether the speech data includes (or may include) a keyword or phrase. In some aspects, the keyword or phrase may be from a predefined (static) list, or may be a user-specific custom word or phrase. As described above, this initial detection can be performed using a relatively lightweight model that requires fewer resources (e.g., less memory usage, reduced processor time, reduced latency, etc.) compared to downstream keyword detection.

[0100] If at box 410, the user device determines that no keyword is detected in the voice data, the method 400 continues to box 435, where the user device discards the utterance (e.g., voice data). Then, the method 400 returns to box 405. In this way, the user device can selectively process the voice data using a downstream model (which is typically more complex and computationally expensive), thereby reducing the computing resources required to process the voice data and improving the operation of the user device. In one aspect, discarding the utterance / voice data may generally include preventing the utterance / voice data from being stored or further processed, deleting the utterance / voice data from a storage device or memory, and the like.

[0101] If at block 410, the user device determines that the keyword or phrase is present in the voice data, then method 400 proceeds to block 415. At block 415, the user device generates a verification score for the voice data by processing the data using the trained machine learning model. Figure 1 and Figure 2 In the verification component 120 of the user device, the user device can process the data to determine whether the voice data includes the voice of an authorized user of the user device. That is, the user device can use the trained / customized user verification voice model to determine whether the spoken keyword is spoken by the authorized user. In some aspects, the verification score is a continuous value indicating the probability or confidence that the voice belongs to the user. In some aspects, the verification score is a binary classification.

[0102] At block 420, the user device generates a keyword score by processing the speech data using the trained machine learning model. Figure 1 and Figure 2The keyword component 110B of the user device can process the voice data to determine whether a keyword or phrase is spoken in the voice data. As described above, this second stage of keyword recognition can use a relatively more robust model (compared to the initial stage, typically requiring more computing resources) and generally provides improved accuracy (e.g., fewer true and false positives). In some aspects, the keyword score is a continuous value indicating the probability or confidence that the voice data includes the utterance of the keyword or phrase. In some aspects, the keyword score is a binary value.

[0103] At block 425, the user device generates a quality score for the voice data. For example, using Figure 1 and Figure 2 The quality component 115 of the user device can generate an overall quality score and / or a composite score containing sub-elements based on the quality of the voice data. As an example, the user device can determine whether the audio meets a minimum duration requirement, whether the clipping is below a threshold, whether the SNR is above a threshold, whether the keyword ratio exceeds a threshold, etc. In some aspects, the quality score is a binary classification indicating whether the voice data is of sufficiently high quality (e.g., whether the quality components used by the user device meet or exceed the required standards). In some aspects, the quality score is a continuous value indicating the quality of the data (where higher values ​​correspond to higher quality) and / or indicating the probability or confidence that the voice data is of sufficiently high quality.

[0104] Although depicted as sequential operations for conceptual clarity, in some aspects, blocks 415, 420, and 425 may be performed substantially in parallel to reduce latency of method 400. In other aspects, operations may be performed sequentially and / or conditionally (in any order) to reduce computational overhead. For example, a user device may first determine whether a verification score meets a threshold before continuing to generate a keyword score or a quality score.

[0105] Once the verification score, keyword score, and quality score have been generated, method 400 continues to block 430. At block 430, the user device determines whether the specified speech criteria are met. In some aspects, the speech criteria may include a determination as to whether the user device is even in the process of enrolling or fine-tuning a new verification model. In some aspects, as described above, the speech criteria may include minimum thresholds for the verification score, keyword score, and / or quality score. That is, the user device may determine whether the speech data is of sufficiently high quality, includes utterances of keywords or phrases (with sufficiently high probability or confidence), and / or whether the utterances are in the voice of an authorized user (with sufficiently high probability or confidence).

[0106] If the speech criteria are not met (e.g., if any score does not meet the criteria), the method 400 continues to block 435, where the user device discards the utterance / voice data. The method 400 then returns to block 405. In this manner, the user device may selectively store the speech data for future training (which may depend on high quality and accurate data), thereby improving model accuracy, simplifying or easing the training process, and improving the operation of the user device. In one aspect, discarding the utterance / voice data may generally include preventing the storage or further processing of the utterance / voice data, deleting the utterance / voice data from a storage device or memory, and the like.

[0107] If at block 430, the user device determines that the speech criteria are met, the method 400 proceeds to block 440, where the user device stores the utterance / speech data. In some aspects, as described above, storing the utterance may include storing raw speech data (e.g., PCM or MFCC data), or extracting and storing features from the speech data. In some aspects, the speech data is stored in a secure repository (e.g., a repository that is accessible only to specified operations or processes of the user device, not to all applications).

[0108] At block 445, the user device determines whether training criteria are met. In some aspects, the training criteria may include a minimum number of example utterances stored during block 440 (e.g., a minimum of five examples). In some aspects, the training criteria may include considerations related to the workload of the user device, such as whether any computationally expensive processes are ongoing, whether the user is currently using the user device, whether sufficient memory or processing power is available for training, etc. In at least one aspect, the training criteria include time of day and / or day of week (e.g., where training is postponed to a particular time window, such as overnight).

[0109] If the training criteria are not met, method 400 returns to block 405. In this manner, the user device may continue to collect, verify, and store or discard voice data during normal operation of the user device (e.g., when the user is using the device normally, including using keyword detection and / or user authentication functionality).

[0110] If the user device determines that the training criteria are met, method 400 continues to box 450, where the user device trains, refines, or fine-tunes the user verification model. For example, as described above, the user device can use the stored utterances to perform registration, fine-tuning, and / or customization of a global verification model for a particular user. In this way, method 400 allows new models and architectures to be dynamically deployed to user devices without interfering with normal operations or requiring any manual action by the user. Although not included in the illustrated example, in some aspects, the customized verification model can then be deployed by the user device (e.g., for use in verifying the identity of the user against future voice data).

[0111] Although not included in the illustrated example, in some aspects, after training is completed, the user device may discard or delete the voice data (or extracted features). This can reduce the memory / storage requirements of method 400 and protect user privacy. In at least one aspect, the user device may transmit the voice data to a secure area, repository or repository portion where the voice data can be maintained in a highly secure manner, rather than deleting the voice data. For example, a secure repository may be protected using encryption of its contents (which may allow it to be stored in any location), may be accessed only by a specified operation or process on the user device (e.g., may be accessed by the kernel), may be inaccessible to third-party applications (or even the operating system of the user device), may be completely stored in different storage devices (e.g., with a hardware boundary between the secure storage device and the remaining memory or storage device used by the device), etc. In some aspects, this may allow the user device to selectively retain speech for future use without compromising user privacy.

[0112] Example method for continuous learning of machine learning models

[0113] Figure 5 is a flow chart depicting an example method 500 for continuous learning of a machine learning model. In the illustrated example, the method 500 can be used to continuously learn or refine a user verification model using voice data. However, as described above, aspects of the present disclosure can be readily applied to a wide variety of machine learning tasks. In some aspects, the method 500 is performed by an edge device. In at least one aspect, the method 500 is Figure 2 Workflow 200 provides additional details.

[0114] In some aspects, method 500 is similar to Figure 4 400. For example, blocks 505, 510, 520, 525, 530, and 535 may include operations or processes similar to blocks 405, 410, 420, 425, 430, and 435, respectively.

[0115] At block 505, a user device receives voice data (eg, Figures 1 to 3 As described above (for example, referring to Figure 4 Box 405 and / or reference Figures 1 to 3 ), the voice data may generally correspond to recorded audio (e.g., via one or more microphones of a user device).

[0116] At block 510, the user device evaluates the speech data to determine whether it contains one or more defined keywords or phrases. Figure 4 As discussed in block 410 of , the user device may use a lightweight initial machine learning model (e.g., Figure 1 and Figure 2 The keyword component 110A) determines whether the speech data includes (or may include) a keyword or phrase.

[0117] If at block 510 , the user device determines that no keyword is detected in the voice data, the method 500 proceeds to block 535 , where the user device discards the utterance (eg, voice data). The method 500 then returns to block 505 .

[0118] If at block 510, the user device determines that a keyword or phrase is present in the voice data, then method 500 proceeds to block 520. At block 520, the user device generates a keyword score by processing the voice data using the trained machine learning model. For example, as described above with reference to Figure 4 As discussed in block 420, use Figure 1 and Figure 2 With the keyword component 110B, the user device may process the voice data to determine whether a keyword or phrase is spoken in the voice data.

[0119] At block 525, the user device generates a quality score for the voice data. Figure 4 As discussed in Box 425, use Figure 1 and Figure 2 The user device may generate an overall quality score and / or a comprehensive score including sub-elements based on the quality of the voice data using the quality component 115 .

[0120] Although depicted as sequential operations for conceptual clarity, in some aspects, blocks 520 and 525 may be performed substantially in parallel to reduce latency of method 500. In other aspects, operations may be performed sequentially and / or conditionally (in any order) to reduce computational overhead. For example, a user device may first determine whether a keyword score satisfies a threshold before continuing to generate a quality score.

[0121] Once the keyword score and quality score have been generated, the method 500 continues to block 530. At block 530, the user device determines whether the specified speech criteria are met. In some aspects, as described above, the speech criteria may include minimum thresholds for the keyword score and / or quality score. That is, the user device may determine whether the speech data is of sufficiently high quality and / or includes utterances of keywords or phrases (with sufficiently high probability or confidence).

[0122] If the speech criteria are not met (eg, if any of the scores do not meet the criteria), the method 500 continues to block 535 where the user device discards the speech / speech data. The method 500 then returns to block 505.

[0123] If at box 530, the user device determines that the speech criteria are met, the method 500 continues to box 540, where the user device can extract features from the speech data and / or mark the speech data (or the extracted features). In some aspects, as described above, the user device marks the data based on whether the recorded speech corresponds to an authorized user (e.g., as determined using a user verification model). In some aspects, the user device extracts and marks the features and stores them for subsequent training. In other aspects, the user device can directly mark the speech data and store it for subsequent training, as described above.

[0124] At block 545, the user device determines whether the refinement criteria are met. In some aspects, the refinement criteria may include a minimum number of example utterances marked / stored during block 540, a minimum number of positive examples, a minimum number of negative examples, a specific ratio (or range of ratios) of positive examples to negative examples, etc. In some aspects, the refinement criteria may include considerations of the speed of sound related to the workload of the user device, whether the user is currently using the user device, whether sufficient memory or processing power is available for training, etc. In at least one aspect, the refinement criteria include time of day and / or day of week.

[0125] If the refinement criteria are not met, method 500 returns to block 505. In this manner, the user device may continue to collect, verify, and store or discard voice data during normal operation of the user device (e.g., when the user is using the device normally, including using keyword detection and / or user authentication functionality).

[0126] If the user device determines that the refinement criteria are met, method 500 continues to box 550, where the user device trains, refines, or fine-tunes the user verification model. For example, as described above, the user device can use the stored utterances (or features) to perform continuous online updates or learning of the user verification model for a particular user. In this way, method 500 allows the verification model to be continuously updated and refined without interfering with normal operation or requiring any manual operation by the user. Although not included in the illustrated example, in some aspects, the updated verification model can then be deployed by the user device (e.g., for use in verifying the identity of the user for future voice data).

[0127] Although not included in the illustrated example, in some aspects, after refinement is complete, the user device may discard or delete the voice data (or the extracted features). This can reduce the memory / storage requirements of method 500 and protect user privacy. In at least one aspect, the user device may transfer the voice data to a secure area, repository, or repository portion where the voice data can be maintained in a secure manner, rather than deleting the voice data.

[0128] Additionally, in at least one aspect, the user device can securely store the voice data and / or features and automatically delete the voice data and / or features when defined criteria occur (e.g., age of the data, number of stored examples, etc.). For example, the user device can delete any examples older than six months, delete the oldest examples when the total number of stored examples meets or exceeds a threshold, etc.

[0129] Example method for improved federated learning of machine learning models

[0130] Figure 6 6 is a flow chart depicting an example method 600 for improved federated learning of a machine learning model. In the illustrated example, the method 600 can be used for initial training and / or continuous learning or refinement of a model using voice data from multiple devices and / or users. However, as described above, aspects of the present disclosure can be readily applied to a wide variety of machine learning tasks. In some aspects, the method 500 is performed by an edge device. In at least one aspect, the method 600 is Figure 3 Workflow 300 provides additional details.

[0131] In some aspects, method 600 is similar to Figure 4 Method 400 and / or Figure 5 For example, blocks 605, 610, 630, and 635 may include operations or processes similar to blocks 405, 410, 430, and 435, respectively. Similarly, block 620 may include operations or processes similar to Figure 4 Similar operations may be performed at blocks 415, 420 and / or 425.

[0132] At block 605, the user device receives voice data (eg, Figures 1 to 3 As described above (for example, referring to Figure 4 Frame 405, Figure 5 Box 505 and / or reference Figures 1 to 3 ), the voice data may generally correspond to recorded audio (e.g., via one or more microphones of a user device).

[0133] At block 610, the user device evaluates the speech data to determine whether it contains one or more defined keywords or phrases. Figure 4 As discussed in block 410 of , the user device may use a lightweight initial machine learning model (e.g., Figure 1 and Figure 2 The keyword component 110A and / or Figure 3 The ML component 310 of the embodiment determines whether the speech data includes (or may include) a keyword or phrase.

[0134] If at block 610 , the user device determines that no keyword is detected in the voice data, the method 600 proceeds to block 635 , where the user device discards the utterance (eg, voice data). The method 600 then returns to block 605 .

[0135] If at block 610, the user device determines that the keyword or phrase is present in the voice data, then method 600 proceeds to block 615. At block 620, the user device generates a verification score, a keyword score, and / or a quality score for the voice data. In some aspects, as described above, the user device generates a verification score, a keyword score, and / or a quality score for the voice data using a trained machine learning model (e.g., Figure 1 and Figure 2 The verification component 120 and / or Figure 3 The ML component 310 of the user device processes the data to generate a verification score to determine whether the voice data includes the voice of an authorized user of the user device.

[0136] In some aspects, as described above, by using a trained machine learning model (e.g., Figure 1 and Figure 2 The keyword component 110B and / or Figure 3 The ML component 310 of the embodiment of the present invention processes the speech data to generate a keyword score to determine whether a keyword or phrase is spoken in the speech data. As described above, this second stage of keyword spotting can use a relatively more robust model (compared to the initial stage, typically requiring more computing resources) and generally provides improved accuracy (e.g., fewer true and false positives).

[0137] In some aspects, as described above, the user device generates a quality score (e.g., using Figure 1 and Figure 2 Quality components 115 and / or Figure 3 ML component 310) to generate an overall quality score and / or a composite score containing sub-elements based on the quality of the speech data.

[0138] Once the verification score, keyword score, and / or quality score have been generated, the method 600 continues to block 630. At block 630, the user device determines whether the specified speech criteria are met. In some aspects, as described above, the speech criteria may include minimum thresholds for the verification score, keyword score, and / or quality score. That is, the user device may determine whether the speech data is of sufficiently high quality, includes utterances of keywords or phrases (with sufficiently high probability or confidence), and / or whether the utterances are in the voice of an authorized user (with sufficiently high probability or confidence).

[0139] If the speech criteria are not met (eg, if any of the scores do not meet the criteria), the method 600 continues to block 635 where the user device discards the utterance / speech data. The method 600 then returns to block 605.

[0140] If at block 630, the user device determines that the speech criteria are met, the method 600 proceeds to block 640, where the user device may extract features from the speech data and / or tag the speech data (or the extracted features). In some aspects, as described above, the user device tags the data based on whether the recorded speech corresponds to an authorized user (e.g., as determined using a user verification model). In some aspects, the user device tags the data based on whether the speech data includes defined keywords or phrases (e.g., as determined using a keyword model). In the illustrated example, the user device may then send the tagged data (or features) to another system, such as in a federated learning system.

[0141] As mentioned above Figure 3 As discussed, other systems to which features are sent can similarly receive labeled features from any number of user devices. Using this aggregated data set, the system can train or refine various machine learning models (e.g., using a federated learning architecture). In the illustrated example, after sending the labeled features, method 600 returns to box 605. In this way, method 600 allows the user device to continuously collect training data for training and / or refining the model without interfering with normal operation or requiring any manual operation by the user. Although not included in the illustrated example, in some aspects, the updated model can then be deployed to the user device by the remote training system (e.g., to verify the user's identity for future voice data, or to perform future keyword detection).

[0142] Although not included in the illustrated example, in some aspects, after sending the features, the user device may discard or delete the voice data (or the extracted features). This can reduce the memory / storage requirements of method 600 and protect user privacy. In at least one aspect, the user device can transmit the voice data to a secure area, repository, or repository portion where the voice data can be maintained in a highly secure manner, rather than deleting the voice data.

[0143] Additionally, in at least one aspect, the user device can securely store the voice data and / or features and automatically delete the voice data and / or features when defined criteria occur (e.g., age of the data, number of stored examples, etc.). For example, the user device can delete any examples older than six months, delete the oldest examples when the total number of stored examples meets or exceeds a threshold, etc.

[0144] Example method for updating a user authentication model

[0145] Figure 7 is a flow chart depicting an example method 700 for updating a user authentication model. In some aspects, the method 700 is performed by an edge device (also referred to as an end device or user device), as described above.

[0146] At block 705, voice data is received from a first user.

[0147] At block 710 , in response to determining that the speech data includes an utterance of a defined keyword, a user verification score is generated by processing the speech data using a first user-verified machine learning (ML) model, and a quality of the speech data is determined.

[0148] At block 715 , in response to determining that the user authentication score and the determined quality satisfy the one or more defined criteria, a second user authentication ML model is updated based on the speech data.

[0149] In some aspects, determining that the speech data includes an utterance of the defined keyword includes processing the speech data using a first keyword identification ML model, and method 700 also includes confirming that the speech data includes an utterance of the defined keyword by processing the speech data using a second keyword identification ML model, and the second keyword identification ML model is more accurate than the first keyword identification ML model.

[0150] In some aspects, determining the quality of the speech data includes at least one of determining a signal-to-noise ratio (SNR) of the speech data, determining a clipping ratio of the speech data, or determining a duration of the speech data.

[0151] In some aspects, method 700 also includes storing the speech data as training examples, wherein updating the second user-verified ML model is performed based further in response to determining that a number of the stored training examples satisfies one or more defined criteria.

[0152] In some aspects, method 700 further includes deleting the stored training examples after updating the second user-verified ML model.

[0153] In some aspects, method 700 further includes, after updating the second user-verified ML model, storing the training examples in a storage location that satisfies one or more defined security criteria.

[0154] In some aspects, method 700 also includes processing subsequent speech data using a second user-verified ML model.

[0155] In some aspects, updating the second user-verified ML model based on the speech data includes extracting one or more features of the speech data, labeling the one or more features of the speech data based on the user-verification score, and storing the one or more features and the labels as training examples.

[0156] In some aspects, updating the second user-verified ML model based on the speech data is performed based on further responding to determining that the number of stored training examples satisfies one or more defined criteria, wherein the one or more defined criteria indicate at least one of the following items: a minimum number of stored positive examples corresponding to utterances spoken by the first user, a minimum number of stored negative examples corresponding to utterances not spoken by the first user, or a ratio of stored positive examples to stored negative examples.

[0157] In some aspects, updating the second user-verified ML model is performed using a federated learning operation.

[0158] In some aspects, the federated learning operation includes sending the training examples to a host system that performs an update of the second user-verified ML model.

[0159] In some aspects, the federated learning operation includes sending updated parameters of the second user-verified ML model to a host system, wherein the host system aggregates the updated parameters to update a global version of the second user-verified ML model.

[0160] Example Method for Evaluating Speech Data for Improved Machine Learning

[0161] Figure 8 8 is a flow chart depicting an example method 800 for evaluating speech data for improved machine learning. In some aspects, the method 800 is performed by an edge device (also referred to as an end device or user device), as described above.

[0162] At block 805, voice data is received from a first user.

[0163] At block 810 , in response to determining that the speech data includes an utterance of a defined keyword, a user verification score is generated by processing the speech data using a first user-verified machine learning (ML) model, and a quality of the speech data is determined.

[0164] At block 815 , in response to determining that the user verification score and the determined quality meet the one or more defined criteria, the speech data is stored as a training example.

[0165] In some aspects, determining that the speech data includes an utterance of the defined keyword includes processing the speech data using a first keyword identification ML model, and method 800 also includes confirming that the speech data includes an utterance of the defined keyword by processing the speech data using a second keyword identification ML model, and the second keyword identification ML model is more accurate than the first keyword identification ML model.

[0166] In some aspects, method 800 also includes updating the second user-verified ML model based further in response to determining that the number of stored training examples satisfies the one or more defined criteria.

[0167] In some aspects, method 800 further includes deleting the stored training examples after updating the second user-verified ML model.

[0168] In some aspects, method 800 further includes, after updating the second user-verified ML model, storing the training examples in a storage location that satisfies one or more defined security criteria.

[0169] In some aspects, method 800 also includes processing subsequent speech data using a second user-verified ML model.

[0170] Example Processing System

[0171] In some respects, reference Figures 1 to 8 The described workflows, techniques, and methods may be implemented on one or more devices or systems. Fig. 9 An example processing system 900 is depicted that is configured to perform various aspects of the present disclosure, including, for example, Figures 1 to 8The techniques and methods described herein. In one aspect, the processing system 900 may correspond to an edge device, such as a smart phone, a smart watch, a laptop, a smart speaker, or a digital assistant. In some aspects, the processing system 900 includes a training system (e.g., for federated learning). Although depicted as a single system for conceptual clarity, in at least some aspects, as discussed above, the operations described below with respect to the processing system 900 may be distributed across any number of devices. For example, a first system may train a model, while a second system evaluates speech data using the trained model.

[0172] The processing system 900 includes a central processing unit (CPU) 902, which may be a multi-core CPU in some examples. Instructions executed at the CPU 902 may be loaded, for example, from a program memory associated with the CPU 902, or may be loaded from a memory partition (e.g., a memory partition in the memory 924).

[0173] The processing system 900 also includes additional processing components tailored for specific functions, such as a graphics processing unit (GPU) 904 , a digital signal processor (DSP) 906 , a neural processing unit (NPU) 908 , a multimedia processing unit 910 , and a wireless connectivity component 912 .

[0174] An NPU, such as NPU 908, is typically a dedicated circuit configured to implement control and arithmetic logic for executing machine learning algorithms, such as algorithms for processing artificial neural networks (ANNs), deep neural networks (DNNs), random forests (RFs), etc. An NPU is sometimes alternatively referred to as a neural signal processor (NSP), a tensor processing unit (TPU), a neural network processor (NNP), an intelligence processing unit (IPU), a vision processing unit (VPU), or a graphics processing unit.

[0175] NPUs, such as NPU 908, are configured to accelerate the execution of common machine learning tasks, such as image classification, machine translation, object detection, and various other predictive models. In some examples, multiple NPUs may be instantiated on a single chip, such as a system on a chip (SoC), while in other examples, the NPU may be part of a dedicated neural network accelerator.

[0176] The NPU can be optimized for training or inference, or in some cases configured to balance performance between the two. For NPUs that can perform both training and inference, the two tasks can generally still be performed independently.

[0177] NPUs designed to accelerate training are generally configured to accelerate the optimization of new models, which is a highly computationally intensive operation that involves inputting an existing data set (usually labeled or tagged), iterating on the data set, and then adjusting model parameters (such as weights and biases) to improve model performance. Typically, optimization based on error predictions involves passing back through the layers of the model and determining the gradient to reduce the prediction error.

[0178] NPUs designed to accelerate inference are generally configured to operate on complete models. Thus, such NPUs can be configured to input new data slices and quickly process them through the trained model to generate model outputs (e.g., inferences).

[0179] In one implementation, the NPU 908 is part of one or more of the CPU 902 , the GPU 904 , and / or the DSP 906 .

[0180] In some examples, wireless connectivity component 912 may include, for example, subcomponents for third generation (3G) connectivity, fourth generation (4G) connectivity (e.g., 4G LTE), fifth generation connectivity (e.g., 5G or NR), Wi-Fi connectivity, Bluetooth connectivity, and other wireless data transmission standards. Wireless connectivity component 912 is further connected to one or more antennas 914.

[0181] The processing system 900 may also include one or more sensor processing units 916 associated with any manner of sensors, one or more image signal processors (ISPs) 918 associated with any manner of image sensors, and / or a navigation processor 920, which may include satellite-based positioning system components (e.g., GPS or GLONASS) and inertial positioning system components.

[0182] The processing system 900 may also include one or more input and / or output devices 922, such as a screen, a touch-sensitive surface (including a touch-sensitive display), physical buttons, speakers, microphones, and the like.

[0183] In some examples, one or more of the processors of processing system 900 may be based on the ARM or RISC-V instruction set.

[0184] The processing system 900 also includes a memory 924, which represents one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, etc. In this example, the memory 924 includes computer-executable components that can be executed by one or more of the aforementioned processors of the processing system 900.

[0185] Specifically, in this example, the memory 924 includes a keyword component 924A, a verification component 924B, a quality component 924C, a feature component 924D, and a training component 924E. The memory 924 also includes a collection of training examples 924F and model parameters 924G. Although the examples are not shown in the figure for the sake of conceptual clarity, the examples are shown in the figure below. Fig. 9 Although depicted as discrete components in the drawings, the illustrated components (and other components not depicted) can be implemented together or separately in various aspects.

[0186] The training examples 924F may generally correspond to stored speech data and / or features, as described above. For example, the training examples 924F may include PCM / MFCC data and / or extracted features (e.g., from Figures 1 to 3 The model parameters 924G may generally correspond to parameters of all or a portion of one or more machine learning models (such as one or more keyword detection models, user authentication models, etc.).

[0187] Processing system 900 also includes keyword circuit 926, verification circuit 927, quality circuit 928, feature circuit 929, and training circuit 930. The depicted circuits, as well as other circuits not depicted, may be configured to perform various aspects of the techniques described herein.

[0188] For example, the keyword component 924A and the keyword circuit 926 (which may correspond to Figure 1 to Figure 2 The keyword components 110A and / or 110B) may be used to process the speech data to detect the presence of a keyword or phrase, as described above. The verification component 924B and the verification circuit 927 (which may correspond to Figure 1 to Figure 2 The verification component 120 of the example embodiment may be used to process the speech data to determine whether the speaker (if any) is an authorized user, as described above. The quality component 924C and the quality circuit 928 (which may correspond to Figure 1 to Figure 2 The quality component 115 of the embodiment of the present invention can be used to evaluate and quantify the audio quality of the recorded voice data, as described above. The feature component 924D and the feature circuit 929 (which can correspond to Figure 2 The marking component 227 and / or Figure 3 The labeling component 327 of the embodiment can be used to extract features from the speech data, label features, etc. The training component 924E and the training circuit 930 (which can correspond to Figure 1 Training components 135, Figure 2 Training component 235 and / or Figure 3 The training component 335) can be used to calculate losses and / or refine the machine learning model (e.g., a user verification model), as described above.

[0189] Although for the sake of clarity Fig. 9 Although depicted as separate components and circuits, keyword circuit 926, verification circuit 927, quality circuit 928, feature circuit 929, and training circuit 930 may be implemented collectively or individually in other processing devices of processing system 900, such as in CPU 902, GPU 904, DSP 906, NPU 908, etc.

[0190] Generally, the processing system 900 and / or its components may be configured to perform the methods described herein.

[0191] It is worth noting that in other aspects, aspects of the processing system 900 may be omitted, such as where the processing system 900 is a server computer, etc. For example, in other aspects, the multimedia processing unit 910, the wireless connectivity component 912, the sensor processing unit 916, the ISP 918, and / or the navigation processor 920 may be omitted. Furthermore, aspects of the processing system 900 may be distributed among multiple devices.

[0192] Sample Clauses

[0193] Specific implementation examples are described in the following numbered clauses:

[0194] Item 1: A method for training a machine learning model for user verification, the method comprising: receiving voice data from a first user; in response to determining that the voice data includes an utterance of a defined keyword: generating a user verification score by processing the voice data using a first user verification machine learning (ML) model; and determining a quality of the voice data; and in response to determining that the user verification score and the determined quality meet one or more defined criteria, updating a second user verification ML model based on the voice data.

[0195] Clause 2: A method according to clause 1, wherein: determining that the speech data includes the utterance of the defined keyword includes processing the speech data using a first keyword identification ML model, and the method also includes confirming that the speech data includes the utterance of the defined keyword by processing the speech data using a second keyword identification ML model, and the second keyword identification ML model is more accurate than the first keyword identification ML model.

[0196] Clause 3: A method according to any one of clauses 1 to 2, wherein determining the quality of the speech data comprises at least one of the following: determining a signal-to-noise ratio (SNR) of the speech data; determining a clipping ratio of the speech data; or determining a duration of the speech data.

[0197] Clause 4: A method according to any one of clauses 1 to 3, further comprising storing the speech data as training examples, wherein updating the second user-verified ML model is performed based on further response to determining that the number of stored training examples meets one or more defined criteria.

[0198] Clause 5: A method according to any one of clauses 1 to 4, further comprising deleting the stored training examples after updating the second user-verified ML model.

[0199] Clause 6: A method according to any one of clauses 1 to 5, further comprising storing the training examples in a storage location that meets one or more defined security criteria after updating the second user-verified ML model.

[0200] Clause 7: A method according to any one of clauses 1 to 6, further comprising processing subsequent speech data using the second user-verified ML model.

[0201] Clause 8: A method according to any one of clauses 1 to 7, wherein updating the second user-verified ML model based on the voice data comprises: extracting one or more features of the voice data; labeling the one or more features of the voice data based on the user verification score; and storing the one or more features and the labels as training examples.

[0202] Clause 9: A method according to any one of clauses 1 to 8, wherein updating the second user-verified ML model based on the speech data is performed based on further responding to determining that the number of stored training examples satisfies one or more defined criteria, wherein the one or more defined criteria indicate at least one of the following items: a minimum number of stored positive examples corresponding to utterances spoken by the first user; a minimum number of stored negative examples corresponding to utterances not spoken by the first user; or a ratio of stored positive examples to stored negative examples.

[0203] Clause 10: A method according to any one of clauses 1 to 9, wherein updating the second user-verified ML model is performed using a federated learning operation.

[0204] Clause 11: The method of any one of clauses 1 to 10, wherein the federated learning operation comprises sending the training examples to a host system that executes the update of the second user-verified ML model.

[0205] Clause 12: A method according to any one of clauses 1 to 11, wherein the federated learning operation includes sending updated parameters of the second user-verified ML model to a host system, wherein the host system aggregates the updated parameters to update a global version of the second user-verified ML model.

[0206] Clause 13: A method for performing user verification using machine learning, the method comprising: receiving voice data from a first user; in response to determining that the voice data includes utterance of a defined keyword: generating a user verification score by processing the voice data using a first user verification machine learning (ML) model; and determining a quality of the voice data; and in response to determining that the user verification score and the determined quality meet one or more defined criteria, storing the voice data as a training example.

[0207] Clause 14: A method according to clause 13, wherein: determining that the speech data includes the utterance of the defined keyword includes processing the speech data using a first keyword identification ML model, and the method also includes confirming that the speech data includes the utterance of the defined keyword by processing the speech data using a second keyword identification ML model, and the second keyword identification ML model is more accurate than the first keyword identification ML model.

[0208] Clause 15: The method of any of clauses 13 to 14, further comprising updating the second user-verified ML model based further in response to determining that the number of stored training examples satisfies one or more defined criteria.

[0209] Clause 16: A method according to any one of clauses 13 to 15, further comprising deleting the stored training examples after updating the second user-verified ML model.

[0210] Clause 17: A method according to any one of clauses 13 to 16, further comprising storing the training examples in a storage location that meets one or more defined security criteria after updating the second user-verified ML model.

[0211] Clause 18: A method according to any one of clauses 13 to 17, further comprising processing subsequent speech data using the second user-verified ML model.

[0212] Clause 19: A processing system, the processing system comprising: a memory, the memory comprising computer-executable instructions; one or more processors, the one or more processors configured to execute the computer-executable instructions and cause the processing system to perform a method according to any one of clauses 1 to 18.

[0213] Clause 20: A processing system comprising means for performing the method of any one of clauses 1 to 18.

[0214] Clause 21: A non-transitory computer readable medium comprising computer executable instructions which, when executed by one or more processors of a processing system, cause the processing system to perform the method of any one of clauses 1 to 18.

[0215] Clause 22: A computer program product embodied on a computer readable storage medium, the computer readable storage medium comprising code for performing the method according to any one of clauses 1 to 18.

[0216] Additional considerations

[0217] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. The examples discussed herein are not limited to the scope, applicability or aspects set forth in the claims. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. For example, without departing from the scope of the present disclosure, the functions and arrangements of the elements discussed may be changed. Various examples may omit, replace or add various procedures or components as appropriate. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted or combined. In addition, the features described for some examples may be combined in some other examples. For example, any number of aspects set forth herein may be used to implement a device or practice method. In addition, the scope of the present disclosure is intended to cover such devices or methods practiced using other structures, functionality or structures and functionality that are supplementary or alternative to the various aspects of the present disclosure set forth herein. It should be understood that any aspect of the disclosure disclosed herein may be embodied by one or more elements of the present invention.

[0218] As used herein, the word “exemplary” means “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects.

[0219] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items (including single members). As an example, "at least one of a, b, or c" is intended to cover: a, b, c, ab, ac, bc, and abc, as well as any combination with multiple identical elements (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or any other ordering of a, b, and c).

[0220] As used herein, the term "determining" encompasses a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, searching (e.g., searching in a table, database, or other data structure), ascertaining, and the like. Furthermore, "determining" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and the like. Furthermore, "determining" may include resolving, selecting, choosing, establishing, and the like.

[0221] The method disclosed herein includes one or more steps or actions for implementing the method. The steps and / or actions of the method can be interchangeable with each other without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions can be modified without departing from the scope of the claims. In addition, the various operations of the method described above can be performed by any appropriate component capable of performing the corresponding function. The component may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs) or processors. Typically, in the case of operations illustrated in the accompanying drawings, those operations may have corresponding corresponding components with similar numbers plus functional components.

[0222] The following claims are not intended to be limited to the various aspects shown herein, but should be given the full scope consistent with the language of the claims. Within the claims, unless otherwise specified, the reference to the singular element is not intended to mean "one and only one", but "one or more". Unless otherwise specified, the term "some" refers to one or more. No claim element should be interpreted according to the provisions of 35 U.S.C. § 112 (f), unless the phrase "parts for..." is used to expressly record the element, or in the case of a method claim, the phrase "step for..." is used to record the element. All structural and functional equivalents of the elements of the various aspects described throughout the present disclosure that are known or will be known later to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be covered by the claims. In addition, nothing disclosed herein is intended to be dedicated to the public, regardless of whether such disclosure is explicitly recorded in the claims.

Claims

1. A computer-implemented method for training a machine learning model for user verification, comprising: receiving voice data from a first user; In response to determining that the speech data includes an utterance of the defined keyword: generating a user verification score by processing the speech data using a first user-verified machine learning (ML) model; as well as determining the quality of the voice data; as well as In response to determining that the user authentication score and the determined quality meet one or more defined criteria, updating a second user authentication ML model based on the speech data.

2. The computer-implemented method of claim 1 , wherein: determining that the speech data includes the utterance of the defined keyword comprises processing the speech data using a first keyword identification ML model, The method also includes confirming that the speech data includes the utterance of the defined keyword by processing the speech data using a second keyword identification ML model, and the second keyword identification ML model is more accurate than the first keyword identification ML model.

3. The computer-implemented method of claim 1 , wherein determining the quality of the speech data comprises at least one of: determining a signal-to-noise ratio (SNR) of the speech data; determining a clipping ratio of the speech data; or A duration of the voice data is determined.

4. The computer-implemented method of claim 1 , further comprising storing the speech data as training examples, wherein updating the second user-verified ML model is performed based further in response to determining that a number of the stored training examples satisfies one or more defined criteria.

5. The computer-implemented method of claim 4, further comprising deleting the stored training examples after updating the second user-verified ML model.

6. The computer-implemented method of claim 4, further comprising, after updating the second user-verified ML model, storing the training examples in a storage location that satisfies one or more defined security criteria.

7. The computer-implemented method of claim 1 , further comprising processing subsequent speech data using the second user-verified ML model.

8. The computer-implemented method of claim 1 , wherein updating the second user-verified ML model based on the speech data comprises: extracting one or more features of the speech data; tagging the one or more features of the speech data based on the user verification score; as well as The one or more features and the labels are stored as training examples.

9. The computer-implemented method of claim 8, wherein updating the second user-verified ML model based on the speech data is performed in further response to determining that the number of stored training examples satisfies one or more defined criteria, wherein the one or more defined criteria indicate at least one of the following: a minimum number of stored positive examples corresponding to an utterance spoken by the first user; a minimum number of stored negative examples corresponding to utterances not uttered by the first user; or The ratio of stored positive examples to stored negative examples.

10. The computer-implemented method of claim 8, wherein updating the second user-verified ML model is performed using a federated learning operation.

11. The computer-implemented method of claim 10, wherein the federated learning operation comprises sending the training examples to a host system that executes the update of the second user-verified ML model.

12. The computer-implemented method of claim 10, wherein the federated learning operation comprises sending updated parameters of the second user-verified ML model to a host system, wherein the host system aggregates the updated parameters to update a global version of the second user-verified ML model.

13. A computer-implemented method for performing user verification using machine learning, comprising: receiving voice data from a first user; In response to determining that the speech data includes an utterance of the defined keyword: generating a user verification score by processing the speech data using a first user-verified machine learning (ML) model; as well as determining the quality of the voice data; as well as In response to determining that the user verification score and the determined quality meet one or more defined criteria, the speech data is stored as a training example.

14. The computer-implemented method of claim 13, wherein: determining that the speech data includes the utterance of the defined keyword comprises processing the speech data using a first keyword identification ML model, The method also includes confirming that the speech data includes the utterance of the defined keyword by processing the speech data using a second keyword identification ML model, and the second keyword identification ML model is more accurate than the first keyword identification ML model.

15. The computer-implemented method of claim 13, further comprising updating a second user-verified ML model based further in response to determining that the number of stored training examples satisfies the one or more defined criteria.

16. The computer-implemented method of claim 15, further comprising deleting the stored training examples after updating the second user-verified ML model.

17. The computer-implemented method of claim 15, further comprising, after updating the second user-verified ML model, storing the training examples in a storage location that satisfies one or more defined security criteria.

18. The computer-implemented method of claim 15, further comprising processing subsequent speech data using the second user-verified ML model.

19. A processing system comprising: a memory including computer executable instructions; and one or more processors configured to execute the computer-executable instructions and cause the processing system to perform operations including: receiving voice data from a first user; In response to determining that the speech data includes an utterance of the defined keyword: generating a user verification score by processing the speech data using a first user-verified machine learning (ML) model; as well as determining the quality of the voice data; as well as In response to determining that the user authentication score and the determined quality meet one or more defined criteria, updating a second user authentication ML model based on the speech data.

20. The processing system of claim 19, wherein: determining that the speech data includes the utterance of the defined keyword comprises processing the speech data using a first keyword identification ML model, The operations also include confirming that the speech data includes the utterance of the defined keyword by processing the speech data using a second keyword identification ML model, and the second keyword identification ML model is more accurate than the first keyword identification ML model.

21. The processing system of claim 19, wherein determining the quality of the speech data comprises at least one of: determining a signal-to-noise ratio (SNR) of the speech data; determining a clipping ratio of the speech data; or A duration of the voice data is determined.

22. The processing system of claim 19, the operations further comprising storing the speech data as training examples, wherein updating the second user-verified ML model is performed based further in response to determining that a number of the stored training examples satisfies one or more defined criteria.

23. The processing system of claim 22, the operations further comprising deleting the stored training examples after updating the second user-verified ML model.

24. The processing system of claim 22, the operations further comprising, after updating the second user-verified ML model, storing the training examples in a storage location that satisfies one or more defined security criteria.

25. The processing system of claim 19, the operations further comprising processing subsequent speech data using the second user-verified ML model.

26. The processing system of claim 19, wherein updating the second user-verified ML model based on the speech data comprises: extracting one or more features of the speech data; tagging the one or more features of the speech data based on the user verification score; as well as The one or more features and the labels are stored as training examples.

27. The processing system of claim 26, wherein updating the second user-verified ML model based on the speech data is performed in further response to determining that the number of stored training examples satisfies one or more defined criteria, wherein the one or more defined criteria indicate at least one of the following: a minimum number of stored positive examples corresponding to an utterance spoken by the first user; a minimum number of stored negative examples corresponding to utterances not uttered by the first user; or The ratio of stored positive examples to stored negative examples.

28. The processing system of claim 26, wherein updating the second user-verified ML model is performed using a federated learning operation.

29. The processing system of claim 28, wherein the federated learning operation comprises sending the training examples to a host system that executes the update of the second user-verified ML model.

30. A processing system comprising: means for receiving voice data from a first user; means for performing the following operations in response to determining that the speech data includes an utterance of the defined keyword: generating a user verification score by processing the speech data using a first user-verified machine learning (ML) model; as well as determining the quality of the voice data; as well as means for updating a second user-verified ML model based on the speech data in response to determining that the user-verification score and the determined quality satisfy one or more defined criteria.