A method of training a
machine learning model to conduct
authentication transactions is provided that includes the steps of obtaining, by an electronic device, a training dataset of audio signals. Each
audio signal includes voice
biometric data of a user and information for a
passphrase spoken by the respective user and belongs to a same or different
data class. Each
data class includes a user identity and a
passphrase identifier. Moreover, the method includes the steps of creating, using a
machine learning model being trained, at least one embedding for each
audio signal. The
machine learning model includes parameters. Furthermore, the method includes calculating, by a
machine learning algorithm using the embeddings, a loss, and updating parameters of the
machine learning model based on the calculated loss. In response to determining criteria defining an end of training have been satisfied, deeming the
machine learning model to be operable for use in simultaneously successfully verifying the identity of a user based on voice
biometric data and verifying a
passphrase spoken by the user matches a secret passphrase during
authentication transactions.