System and method of speaker-independent embedding for identification and verification from audio

By using machine learning models to generate DP vectors from speaker-independent speech characteristics, the system addresses voice spoofing and noise issues in ASV, enhancing call source authentication and security.

JP2025169252APending Publication Date: 2025-11-12PINDROP SECURITY INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2025121219
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-03-05
Filing Date
2025-07-18
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

Existing automatic speaker verification (ASV) systems are susceptible to voice spoofing, synthetic voices, and background noise, making them unreliable for authenticating call sources, as telephone numbers and caller IDs can be intentionally altered.

Method used

Implementing machine learning models, including Gaussian mixture models and neural networks, to evaluate speaker-independent characteristics of speech signals, generating deep phonprint (DP) vectors that capture device type, microphone type, geographic location, and other attributes independent of the speaker's voice.

Benefits of technology

DP vectors enhance voice authentication by providing reliable verification of audio sources, detecting spoofed services, and identifying device and network types, complementing voice biometrics for improved security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025169252000001_ABST
    Figure 2025169252000001_ABST
Patent Text Reader

Abstract

To provide audio processing operations that evaluate characteristics of an audio signal independent of a speaker's voice.SOLUTION: A neural network architecture trains and applies a discriminative neural network tasked with modeling and classifying speaker-independent characteristics. Task-specific models generate or extract feature vectors from input audio data based on trained embedding extraction models. Embeddings from the task-specific models are concatenated to form a deep-phoneprint vector for an input audio signal. The DP vector is a low-dimensional representation of each speaker-independent characteristic of the audio signal and is applied to various downstream operations.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 985,757, filed March 5, 2020, the entire contents of which are incorporated by reference.

[0002] This application is generally related to U.S. Patent Application No. 17 / 066,210, filed October 8, 2020, U.S. Patent Application No. 17 / 079,082, filed October 23, 2020, and U.S. Patent Application No. 17 / 155,851, filed January 22, 2021, each of which is incorporated by reference in its entirety.

[0003] This application relates generally to systems and methods for training and deploying speech processing machine learning models. [Background technology]

[0004] Today, there are various forms of communication channels and devices available for voice communication, including Internet of Things (IoT) devices for communication over computing networks or various forms of telephone calls, such as landline calls, mobile phone calls, and Voice over IP (VoIP) calls, among others. In telephony systems, due to the introduction of virtual phone numbers, telephone numbers, automatic number identification (ANI), or caller identification (caller ID) are no longer uniquely tied to individual subscribers or telephone lines. Some VoIP services allow intentional spoofing of such identifiers (e.g., telephone numbers, ANI, caller ID), allowing callers to intentionally alter information transmitted to the recipient's display to disguise their identity. As a result, telephone numbers and similar telephony identifiers are no longer reliable for verifying the audio source of a call.

[0005] As caller ID services become less reliable, automatic speaker verification (ASV) systems are becoming necessary to authenticate the origin of telephone calls. However, ASV systems have strict net speech requirements and are susceptible to voice spoofing, such as voice modulation, synthetic voices (e.g., deepfakes), and replay attacks. ASV is also affected by background noise often experienced in telephone calls. Therefore, what is needed is a means to evaluate other attributes of a speech signal that are independent of the speaker's voice in order to verify the legitimate source of the speech signal. Summary of the Invention

[0006] Disclosed herein are systems and methods that can address the above-mentioned shortcomings, which may also provide any number of additional or alternative benefits and advantages. The embodiments described herein provide speech processing operations that evaluate characteristics of a speech signal that are speaker-independent or complementary to evaluating speaker-dependent characteristics. Computer-implemented software executes one or more machine learning models, which may include Gaussian mixture models (GMMs) and / or neural network architectures with discriminative neural networks, such as convolutional neural networks (CNNs), deep neural networks (DNNs), and recurrent neural networks (RNNs), referred to herein as "task-specific machine learning models" or "task-specific models," each tasked with and configured to model and / or classify corresponding speaker-independent characteristics.

[0007] Task-specific machine learning models are trained for each speaker-independent characteristic of the input speech signal using the input speech data and metadata associated with the speech data or speech source. Discriminant models are trained and developed to distinguish between classifications of speech characteristics. One or more modeling layers (or modeling operations) generate or extract feature vectors, sometimes called "embeddings" or sometimes combined to form embeddings, based on the input speech data. Specific post-modeling operations (or post-modeling layers) incorporate embeddings from task-specific models and train task-specific models. The post-modeling layers (or post-modeling operations) concatenate the speaker-independent embeddings to form deep-phoneprint (DP) vectors of the input speech signal. The DP vectors are low-dimensional representations of each of the various speaker-independent characteristics of speech signal aspects. Non-limiting examples of additional or alternative post-modeling operations or post-modeling layers of task-specific models may include classification operations / layers, fully connected layers, loss functions / layers, and regression operations / layers (e.g., probabilistic linear discriminant analysis (PLDA)).

[0008] The DP vectors may be used for various downstream operations or tasks such as creating audio-based exclusion / allow lists, enforcing audio-based exclusion / allow lists, authenticating registered legitimate audio sources, determining the device type, determining the microphone type, determining the geographic location of the source of the audio, determining the codec, determining the carrier, determining the network type involved in transmitting the audio, detecting spoofed services that have spoofed device identifiers, recognizing spoofed services, and recognizing audio events occurring in audio signals.

[0009] DP vectors can be used in voice-based authentication operations either alone or complementary to voice biometric features. Additionally or alternatively, DP vectors can be used for voice quality measurement purposes and can be combined with voice biometric systems for various downstream operations or tasks.

[0010] In one embodiment, a computer-implemented method includes applying, by a computer, a plurality of task-specific machine learning models to an inbound speech signal having one or more speaker-independent characteristics to extract a plurality of speaker-independent embeddings for the inbound speech signal; extracting, by the computer, a deep phonprint (DP) vector for the inbound speech signal based on the plurality of speaker-independent embeddings extracted for the inbound speech signal; and applying, by the computer, one or more post-modeling operations to the plurality of speaker-independent embeddings extracted for the inbound speech signal to generate one or more post-modeling outputs for the inbound speech signal.

[0011] In another embodiment, the database comprises a non-transitory memory configured to store a plurality of training speech signals having one or more speaker-independent characteristics. The server comprises a processor configured to: apply a plurality of task-specific machine learning models to the inbound speech signal having the one or more speaker-independent characteristics to extract a plurality of speaker-independent embeddings for the inbound speech signal; extract a deep phonprint (DP) vector for the inbound speech signal based on the plurality of speaker-independent embeddings extracted for the inbound speech signal; and apply one or more post-modeling operations to the plurality of speaker-independent embeddings extracted for the inbound speech signal to generate one or more post-modeling outputs for the inbound speech signal.

[0012] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are intended to provide further explanation of the invention as claimed. [Brief explanation of the drawings]

[0013] The present disclosure may be better understood by reference to the following drawings, in which components are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the present disclosure, in which reference characters refer to corresponding parts throughout the various views.

[0014] [Figure 1] 1 illustrates components of a system for receiving and analyzing audio signals from an end user, according to one embodiment. [Figure 2] 1 illustrates method steps for implementing a task-specific model for processing speaker-independent aspects of a speech signal, according to one embodiment. [Figure 3] 1 illustrates the execution steps of a method for training or enrolling operations of a neural network architecture for speaker-independent embedding, according to one embodiment. [Figure 4] 1 illustrates the execution steps of a multi-task learning method for training a neural network architecture for speaker-independent embedding, according to one embodiment. [Figure 5] 1 illustrates the execution steps of a method for applying a task-specific model of a neural network architecture and extracting a speaker-independent embedding and a DP vector of an input speech signal according to one embodiment. [Figure 6] FIG. 1 shows a diagram illustrating data flow between layers of a neural network architecture that uses DP vectors as a complement to speaker-dependent embeddings, according to one embodiment. [Figure 7] FIG. 1 shows a diagram illustrating data flow between layers of a neural network architecture that uses DP vectors as a complement to speaker-dependent embeddings, according to one embodiment. [Figure 8] FIG. 1 shows a diagram illustrating data flow between layers of a neural network architecture that uses DP vectors to authenticate audio sources according to an exclusion list and / or an allow list, according to one embodiment. [Figure 9] FIG. 1 shows a diagram illustrating data flow between layers of a neural network architecture using DP vectors to authenticate a device using a device identifier, according to one embodiment. [Figure 10] FIG. 1 shows a diagram illustrating data flow between layers of a neural network architecture using deep phone printing for dynamic registration, according to one embodiment. [Figure 11] FIG. 1 shows a diagram illustrating the data flow of a neural network architecture for using DP vectors to detect replay attacks, according to one embodiment. [Figure 12] FIG. 1 shows a diagram illustrating data flow between layers of a neural network architecture that uses DP vectors to identify spoof services associated with an audio signal, according to one embodiment. [Figure 13] FIG. 1 shows a diagram illustrating data flow between layers of a neural network architecture for identifying spoof services associated with an audio signal according to a label mismatch approach using DP vectors, according to one embodiment. [Figure 14] FIG. 1 shows a diagram illustrating data flow between layers of a neural network architecture that uses DP vectors to determine the microphone type associated with an audio signal, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0015] Reference will now be made to the exemplary embodiments illustrated in the drawings, and specific language will be used herein to describe the same. It will be understood that no limitation of the scope of the invention is thereby intended. Alterations and further modifications of the features of the invention illustrated herein, and further applications of the principles of the invention as illustrated herein, which will occur to those skilled in the art and in possession of this disclosure, are to be considered within the scope of the invention.

[0016] Described herein are systems and methods for processing audio signals with samples of a speaker's voice and using the results in any number of downstream computations or tasks. A computing device (e.g., a server) of the system executes software programming that implements various machine learning algorithms, including various types of variations of neural networks, such as Gaussian mixture models (GMMs) or convolutional neural networks (CNNs), deep neural networks (DNNs), and recurrent neural networks (RNNs). The software programming trains models to recognize and evaluate speaker-independent characteristics of audio signals received from audio sources (e.g., caller users, speaker users, originating locations, originating systems). Various characteristics of the audio signals are independent of a particular speaker of the audio signal, as opposed to speaker-dependent characteristics associated with a particular speaker's voice.

[0017] Non-limiting examples of speaker-independent characteristics include, among others, the device type from which the audio originates and is recorded (e.g., landline, mobile phone, computing device, Internet of Things (IoT) / edge device), the microphone type used to capture the audio (e.g., speakerphone, headset, wired or wireless headset, IoT device), the carrier over which the audio is transmitted (e.g., AT&T, Sprint, T-MOBILE, Google voice), codecs applied to compress and decompress the audio for transmission or storage, geographic location associated with the audio source (e.g., continent, country, state / province, county / city), determining if an identifier associated with the audio source has been spoofed, determining spoofing services (e.g., Zang, Tropo, Twilio) that may be used to change the source identifier associated with the audio, the type of network over which the audio is transmitted, audio events occurring in the input audio signal (e.g., cellular network, landline communication network, VOIP device), audio events occurring in the input audio signal (e.g., background noise, traffic sounds, television noise, music, crying baby, train or factory whistle, laughter), and the communication channel over which the audio was received.

[0018] The device type may be more granular, for example, to reflect the device manufacturer (e.g., Samsung, Apple) or model (e.g., Galaxy S10, iPhone X). The codec classification may further indicate, for example, a single codec or multiple cascaded codecs. The codec classification may also indicate or be used to determine other information. For example, an audio signal such as a phone call may originate from different source device types (e.g., landline, mobile phone, VoIP device), and therefore different audio codecs are applied to the audio of the call (e.g., SILK codec on Skype, G722 on WhatsApp, PSTN, GSM codec).

[0019] The system's server (or other computing device) executes one or more machine learning models and / or neural network architectures, including machine learning modeling layers or neural network layers for performing various operations, including layers of discriminative neural network models (e.g., DNN, CNN, RNN). Each particular machine learning model is trained to respond to aspects of the audio signal using audio data and / or metadata related to a particular aspect. The neural network learns to distinguish between classification labels for various aspects of the audio signal. One or more fully connected layers of the neural network architecture extract feature vectors or embeddings of the audio signal generated from each of the neural networks and concatenate the respective feature vectors to form a deep phonprint (DP) vector. The DP vector is a low-dimensional representation of different aspects of the audio signal.

[0020] DP vectors are used in various downstream operations. Non-limiting examples may include, among others, creating and enforcing audio-based exclusion lists, authenticating registered or legitimate audio sources, determining the device type, microphone type, geographic location of the audio source, codec, carrier, and / or network type involved in transmitting audio, detecting spoofed identifiers associated with audio signals, such as spoofed caller identifiers (Caller IDs), spoofed Automatic Number Identifiers (ANIs), or spoofed phone numbers, or recognizing spoofed services. Additionally or alternatively, DP vectors are complementary to voice biometric features. For example, DP vectors can be used with complementary voiceprints (e.g., audio-based speaker vectors or embeddings) to perform voice quality measurements and voice enhancements or to perform authentication operations with voice biometric systems.

[0021] For ease of explanation and understanding, the embodiments described herein involve a neural network architecture including any number of task-specific machine learning models configured to model and classify particular aspects of an audio signal, with each task corresponding to modeling and classifying a particular characteristic of the audio signal. For example, the neural network architecture may include a device-type neural network and a carrier neural network, where the device-type neural network models and classifies the type of device that emitted the input audio signal, and the carrier neural network models and classifies the particular communications carrier associated with the input audio signal. However, the neural network architecture need not include each task-specific machine learning model. For example, the server may run the task-specific machine learning models individually as separate neural network architectures, or may run any number of neural network architectures including any combination of task-specific machine learning models. The server then models or clusters the resulting outputs of each task-specific machine learning model to generate a DP vector.

[0022] Although the system is described herein as implementing a neural network architecture with any number of machine learning model layers or neural network layers, any number of combinations or architectural configurations of machine learning architectures are possible. For example, a shared machine learning model may use a shared GMM operation to jointly model an input audio signal and extract one or more feature vectors, and then for each task, a separate fully connected layer (of a separate fully connected neural network) may be implemented that performs various pooling and statistical operations for the specific task. Generally, the architecture includes a modeling layer, a pre-modeling layer, and a post-modeling layer. The modeling layer includes layers for performing audio processing operations, such as extracting feature vectors or embeddings from various types of features extracted from the input audio signal or metadata. The pre-modeling layer performs pre-processing operations, transformation operations, etc. to take the input audio signal and metadata and prepare them for the modeling layer, such as extracting features from the audio signal or metadata. The post-modeling layer performs operations that use the output of the modeling layer, such as training operations, loss functions, classification operations, and regression functions. The boundaries and functions of the layer types may vary in different implementations.

[0023] System Architecture FIG. 1 illustrates components of a system 100 for receiving and analyzing audio signals from end users. The system 100 includes an analysis system 101, service provider systems 110 for various types of entities (e.g., businesses, government agencies, universities), and end-user devices 114. The analysis system 101 includes an analysis server 102, an analysis database 104, and a management device 103. The service provider system 110 includes a provider server 111, a provider database 112, and an agent device 116. An embodiment may include additional or alternative components, or omit certain components from those illustrated in FIG. 1, and still fall within the scope of the present disclosure. For example, it may be common to include multiple service provider systems 110, or for the analysis system 101 to have multiple analysis servers 102. An embodiment may include any number of devices capable of performing the various features and tasks described herein, or may otherwise implement them. For example, FIG. 1 illustrates the analysis server 102 as a separate computing device from the analysis database 104. In some embodiments, the analytical database 104 may be integrated into the analytical server 102 .

[0024] The embodiment described with respect to FIG. 1 is merely an example of using speaker-independent embedding and Deep Phone Printing and does not necessarily limit other potential embodiments. The description of FIG. 1 refers to a situation in which an end user calls the service provider system 110 through various communication channels to contact and / or interact with services provided by the service provider. However, the operations and features of various Deep Phone Printing implementations described herein may be applicable to many situations for evaluating speaker-independent aspects of a speech signal. For example, the Deep Phone Printing speech processing operations described herein may be implemented within various types of devices and need not be implemented within a larger infrastructure. As one example, an IoT device 114d may implement various processes described herein when capturing an input speech signal from an end user or when receiving an input speech signal from another end user over a TCP / IP network. As another example, an end user device 114 may execute locally installed software that implements the Deep Phone Printing processes described herein, e.g., that enables the Deep Phone Printing processes in user-to-user interactions. Smartphone 114b may execute deep phone printing software when it receives an inbound call from another end user to perform certain downstream operations, such as verifying the identity of the other end user or indicating whether the other end user is using a spoofing service.

[0025] Various hardware and software components of one or more public or private networks may interconnect the various components of system 100. Non-limiting examples of such networks may include a local area network (LAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a wide area network (WAN), and the Internet. Communications over the networks may be performed according to various communication protocols, such as Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communication protocols. Similarly, end-user devices 114 may communicate with called parties (e.g., provider system 110) via telecommunications protocols, hardware, and software capable of hosting, transporting, and exchanging telephone communications and voice data associated with telephone calls. Non-limiting examples of telecommunications hardware may include switches and trunks, among other additional or alternative hardware used to host, route, or manage telephone calls, circuits, and signaling. Non-limiting examples of software and protocols for telecommunications may include SS7, SIGTRAN, SCTP, ISDN, and DNIS, among other additional or alternative software and protocols used to host, route, or manage telephone calls, circuits, and signaling. Components for telecommunications may be organized or managed by a variety of different entities, such as carriers, switches, and networks, among others.

[0026] The end user device 114 may be any communication or computing device that operates to allow callers to access the services of the service provider system 110 through various communication channels. For example, an end user may place a call to the service provider system 110 through a telephony network or through a software application executed by the end user device 114. Non-limiting examples of the end user device 114 may include a landline phone 114a, a mobile phone 114b, a calling computing device 114c, or an edge device 114d. The landline phone 114a and the mobile phone 114b are telecommunication-oriented devices (e.g., telephones) that communicate through telecommunication channels. The end user device 114 is not limited to telecommunication-oriented devices or channels. For example, in some cases, the mobile phone 114b may communicate through a computing network channel (e.g., the Internet). The end-user devices 114 may also include electronic devices with processors and / or software, such as a call computing device 114c or an edge device 114d, that implement voice-over-IP (VoIP) communications, data streaming over a TCP / IP network, or other computing network channels. The edge device 114d may include any IoT device or other electronic device for computing network communications. The edge device 114d may be any smart device capable of running software applications and / or performing voice interface operations. Non-limiting examples of the edge device 114d may include a voice assistant device, an automobile, a smart appliance, etc.

[0027] The service provider system 110 comprises various hardware and software components that capture and store various types of voice signal data or metadata associated with a caller's contact with the service provider system 110. This voice data may include, for example, an audio recording of the call and metadata associated with the software and various protocols used for a particular communication channel. Speaker-independent characteristics of the voice signal, such as voice quality or sampling rate, can represent (and be used to evaluate) various speaker-independent aspects such as the codec, type of end-user device 114, or carrier, among others.

[0028] The analysis system 101 and the provider system 110 represent network infrastructures 101, 110 comprising physically and logically related software and electronic devices managed or operated by various business organizations. Each network system infrastructure 101, 110 device is configured to provide the intended services of a particular business organization.

[0029] The analysis server 102 of the analysis system 101 may be any computing device equipped with one or more processors and software and capable of performing the various processes and tasks described herein. The analysis server 102 may host or communicate with the analysis database 104 and receive and process audio signal data (e.g., audio recordings, metadata) received from one or more provider systems 110. While FIG. 1 shows only a single analysis server 102, the analysis server 102 may include any number of computing devices. In some cases, a computing device of the analysis server 102 may perform all or some of the processes and services of the analysis server 102. The analysis server 102 may comprise a computing device operating in a distributed computing configuration or a cloud computing configuration and / or in a virtual machine configuration. In some embodiments, the functionality of the analysis server 102 may be partially or fully performed by a computing device of a provider system 110 (e.g., the provider server 111).

[0030] The analysis server 102 executes speech processing software including one or more neural network architectures with neural network layers for deep phoneprinting operations (e.g., extracting speaker-independent embeddings, extracting DP vectors) and any number of downstream speech processing operations. For ease of explanation, the analysis server 102 is described as executing a single neural network architecture for implementing deep phoneprinting, including neural network layers for extracting speaker-independent embeddings and deep phoneprint vectors (DP vectors), although in some embodiments, multiple neural network architectures may be employed. The analysis server 102 and neural network architectures logically operate in several operational phases, including a training phase, an enrollment phase, and a deployment phase (sometimes referred to as a "test" phase or an "inference" phase), although some embodiments need not perform an enrollment phase. Input speech signals processed by the analysis server 102 and neural network architectures include a training speech signal, an enrollment speech signal, and an inbound speech signal (processed during the deployment phase). The analysis server 102 applies a neural network architecture to each type of input audio signal during a corresponding computation phase.

[0031] The analysis server 102 or other computing devices (e.g., provider server 111) of system 100 can perform various preprocessing and / or data augmentation operations on input speech signals (e.g., training speech signals, enrollment speech signals, inbound speech signals). While the analysis server 102 may perform preprocessing and data augmentation operations when executing a particular neural network layer, the analysis server 102 may also perform a particular preprocessing or data augmentation operation as a separate operation from the neural network architecture (e.g., before feeding the input speech signal into the neural network architecture).

[0032] Optionally, the analysis server 102 performs any number of preprocessing operations before feeding the audio data to the neural network. The analysis server 102 may perform various preprocessing operations during one or more of the operation phases (e.g., training phase, enrollment phase, deployment phase), although the specific preprocessing operations performed may vary between operation phases. The analysis server 102 may perform various preprocessing operations separately from the neural network architecture or when executing an inner network layer of the neural network architecture. Non-limiting examples of preprocessing operations performed on the input audio signal include running voice activity detection (VAD) software or a VAD neural network layer, extracting features (e.g., one or more spectrotemporal features) from a portion (e.g., a frame, a segment) or from substantially all of a particular input audio signal, and converting the extracted features from a time-domain representation to a frequency-domain representation by performing a short-time Fourier transform (SFT) operation and / or a fast Fourier transform (FFT) operation, among other preprocessing operations. The features extracted from the input audio signal may include, for example, Mel Frequency Cepstral Coefficients (MFCCs), Mel filter banks, linear filter banks, bottleneck features, and communication protocol metadata fields, among other types of data. Preprocessing operations may also include parsing the audio signal into frames or subframes and performing various normalization or scaling operations.

[0033] As an example, the neural network architecture may include a neural network layer for VAD computation that parses a set of speech portions and a set of non-speech portions from each particular input audio signal. The analysis server 102 may train a classifier for VAD separately (as a separate neural network architecture) or together with the neural network architecture (as part of the same neural network architecture). When VAD is applied to features extracted from the input audio signal, it may output a binary result (e.g., speech detected, no speech detected) or an argument value (e.g., probability that speech occurs) for each window of the input audio signal, thereby indicating whether a speech portion occurs in a given window. The server may store the one or more sets of speech portions, the one or more sets of non-speech portions, and the input audio signal in memory storage locations, including short-term RAM, a hard disk, or one or more databases 104, 112.

[0034] As mentioned, the analysis server 102 or other computing devices (e.g., provider server 111) of the system 100 may perform various augmentation operations on input speech signals (e.g., training speech signals, enrollment speech signals, inbound speech signals). Non-limiting examples of augmentation operations include frequency augmentation, speech clipping, and duration augmentation, among others. The augmentation operations generate various types of distortions or degradations of the input speech signals so that the resulting speech signals are captured by, for example, convolution operations in a modeling layer that generate feature vectors or speaker-independent embeddings. The analysis server 102 may perform the various augmentation operations as separate operations from the neural network architecture or as an in-network augmentation layer. The analysis server 102 may perform the various augmentation operations in one or more of the computation phases, although the specific augmentation operations performed may vary between computation phases.

[0035] As described in detail herein, the neural network architecture includes any number of task-specific models configured to model and classify particular speaker-independent aspects of the input audio signal, with each task corresponding to modeling and classifying a particular speaker-independent characteristic of the input audio signal. For example, the neural network architecture may include a device-type neural network that models and classifies the type of device that emitted the input audio signal and a carrier neural network that models and classifies the particular communications carrier associated with the input audio signal. As noted, the neural network architecture need not include each task-specific model. For example, the server may execute the task-specific models individually as separate neural network architectures, or may execute any number of neural network architectures including any combination of task-specific models.

[0036] The neural network architecture includes task-specific models configured to extract corresponding speaker-independent embeddings and DP vectors based on the speaker-independent embeddings. The server applies the task-specific models to the speech portions (or the speech-only abridged speech signal) and then to the non-speech portions (or the non-speech-only abridged speech signal). The analysis server 102 applies a particular type of task-specific model (e.g., a speech event neural network) to substantially all of the input speech signal (or at least to the portion not parsed by the VAD). For example, the task-specific models of the neural network architecture include a device-type neural network and a speech event neural network. In this example, the analysis server 102 applies the device-type neural network to the speech portions and then again to the non-speech portions to extract speaker-independent embeddings for the speech and non-speech portions. The analysis server 102 then applies the device-type neural network to the input speech signal to extract the entire speech signal embedding.

[0037] During the training phase, the analysis server 102 receives training speech signals with a variety of speaker-independent characteristics (e.g., codecs, carriers, device types, microphone types) from one or more corpora of training speech signals stored in the analysis database 104 or other storage medium. The training speech signals may further include clean speech signals and simulated speech signals, each of which the analysis server 102 uses to train various layers of the neural network architecture.

[0038] The analytic server 102 may generate simulated speech signals by retrieving simulated speech signals from more analytic databases 104 and / or performing various data augmentation operations. In some cases, the data augmentation operations may generate simulated speech signals for a given input speech signal (e.g., training signal, enrollment signal), where the simulated speech signal incorporates manipulated features of the input speech signal that mimic the effects of a particular type of signal degradation or distortion of the input speech signal. The analytic server 102 stores the training speech signals in non-transitory media on the analytic server 102 and / or the analytic database 104 for future reference or operation of the neural network architecture.

[0039] Training speech signals are associated with training labels, which are separate machine-readable data records or metadata coding of the training speech signal data file. The labels may be generated by a user to indicate expected data (e.g., expected classification, expected features, expected feature vectors), or the labels may be generated automatically according to a computer-implemented process used to generate a particular training speech signal. For example, during a noise enhancement operation, the analysis server 102 generates a simulated speech signal by algorithmically combining an input speech signal and a type of noise impairment. The analysis server 102 generates or updates corresponding labels for the simulated speech signal that indicate the expected features, expected feature vectors, or other expected types of data for the simulated speech signal.

[0040] In some embodiments, the analysis server 102 performs an enrollment phase to develop enrollee speaker-independent embeddings and enrollee DP vectors. The analysis server 102 may perform some or all of the preprocessing and / or data augmentation operations on the enrollee speech signals of the enrolled speech sources.

[0041] During the training and enrollment phases in some embodiments, one or more fully connected layers, classification layers, and / or output layers of each task-specific model generate predicted outputs (e.g., predicted classifications, predicted speaker-independent feature vectors, predicted speaker-independent embeddings, predicted DP vectors, predicted similarity scores) for the training speech signals (or enrollment speech signals). Loss layers execute various types of loss functions to evaluate distances (e.g., differences, similarities) between predicted outputs (e.g., predicted classifications) to determine the level error between the predicted outputs and the corresponding expected outputs indicated by the training labels associated with the training speech signals (or enrollment speech signals). The loss layers, or other functions performed by the analysis server 102, tune or adjust hyperparameters of the neural network architecture until the distance between the predicted outputs and the expected outputs meets a training threshold.

[0042] During the enrollment operation phase, a registered voice source (e.g., end-user device 114, registered organization, enrollee user), such as a registered user of the service provider system 110, provides (to the analysis system 101) several enrollment voice signals containing examples of speaker-independent characteristics. In some embodiments, the registered voice source further includes examples of the enrolled user's speech. The enrolled user may provide the enrollee voice signal via and / or using any number of channels. The analysis server 102 or the provider server 111 actively or passively captures the enrollment voice signal. In active enrollment, the enrollee responds to voice or GUI prompts to provide the enrollee voice signal to the provider server 111 or the analysis server 102. As one example, the enrollee may respond to various interactive voice response (IVR) prompts in IVR software executed by the provider server 111 via a telephone channel. As another example, the enrollee may respond to various prompts generated by the provider server 111 and exchanged with a software application on the edge device 114d via a corresponding data communication channel. As another example, a registrant may upload a media file (e.g., WAV, MP3, MP4, MPEG) containing audio data to the provider server 111 or the analysis server 102 via a computing network channel (e.g., the Internet, TCP / IP). In passive enrollment, the provider server 111 or the analysis server 102 collects the enrollment audio signal through one or more communication channels without the registrant's knowledge and / or continuously over time. For embodiments in which the provider server 111 receives or otherwise collects the enrollment audio signal, the provider server 111 forwards (or otherwise transmits) the authentic enrollment audio signal to the analysis server 102 via one or more networks.

[0043] The analysis server 102 feeds each enrollment speech signal to a VAD, which parses the particular enrollment speech signal into a speech portion (or a reduced speech signal with only speech) and a non-speech portion (or a reduced speech signal with only no speech). For each enrollment speech signal, the analysis server 102 applies a trained neural network architecture, including a trained task-specific model, to the set of speech portions and again to the set of non-speech portions. The task-specific model generates an enrollment speaker-independent feature vector for the enrollment speech signal based on features extracted from the enrollment speech signal. The analysis server 102 algorithmically combines the enrollment feature vectors generated from the entire enrollment speech signal to extract a spoken speaker-independent enrollment embedding (for the speech portion) and a non-speech speaker-independent enrollment embedding (for the non-speech portion). The analysis server 102 then applies the full-speech task-specific model to generate a full-speech enrollment feature vector for each of the enrollment speech signals. The analysis server 102 then algorithmically combines the full-speech enrollment feature vectors to extract a full-speech speaker-independent enrollment embedding. The analysis server 102 then extracts an enrollment DP vector for the enrollee speech source by algorithmically combining each of the speaker-independent embeddings. The speaker-independent embeddings and / or DP vectors are sometimes referred to as "DeepPhonPrints."

[0044] The analysis server 102 stores the extracted enrollment speaker-independent embeddings and the extracted DP vectors for each of the various enrolled speech sources. In some embodiments, the analysis server 102 may also store the extracted enrollment speaker-dependent embeddings (sometimes called "voiceprints" or "enrollment voiceprints"). The enrolled speaker-independent embeddings are stored in the analysis database 104 or the provider database 112. Examples of neural networks for speaker verification are described in U.S. Patent Application Nos. 17 / 066,210 and 17 / 079,082, which are incorporated herein by reference.

[0045] Optionally, certain end user devices 114 (e.g., computing device 114c, edge device 114d) execute software programming associated with analysis system 101. The software program generates enrollment feature vectors by locally capturing enrollment audio signals and / or applies a trained neural network architecture locally (on-device) to each of the enrollment audio signals. The software program then transmits the enrollment feature vectors to provider server 111 or analysis server 102.

[0046] Following the training and / or enrollment phases, the analytic server 102 stores the trained or developed neural network architecture in the analytic database 104 or the provider database 112. The analytic server 102 places the neural network architecture in the training or enrollment phase, which may include enabling or disabling specific layers of the neural network architecture. In some implementations, a device in the system 100 (e.g., the provider server 111, the agent device 116, the management device 103, or the end-user device 114) instructs the analytic server 102 to move to the enrollment phase to develop the neural network architecture by extracting various types of embeddings for enrollment speech sources. The analytic server 102 then stores the extracted enrollment embeddings and the trained neural network architecture in one or more databases 104, 112 for later reference during the deployment phase.

[0047] During the deployment phase, the analysis server 102 receives an inbound speech signal from an inbound speech source originating from an end-user device 114 received through a particular communication channel. The analysis server 102 applies a trained neural network architecture to the inbound speech signal to generate a set of speech portions and a set of non-speech portions, extracts features from the inbound speech signal, and extracts an inbound speaker-independent embedding and an inbound DP vector for the inbound speech source. The analysis server 102 may use the extracted embedding and / or DP vector in various downstream operations. For example, the analysis server 102 may determine a similarity score based on the distance, difference / similarity between the enrollment DP vector and the inbound DP vector, where the similarity score indicates the likelihood that the enrollment DP vector originated from the same speech source as the inbound DP vector. As described herein, deep phoneprinting outputs generated by machine learning models, such as speaker-independent embeddings and DP vectors, can be used in various downstream operations.

[0048] The analysis database 104 and / or the provider database 112 may be hosted on a computing device (e.g., a server, a desktop computer) that includes hardware and software components capable of performing the various processes and tasks described herein, such as a non-transitory machine-readable storage medium and database management software (DBMS). The analysis database 104 and / or the provider database 112 contain any number of corpora of training speech signals that are accessible to the analysis server 102 via one or more networks. In some embodiments, the analysis server 102 trains a neural network using supervised training, and the analysis database 104 and / or the provider database 112 contain labels associated with the training or enrollment speech signals. The labels indicate, for example, expected data for the training or enrollment speech signals. The analysis server 102 may also query an external database (not shown) to access a third-party corpus of training speech signals. An administrator may configure the analysis server 102 to select training speech signals with various types of speaker-independent characteristics.

[0049] The provider server 111 of the provider system 110 executes software processes for interacting with end users through various channels. The processes may include, for example, routing calls to the appropriate agent device 116 based on the inbound caller's comments, instructions, IVR inputs, or other inputs submitted during the inbound call. The provider server 111 can capture, query, or generate various types of information about the inbound voice signals, the caller, and / or the end user device 114 and forward the information to the agent device 116. A graphical user interface (GUI) of the agent device 116 displays the information to the service provider's agents. The provider server 111 also sends information about the inbound voice signals to the analysis system 101 to perform various analysis processes on the inbound voice signals and any other voice data. The provider server 111 may send the information and voice data based on a preconfigured trigger condition (e.g., receiving an inbound telephone call), an instruction or query received from another device in the system 100 (e.g., the agent device 116, the management device 103, the analysis server 102), or as part of a batch sent at regular intervals or at predetermined times.

[0050] The management device 103 of the analysis system 101 is a computing device that enables personnel of the analysis system 101 to perform various administrative tasks or user-prompted analytical operations. The management device 103 may be any computing device equipped with a processor and software and capable of performing the various tasks and processes described herein. Non-limiting examples of the management device 103 may include a server, a personal computer, a laptop computer, a tablet computer, etc. In operation, a user uses the management device 103 to configure the operation of various components of the analysis system 101 or the provider system 110 and to issue queries and commands to such components.

[0051] The agent device 116 of the provider system 110 may allow an agent or other user of the provider system 110 to configure the operation of the device of the provider system 110. For calls made to the provider system 110, the agent device 116 receives and displays some or all of the information associated with the inbound voice signal routed from the provider server 111.

[0052] Example Operations Calculation Phase 2 illustrates steps of a method 200 for implementing a task-specific model for processing speaker-independent aspects of a speech signal. Embodiments may include additional, fewer, or different operations than those described in method 200. Method 200 is performed by a server executing machine-readable software code of a neural network architecture, including any number of neural network layers and neural networks, although various operations may be performed by one or more computing devices and / or processors. While a server is described as generating and evaluating enrollee embeddings, a server need not generate and evaluate enrollee embeddings in all embodiments.

[0053] In step 202, the server places the neural network architecture and task-specific model in a training phase. The server applies the neural network architecture to any number of training speech signals to train the task-specific model. The task-specific model includes a modeling layer. The modeling layer may include, for example, an "embedding extraction layer" or a "hidden layer" that generates feature vectors for the input speech signals.

[0054] During the training phase, the server applies modeling layers to the training audio signals to generate training feature vectors. Various post-modeling layers, such as a fully connected layer and a classification layer (sometimes referred to as a "classifier layer" or "classifier") of each task-specific model, determine task-related classifications based on the training feature vectors. For example, a task-specific model may include a device-type neural network or a carrier neural network. The classifier layer of the device-type neural network outputs a predicted brand classification of the device of the audio source, and the classifier layer of the carrier neural network outputs a predicted carrier classification associated with the audio source.

[0055] The task-specific model generates a predicted output (e.g., a predicted training feature vector, a predicted classification) for a particular training audio signal. The post-embedding modeling layers (e.g., a classification layer, a fully connected layer, a loss layer) of the neural network architecture execute a loss function according to the predicted output of the training signal and the label associated with the training audio signal. The server executes the loss function to determine the level of error of the training feature vector generated by the modeling layer of the particular task-specific model. The classifier layer (or other layer) adjusts hyperparameters of the task-specific model and / or other layers of the neural network architecture until the training feature vector converges to the expected feature vector indicated by the label associated with the training audio signal. Once the training phase is complete, the server stores the hyperparameters in a server memory location or other memory location. The server may also disable one or more layers of the neural network architecture during later computation phases to keep the hyperparameters fixed.

[0056] A particular type of task-specific model is trained on the entire training speech signal, such as a speech event neural network. The server applies these task-specific models to the entire training speech signal and outputs a predicted classification (e.g., a speech event classification). Similarly, a particular task-specific model is trained on the speech portion and separately on the non-speech portion of the training signal. For example, one or more layers of the neural network architecture define a VAD layer that the server applies to the training speech signal to parse the training speech signal into a set of speech portions and a set of non-speech portions. The server then applies each task-specific model to the speech portions to train a speech task-specific model, and again to the non-speech portions to train a non-speech task-specific model.

[0057] The server can train each task-specific model individually and / or sequentially, sometimes referred to as a "single-task" configuration, where each task-specific model includes a separate modeling layer for extracting a separate embedding. Each task-specific model outputs a separate predicted output (e.g., predicted feature vector, predicted classification). A loss function evaluates the level of error for each task-specific model based on the relative distance (e.g., similarity or difference) between the predicted output and the expected output, as indicated by the label. The loss function then adjusts hyperparameters or other aspects of the neural network architecture to minimize the level of error. Once the level error meets a training threshold, the server fixes (e.g., preserves, does not disturb) the hyperparameters or other aspects of the neural network architecture.

[0058] The server can jointly train task-specific models, sometimes referred to as a "multi-task" configuration, where multiple task-specific models share the same hidden modeling and loss layers but have specific, separate post-modeling layers (e.g., fully connected layers, classification layers). The server feeds training audio signals into a neural network architecture and applies each of the task-specific models. The neural network architecture includes a shared hidden layer for generating a joint feature vector for the input audio signals. Each of the task-specific models includes a separate post-modeling layer that takes the joint feature vector and generates, for example, a task-specific predicted output (e.g., predicted feature vector, predicted classification). In some cases, a shared loss layer is applied to each of the predicted outputs and labels associated with the training audio signals to tune hyperparameters and minimize error levels. Additional post-modeling layers may algorithmically combine or concatenate the predicted feature vectors of specific training audio signals to output a predicted DP vector, among other potential predicted outputs (e.g., predicted classification). As described above, the server executes a shared loss function that evaluates the level of error between predicted and expected joint outputs according to one or more labels associated with the training speech signals, and adjusts one or more hyperparameters to minimize the level of error. If the level error meets a training threshold, the server fixes (e.g., preserves, does not disturb) the hyperparameters or other aspects of the neural network architecture.

[0059] In step 204, the server places the neural network architecture and task-specific models in an enrollment phase and extracts enrolled embeddings for the enrolled speech sources. In some implementations, the server may enable and / or disable particular layers of the neural network architecture during the enrollment phase. For example, the server typically enables and applies each of the layers during the enrollment phase, but the server disables the classification layer. While the enrolled embeddings include speaker-independent embeddings, in some embodiments, the speaker modeling neural network may extract one or more speaker-dependent embeddings for the enrolled speech sources.

[0060] During the enrollment phase, the server receives enrollment speech signals of enrolled speech sources and applies task-specific models to extract speaker-independent embeddings. The server applies the task-specific models to the enrollment speech signals to generate enrollment feature vectors for each of the enrollment speech signals as described for the training phase (e.g., single-task configuration, multi-task configuration). For each task-specific model, the server statistically or algorithmically combines each of the enrollment feature vectors to extract a task-specific enrollment embedding.

[0061] Specific task-specific models are applied separately to the speech portions of the enrollment speech signal and again to the non-speech portions. The server applies VAD to each specific enrollment speech signal to parse the enrollment speech signal into a set of speech portions and a set of non-speech portions. For these task-specific models, a neural network architecture extracts two speaker-independent embeddings: a task-specific embedding of the speech portions and a task-specific embedding of the non-speech portions. Similarly, a specific task-specific model is applied to the entire enrollment speech signal (e.g., a speech event neural network) to generate an enrollment feature vector and extract the corresponding task-specific enrollment embedding.

[0062] The neural network architecture extracts enrollment DP vectors for the speech source. One or more post-modeling or output layers of the neural network architecture concatenate or algorithmically combine the various speaker-independent embeddings. The server then stores the enrollment DP vectors in memory.

[0063] In step 206, the server places the neural network architecture in a deployment phase (sometimes called the "inference" or "test" phase) when the neural network architecture generates inbound embeddings and inbound DP vectors for the inbound speech source. The server may enable and / or disable specific layers and task-specific models of the neural network architecture during the deployment phase. For example, the server typically enables and applies each of the layers during the deployment phase, but the server disables the classification layer. In the current step 206, the server receives an inbound speech signal of an inbound speaker and feeds the inbound speech signal to the neural network architecture.

[0064] In step 208, during the deployment phase, the server applies the neural network architecture and task-specific model to the inbound speech signal to extract an inbound embedding and an inbound DP vector. The neural network architecture then generates one or more similarity scores based on the relative distance (e.g., similarity, difference) between the inbound DP vector and one or more enrolled DP vectors. The server applies the task-specific model to the inbound speech signal to generate an inbound feature vector and extract an inbound speaker-independent embedding and an inbound DP vector for the inbound speech source as described for the training and enrollment phases (e.g., single-task configuration, multi-task configuration).

[0065] As an example, the neural network architecture extracts an inbound DP vector and outputs a similarity score indicating the distance (e.g., similarity, difference) between the inbound DP vector and the enrollee DP vector. A larger distance may indicate a lower likelihood that the inbound speech signal originated from the enrollee speech source that originated the enrollee DP vector due to less / lower similarity between the speaker-independent aspects of the inbound speech signal and the enrollee speech signal. In this example, the server determines that the inbound speech signal originated from the enrollee speech source when the similarity score meets a speech source matching threshold. The task-specific models and DP vectors may be used in any number of downstream operations, as described in various embodiments herein.

[0066] Single-task configuration training and enrollment 3 illustrates the execution steps of a method 300 for training or enrolling a neural network architecture for speaker-independent embedding. Embodiments may include operations in addition to, fewer than, or different from those described in method 300. Although method 300 is performed by a server executing machine-readable software code of the neural network architecture, various operations may be performed by one or more computing devices and / or processors. Embodiments may include operations in addition to, fewer than, or different from those described in method 300.

[0067] In step 302, the input layer of the neural network architecture takes in an input speech signal 301, which may be a training speech signal from a training phase or an enrollment speech signal from an enrollment phase. The input layer performs various pre-processing operations (e.g., training speech signal, enrollment speech signal) on the input speech signal 301 before feeding it to various other layers of the neural network architecture. Pre-processing operations may include, for example, applying VAD operations, extracting low-level spectro-temporal features, and performing data transformation operations.

[0068] In some embodiments, the input layer performs various data augmentation operations during the training or enrollment phase. The data augmentation operations may generate or obtain specific training speech signals, including clean speech signals and noise samples. The server may receive or request clean speech signals from one or more corpus databases. The clean speech signals may include speech signals originating from various types of speech sources with diverse speaker-independent characteristics. The clean speech signals may be stored in a non-transitory storage medium accessible to the server or received via a network or other data source. The data augmentation operations may also receive simulated speech signals from one or more databases or generate simulated speech signals based on the clean speech signal or the input speech signal 301 by applying various forms of data augmentation to the input speech signal 301 or the clean speech signal. Examples of data augmentation techniques are described in U.S. Patent Application No. 17 / 155,851, which is incorporated by reference in its entirety.

[0069] In step 304, the server applies a VAD to the input speech signal 301. A neural network layer of the VAD detects the occurrence of speech or non-speech windows in the input speech signal 301. The VAD parses the input speech signal 301 into speech portions 303 and non-speech portions 305. The VAD includes a classification layer (e.g., a classifier) ​​that the server trains separately or jointly with other layers of the neural network architecture. The VAD may directly output a binary result (e.g., speech, non-speech) for a portion of the input signal 301, or may generate an argument value (e.g., probability) for each portion of the input speech signal 301 that the server evaluates against a speech detection threshold to output a binary result for a given portion. From the input speech signal 301, the VAD generates a set of speech portions 303 and a set of non-speech portions 305 parsed by the VAD.

[0070] In step 306, the server extracts features from the set of speech portions 303, and in step 308, the server extracts corresponding features from the set of non-speech portions 305. The server performs various pre-processing operations on the portions 303, 305 of the input audio signal 301, for example to extract low-level features from the portions 303, 305 and convert such features from a time-domain representation to a frequency-domain representation by performing a short-time Fourier transform (SFT) and / or a fast Fourier transform (FFT).

[0071] While the server and neural network architecture of method 300 are configured to extract features related to speaker-independent characteristics of the input speech signal 301, in some embodiments the server is further configured to extract features related to utterance-dependent speaker-dependent characteristics.

[0072] In step 310, the server applies the task-specific model to the speech portion 303 to separately train a specific neural network for the speech portion 303. The server receives the input speech signal 301 (e.g., training speech signal, enrollment speech signal) along with labels. The labels indicate certain expected speaker-independent characteristics of the input speech signal 301, such as expected classification, expected features, expected feature vectors, expected metadata, and the type or degree of degradation present in the input speech signal 301, among other speaker-independent characteristics.

[0073] Each task-specific model includes one or more embedding and extraction layers to model specific characteristics of the input speech signal 301. The embedding and extraction layers generate feature vectors, or embeddings, based on features extracted from the speech portion 303. During the training phase (and possibly during the enrollment phase), the classifier layer of the task-specific model determines a predicted classification of the speech portion 303 based on the feature vectors. The server executes a loss function to determine the level of error based on the difference between the predicted speech output (e.g., predicted feature vector, predicted classification) and the expected speech output (e.g., expected feature vector, expected classification) according to the label associated with the particular input speech signal 301. The loss function or other computational layer of the neural network architecture adjusts the hyperparameters of the task-specific model until the level of error meets a threshold error degree.

[0074] During the enrollment phase, the task-specific models output feature vectors and / or classifications. A neural network architecture statistically or algorithmically combines the feature vectors generated for the speech portions of the enrollment speech signal to extract enrolled speaker-independent embeddings of the speech portions.

[0075] In step 312, the server similarly applies the task-specific model for the non-speech portion 305 to the specific neural network for the non-speech portion 303. The embedding extraction layer generates a feature vector or embedding based on features extracted from the non-speech portion 305. During the training phase (and possibly during the enrollment phase), the classifier layer of the task-specific model determines a predicted classification for the non-speech portion 305 based on the feature vector. The server executes a loss function to determine an error level based on the difference between the predicted non-speech output (e.g., predicted feature vector, predicted classification) and the expected non-speech output (e.g., expected feature vector, expected classification) according to the labels associated with the particular input audio signal 301. The loss function or other computational layer of the neural network architecture adjusts the hyperparameters of the task-specific model until the error level meets a threshold error metric.

[0076] During the enrollment phase, the task-specific models output feature vectors and / or classifications. A neural network architecture statistically or algorithmically combines the feature vectors generated for the non-speech portions of the enrollment speech signal to extract enrolled speaker-independent embeddings of the non-speech portions.

[0077] In step 314, the server extracts features from the entire input audio signal 301. The server performs various pre-processing operations on the input audio signal 301, for example, to extract features from the input audio signal 301 and convert one or more extracted features from a time domain representation to a frequency domain representation by an SFT or FFT operation.

[0078] In step 316, the server applies a task-specific model (e.g., an audio event neural network) to the entire input audio signal 301. The embedding extraction layer generates a feature vector based on features extracted from the input audio signal 301. During the training phase (and possibly during the enrollment phase), the classifier layer of the task-specific model determines a predicted classification of the input audio signal 301 based on the feature vector. The server executes a loss function to determine the level of error based on the difference between the predicted output (e.g., predicted feature vector, predicted classification) and the expected output (e.g., expected feature vector, expected classification) according to the label associated with the particular input audio signal 301. The loss function or other computational layer of the neural network architecture adjusts the hyperparameters of the task-specific model (e.g., an audio event neural network) until the level of error meets a threshold error metric.

[0079] During the enrollment phase, the task-specific models of the current step 316 output feature vectors and / or classifications. The neural network architecture statistically or algorithmically combines the feature vectors generated for all (or substantially all) of the enrollment speech signals to extract speaker-independent embeddings based on the feature vectors generated for the enrollment speech signals.

[0080] The neural network architecture further extracts DP vectors for the enrolled speech sources. For each task-specific model that evaluates the speech portion 303 separately from the non-speech portion 305, the neural network architecture extracts a pair of speaker-independent embeddings: a speech embedding and a non-speech embedding. Furthermore, the neural network architecture extracts a single speaker-independent embedding for each task-specific model that evaluates the entire input speech signal 301. One or more post-modeling layers of the neural network architecture concatenate or algorithmically combine the speaker-independent embeddings to extract enrolled DP embeddings for the enrolled speech sources.

[0081] Training and Enrolling Multitask Configurations 4 illustrates the steps performed in a multi-task learning method 400 for training a neural network architecture for speaker-independent embedding. Embodiments may include operations in addition to, fewer than, or different from those described in method 400. Method 400 is performed by a server executing machine-readable software code of the neural network architecture, although various operations may be performed by one or more computing devices and / or processors.

[0082] In the multi-task learning method 400, the server trains only two task-specific models (e.g., speech, non-speech) or three task-specific models (e.g., speech, non-speech, and the entire audio signal) for multi-task learning, rather than separately training and developing task-specific models (as in the signal-task configuration of method 300 of FIG. 3 ). In the multi-task learning method 400, the neural network architecture includes a shared hidden layer (e.g., an embedding extraction layer) shared by (and common with) the multiple task-specific models, such that the hidden layer generates a feature vector for a given portion 403, 405 of the audio signal. In addition, the input audio signal 401 (or its speech portion 403 and non-speech portion 405) is shared by the task-specific models. The neural network architecture includes a single final loss function that is a weighted sum of the single-task losses. The shared hidden layer is followed by separate task-specific fully connected (FC) layers and task-specific output layers. In the multi-task learning method 400, the server runs a single loss function across all task-specific models.

[0083] In step 402, an input layer of the neural network architecture takes in an input audio signal 401, which may be a training audio signal from a training phase or an enrollment audio signal from an enrollment phase. The input layer performs various pre-processing operations (e.g., training audio signal, enrollment audio signal) on the input audio signal 401 before feeding it to various other layers of the neural network architecture. Pre-processing operations include, for example, applying VAD operations, extracting various types of features (e.g., MFCCs, metadata), and performing data transformation operations.

[0084] In some embodiments, the input layer performs various data augmentation operations during the training or enrollment phase. The data augmentation operations may generate or obtain specific training speech signals, including clean speech signals and noise samples. The server may receive or request clean speech signals from one or more corpus databases. The clean speech signals may include speech signals from various types of speech sources with diverse speaker-independent characteristics. The clean speech signals may be stored in a non-transitory storage medium accessible to the server or received via a network or other data source. The data augmentation operations may also receive simulated speech signals from one or more databases or generate simulated speech signals based on the clean speech signal or the input speech signal 401 by applying various forms of data augmentation to the input speech signal 401 or the clean speech signal. Examples of data augmentation techniques are described in U.S. Patent Application No. 17 / 155,851, which is incorporated by reference in its entirety.

[0085] In step 404, the server applies a VAD to the input speech signal 401. The VAD's neural network layer detects the occurrence of speech and non-speech windows in the input speech signal 401. The VAD parses the input speech signal 401 into speech portions 403 and non-speech portions 405. The VAD includes a classification layer (e.g., a classifier) ​​that the server trains separately or jointly with other layers of the neural network architecture. The VAD may directly output a binary result (e.g., speech, non-speech) for a portion of the input signal 401, or may generate an argument value (e.g., probability) for each portion of the input speech signal 401 that the server evaluates against a speech detection threshold to output a binary result for a given portion. From the input speech signal 401, the VAD generates a set of speech portions 403 and a set of non-speech portions 405 parsed by the VAD.

[0086] In step 406, the server extracts features from the set of speech portions 403, and in step 408, the server extracts corresponding features from the set of non-speech portions 405. The server performs various pre-processing operations on the portions 403, 405 of the input audio signal 401, for example to extract various types of features from the portions 403, 405 and to convert one or more extracted features from a time domain representation to a frequency domain representation by performing an SFT or FFT operation.

[0087] While the server and neural network architecture of method 400 are configured to extract features related to speaker-independent characteristics of the input speech signal 401, in some embodiments the server is further configured to extract features related to utterance-dependent speaker-dependent characteristics.

[0088] In step 410, the server applies task-specific models to the utterance portion 403 to jointly train a neural network architecture for the utterance portion 403. The task-specific models share hidden layers for modeling and generating a joint feature vector based on the utterance portion of a particular input audio signal 401. Each task-specific model includes a fully connected layer and an output layer that are distinct from the fully connected layers and output layers of other task-specific models and independently influence a particular loss function. In the current step 410, the shared loss function evaluates the output associated with the utterance portion 403. For example, the loss function can be a sum, concatenation, or other algorithmic combination of several output layers processing the utterance portion 403.

[0089] In particular, the server receives an input speech signal 401 (e.g., a training speech signal, an enrollment speech signal) along with one or more labels that indicate particular aspects of the particular input speech signal 401, such as, among other aspects, an expected classification, expected features, expected feature vectors, expected metadata, and the type or degree of impairment present in the input speech signal 401. A shared hidden layer (e.g., an embedding extraction layer) generates a feature vector based on features extracted from the speech portion 403. During the training phase (and, optionally, during the enrollment phase), the fully connected layer and output layer (e.g., a classifier layer) of each task-specific model determine a predicted classification of the speech portion 403 based on the common feature vector generated by the shared hidden layer for the speech portion 403. The server executes a common loss function to determine a level of error, for example, according to the difference between one or more predicted speech outputs (e.g., predicted feature vectors, predicted classifications) and one or more expected speech outputs (e.g., expected feature vectors, expected classifications) indicated by the labels associated with the particular input speech signal 401. A shared loss function or other computational layer of the neural network architecture adjusts one or more hyperparameters of the neural network architecture until the level of error meets a threshold error measure.

[0090] Similarly, in step 412, the server applies a shared hidden layer (e.g., an embedding extraction layer) and a task-specific model to the non-speech portion 405. The embedding extraction layer generates a non-speech feature vector based on features extracted from the non-speech portion 405. During the training phase (and possibly during the enrollment phase), the fully connected layer and output layer (e.g., a classification layer) of each task-specific model generate a predicted non-speech output (e.g., a predicted feature vector, a predicted classification) for the non-speech portion 405 based on the non-speech feature vector generated by the shared hidden layer. The server executes a shared loss function to determine an error level based on the difference between the predicted non-speech output (e.g., a predicted feature vector, a predicted classification) and the expected non-speech output (e.g., an expected feature vector, an expected classification) according to the label associated with the particular input speech signal 401. The shared loss function or other computational layer of the neural network architecture adjusts one or more hyperparameters of the neural network architecture until the error level meets a threshold error metric.

[0091] In step 414, the server extracts features from the entire input audio signal 401. The server performs various pre-processing operations on the input audio signal 401, for example, to extract different types of features and convert certain extracted features from a time domain representation to a frequency domain representation by an SFT or FFT function.

[0092] In step 416, the server applies a hidden layer (e.g., an embedding extraction layer) to the features extracted (in step 414). The hidden layer is shared by task-specific models that evaluate the entire input audio signal 401 (e.g., an audio event neural network). The embedding extraction layer generates a full audio feature vector based on the features extracted from the input audio signal 401. During the training phase (and possibly during the enrollment phase), a fully connected layer and an output layer (e.g., a classification layer) specific to each task-specific model generate a predicted full audio output (e.g., a predicted classification, a predicted full audio feature vector) according to the full audio feature vector. The server executes a shared loss function to determine an error level based on the difference between the predicted full audio output (e.g., a predicted feature vector, a predicted classification) and the expected full audio output (e.g., an expected feature vector, an expected classification) indicated by the label associated with the particular input audio signal 401. The shared loss function or other computational layer of the neural network architecture adjusts one or more hyperparameters of the neural network architecture until the error level meets a threshold error level.

[0093] Speaker-independent embedding and extraction of DP vectors 5 illustrates the execution steps of a method 500 for applying a task-specific model of a neural network architecture and extracting a speaker-independent embedding and DP vector of an input speech signal. Embodiments may include operations in addition to, fewer than, or different from those described in method 500. Method 500 is performed by a server executing machine-readable software code of the neural network architecture, although various operations may be performed by one or more computing devices and / or processors.

[0094] In step 502, the server receives an input speech signal (e.g., a training speech signal, an enrollment speech signal, an inbound speech signal) from a particular speech source. The server may perform any number of pre-processing and / or data augmentation operations on the input speech signal before feeding it into the neural network architecture.

[0095] In step 504, the server optionally applies a VAD to the input speech signal. A neural network layer of the VAD detects the occurrence of speech or non-speech windows in the input speech signal. The VAD parses the input speech signal into speech and non-speech portions. The VAD includes a classification layer (e.g., a classifier) ​​that the server trains separately or jointly with other layers of the neural network architecture. The VAD may directly output a binary result (e.g., speech, non-speech) for a portion of the input signal, or may generate an argument value (e.g., probability) for each portion of the input speech signal that the server evaluates against a speech detection threshold to output a binary result for a given portion. The VAD generates a set of speech portions and a set of non-speech portions from the input speech signal that are parsed by the VAD.

[0096] In step 506, the server extracts various features from the input audio signal. The extracted features may include low-level spectrotemporal features (e.g., MFCCs) and communication metadata, among other types of data. These features are associated with a set of speech portions, a set of non-speech portions, and the entire audio signal.

[0097] In step 508, the server applies a neural network architecture to the input speech signal. In particular, the server applies the task-specific models to the features extracted from the speech portions to generate a first set of speaker-independent feature vectors corresponding to each of the task-specific models. The server again applies the task-specific models to the features extracted from the non-speech portions to generate a second set of speaker-independent feature vectors corresponding to each of the task-specific models. The server also applies an entire speech task-specific model (e.g., a speech event neural network) to the entire input speech signal.

[0098] In a single-task configuration, each task-specific model includes a separate modeling layer for generating feature vectors. In a multi-task configuration, a shared hidden layer generates feature vectors and a shared loss layer tunes hyperparameters, but each task-specific model includes a separate set of post-modeling layers, such as a fully connected layer, a classification layer, and / or an output layer. In either configuration, the fully connected layer of each particular task-specific model statistically or algorithmically combines feature vectors to extract one or more corresponding speaker-independent embeddings. For example, a speech event neural network extracts a full speech embedding based on one or more full speech feature vectors, while a device neural network extracts an utterance embedding and then extracts a non-speech feature vector based on one or more utterance feature vectors and one or more non-speech feature vectors.

[0099] Then, after extracting speaker-independent embeddings from the input speech signal, the server extracts DP vectors for the speech source based on the extracted speaker-independent embeddings. In particular, the server algorithmically combines the extracted embeddings from each of the task-specific models to generate the DP vectors.

[0100] As shown in FIG. 5, the neural network architecture includes nine task-specific models that extract speaker-independent embeddings. The server applies the speech event neural network to the entire input speech signal to extract corresponding speaker-independent embeddings. The server also applies each of the remaining eight task-specific models to the speech portion and the non-speech portion. The server extracts the following pair of speaker-independent embeddings output by each task-specific model: a first speaker-independent embedding for the speech portion and a second speaker-independent embedding for the non-speech portion. In this example, the server extracts eight of the first speaker-independent embeddings and eight of the second speaker-independent embeddings. The server then extracts a DP vector for the input speech signal by concatenating the speaker-independent embeddings. In the example of FIG. 5, the DP vector is generated by concatenating 17 speaker-independent embeddings output by the nine task-specific models.

[0101] Additional Embodiments Speaker-independent DP vectors complementary to voice biometrics Deep phone printing systems can be used for authentication computation of audio features and / or metadata associated with audio signals. DP vectors or speaker-independent embeddings can be evaluated as a complement to voice biometrics.

[0102] In a voice biometric authentication system, during the training phase, an input audio signal is preprocessed using a VAD to parse the voice-only portion of the audio. A discriminative DNN model is trained on the voice-only portion of the input audio signal using one or more speaker labels. During feature extraction, the audio signal is passed through a VAD to extract one or more features of the voice-only portion of the audio. The audio is taken as input to the DNN model to extract a speaker embedding vector (e.g., a voiceprint) representing speaker characteristics of the input audio signal.

[0103] In a DP system, during the training phase, audio is processed using a VAD to parse both the voice-only and non-voiced portions of the input audio signal. The voice-only portions of the audio are used to train a set of DNN models using a set of metadata labels. The non-speech portions of the audio are then used to train another set of DNN models using the same set of metadata labels. In some implementations, the metadata labels are unrelated to speaker labels. Because of these differences, the DP vectors generated by the DP system add complementary information about speaker-independent aspects of audio to a voice biometric authentication system. The DP vectors capture information complementary to features extracted in a voice biometric authentication system, allowing a neural network architecture to use speaker embeddings (e.g., voiceprints) and the DP vectors in various voice biometric authentication operations, such as authentication. The fusion system described herein can be used for various downstream operations, such as authenticating the audio source associated with the input audio signal or creating and implementing a voice-based exclusion list.

[0104] FIG. 6 illustrates the data flow between layers of a neural network architecture 600 that uses DP vectors as a complement to speaker-dependent embeddings. DP vectors can be fused with speaker-dependent embeddings at either the “feature level” or the “score level.” In a feature-level fusion embodiment, the DP vectors are concatenated or otherwise algorithmically combined with speaker-dependent embeddings (e.g., voiceprints), thereby generating a joint embedding. The concatenated joint feature vector is then used to train a machine-learning classifier for a classification task and / or for use in any number of downstream operations. While neural network architecture 600 is executed by a server during the training phase and optional enrollment and deployment phases, neural network architecture 600 may be executed by any computing device that includes a processor capable of performing the operations of neural network architecture 600, and by any number of such computing devices.

[0105] The speech intake layer 602 receives one or more input speech signals (e.g., a training speech signal, an enrollment speech signal, an inbound speech signal) and performs various preprocessing operations, including applying a VAD to the input speech signals and extracting various types of features. The VAD detects speech and non-speech portions and outputs a speech-only speech signal and a non-speech-only speech signal. The intake layer 602 may also extract various types of features from the input speech signal, the speech-only speech signal, and / or the non-speech-only speech signal.

[0106] A speaker-dependent embedding layer 604 is applied to the input speech signal to extract a speaker-dependent embedding (e.g., a voiceprint). The speaker-dependent embedding layer 604 includes an embedding extraction layer for generating a feature vector according to the extracted features used to model speaker-dependent aspects of the input speech signal. Examples of extracting such speaker-dependent embeddings are described in U.S. Patent Application Nos. 17 / 066,210 and 17 / 079,082, which are incorporated herein by reference.

[0107] Sequentially or simultaneously, a DP vector layer 606 is applied to the input speech signal. The DP vector layer includes task-specific models that extract various speaker-independent embeddings from the input speech signal. The DP vector layer 606 then extracts a DP vector for the input speech signal by concatenating (or otherwise combining) the various speaker-independent embeddings.

[0108] The neural network layers defining the joint classifier 608 are applied to the joint embedding to output a classification or other output indicative of the classification. The DP vector and speaker-dependent embedding are concatenated or otherwise algorithmically combined to generate the joint embedding. The concatenated joint embedding is then used to train the joint classifier for classification tasks or any number of downstream operations that use classification decisions. During the training phase, the joint classifier outputs predicted outputs (e.g., predicted classifications, predicted feature vectors) for the training speech signals, which are used to determine the level of error according to training labels indicative of the expected output. The hyperparameters of the joint classifier 608 and other layers of the neural network architecture 600 are adjusted to minimize the level of error.

[0109] 7 illustrates the data flow between layers of a neural network architecture 700 that uses DP vectors as complements to speaker-dependent embeddings. Although neural network architecture 700 is executed by a server during the training phase and optional enrollment and deployment phases, neural network architecture 700 may be executed by any computing device that includes a processor capable of performing the operations of neural network architecture 700, and by any number of such computing devices.

[0110] In a score-level fusion embodiment, such as in Figure 7, one or more DP vectors are used to train a speaker-independent classifier 708, and speaker-dependent embeddings are used to train a speaker-dependent classifier 710. The two types of classifiers 708, 710 independently generate predicted classifications (e.g., predicted classification labels, predicted classification probabilities) for corresponding types of embeddings. A classification score fusion layer 712 algorithmically combines the predicted labels or predicted probabilities, for example, using an ensemble algorithm (e.g., a logistic regression model), to output a joint prediction.

[0111] The speech intake layer 702 receives one or more input speech signals (e.g., a training speech signal, an enrollment speech signal, an inbound speech signal) and performs various preprocessing operations, including applying a VAD to the input speech signals and extracting various types of features. The VAD detects speech and non-speech portions and outputs a speech-only speech signal and a no-speech-only speech signal. The intake layer 702 may also extract various types of features from the input speech signal, the speech-only speech signal, and / or the no-speech-only speech signal.

[0112] A speaker-dependent embedding layer 704 is applied to the input speech signal to extract a speaker-dependent embedding (e.g., a voiceprint). The speaker-dependent embedding layer 704 includes an embedding extraction layer for generating a feature vector according to the extracted features used to model speaker-dependent aspects of the input speech signal. Examples of extracting such speaker-dependent embeddings are described in U.S. Application Nos. 17 / 066,210 and 17 / 079,082, which are incorporated herein by reference.

[0113] Sequentially or simultaneously, a DP vector layer 706 is applied to the input speech signal. The DP vector layer includes task-specific models that extract various speaker-independent embeddings from the input speech signal. The DP vector layer 706 then extracts a DP vector for the input speech signal by concatenating (or otherwise combining) the various speaker-independent embeddings.

[0114] The neural network layers defining the speaker-dependent classifier 708 are applied to the speaker-dependent embeddings to output a speaker classification (e.g., genuine, dishonest) or other output indicative of the speaker classification. During the training phase, the speaker-dependent classifier 708 generates predicted outputs (e.g., predicted classifications, predicted feature vectors, predicted DP vectors) for the training speech signals and uses these to determine the level of error according to training labels indicative of the expected output. The hyperparameters of the speaker-dependent classifier 708 or other layers of the neural network architecture 700 are adjusted by a server loss function or other function to minimize the level of error of the speaker-dependent classifier 708.

[0115] The neural network layers defining the speaker-independent classifier 710 are applied to the DP vectors to output a speaker classification (e.g., genuine, dishonest) or other output indicative of the speaker classification. During the training phase, the speaker-independent classifier 710 generates predicted outputs (e.g., predicted classifications, predicted feature vectors, predicted DP vectors) for the training speech signals and uses these to determine the level of error according to training labels indicative of the expected output. A loss function or other function executed by the server adjusts hyperparameters of the speaker-independent classifier 710 or other layers of the neural network architecture 700 to minimize the level of error of the speaker-independent classifier 710.

[0116] The server performs a classification score fusion operation 712 based on the output classification scores or decisions generated by the speaker-dependent classifier 708 and the speaker-independent classifier 710. In particular, each classifier 708, 710 independently generates a respective classification output (e.g., classification label, classification probability value). The classification score fusion operation 712 algorithmically combines the predicted outputs, for example, using an ensemble algorithm (e.g., a logistic regression model), to output a joint embedding.

[0117] Speaker-independent DP vectors of the exclusion list 8 illustrates the data flow between layers of a neural network architecture 800 that uses DP vectors to authenticate audio sources according to an exclusion list and / or an allow list. The exclusion list may operate as a rejection list (sometimes called a "blacklist") and / or an allow list (sometimes called a "whitelist"). In an authentication operation, the exclusion list is used to exclude and / or allow particular audio sources. While neural network architecture 800 is executed by a server during a training phase and optional enrollment and deployment phases, neural network architecture 800 may be executed by any computing device that includes a processor capable of performing the operations of neural network architecture 800, and by any number of such computing devices.

[0118] In particular, DP vectors may be referenced to create and / or implement the exclusion list. DP vectors are extracted from input speech signals (e.g., training speech signals, enrollment speech signals, inbound speech signals) for various computation phases. DP vectors are associated with speech sources that are part of the exclusion list and with speech associated with a normal population. A machine learning model is trained on the DP vectors using the labels. At test time, the model predicts new sample speech to determine whether the speech belongs to the exclusion list.

[0119] The speech intake layer 802 receives one or more input speech signals (e.g., a training speech signal, an enrollment speech signal, an inbound speech signal) and performs various preprocessing operations, including applying a VAD to the input speech signals and extracting various types of features. The VAD detects speech and non-speech portions and outputs a speech-only speech signal and a non-speech-only speech signal. The intake layer 802 may also extract various types of features from the input speech signal, the speech-only speech signal, and / or the non-speech-only speech signal.

[0120] A DP vector layer 804 is applied to the input speech signal. The DP vector layer includes task-specific models that extract various speaker-independent embeddings from the input speech signal. The DP vector layer 804 then extracts a DP vector for the input speech signal by concatenating (or otherwise combining) the various speaker-independent embeddings.

[0121] The exclusion list modeling layer 806 is applied to the DP vector. The exclusion list modeling layer 806 is trained to determine whether the audio source that emitted the input audio signal is on the exclusion list. The exclusion list modeling layer 806 determines a similarity score based on the relative distance (e.g., similarity, difference) between the DP vector of the input audio signal and the DP vector of the audio source in the exclusion list. During the training phase, the exclusion list modeling layer 806 generates predicted outputs (e.g., predicted classifications, predicted similarity scores) for the training audio signals and uses these to determine the level of error according to the training labels that indicate the expected output. The hyperparameters of the exclusion list modeling layer 806 or other layers of the neural network architecture 800 minimize the level of error of the exclusion list modeling layer 806.

[0122] As an example, the neural network architecture 800 uses a fraudulent exclusion list that is applied to an inbound audio source. n} are the corresponding training labels, Y={y1,y2,y3,...y n}, where each label (y i ) is the corresponding training speech signal (x i ) classification {illegible, genuine}. DP vector (v i ) is the input audio signal (x i ) and the corresponding DP vector, V={v1,v2,v3,...v n}. The modeling layer 806 is trained on the DP vectors using the labels. At test time, the trained modeling layer 806 determines a classification score based on a similarity score or likelihood score that the inbound speech signal is within a threshold distance to the DP vector of the inbound speech signal.

[0123] DP Vector for Authentication 9 is a diagram illustrating data flow between layers of a neural network architecture 900 that uses DP vectors to authenticate a device using a device identifier. Although neural network architecture 900 is executed by a server during a training phase and optional enrollment and deployment phases, neural network architecture 900 may be executed by any computing device that includes a processor capable of performing the operations of neural network architecture 900, and by any number of such computing devices.

[0124] An authentication embodiment may register an audio source entity with a device ID or source ID, and the DP vector used to register the audio source is stored in a database as a stored DP vector. Upon testing, each audio source associated with a device ID should be authenticated (or rejected) based on an algorithmic comparison between its inbound DP vector and the stored DP vector associated with the device ID of the inbound audio device.

[0125] The speech intake layer 902 receives one or more input speech signals (e.g., a training speech signal, an enrollment speech signal, an inbound speech signal) and performs various preprocessing operations, including applying a VAD to the input speech signals and extracting various types of features. The VAD detects speech and non-speech portions and outputs a speech-only speech signal and a non-speech-only speech signal. The intake layer 902 may also extract various types of features from the input speech signal, the speech-only speech signal, and / or the non-speech-only speech signal.

[0126] The audio intake layer 902 includes identifying the input audio signal and a device identifier (Device ID 903) associated with the audio source. The Device ID 903 can be any identifier associated with the originating device or audio source. For example, if the originating device is a phone, the Device ID 903 can be an ANI or a phone number. If the originating device is an IoT device, the Device ID 903 can be a MAC address, IP address, computer name, etc.

[0127] The audio intake layer 902 also performs various preprocessing operations (e.g., feature extraction, VAD operations) on the input audio signal to output preprocessed signal data 905. The preprocessed signal data 905 includes, for example, a speech-only signal, a no-speech-only signal, and various types of extracted features.

[0128] A DP vector layer 904 is applied to the input speech signal. The DP vector layer includes preprocessed signal data 905 and a task-specific model that extracts various speaker-independent embeddings from substantially all of the input speech signal. The DP vector layer 904 then extracts DP vectors for the input speech signal by concatenating (or otherwise combining) the various speaker-independent embeddings.

[0129] Additionally, the server uses the device ID 903 to query a database containing stored DP vectors 906. The server identifies one or more stored DP vectors 906 in the database associated with the device ID 903.

[0130] The classifier layer 908 determines a similarity score based on the relative distance (e.g., similarity, difference) between the DP vector extracted for the inbound voice signal and the stored DP vector 906. In response to determining that the similarity score meets the authentication threshold score, the classifier layer 908 authenticates the calling device as a genuine registered device. Any type of metric may be used to calculate the similarity between the stored vector and the inbound DP vector. Non-limiting examples of potential similarity calculations may include probabilistic linear discriminant analysis (PLDA) similarity, the inverse of Euclidean distance, or the inverse of Manhattan distance, among others.

[0131] The server performs one or more authentication output operations 910 based on the similarity scores generated by the classifier layer 908. The server may, for example, in response to authenticating the calling device, connect the calling device to a destination device (e.g., a provider server) or an agent phone.

[0132] DP Vectors for Speech Quality Measurement A DP vector can be created by concatenating embeddings of different speaker-independent aspects of a speech signal. Speaker-independent aspects such as the microphone type used to capture the speech signal, the codec used to compress and decompress the speech signal, and speech events (e.g., background noise present in the speech) provide useful information about the quality of the speech. The DP vector can be used to represent speech quality measurements based on speaker-independent aspects, which can be used in various downstream operations, such as authentication systems.

[0133] For example, in authentication systems (e.g., voice biometric authentication systems), it is important to enroll speaker embeddings extracted from good quality speech. If the speech signal contains a high level of noise, the quality of the speaker embedding may be adversely affected. Similarly, speech-related artifacts generated by the device and microphone may degrade the quality of the speech for enrollment embedding. Furthermore, high-energy speech events (e.g., overwhelming noise) may be present in the input speech signal, such as a crying baby, a barking dog, or music. These high-energy speech events are evaluated as voice parts by the VAD applied to the speech signal. This adversely affects the quality of the speaker embedding extracted from the speech signal for enrollment purposes. If a poor-quality speaker embedding is enrolled and stored in the authentication system, the enrolled speaker embedding may adversely affect the performance of the entire authentication system. DP vectors can be used to characterize the quality of speech either using rule-based thresholding techniques or using machine learning models as binary classifiers.

[0134] 10 illustrates the data flow between layers of a neural network architecture 1000 that uses deep phone printing for dynamic registration, where a server determines whether to register based on voice quality. Although neural network 1000 is executed by a server during a training phase and optional registration and deployment phases, neural network architecture 1000 may be executed by any computing device with a processor capable of performing the operations of neural network architecture 1000, and by any number of such computing devices.

[0135] The audio intake layer 1002 receives one or more enrollment audio signals and performs various preprocessing operations, including applying a VAD to the input audio signals and extracting various types of features. The VAD detects speech and non-speech portions and outputs a speech-only audio signal and a no-speech-only audio signal. The intake layer 1002 may also extract various types of features from the v-speech signal, the speech-only audio signal, and / or the no-speech-only audio signal.

[0136] A DP vector layer 1006 is applied to the enrollment speech signal. The DP vector layer includes a task-specific model that extracts various speaker-independent embeddings from the enrollment speech signal. The DP vector layer 1006 then extracts a DP vector for the enrollment speech signal by concatenating (or otherwise combining) the various speaker-independent embeddings. The DP vector represents the quality of the enrollment speech signal. The server applies a layer of speech quality classifier, which generates a quality score and makes a binary classification decision of whether to use the particular enrollment speech signal to generate an enrolled speaker embedding based on an enrollment quality threshold.

[0137] In response to determining that the quality meets the enrollment quality threshold, the server performs various enrollment operations 1004, such as generating, storing, or updating enrollment speaker embeddings.

[0138] DP Vectors for Replay Attack Detection 11 illustrates a data flow diagram for a neural network architecture 1100 for detecting replay attacks using DP vectors. While neural network architecture 1100 is executed by a server during a training phase and optional enrollment and deployment phases, neural network 1100 may be executed by any computing device, and by any number of such computing devices, that includes a processor capable of performing the operations of neural network architecture 1100. Neural network architecture 1100 need not always perform the operations of the enrollment phase. Thus, in some embodiments, neural network architecture 1100 includes a training phase and a deployment phase.

[0139] Automatic Speaker Verification (ASV) systems are commonly used as a reliable solution for person authentication. However, ASV systems are vulnerable to voice spoofing. Voice spoofing can be either local access (LA) or physical access (PA). Text-to-speech (TTS) functions that generate synthetic speech and voice conversions are under the protection of LA. A replay attack is an example of PA spoofing. In a PA scenario, speech data is assumed to be captured by a microphone in a physical, reverberant space. A replace spoofing attack is a recording of authentic speech that is assumed to be captured and then re-presented to the microphone of the ASV system using a replay device. The replayed speech is assumed to be first captured by a recording device before being replayed using a nonlinear replay device.

[0140] The speech intake layer 1102 receives one or more input speech signals (e.g., training speech signals, enrollment speech signals, inbound speech signals) and performs various preprocessing operations, including applying a VAD to the input speech signals and extracting various types of features. The VAD detects speech and non-speech portions and outputs speech-only and non-speech-only speech signals. The intake layer 1102 may also extract various types of features from the input speech signals, speech-only and / or non-speech-only speech signals.

[0141] A DP vector layer 1104 is applied to the input speech signal. The DP vector layer includes task-specific models that extract various speaker-independent embeddings from the input speech signal. The DP vector layer 1104 then extracts a DP vector for the input speech signal by concatenating (or otherwise combining) the various speaker-independent embeddings.

[0142] The device type and microphone type are important speaker-independent aspects used by the deep printing system to detect replay attacks using DP vectors. Therefore, during the training phase, DP vectors are extracted from one or more corpora of training speech signals that are also used for training speaker embedding, including simulated speech signals with various forms of impairments (e.g., additive noise, reverberation). The DP vector layer 1104 is applied to the training speech signals, including the simulated speech signals.

[0143] The replay attack detection classifier 1106 determines a detection score that indicates the likelihood that the inbound audio signal is a replay attack. The binary classification model of the replay attack detection classifier 1106 is trained on the training data and training labels. During testing, the replay attack detection classifier 1106 model outputs an indication of whether the neural network architecture 1100 detected a replay attack or a genuine audio signal in the inbound audio.

[0144] DP Vectors for Spoof Detection In recent years, voice over IP (VoIP) services have risen due to their cost-effectiveness. VoIP services, such as Google Voice or Skype, use virtual phone numbers, also known as direct inward dialing (DID) or access numbers, which are phone numbers not directly associated with a telephone line. These phone numbers are also called gateway numbers. When a user makes a call through a VoIP service, the VoIP service may present a common gateway number or a number from a set of gateway numbers as a caller ID that is not necessarily unique to the device or caller. Some VoIP services allow loopholes, such as intentional caller ID / ANI spoofing, which allows callers to intentionally alter the information sent to the recipient's caller ID display to disguise their identity. There are several caller ID / ANI spoofing services available in the form of Android and iOS applications.

[0145] The deep phoneprinting system has a component for spoof service classification. Embeddings extracted from a DNN model trained for the ANI spoof service classification task can be used for spoof service recognition for new voices. These embeddings can also be used for spoof detection by considering all ANI spoof services as a spoofed class.

[0146] 12 illustrates the data flow between layers of a neural network architecture 1200 that uses DP vectors to identify spoofing services associated with an audio signal. While neural network architecture 1200 is executed by a server during a training phase and optional enrollment and deployment phases, neural network architecture 1200 may be executed by any computing device, and by any number of computing devices, that includes a processor capable of performing the operations of neural network architecture 1200. Neural network architecture 1200 need not always perform the operations of the enrollment phase. Thus, in some embodiments, neural network architecture 1200 includes a training phase and a deployment phase.

[0147] The speech intake layer 1202 receives one or more input speech signals (e.g., a training speech signal, an enrollment speech signal, an inbound speech signal) and performs various preprocessing operations, including applying a VAD to the input speech signals and extracting various types of features. The VAD detects speech and non-speech portions and outputs a speech-only speech signal and a non-speech-only speech signal. The intake layer 1202 may also extract various types of features from the input speech signal, the speech-only speech signal, and / or the non-speech-only speech signal.

[0148] A DP vector layer 1204 is applied to the input speech signal. The DP vector layer includes task-specific models that extract various speaker-independent embeddings from the input speech signal. The DP vector layer 1204 then extracts a DP vector for the input speech signal by concatenating (or otherwise combining) the various speaker-independent embeddings.

[0149] The spoof detection layer 1206 defines a binary classifier trained to determine a spoof detection likelihood score. The spoof detection layer 1206 is trained to determine whether an audio source is spoofing an input audio signal (spoof detection) or whether no spoofing is detected. The spoof detection layer 1206 determines the likelihood score based on the relative distance (e.g., similarity, difference) between the DP vector of the input audio signal and the DP vector of the trained spoof detection classification or stored spoof detection DP vector. During the training phase, the spoof detection layer 1206 generates predicted outputs (e.g., predicted classifications, predicted similarity scores) for the training audio signals and uses these to determine the level of error according to training labels indicating the expected output. A loss function or other server-implemented process adjusts the hyperparameters of the spoof detection layer 1206 or other layers of the neural network architecture 1200 to minimize the level of error of the spoof detection layer 1206. The spoof detection layer 1206 determines whether a given inbound voice signal, when tested, has a spoof detection likelihood score that meets the spoof detection threshold.

[0150] If the neural network architecture 1200 detects a spoof in the inbound voice signal, the neural network architecture applies a spoof service recognition layer 1208. The spoof service recognition layer 1208 is a multi-class classifier trained to determine the likely spoof service used to generate the spoofed inbound signal. The spoof service recognition layer 1208 determines a spoof service likelihood score based on the relative distance (e.g., similarity, difference) between the DP vector of the input voice signal and the DP vector of the trained spoof service recognition classification or a stored spoof service recognition DP vector. During the training phase, the spoof service recognition layer 1208 generates predicted outputs (e.g., predicted classifications, predicted similarity scores) for the training voice signal and uses these to determine the level of error according to training labels indicating the expected output. The hyperparameters of the spoof service awareness layer 1208 or other layers of the neural network architecture 1200 minimize the level of error of the spoof service awareness layer 1208. The spoof service awareness layer 1208 determines the specific spoof service to apply to generate a spoofed inbound voice signal at test time if the spoof service likelihood score meets the spoof detection threshold.

[0151] 13 illustrates the data flow between layers of a neural network architecture 1300 for identifying spoof services associated with an audio signal according to a label mismatch approach using DP vectors. While neural network architecture 1300 is executed by a server during a training phase and optional enrollment and deployment phases, neural network architecture 1300 may be executed by any computing device, and by any number of such computing devices, that includes a processor capable of performing the operations of neural network architecture 1300. Neural network architecture 1300 need not always perform the operations of the enrollment phase. Thus, in some embodiments, neural network architecture 1300 includes a training phase and a deployment phase.

[0152] The audio intake layer receives one or more input audio signals 1301 and performs various pre-processing operations, including applying a VAD to the input audio signals and extracting various types of features. The VAD detects speech and non-speech portions and outputs speech-only and non-speech-only audio signals. The intake layer may also extract various types of features from the input audio signals 1301, the speech-only audio signals, and / or the non-speech-only audio signals.

[0153] A DP vector layer 1302 is applied to the input speech signal 1301. The DP vector layer includes task-specific models that extract various speaker-independent embeddings from the input speech signal. The DP vector layer 1302 then extracts a DP vector for the input speech signal 1301 by concatenating (or otherwise combining) the various speaker-independent embeddings.

[0154] A deep phone printing system implementing the neural network architecture 1300 has components based on speaker-independent aspects such as carrier, network type, and geography, among others. The DP vectors can be used to train one or more label classification models 1304, 1306, and 1308 of speaker-independent embeddings for each of these speaker-independent aspects. During testing, the DP vectors extracted from an inbound speech signal are provided as inputs to the classification models 1304, 1306, and 1308 to predict various types of labels, such as carrier, network type, and geography, among others. For example, for an inbound speech signal received over a telephone channel, the telephone number or ANI of the originating speech source can be used to collect or determine specific metadata, such as carrier, network type, or geography, among others. The classification models 1304, 1306, and 1308 can detect discrepancies between the classification labels predicted based on the DP vectors and the metadata associated with the ANI. The output layers 1310 , 1312 , 1314 feed the corresponding classifications of the classification models 1304 , 1306 , 1308 as inputs to a spoof detection layer 1316 .

[0155] In some implementations, the classification models 1304, 1306, 1308 may be classifiers or other types of neural network layers (e.g., fully connected layers) of the respective task-specific models. Similarly, in some implementations, the output layers 1310, 1312, 1314 may be layers of the respective task-specific models.

[0156] The spoof detection layer 1316 defines a binary classifier trained to determine a spoof detection likelihood score based on the classification scores or labels received from a particular output layer 1310, 1312, 1314. The spoof detection layer 1316 is trained to determine whether an audio source is spoofing the input audio signal 1310 (spoof detection) or whether no spoof is detected. The spoof detection layer 1316 determines the likelihood score based on the relative distance (e.g., similarity, difference) between the DP vector 1302 of the input audio signal 1301 and the DP vector of the trained spoof detection classification, score, cluster, or stored spoof detection DP vector. During the training phase, the spoof detection layer 1316 generates predicted outputs (e.g., predicted classifications, predicted similarity scores) for the training audio signal and uses these to determine the level of error according to the training labels indicating the expected output. The loss function (or other function executed by the server) adjusts the hyperparameters of the spoof detection layer 1316 or other layers of the neural network architecture 1300 to minimize the level of error of the spoof detection layer 1316. The spoof detection layer 1316 determines whether a given inbound voice signal 1301, when tested, has a spoof detection likelihood score that meets the spoof detection threshold.

[0157] DP Vectors for Speaker-Independent Feature Classification. 14 illustrates data flow between layers of a neural network architecture 1400 that uses DP vectors to determine the microphone type associated with an audio signal. The neural network architecture 1400 is executed by one or more computing devices, such as a server, using various types of input audio signals (e.g., training, audio signals, enrollment audio signals, inbound audio signals) and various other types of data, such as training labels. The neural network architecture 1400 includes an audio intake layer, a DP vector layer 1404, and a classifier layer 1406.

[0158] The speech intake layer 1402 takes in an input speech signal and performs one or more preprocessing operations on the input speech signal. The DP vector layer 1404 extracts multiple speaker-independent embeddings for various types of speaker-independent features and extracts a DP vector for the inbound speech signal based on the speaker-independent embeddings. The classifier 1406 generates a classification score (or other type of score) based on the DP vector of the inbound signal and various types of stored data used for classification (e.g., stored embeddings, stored DP vectors), along with any number of additional outputs related to the classification results.

[0159] The DP vector can be used to recognize the type of microphone used to capture an input audio signal at a calling device. A microphone is a transducer that converts sound impulses into an electrical signal. There are several types of microphones used for various purposes, such as dynamic microphones, condenser microphones, piezoelectric microphones, microelectromechanical systems (MEMS) microphones, ribbon microphones, carbon microphones, etc. Microphones respond to sound differently based on direction. Microphones have different polar patterns, such as omnidirectional, bidirectional, and unidirectional, and have proximity effects. The distance from the sound source to the microphone and the direction of the sound source have a significant impact on the microphone's response. For telephone communication, while making a voice phone call, a user may use different types of microphones, such as a traditional landline receiver, a speakerphone, a mobile phone microphone, a wired headset with a microphone, a wireless microphone, etc. For IoT devices (e.g., Alexa®, Google Home®), microphone arrays with built-in direction detection, noise cancellation, and echo cancellation capabilities are used. The microphone array uses beamforming principles to improve the quality of the captured audio. The neural network architecture 1400 can recognize, or can be trained to recognize, microphone types by evaluating the DP vectors of the inbound audio signal against trained classifications, clusters, pre-stored embeddings, etc.

[0160] The DP vector can be used to recognize the type of device that captured the input audio signal. For telephony, there are several types of devices used. There are several telephony device manufacturers that produce different types of phones. For example, Apple produces the iPhone, which has several models such as iPhone 8, iPhone X, and iPhone 11, to name a few; Samsung produces the Samsung Galaxy series, etc.

[0161] For IoT, there is a wide variety of devices involved, starting from voice-based assistance devices such as Amazon Echo with Alexa, Google Home, smartphones, smart refrigerators, smart watches, smart fire alarms, smart door locks, smart bicycles, medical sensors, fitness trackers, smart security systems, etc. Recognizing the type of device from which audio is captured adds useful information about the audio.

[0162] The DP vector can be used to recognize the type of codec used to capture the input voice signal at the originating device. An audio codec is a device or computer program that can encode and decode an audio data stream or audio signal. A codec is used to compress the audio at one end before transmission or storage, and then decompress the audio at the receiving end. There are several types of codecs used in telephone audio, such as G.711, G.721, G729, Adaptive Multi-Rate (AMR), Enhanced Variable Rate Codec (EVRC) for CDMA networks, GSM used in GSM-based mobile phones, SILK used in Skype, WhatsApp, Speex used in VoIP apps, iLBC used in open source VoIP apps, etc.

[0163] The DP vector can be used to identify the carrier (e.g., AT&T, Verizon Wireless, T-Mobile, Sprint) associated with the transmission of the input voice signal and / or associated with the originating device. A carrier is a component of a telecommunications system that transmits information such as voice signals and / or protocol metadata. Carrier networks deliver large amounts of data over long distances. Carrier systems typically use various forms of multiplexing to simultaneously transmit multiple communication channels over a shared medium.

[0164] The DP vector can be used to recognize the geography associated with the originating device. The DP vector can be used to recognize the geographic location of the audio source. The geographic location can be a broad classification such as a specific continent, country, state / province, city / county, national or international.

[0165] The DP vector can be used to recognize the type of network associated with the transmission of the input voice signal and / or associated with the originating device. Generally, there are three classes of telephone networks: public switched telephone network (PSTN), cellular network, and voice over Internet Protocol (VoIP) network. PSTN is a traditional circuit-switched telecommunications system. Like PSTN systems, cellular networks have a circuit-switched core, parts of which are now replaced by IP links. These networks, for example, may deploy different technologies over the air interface, but the core technology and protocols of cellular networks are similar to PSTN networks. Finally, VoIP networks run on top of IP links and generally share the same paths as other Internet-based traffic.

[0166] The DP vectors can be used to recognize audio events in an input audio signal. Audio signals contain several types of audio events. These audio events carry information about everyday environments and physical events occurring within these environments. Recognizing such audio events and specific classes and detecting their precise locations in an audio stream has downstream benefits provided by a deep phone printing system, such as audio event-based multimedia searching, context-aware IoT devices (e.g., mobile, automotive), and intelligent monitoring in audio event-based security systems. For example, during the training phase, a training audio signal may be received along with labels or other metadata indicators indicating specific audio events in the training audio signal. The audio event indicators are used to train audio event task-specific machine learning models according to classifiers, fully connected layers, regression algorithms, and / or loss functions that evaluate the distance between the predicted output of the training audio signal and the expected output indicated by the labels.

[0167] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.

[0168] Computer software-implemented embodiments may be implemented in software, firmware, middleware, microcode, hardware description languages, or any combination thereof. A code segment or machine-executable instruction may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0169] The actual software code or specialized control hardware used to implement these systems and methods is not a limitation of the present invention. Accordingly, the operation and behavior of the systems and methods have been described without reference to specific software code, with the understanding that software and control hardware may be designed to implement the systems and methods based on the description herein.

[0170] If implemented in software, the functions may be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of a method or algorithm disclosed herein may be embodied in a processor-executable software module, which may reside on a computer-readable or processor-readable storage medium. Non-transitory computer-readable or processor-readable media includes both computer storage media and tangible storage media that facilitate transfer of a computer program from one place to another. Non-transitory processor-readable storage media may be any available medium that can be accessed by a computer. By way of example, and not limitation, such non-transitory processor-readable media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other tangible storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer or processor. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically while discs reproduce data optically with a laser. Combinations of the above are also included within the scope of computer-readable media. Additionally, the operations of a method or algorithm may reside as one or any combination or set of code and / or instructions on a non-transitory processor-readable medium and / or computer-readable medium, which may be incorporated into a computer program product.

[0171] The previous description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.

[0172] While various aspects and embodiments have been disclosed, other aspects and embodiments are contemplated. The various aspects and embodiments disclosed are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.

Claims

1. 1. A computer-implemented method comprising: applying, by a computer, a plurality of task-specific machine learning models to an inbound speech signal having one or more speaker-independent characteristics to extract a plurality of speaker-independent embeddings for the inbound speech signal; extracting, by the computer, a deep phonprint (DP) vector of the inbound speech signal based on the plurality of speaker-independent embeddings extracted for the inbound speech signal; applying, by the computer, one or more post-modeling operations to the plurality of speaker-independent embeddings extracted for the inbound speech signal to generate one or more post-modeling outputs for the inbound speech signal.

2. extracting, by the computer, a plurality of features from the training speech signal, the plurality of features including at least one of spectro-temporal features and metadata associated with the training speech signal; The method of claim 1 , further comprising: extracting, by the computer, the plurality of features from the inbound speech signal.

3. the metadata includes at least one of: a microphone type used to capture the training audio signal; a device type from which the training audio signal originated; a codec type applied to compress and decompress the training audio signal for transmission; a carrier from which the training audio signal originated; a spoofing service used to change a source identifier associated with the training audio signal; a geography associated with the training audio signal; a network type from which the training audio signal originated; and an audio event indicator associated with the training audio signal; 3. The method of claim 2, wherein the one or more classifications of the one or more post-modeling outputs include at least one of the microphone type, the device type, the codec type, the carrier, the spoofing service, the geography, the network type, and the audio event indicator.

4. The method of claim 1 , wherein a post-modeling operation of the one or more post-modeling operations comprises at least one of a classification operation and a regression operation.

5. The method of claim 1 , wherein the task-specific machine learning model comprises at least one of a neural network and a Gaussian mixture model.

6. 10. The method of claim 1, further comprising: training, by the computer, the plurality of task-specific machine learning models by applying each of the plurality of task-specific machine learning models to a plurality of training speech signals having the one or more speaker-independent characteristics.

7. Training task-specific machine learning models extracting, by the computer, one or more first spectrotemporal features of speech portions of the training speech signal and one or more second spectrotemporal features of non-speech portions of the training speech signal; applying, by the computer, one or more modeling layers of the task-specific machine learning model to the one or more first spectrotemporal features and the one or more second spectrotemporal features to generate first and second speaker-independent embeddings; generating, by the computer, predicted output data for the training speech signal by applying one or more post-modeling layers of the task-specific machine learning model to the first speaker-independent embedding and to the second speaker-independent embedding; and tuning, by the computer, one or more hyperparameters of the one or more modeling layers of the task-specific machine learning model by running a loss function using the predicted output data and expected output data indicated by one or more labels associated with the training audio signals.

8. Training task-specific machine learning models extracting, by the computer, one or more first spectrotemporal features of speech portions of the training speech signal and one or more second spectrotemporal features of non-speech portions of the training speech signal; applying, by the computer, one or more shared modeling layers of one or more task-specific machine learning models to the one or more first spectrotemporal features and the one or more second spectrotemporal features to generate first and second speaker-independent embeddings; generating, by the computer, predicted output data for the training speech signal by applying one or more post-modeling layers of the task-specific machine learning model to the first speaker-independent embedding and to the second speaker-independent embedding; and tuning, by the computer, one or more hyperparameters of the task-specific machine learning models by running a shared loss function of the one or more task-specific machine learning models using the predicted output data and expected output data indicated by one or more labels associated with the training audio signals.

9. training the plurality of task-specific machine learning models, extracting, by the computer, one or more spectrotemporal features of a substantial portion of a training speech signal; applying, by the computer, a task-specific machine learning model to the one or more spectrotemporal features of the training speech signal to generate a speaker-independent embedding; generating, by the computer, predicted output data for the training speech signal by applying one or more post-modeling layers of the task-specific machine learning model to the speaker-independent embedding of the substantial portion of the training speech signal; and tuning, by the computer, one or more hyperparameters of the task-specific machine learning model by running a loss function using the predicted output data and expected output data indicated by one or more labels associated with the training audio signals.

10. The method of claim 1 , wherein the task-specific machine learning model comprises at least one of a convolutional neural network, a recurrent neural network, and a fully connected neural network.

11. performing, by the computer, a voice activity detection (VAD) operation on a training speech signal to thereby generate one or more speech portions of the training speech signal and one or more non-speech portions of the training speech signal; 2. The method of claim 1, further comprising: performing, by the computer, the VAD operation on the inbound speech signal, thereby generating the one or more speech portions of the inbound speech signal and the one or more non-speech portions of the inbound speech signal.

12. extracting, by the computer, a plurality of enrolled speaker-independent embeddings and DP vectors for an enrolled speech source by applying the plurality of task-specific machine learning models to a plurality of enrollment speech signals associated with the enrolled speech source; extracting, by the computer, the DP vectors of the enrolled speech sources based on the plurality of enrolled speaker-independent embeddings; The method of claim 1 , wherein the computer uses the DP vectors of the registered speech sources and the DP vectors of the inbound speech signal to generate the one or more post-modeling outputs.

13. 13. The method of claim 12, further comprising determining, by the computer, an authentication classification of the inbound speech signal based on one or more similarity scores between the DP vector of the enrolled speech source and the DP vector of the inbound speech signal.

14. 1. A system comprising: a database comprising a non-transitory memory configured to store a plurality of speech signals having one or more speaker-independent characteristics; a server including a processor, wherein the processor: applying a plurality of task-specific machine learning models to the inbound speech signal having the one or more speaker-independent characteristics to extract a plurality of speaker-independent embeddings for the inbound speech signal; extracting a deep phonprint (DP) vector of the inbound speech signal based on the plurality of speaker-independent embeddings extracted for the inbound speech signal; and applying one or more post-modeling operations to the plurality of speaker-independent embeddings extracted for the inbound speech signal to generate one or more post-modeling outputs for the inbound speech signal.

15. The server: extracting a plurality of features from a training speech signal, the plurality of features including at least one of spectro-temporal features and metadata associated with the training speech signal; The system of claim 14 , further configured to: extract the plurality of features from the inbound speech signal.

16. the metadata includes at least one of: a microphone type used to capture the training audio signal; a device type from which the training audio signal originated; a codec type applied to compress and decompress the training audio signal for transmission; a carrier from which the training audio signal originated; a spoofing service used to change a source identifier associated with the training audio signal; a geography associated with the training audio signal; a network type from which the training audio signal originated; and an audio event indicator associated with the training audio signal; 16. The system of claim 15, wherein the one or more classifications of the one or more post-modeling outputs include at least one of the microphone type, the device type, the codec type, the carrier, the spoofing service, the geography, the network type, and the audio event indicator.

17. The system of claim 14 , wherein a post-modeling operation of the one or more post-modeling operations comprises at least one of a classification operation and a regression operation.

18. The system of claim 14 , wherein the task-specific machine learning model comprises at least one of a neural network and a Gaussian mixture model.

19. 15. The system of claim 14, wherein the server is further configured to train the plurality of task-specific machine learning models by applying each of the plurality of task-specific machine learning models to a plurality of training speech signals having the one or more speaker-independent characteristics.

20. When training a task-specific machine learning model, the server: extracting one or more first spectrotemporal features of a speech portion of a training speech signal and one or more second spectrotemporal features of a non-speech portion of the training speech signal; applying one or more modeling layers of the task-specific machine learning model to the one or more first spectrotemporal features and the one or more second spectrotemporal features to generate first and second speaker-independent embeddings; and generating predicted output data for the training speech signal by applying one or more post-modeling layers of the task-specific machine learning model to the first speaker-independent embedding and to the second speaker-independent embedding; 20. The system of claim 19, further configured to: tune one or more hyperparameters of the one or more modeling layers of the task-specific machine learning model by running a loss function using the predicted output data and expected output data indicated by one or more labels associated with the training audio signals.

21. When training a task-specific machine learning model, the server: extracting one or more first spectrotemporal features of a speech portion of a training speech signal and one or more second spectrotemporal features of a non-speech portion of the training speech signal; applying one or more shared modeling layers of one or more task-specific machine learning models on a plurality of features to the one or more first spectrotemporal features and the one or more second spectrotemporal features to generate first and second speaker-independent embeddings; generating predicted output data for the training speech signal by applying one or more post-modeling layers of the task-specific machine learning model to the first speaker-independent embedding and to the second speaker-independent embedding; 20. The system of claim 19, further configured to: tune one or more hyperparameters of the task-specific machine learning models by running a shared loss function of the one or more task-specific machine learning models using the predicted output data and expected output data indicated by one or more labels associated with the training audio signals.

22. When training the plurality of task-specific machine learning models, the server: one or more spectrotemporal features of a substantial portion of a training speech signal; applying a task-specific machine learning model to the one or more spectrotemporal features of the training speech signal to generate a speaker-independent embedding; generating predicted output data for the training speech signal by applying one or more post-modeling layers of the task-specific machine learning model to the speaker-independent embedding of the substantial portion of the training speech signal; 20. The system of claim 19, further configured to: tune one or more hyperparameters of the task-specific machine learning model by running a loss function using the predicted output data and expected output data indicated by one or more labels associated with the training audio signals.

23. 15. The system of claim 14, wherein the task-specific machine learning model comprises at least one of a convolutional neural network, a recurrent neural network, and a fully connected neural network.

24. The server: performing a voice activity detection (VAD) operation on a training speech signal, thereby generating one or more speech portions of the training speech signal and one or more non-speech portions of the training speech signal; 15. The system of claim 14, further configured to: perform the VAD operation on the inbound speech signal, thereby generating the one or more speech portions of the inbound speech signal and the one or more non-speech portions of the inbound speech signal.

25. The server: extracting a plurality of enrolled speaker-independent embeddings and DP vectors for enrolled speech sources by applying the plurality of task-specific machine learning models to a plurality of enrollment speech signals associated with the enrolled speech sources; extracting the DP vectors of the enrolled speech sources based on the plurality of enrolled speaker-independent embeddings; The system of claim 14 , wherein the server uses the DP vectors of the registered audio sources and the DP vectors of the inbound audio signal to generate the one or more post-modeling outputs.

26. The server:

26. The system of claim 25, further configured to determine an authentication classification of the inbound speech signal based on one or more similarity scores between the DP vector of the enrolled speech source and the DP vector of the inbound speech signal.

Citation Information

Patent Citations

  • Fraud detection database

    US20150269946A1

  • Systems and methods for detecting call provenance from call audio

    US20170126884A1

  • Improved fixed point integer implementations for neural networks

    US20170220929A1

  • Dimensionality reduction of baum-welch statistics for speaker recognition

    US20180082691A1

  • Classification Training Techniques to Map Datasets to a Standardized Data Model

    US20190034801A1