System and method for speaker-independent embedding for identification and verification from speech

By using machine learning models to generate deep phoneprint vectors from speaker-independent characteristics, the system addresses voice spoofing and noise issues in ASV, providing reliable voice-based authentication and spoofing detection.

JP7716420B2Active Publication Date: 2025-07-31PINDROP SECURITY INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022552583
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-05
Filing Date
2021-03-04
Publication Date
2025-07-31
Estimated Expiration
2041-03-04

AI Technical Summary

Technical Problem

Existing automatic speaker verification (ASV) systems are susceptible to voice spoofing, such as voice modulation and synthetic voices, and are affected by background noise, making them unreliable for authenticating the origin of telephone calls.

Method used

Implementing machine learning models, including Gaussian mixture models and neural networks, to evaluate speaker-independent characteristics of speech signals, generating deep phoneprint (DP) vectors that can be used for voice-based authentication and exclusion/permission lists, and detecting spoofing services.

Benefits of technology

The system provides reliable authentication by evaluating speaker-independent characteristics, enhancing voice-based verification and detection of spoofing, while being resilient to background noise and synthetic voices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007716420000001
    Figure 0007716420000001
  • Figure 0007716420000002
    Figure 0007716420000002
  • Figure 0007716420000003
    Figure 0007716420000003
Patent Text Reader

Abstract

The embodiments described herein provide speech processing operations that evaluate characteristics of a speech signal that are independent of the speaker's voice. A neural network architecture trains and applies a discriminative neural network whose task is to model and classify speaker-independent characteristics. A task-specific model generates or extracts feature vectors from input speech data based on a trained embedding extraction model. The embeddings from the task-specific models are concatenated to form a deep phonprint vector of the input speech signal. The DP vector is a low-dimensional representation of each of the speaker-independent characteristics of the speech signal and is applied to various downstream operations.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 62 / 985,757, filed March 5, 2020, the entire contents of which are incorporated by reference.

[0002] This application is generally related to U.S. Patent Application No. 17 / 066,210, filed October 8, 2020, U.S. Patent Application No. 17 / 079,082, filed October 23, 2020, and U.S. Patent Application No. 17 / 155,851, filed January 22, 2021, each of which is incorporated by reference in its entirety.

[0003] This application relates generally to systems and methods for training and deploying speech processing machine learning models. [Background technology]

[0004] Today, there are various forms of communication channels and devices available for voice communication, including Internet of Things (IoT) devices for communication over computing networks or various forms of telephone calls, such as landline calls, mobile phone calls, and Voice over IP (VoIP) calls, among others. In telephony systems, due to the introduction of virtual phone numbers, telephone numbers, automatic number identification (ANI), or caller identification (caller ID) are no longer uniquely tied to individual subscribers or telephone lines. Some VoIP services allow intentional spoofing of such identifiers (e.g., telephone numbers, ANI, caller ID), allowing callers to intentionally alter information transmitted to the recipient's display to disguise their identity. As a result, telephone numbers and similar telephony identifiers are no longer reliable for verifying the audio source of a call.

[0005] As caller ID services become less reliable, automatic speaker verification (ASV) systems are becoming necessary to authenticate the origin of telephone calls. However, ASV systems have strict net speech requirements and are susceptible to voice spoofing, such as voice modulation, synthetic voices (e.g., deepfakes), and replay attacks. ASV is also affected by background noise often experienced in telephone calls. Therefore, what is needed is a means to evaluate other attributes of a speech signal that are independent of the speaker's voice in order to verify the legitimate source of the speech signal. Summary of the Invention

[0006] Disclosed herein are systems and methods that can address the above-mentioned shortcomings, which may also provide any number of additional or alternative benefits and advantages. The embodiments described herein provide speech processing operations that evaluate characteristics of a speech signal that are speaker-independent or complementary to evaluating speaker-dependent characteristics. Computer-implemented software executes one or more machine learning models, which may include Gaussian mixture models (GMMs) and / or neural network architectures with discriminative neural networks, such as convolutional neural networks (CNNs), deep neural networks (DNNs), and recurrent neural networks (RNNs), referred to herein as "task-specific machine learning models" or "task-specific models," each tasked with and configured to model and / or classify corresponding speaker-independent characteristics.

[0007] The task-specific machine learning model is trained for each speaker-independent characteristic of the input audio signal using the input audio data and metadata associated with this audio data or the audio source. The discriminant model is trained and developed to distinguish the classification of the characteristics of the voice. One or more modeling layers (or modeling operations) generate or extract a feature vector, which may also be called an "embedding" or may be combined to form an embedding, based on the input audio data. A specific post-modeling operation (or post-modeling layer) takes the embedding from the task-specific model and trains the task-specific model. The post-modeling layer (or post-modeling operation) concatenates the speaker-independent embeddings to form a deep-phoneprint (DP) vector of the input audio signal. The DP vector is a low-dimensional representation of each of the various speaker-independent characteristics of the audio signal aspect of the voice. Non-limiting examples of additional or alternative post-modeling operations or post-modeling layers of the task-specific model may include classification operations / layers, fully connected layers, loss functions / layers, and regression operations / layers (e.g., probabilistic linear discriminant analysis (PLDA)).

[0008] The DP vector may be used for various downstream operations or downstream tasks such as creating a voice-based exclusion / permission list, implementing a voice-based exclusion / permission list, authenticating a registered legitimate audio source, determining the device type, determining the microphone type, determining the geographical location of the source of the voice, determining the codec, determining the carrier, determining the network type involved in the transmission of the voice, detecting a spoofing service that spoofs the device identifier, recognizing the spoofing service, and recognizing voice events occurring in the audio signal.

[0009] DP vectors can be used in voice-based authentication operations either alone or complementary to voice biometric features. Additionally or alternatively, DP vectors can be used for voice quality measurement purposes and can be combined with voice biometric systems for various downstream operations or tasks.

[0010] In one embodiment, a computer-implemented method includes applying, by a computer, a plurality of task-specific machine learning models to an inbound speech signal having one or more speaker-independent characteristics to extract a plurality of speaker-independent embeddings for the inbound speech signal; extracting, by the computer, a deep phonprint (DP) vector for the inbound speech signal based on the plurality of speaker-independent embeddings extracted for the inbound speech signal; and applying, by the computer, one or more post-modeling operations to the plurality of speaker-independent embeddings extracted for the inbound speech signal to generate one or more post-modeling outputs for the inbound speech signal.

[0011] In another embodiment, the database comprises a non-transitory memory configured to store a plurality of training speech signals having one or more speaker-independent characteristics. The server comprises a processor configured to: apply a plurality of task-specific machine learning models to the inbound speech signal having the one or more speaker-independent characteristics to extract a plurality of speaker-independent embeddings for the inbound speech signal; extract a deep phonprint (DP) vector for the inbound speech signal based on the plurality of speaker-independent embeddings extracted for the inbound speech signal; and apply one or more post-modeling operations to the plurality of speaker-independent embeddings extracted for the inbound speech signal to generate one or more post-modeling outputs for the inbound speech signal.

[0012] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are intended to provide further explanation of the invention as claimed.

Brief Description of the Drawings

[0013] The present disclosure can be better understood by referring to the following drawings. The components in the drawings are not necessarily to scale, and instead, emphasis is placed on illustrating the principles of the present disclosure. In the drawings, reference numerals refer to corresponding parts throughout the various drawings.

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

[0015] Reference will now be made to the exemplary embodiments illustrated in the drawings, and specific language will be used herein to describe the same. It will be understood that no limitation of the scope of the invention is thereby intended. Alterations and further modifications of the features of the invention illustrated herein, and further applications of the principles of the invention as illustrated herein, which will occur to those skilled in the art and in possession of this disclosure, are to be considered within the scope of the invention.

[0016] This specification describes systems and methods for processing an audio signal with a sample of a speaker's voice and using the result in any number of downstream operations or downstream tasks. The computing device (e.g., a server) of the system executes software programming that implements various machine learning algorithms, including various types of variants of neural networks such as Gaussian mixture models (GMMs), convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs). The software programming trains a model for recognizing and evaluating speaker-independent characteristics of an audio signal received from an audio source (e.g., a calling user, a speaking user, a source location, a source system). Various characteristics of the audio signal are independent of a particular speaker of the audio signal, as opposed to speaker-dependent characteristics associated with the voice of a particular speaker.

[0017] Non-limiting examples of speaker-independent features include, among others, the device type on which the voice is uttered and recorded (e.g., landline phone, mobile phone, computing device, Internet of Things (IoT) / edge device), the microphone type used to capture the voice (e.g., speakerphone, headset, wired or wireless headset, IoT device), the carrier over which the voice is transmitted (e.g., AT&T, Sprint, T-Mobile, Google Voice), the codec applied for compression and decompression of the voice for transmission or storage, the geographical location associated with the voice source (e.g., continent, country, state / province, county / city), determining whether the identifier associated with the voice source is spoofed, determining spoofing services (e.g., Zang, Tropo, Twilio) that can be used to change the source identifier associated with the voice, the type of network over which the voice is transmitted, voice events occurring in the input voice signal (e.g., cellular network, landline communication network, VOIP device), voice events occurring in the input voice signal (e.g., background noise, traffic sound, TV noise, music, crying baby, train or factory whistle, laughter), and the communication channel over which the voice is received.

[0018] The device type may be more granular, for example, to reflect the manufacturer of the device (e.g., Samsung, Apple) or the model (e.g., Galaxy S10, iPhone X). The codec classification may further indicate, for example, a single codec or multiple cascaded codecs. The codec classification may further indicate other information or may be further used to determine other information. For example, since voice signals such as telephone calls can be emitted from different source device types (e.g., landline phone, mobile phone, VoIP device), different voice codecs are applied to the voice of the call (e.g., SILK codec on Skype, WhatsApp, G722 for PSTN, GSM codec).

[0019] The system's server (or other computing device) executes one or more machine learning models and / or neural network architectures that include a machine learning modeling layer or neural network layer for performing various operations, including layers of discriminative neural network models (e.g., DNN, CNN, RNN). Each specific machine learning model is trained to correspond to aspects of an audio signal using audio data and / or metadata related to specific aspects. The neural network learns to distinguish classification labels for various aspects of the audio signal. One or more fully connected layers of the neural network architecture extract feature vectors or embeddings of the audio signals generated from each of the neural networks, and concatenate the respective feature vectors to form a deep phone print (DP) vector. The DP vector is a low-dimensional representation of different aspects of the audio signal.

[0020] The DP vector is used in various downstream operations. By way of non-limiting example, among other things, it can be used to create and enforce an audio-based exclusion list, authenticate registered or legitimate audio sources, determine the device type, microphone type, geographical location of the audio source, codec, carrier, and / or network type involved in the transmission of the audio, detect spoofed identifiers associated with the audio signal such as a spoofed caller identifier (caller ID), spoofed automatic number identifier (ANI), or spoofed phone number, or recognize a spoofing service. Additionally or alternatively, the DP vector is complementary to audio biometric features. For example, the DP vector can be used together with complementary voice prints (e.g., audio-based speaker vectors or embeddings) to perform audio quality measurement and audio enhancement, or to perform authentication operations by an audio biometric system.

[0021] For ease of explanation and understanding, embodiments described herein involve a neural network architecture that includes any number of task-specific machine learning models configured to model and classify specific aspects of an audio signal, where each task corresponds to modeling and classifying a specific characteristic of the audio signal. For example, the neural network architecture can include a device type neural network and a carrier neural network, where the device type neural network models and classifies the type of device that emitted the input audio signal, and the carrier neural network models and classifies a specific communication carrier associated with the input audio signal. However, the neural network architecture need not include a machine learning model for each task. The server may, for example, execute the task-specific machine learning models individually as separate neural network architectures, or execute any number of neural network architectures that include any combination of the task-specific machine learning models. The server then models or clusters the outputs resulting from each task-specific machine learning model to generate a DP vector.

[0022] Although the system is described herein as implementing a neural network architecture with any number of machine learning model layers or neural network layers, any number of combinations or architectural configurations of machine learning architectures are possible. For example, a shared machine learning model may use a shared GMM operation to jointly model an input audio signal and extract one or more feature vectors, and then for each task, a separate fully connected layer (of a separate fully connected neural network) may be implemented that performs various pooling and statistical operations for the specific task. Generally, the architecture includes a modeling layer, a pre-modeling layer, and a post-modeling layer. The modeling layer includes layers for performing audio processing operations, such as extracting feature vectors or embeddings from various types of features extracted from the input audio signal or metadata. The pre-modeling layer performs pre-processing operations, transformation operations, etc. to take the input audio signal and metadata and prepare them for the modeling layer, such as extracting features from the audio signal or metadata. The post-modeling layer performs operations that use the output of the modeling layer, such as training operations, loss functions, classification operations, and regression functions. The boundaries and functions of the layer types may vary in different implementations.

[0023] System Architecture FIG. 1 shows the components of a system 100 for receiving and analyzing voice signals from an end user. System 100 includes an analysis system 101, a service provider system 110 of various types of entities (e.g., companies, government agencies, universities), and an end user device 114. The analysis system 101 includes an analysis server 102, an analysis database 104, and a management device 103. The service provider system 110 includes a provider server 111, a provider database 112, and an agent device 116. Embodiments may include additional or alternative components, or may omit certain components from the components of FIG. 1 and still fall within the scope of the present disclosure. For example, it may be common to include multiple service provider systems 110, or for the analysis system 101 to have multiple analysis servers 102. Embodiments may include any number of devices capable of performing the various features and tasks described herein, or may implement them otherwise. For example, FIG. 1 shows the analysis server 102 as a computing device separate from the analysis database 104. In some embodiments, the analysis database 104 may be integrated with the analysis server 102.

[0024] The embodiments described with respect to FIG. 1 are merely examples that use speaker-independent embedding and deep phone printing and do not necessarily limit other potential embodiments. The description of FIG. 1 refers to a situation where an end user calls a service provider system 110 through various communication channels to contact and / or interact with services provided by the service provider. However, the operations and features of the various deep phone printing implementation aspects described herein may be applicable to many situations for evaluating speaker-independent aspects of audio signals. For example, the deep phone printing audio processing operations described herein may be implemented within various types of devices and do not need to be implemented within a larger infrastructure. As an example, the IoT device 114d may implement the various processes described herein when capturing an input audio signal from an end user or receiving an input audio signal from another end user via a TCP / IP network. As another example, the end user device 114 may execute locally installed software that implements the deep phone printing process described herein, for example, enabling the deep phone printing process in an interaction between users. The smartphone 114b may execute deep phone printing software when receiving an inbound call from another end user to perform certain downstream operations such as verifying the identity of the other end user or indicating whether the other end user is using a spoofing service.

[0025] Various hardware and software components of one or more public or private networks may interconnect the various components of system 100. Non-limiting examples of such networks may include a local area network (LAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a wide area network (WAN), and the Internet. Communications over the networks may be performed according to various communication protocols, such as Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communication protocols. Similarly, end-user devices 114 may communicate with called parties (e.g., provider system 110) via telecommunications protocols, hardware, and software capable of hosting, transporting, and exchanging telephone communications and voice data associated with telephone calls. Non-limiting examples of telecommunications hardware may include switches and trunks, among other additional or alternative hardware used to host, route, or manage telephone calls, circuits, and signaling. Non-limiting examples of software and protocols for telecommunications may include SS7, SIGTRAN, SCTP, ISDN, and DNIS, among other additional or alternative software and protocols used to host, route, or manage telephone calls, circuits, and signaling. Components for telecommunications may be organized or managed by a variety of different entities, such as carriers, switches, and networks, among others.

[0026] The end user device 114 may be any communication or computing device that operates to allow callers to access the services of the service provider system 110 through various communication channels. For example, an end user may place a call to the service provider system 110 through a telephony network or through a software application executed by the end user device 114. Non-limiting examples of the end user device 114 may include a landline phone 114a, a mobile phone 114b, a calling computing device 114c, or an edge device 114d. The landline phone 114a and the mobile phone 114b are telecommunication-oriented devices (e.g., telephones) that communicate through telecommunication channels. The end user device 114 is not limited to telecommunication-oriented devices or channels. For example, in some cases, the mobile phone 114b may communicate through a computing network channel (e.g., the Internet). The end-user devices 114 may also include electronic devices with processors and / or software, such as a call computing device 114c or an edge device 114d, that implement voice-over-IP (VoIP) communications, data streaming over a TCP / IP network, or other computing network channels. The edge device 114d may include any IoT device or other electronic device for computing network communications. The edge device 114d may be any smart device capable of running software applications and / or performing voice interface operations. Non-limiting examples of the edge device 114d may include a voice assistant device, an automobile, a smart appliance, etc.

[0027] The service provider system 110 comprises various hardware components and software components that capture and store various types of voice signal data or metadata related to the contact between the caller and the service provider system 110. This voice data can include, for example, the voice recording of the call and the metadata related to the software used for a particular communication channel and various protocols. Speaker-independent features of the voice signal, such as voice quality or sampling rate, can represent (and be used for evaluating) various speaker-independent aspects, such as, inter alia, the codec, the type of end-user device 114, or the carrier.

[0028] The analysis system 101 and the provider system 110 represent network infrastructures 101, 110 that comprise physically and logically related software and electronic devices managed or operated by various enterprise organizations. The devices of each network system infrastructure 101, 110 are configured to provide the intended services of a particular enterprise organization.

[0029] The analysis server 102 of the analysis system 101 can be any computing device that includes one or more processors and software and can execute the various processes and tasks described herein. The analysis server 102 can host the analysis database 104 or communicate with the analysis database 104, and receive and process voice signal data (e.g., voice recordings, metadata) received from one or more provider systems 110. Although FIG. 1 shows only a single analysis server 102, the analysis server 102 may include any number of computing devices. In some cases, the computing devices of the analysis server 102 can execute all or part of the processes and benefits of the analysis server 102. The analysis server 102 can include computing devices that operate in a distributed computing configuration or a cloud computing configuration and / or in a virtual machine configuration. In some embodiments, the functions of the analysis server 102 can be partially or fully executed by the computing devices of the provider system 110 (e.g., provider server 111).

[0030] The analysis server 102 executes speech processing software including one or more neural network architectures with neural network layers for deep phoneprinting operations (e.g., extracting speaker-independent embeddings, extracting DP vectors) and any number of downstream speech processing operations. For ease of explanation, the analysis server 102 is described as executing a single neural network architecture for implementing deep phoneprinting, including neural network layers for extracting speaker-independent embeddings and deep phoneprint vectors (DP vectors), although in some embodiments, multiple neural network architectures may be employed. The analysis server 102 and neural network architectures logically operate in several operational phases, including a training phase, an enrollment phase, and a deployment phase (sometimes referred to as a "test" phase or an "inference" phase), although some embodiments need not perform an enrollment phase. Input speech signals processed by the analysis server 102 and neural network architectures include a training speech signal, an enrollment speech signal, and an inbound speech signal (processed during the deployment phase). The analysis server 102 applies a neural network architecture to each type of input audio signal during a corresponding computation phase.

[0031] The analysis server 102 or other computing devices (e.g., provider server 111) of system 100 can perform various preprocessing and / or data augmentation operations on input speech signals (e.g., training speech signals, enrollment speech signals, inbound speech signals). While the analysis server 102 may perform preprocessing and data augmentation operations when executing a particular neural network layer, the analysis server 102 may also perform a particular preprocessing or data augmentation operation as a separate operation from the neural network architecture (e.g., before feeding the input speech signal into the neural network architecture).

[0032] Optionally, the analysis server 102 performs any number of preprocessing operations before feeding the audio data to the neural network. The analysis server 102 may perform various preprocessing operations during one or more of the operation phases (e.g., training phase, enrollment phase, deployment phase), although the specific preprocessing operations performed may vary between operation phases. The analysis server 102 may perform various preprocessing operations separately from the neural network architecture or when executing an inner network layer of the neural network architecture. Non-limiting examples of preprocessing operations performed on the input audio signal include running voice activity detection (VAD) software or a VAD neural network layer, extracting features (e.g., one or more spectrotemporal features) from a portion (e.g., a frame, a segment) or from substantially all of a particular input audio signal, and converting the extracted features from a time-domain representation to a frequency-domain representation by performing a short-time Fourier transform (SFT) operation and / or a fast Fourier transform (FFT) operation, among other preprocessing operations. The features extracted from the input audio signal may include, for example, Mel Frequency Cepstral Coefficients (MFCCs), Mel filter banks, linear filter banks, bottleneck features, and communication protocol metadata fields, among other types of data. Preprocessing operations may also include parsing the audio signal into frames or subframes and performing various normalization or scaling operations.

[0033] As an example, the neural network architecture can include a neural network layer for a VAD operation that syntax analyzes a set of utterance portions and a set of non-utterance portions from each specific input audio signal. The analysis server 102 can train the VAD classifier separately (as a separate neural network architecture) or together with the neural network architecture (as part of the same neural network architecture). When the VAD is applied to the features extracted from the input audio signal, the VAD outputs a binary result (e.g., speech detection, no speech detection) or a likelihood value (e.g., the probability that speech occurs) for each window of the input audio signal, thereby indicating whether an utterance portion occurs in a given window. The server can store one or more sets of utterance portions, one or more sets of non-utterance portions, and the input audio signal in a memory storage location, including short-term RAM, a hard disk, or one or more databases 104, 112.

[0034] As described above, the analysis server 102 of the system 100 or another computing device (e.g., the provider server 111) can perform various augmentation operations on the input audio signal (e.g., a training audio signal, a registration audio signal, an inbound audio signal). Non-limiting examples of augmentation operations include, among others, frequency augmentation, audio clipping, and duration augmentation. The augmentation operations generate various types of distortion or degradation of the input audio signal such that the resulting audio signal is taken in by, for example, a convolutional operation of a modeling layer that generates a feature vector or a speaker-independent embedding. The analysis server 102 can perform various augmentation operations as an operation separate from the neural network architecture or as an in-network augmentation layer. The analysis server 102 can perform various augmentation operations in one or more of the operational phases, although the specific augmentation operations performed may vary between operational phases.

[0035] As described in detail herein, the neural network architecture includes any number of task-specific models configured to model and classify particular speaker-independent aspects of the input audio signal, with each task corresponding to modeling and classifying a particular speaker-independent characteristic of the input audio signal. For example, the neural network architecture may include a device-type neural network that models and classifies the type of device that emitted the input audio signal and a carrier neural network that models and classifies the particular communications carrier associated with the input audio signal. As noted, the neural network architecture need not include each task-specific model. For example, the server may execute the task-specific models individually as separate neural network architectures, or may execute any number of neural network architectures including any combination of task-specific models.

[0036] The neural network architecture includes task-specific models configured to extract corresponding speaker-independent embeddings and DP vectors based on the speaker-independent embeddings. The server applies the task-specific models to the speech portions (or the speech-only abridged speech signal) and then to the non-speech portions (or the non-speech-only abridged speech signal). The analysis server 102 applies a particular type of task-specific model (e.g., a speech event neural network) to substantially all of the input speech signal (or at least to the portion not parsed by the VAD). For example, the task-specific models of the neural network architecture include a device-type neural network and a speech event neural network. In this example, the analysis server 102 applies the device-type neural network to the speech portions and then again to the non-speech portions to extract speaker-independent embeddings for the speech and non-speech portions. The analysis server 102 then applies the device-type neural network to the input speech signal to extract the entire speech signal embedding.

[0037] During the training phase, the analysis server 102 receives training audio signals having various speaker-independent characteristics (e.g., codec, carrier, device type, microphone type) from one or more corpora of training audio signals stored in the analysis database 104 or other storage media. The training audio signals may further include clean audio signals and simulated audio signals, and the analysis server 102 uses each of them to train various layers of the neural network architecture.

[0038] The analysis server 102 may obtain simulated audio signals from more analysis databases 104 and / or generate simulated audio signals by performing various data augmentation operations. In some cases, the data augmentation operations may generate a simulated audio signal of a given input audio signal (e.g., training signal, registration signal), and the simulated audio signal encompasses the manipulated features of the input audio signal that mimic the effects of a particular type of signal degradation or distortion of the input audio signal. The analysis server 102 stores the training audio signals in a non-transitory medium of the analysis server 102 and / or the analysis database 104 for future reference or operations of the neural network architecture.

[0039] The training audio signals are associated with training labels that are separate machine-readable data records or metadata codings of the training audio signal data files. The labels can be generated by the user to indicate the expected data (e.g., expected classification, expected features, expected feature vectors), or the labels can be automatically generated according to a computer-executed process used to generate a particular training audio signal. For example, during the noise augmentation operation, the analysis server 102 generates a simulated audio signal by algorithmically combining the input audio signal and the type of noise degradation. The analysis server 102 generates or updates the corresponding label of the simulated audio signal indicating the expected features, expected feature vectors, or other expected types of data of the simulated audio signal.

[0040] In some embodiments, the analysis server 102 performs an enrollment phase to develop enrollee speaker-independent embeddings and enrollee DP vectors. The analysis server 102 may perform some or all of the preprocessing and / or data augmentation operations on the enrollee speech signals of the enrolled speech sources.

[0041] During the training and enrollment phases in some embodiments, one or more fully connected layers, classification layers, and / or output layers of each task-specific model generate predicted outputs (e.g., predicted classifications, predicted speaker-independent feature vectors, predicted speaker-independent embeddings, predicted DP vectors, predicted similarity scores) for the training speech signals (or enrollment speech signals). Loss layers execute various types of loss functions to evaluate distances (e.g., differences, similarities) between predicted outputs (e.g., predicted classifications) to determine the level error between the predicted outputs and the corresponding expected outputs indicated by the training labels associated with the training speech signals (or enrollment speech signals). The loss layers, or other functions performed by the analysis server 102, tune or adjust hyperparameters of the neural network architecture until the distance between the predicted outputs and the expected outputs meets a training threshold.

[0042] During the enrollment operation phase, registered sound sources such as enrolled users of the service provider system 110 (e.g., end user device 114, enrolled organization, enrolling user) provide several registered voice signals that encapsulate examples of speaker-independent characteristics (to the analysis system 101). In some embodiments, the registered sound source further includes examples of the utterances of the registered user. The registered user may provide the enrolling voice signal via any number of channels and / or using any number of channels. The analysis server 102 or the provider server 111 actively or passively captures the registered voice signal. In active enrollment, the enroller responds to a voice prompt or a GUI prompt for supplying the enrolling voice signal to the provider server 111 or the analysis server 102. As an example, the enroller may respond to various automated voice response (IVR) prompts of the IVR software executed by the provider server 111 via a telephone channel. As another example, the enroller may respond to various prompts generated by the provider server 111 and exchanged with the software application of the edge device 114d via the corresponding data communication channel. As another example, the enroller may upload a media file (e.g., WAV, MP3, MP4, MPEG) containing voice data to the provider server 111 or the analysis server 102 via a computing network channel (e.g., Internet, TCP / IP). In passive enrollment, the provider server 111 or the analysis server 102 collects the registered voice signal continuously over time without the enroller's awareness and / or through one or more communication channels. In embodiments where the provider server 111 receives or otherwise collects the registered voice signal, the provider server 111 transfers (or otherwise transmits) the authentic registered voice signal to the analysis server 102 via one or more networks.

[0043] The analysis server 102 feeds each registered voice signal to the VAD and syntactically analyzes a specific registered voice signal into an utterance part (or a reduced voice signal of only the utterance) and a non-utterance part (or a reduced voice signal of only no utterance). For each registered voice signal, the analysis server 102 applies a trained neural network architecture including a task-specific model to the set of utterance parts and again to the set of non-utterance parts. The task-specific model generates a registered speaker-independent feature vector of the registered voice signal based on features extracted from the registered voice signal. The analysis server 102 algorithmically combines the registered feature vectors generated from the entire registered voice signal to extract an utterance speaker-independent registration embedding (for the utterance part) and a non-utterance speaker-independent registration embedding (for the non-utterance part). Next, the analysis server 102 applies a full voice task-specific model to generate a full voice registered feature vector for each of the registered voice signals. Next, the analysis server 102 algorithmically combines the full voice registered feature vectors to extract a full voice speaker-independent registration embedding. Next, the analysis server 102 extracts a registration DP vector of the registered sound source of the registrant by algorithmically combining each of the speaker-independent embeddings. The speaker-independent embedding and / or the DP vector may be referred to as a "deep phone print".

[0044] The analysis server 102 stores the extracted registered speaker-independent embedding and the extracted DP vector for each of various registered sound sources. In some embodiments, the analysis server 102 may similarly store the extracted registered speaker-dependent embedding (sometimes referred to as a "voice print" or "registered voice print"). The registered speaker-independent embedding is stored in the analysis database 104 or the provider database 112. Examples of neural networks for speaker verification are described in U.S. Patent Application Nos. 17 / 066,210 and 17 / 079,082, which are incorporated herein by reference.

[0045] Optionally, a particular end-user device 114 (e.g., computing device 114c, edge device 114d) executes software programming associated with the analysis system 101. The software program generates an enrollment feature vector by locally capturing an enrolled audio signal and / or locally applies a trained neural network architecture to each of the enrolled audio signals (on-device). The software program then transmits the enrollment feature vector to the provider server 111 or the analysis server 102.

[0046] Following the training phase and / or the enrollment phase, the analysis server 102 stores the trained neural network architecture or the developed neural network architecture in the analysis database 104 or the provider database 112. The analysis server 102 places the neural network architecture in the training phase or the enrollment phase, which may include enabling or disabling a particular layer of the neural network architecture. In some implementations, the devices of the system 100 (e.g., provider server 111, agent device 116, management device 103, end-user device 114) instruct the analysis server 102 to transition to an enrollment phase for developing a neural network architecture by extracting various types of embeddings of the enrolled audio source. The analysis server 102 then stores the extracted enrollment embeddings and the trained neural network architecture in one or more databases 104, 112 for later reference during the deployment phase.

[0047] During the expansion phase, the analysis server 102 receives an inbound voice signal from an inbound voice source, emitted from the end-user device 114, received through a specific communication channel. The analysis server 102 applies a trained neural network architecture to the inbound voice signal to generate a set of speaking parts and a set of non-speaking parts, extracts features from the inbound voice signal, and extracts an inbound speaker-independent embedding and an inbound DP vector of the inbound voice source. The analysis server 102 may use the extracted embedding and / or DP vector in various downstream operations. For example, the analysis server 102 may determine a similarity score based on the distance, difference / similarity between the registered DP vector and the inbound DP vector, and the similarity score indicates the likelihood that the registered DP vector was emitted from the same voice source as the inbound DP vector. As described herein, deep phone printing outputs generated by a machine learning model, such as speaker-independent embeddings and DP vectors, may be used in various downstream operations.

[0048] The analysis database 104 and / or the provider database 112 may be hosted on a computing device (e.g., a server, a desktop computer) that includes hardware and software components capable of performing the various processes and tasks described herein, such as a non-transitory machine-readable storage medium and database management software (DBMS). The analysis database 104 and / or the provider database 112 contain any number of corpora of training speech signals that are accessible to the analysis server 102 via one or more networks. In some embodiments, the analysis server 102 trains a neural network using supervised training, and the analysis database 104 and / or the provider database 112 contain labels associated with the training or enrollment speech signals. The labels indicate, for example, expected data for the training or enrollment speech signals. The analysis server 102 may also query an external database (not shown) to access a third-party corpus of training speech signals. An administrator may configure the analysis server 102 to select training speech signals with various types of speaker-independent characteristics.

[0049] The provider server 111 of the provider system 110 executes software processes for interacting with end users through various channels. The process may include, for example, routing calls to the appropriate agent device 116 based on comments, commands, IVR inputs, or other inputs submitted during an inbound call by an inbound caller. The provider server 111 can capture, query, or generate various types of information regarding the inbound voice signal, the caller, and / or the end user device 114, and transfer the information to the agent device 116. The graphical user interface (GUI) of the agent device 116 displays the information to the agent of the service provider. The provider server 111 also sends information regarding the inbound voice signal to the analysis system 101 to perform various analysis processes on the inbound voice signal and any other voice data. The provider server 111 can send information and voice data based on preconfigured trigger conditions (e.g., receiving an inbound telephone call), commands or queries received from another device of the system 100 (e.g., the agent device 116, the management device 103, the analysis server 102), or as part of a batch sent at regular intervals or at a predetermined time.

[0050] The management device 103 of the analysis system 101 is a computing device that enables a person in charge of the analysis system 101 to perform various management tasks or analysis operations urged by the user. The management device 103 can be any computing device equipped with a processor and software and capable of performing various tasks and processes described herein. Non-limiting examples of the management device 103 can include servers, personal computers, laptop computers, tablet computers, and the like. In operation, the user uses the management device 103 to configure the operation of various components of the analysis system 101 or the provider system 110, and issue queries and commands to such components.

[0051] The agent device 116 of the provider system 110 may enable an agent of the provider system 110 or other users to configure the operation of the devices of the provider system 110. For a call made to the provider system 110, the agent device 116 receives and displays some or all of the information associated with the inbound voice signal routed from the provider server 111.

[0052] Exemplary operations Operation phase FIG. 2 shows the steps of a method 200 for implementing a task-specific model for processing speaker-independent aspects of a voice signal. Embodiments may include additional operations, fewer operations than those described in method 200, or operations different from those described in method 200. Method 200 is implemented by a server that executes machine-readable software code of a neural network architecture that includes any number of neural network layers and neural networks, although various operations may be executed by one or more computing devices and / or processors. The server is described as generating and evaluating enrollee embeddings, although the server need not generate and evaluate enrollee embeddings in all embodiments.

[0053] In step 202, the server places the neural network architecture and the task-specific model in a training phase. The server applies the neural network architecture to any number of training voice signals to train the task-specific model. The task-specific model includes a modeling layer. The modeling layer may include, for example, an "embedding extraction layer" or a "hidden layer" that generates a feature vector of the input voice signal.

[0054] During the training phase, the server applies modeling layers to the training audio signals to generate training feature vectors. Various post-modeling layers, such as a fully connected layer and a classification layer (sometimes referred to as a "classifier layer" or "classifier") of each task-specific model, determine task-related classifications based on the training feature vectors. For example, a task-specific model may include a device-type neural network or a carrier neural network. The classifier layer of the device-type neural network outputs a predicted brand classification of the device of the audio source, and the classifier layer of the carrier neural network outputs a predicted carrier classification associated with the audio source.

[0055] The task-specific model generates a predicted output (e.g., a predicted training feature vector, a predicted classification) for a particular training audio signal. The post-embedding modeling layers (e.g., a classification layer, a fully connected layer, a loss layer) of the neural network architecture execute a loss function according to the predicted output of the training signal and the label associated with the training audio signal. The server executes the loss function to determine the level of error of the training feature vector generated by the modeling layer of the particular task-specific model. The classifier layer (or other layer) adjusts hyperparameters of the task-specific model and / or other layers of the neural network architecture until the training feature vector converges to the expected feature vector indicated by the label associated with the training audio signal. Once the training phase is complete, the server stores the hyperparameters in a server memory location or other memory location. The server may also disable one or more layers of the neural network architecture during later computation phases to keep the hyperparameters fixed.

[0056] A particular type of task-specific model is trained on the entire training speech signal, such as a speech event neural network. The server applies these task-specific models to the entire training speech signal and outputs a predicted classification (e.g., a speech event classification). Similarly, a particular task-specific model is trained on the speech portion and separately on the non-speech portion of the training signal. For example, one or more layers of the neural network architecture define a VAD layer that the server applies to the training speech signal to parse the training speech signal into a set of speech portions and a set of non-speech portions. The server then applies each task-specific model to the speech portions to train a speech task-specific model, and again to the non-speech portions to train a non-speech task-specific model.

[0057] The server can train each task-specific model individually and / or sequentially, sometimes referred to as a "single-task" configuration, where each task-specific model includes a separate modeling layer for extracting a separate embedding. Each task-specific model outputs a separate predicted output (e.g., predicted feature vector, predicted classification). A loss function evaluates the level of error for each task-specific model based on the relative distance (e.g., similarity or difference) between the predicted output and the expected output, as indicated by the label. The loss function then adjusts hyperparameters or other aspects of the neural network architecture to minimize the level of error. Once the level error meets a training threshold, the server fixes (e.g., preserves, does not disturb) the hyperparameters or other aspects of the neural network architecture.

[0058] The server can jointly train task-specific models, which may be referred to as a "multi-task" configuration. Multiple task-specific models share the same hidden modeling layer and loss layer, but have specific individual post-modeling layers (e.g., fully connected layers, classification layers). The server feeds the training audio signal into a neural network architecture and applies each of the task-specific models. The neural network architecture includes a shared hidden layer for generating a joint feature vector of the input audio signal. Each of the task-specific models includes a separate post-modeling layer that takes in the joint feature vector and generates, for example, a task-specific predicted output (e.g., a predicted feature vector, a predicted classification). Optionally, a shared loss layer is applied to each of the predicted output and the label associated with the training audio signal to adjust hyperparameters and minimize the level of error. An additional post-modeling layer may algorithmically combine or concatenate the predicted feature vectors of specific training audio signals to output, among potential predicted outputs (e.g., predicted classifications), a predicted DP vector in particular. As described above, the server executes a shared loss function that evaluates the level of error between the predicted joint output and the expected joint output according to one or more labels associated with the training audio signal, and adjusts one or more hyperparameters to minimize the level of error. When the level of error meets a training threshold, the server fixes (e.g., saves, does not interfere with) the hyperparameters or other aspects of the neural network architecture.

[0059] In step 204, the server places the neural network architecture and the task-specific model in the registration phase and extracts the registered embeddings of the registered sound sources. In some implementations, the server may enable and / or disable specific layers of the neural network architecture during the registration phase. For example, the server typically enables and applies each layer during the registration phase, but the server disables the classification layer. The registered embeddings include speaker-independent embeddings, but in some embodiments, the speaker modeling neural network may extract one or more speaker-dependent embeddings of the registered sound sources.

[0060] During the registration phase, the server receives the registered audio signals of the registered sound sources and applies the task-specific model to extract speaker-independent embeddings. The server applies the task-specific model to the registered audio signals and generates the registered feature vectors of each of the registered audio signals as described for the training phase (e.g., single-task configuration, multi-task configuration). For each of the task-specific models, the server combines each of the registered feature vectors statistically or algorithmically to extract the task-specific registered embeddings.

[0061] A particular task-specific model is applied separately to the speech portion and again to the non-speech portion of the registered audio signal. The server applies VAD to each particular registered audio signal to parse the registered audio signal into a set of speech portions and a set of non-speech portions. For these task-specific models, the neural network architecture extracts the following two speaker-independent embeddings: the task-specific embedding of the speech portion and the task-specific embedding of the non-speech portion. Similarly, a particular task-specific model is applied to the entire registered audio signal (e.g., audio event neural network) to generate the registered feature vectors and extract the corresponding task-specific registered embeddings.

[0062] The neural network architecture extracts the registration DP vectors of the sound sources. One or more post-modeling layers or output layers of the neural network architecture concatenate or algorithmically combine various speaker-independent embeddings. Then, the server stores the registration DP vectors in memory.

[0063] In step 206, the server places the neural network architecture in the deployment phase (sometimes referred to as the "inference" or "test" phase) when the neural network architecture generates the inbound embedding and the inbound DP vector of the inbound sound source. The server may enable and / or disable specific layers and task-specific models of the neural network architecture during the deployment phase. For example, the server typically enables and applies each of the layers during the deployment phase, but the server disables the classification layer. In the current step 206, the server receives the inbound voice signal of the inbound speaker and feeds the inbound voice signal into the neural network architecture.

[0064] In step 208, during the deployment phase, the server applies the neural network architecture and the task-specific model to the inbound voice signal to extract the inbound embedding and the inbound DP vector. Then, the neural network architecture generates one or more similarity scores based on the relative distance (e.g., similarity, difference) between the inbound DP vector and one or more registered DP vectors. The server applies the task-specific model to the inbound voice signal to generate an inbound feature vector and extracts the inbound speaker-independent embedding and the inbound DP vector of the inbound sound source as described for the training and registration phases (e.g., single-task configuration, multi-task configuration).

[0065] As an example, a neural network architecture extracts an inbound DP vector and outputs a similarity score indicating the distance (e.g., similarity, difference) between the inbound DP vector and the enrollee DP vector. A larger distance may indicate a lower likelihood that the inbound audio signal was emitted from the enrollee audio source that emitted the enrollee DP vector, due to less / lower similarity between the speaker-independent aspects of the inbound audio signal and the enrollee audio signal. In this example, the server determines that the inbound audio signal was emitted from the enrollee audio source if the similarity score meets a threshold for audio source verification. The task-specific model and the DP vector may be used in any number of downstream operations, as described in various embodiments of this specification.

[0066] Training and enrolment for single-task configurations FIG. 3 shows the execution steps of method 300 for a training operation or an enrolment operation of a neural network architecture for speaker-independent embedding. Embodiments may include additional operations to those described in method 300, fewer operations than those described in method 300, or operations different from those described in method 300. Method 300 is implemented by a server executing machine-readable software code of a neural network architecture, although the various operations may be executed by one or more computing devices and / or processors. Embodiments may include additional operations to those described in method 300, fewer operations than those described in method 300, or operations different from those described in method 300.

[0067] In step 302, the input layer of the neural network architecture takes in an input speech signal 301, which may be a training speech signal from a training phase or an enrollment speech signal from an enrollment phase. The input layer performs various pre-processing operations (e.g., training speech signal, enrollment speech signal) on the input speech signal 301 before feeding it to various other layers of the neural network architecture. Pre-processing operations may include, for example, applying VAD operations, extracting low-level spectro-temporal features, and performing data transformation operations.

[0068] In some embodiments, the input layer performs various data augmentation operations during the training or enrollment phase. The data augmentation operations may generate or obtain specific training speech signals, including clean speech signals and noise samples. The server may receive or request clean speech signals from one or more corpus databases. The clean speech signals may include speech signals originating from various types of speech sources with diverse speaker-independent characteristics. The clean speech signals may be stored in a non-transitory storage medium accessible to the server or received via a network or other data source. The data augmentation operations may also receive simulated speech signals from one or more databases or generate simulated speech signals based on the clean speech signal or the input speech signal 301 by applying various forms of data augmentation to the input speech signal 301 or the clean speech signal. Examples of data augmentation techniques are described in U.S. Patent Application No. 17 / 155,851, which is incorporated by reference in its entirety.

[0069] In step 304, the server applies VAD to the input voice signal 301. The neural network layer of VAD detects the occurrence of an active window or a non-active window of the input voice signal 301. VAD parses the input voice signal 301 into an active part 303 and a non-active part 305. VAD includes a classification layer (e.g., a classifier) that the server trains separately or jointly with other layers of the neural network architecture. VAD can directly output the binary result (e.g., active, non-active) of a part of the input signal 301, or generate a likelihood value (e.g., probability) of each part of the input voice signal 301 that the server evaluates against a voice activity detection threshold for outputting the binary result of a given part. VAD generates a set of active parts 303 and a set of non-active parts 305 parsed by VAD from the input voice signal 301.

[0070] In step 306, the server extracts features from the set of active parts 303, and in step 308, the server extracts corresponding features from the set of non-active parts 305. The server extracts, for example, low-level features from parts 303, 305 and performs various preprocessing operations on parts 303, 305 of the input voice signal 301 to convert such features from the time-domain representation to the frequency-domain representation by performing a short-time Fourier transform (SFT) and / or a fast Fourier transform (FFT).

[0071] The server and neural network architecture of method 300 are configured to extract features related to the speaker-independent characteristics of the input voice signal 301, but in some embodiments, the server is further configured to extract features related to speaker-dependent characteristics that depend on the speech.

[0072] In step 310, the server applies a task-specific model to the utterance portion 303 to separately train a specific neural network for the utterance portion 303. The server receives an input audio signal 301 (e.g., a training audio signal, a registration audio signal) along with a label. The label indicates specific expected speaker-independent characteristics of the input audio signal 301, such as, among other speaker-independent characteristics, the expected classification, the expected features, the expected feature vector, the expected metadata, and the type or degree of degradation present in the input audio signal 301.

[0073] Each task-specific model includes one or more embedding extraction layers for modeling specific characteristics of the input audio signal 301. The embedding extraction layer generates a feature vector or an embedding based on the features extracted from the utterance portion 303. During the training phase (and optionally, during the registration phase), the classifier layer of the task-specific model determines the predicted classification of the utterance portion 303 based on the feature vector. The server executes a loss function to determine the level of error based on the difference between the predicted utterance output (e.g., the predicted feature vector, the predicted classification) and the expected utterance output (e.g., the expected feature vector, the expected classification) according to the label associated with the specific input audio signal 301. The loss function or other computational layers of the neural network architecture adjust the hyperparameters of the task-specific model until the level of error meets the threshold error degree.

[0074] During the registration phase, the task-specific model outputs a feature vector and / or a classification. The neural network architecture statistically or algorithmically combines the feature vectors generated for the utterance portion of the registration audio signal to extract the registered speaker-independent embedding of the utterance portion.

[0075] In step 312, the server applies the task-specific model for the non-speaking part 305 to the specific neural network of the non-speaking part 303 as well. The embedding extraction layer generates a feature vector or embedding based on the features extracted from the non-speaking part 305. During the training phase (and optionally, during the registration phase), the classifier layer of the task-specific model determines the predicted classification of the non-speaking part 305 based on the feature vector. The server executes a loss function to determine the level of error based on the difference between the predicted non-speaking output (e.g., the predicted feature vector, the predicted classification) and the expected non-speaking output (e.g., the expected feature vector, the expected classification) according to the label associated with the specific input audio signal 301. The loss function or other operation layer of the neural network architecture adjusts the hyperparameters of the task-specific model until the level of error meets the threshold error degree.

[0076] During the registration phase, the task-specific model outputs a feature vector and / or a classification. The neural network architecture statistically or algorithmically combines the feature vectors generated for the non-speaking parts of the registered audio signals to extract the registered speaker-independent embedding of the non-speaking parts.

[0077] In step 314, the server extracts features from the entire input audio signal 301. The server performs various preprocessing operations on the input audio signal 301, for example, to extract features from the input audio signal 301 and convert one or more of the extracted features from the time-domain representation to the frequency-domain representation by means of an SFT operation or an FFT operation.

[0078] In step 316, the server applies a task-specific model (e.g., a voice event neural network) to the entire input voice signal 301. The embedding extraction layer generates a feature vector based on the features extracted from the input voice signal 301. During the training phase (and optionally, during the registration phase), the classifier layer of the task-specific model determines the predicted classification of the input voice signal 301 based on the feature vector. The server executes a loss function to determine the level of error based on the difference between the predicted output (e.g., the predicted feature vector, the predicted classification) and the expected output (e.g., the expected feature vector, the expected classification) according to the label associated with a specific input voice signal 301. The loss function or other operation layers of the neural network architecture adjust the hyperparameters of the task-specific model (e.g., a voice event neural network) until the level of error meets the threshold error degree.

[0079] During the registration phase, the task-specific model of the current step 316 outputs the feature vector and / or classification. The neural network architecture statistically or algorithmically combines the feature vectors generated for all (or substantially all) of the registered voice signals to extract a speaker-independent embedding based on the feature vectors generated for the registered voice signals.

[0080] The neural network architecture further extracts the DP vector of the registered sound source. For each task-specific model that evaluates the speaking part 303 separately from the non-speaking part 305, the neural network architecture extracts the following pair of speaker-independent embeddings: a speaking embedding and a non-speaking embedding. Further, the neural network architecture extracts a single speaker-independent embedding for each task-specific model that evaluates the entire input voice signal 301. One or more post-modeling layers of the neural network architecture concatenate or algorithmically combine the speaker-independent embeddings to extract the registered DP embedding of the registered sound source.

[0081] Training and Registration of Multitask Configuration Figure 4 shows the execution steps of a multi-task learning method 400 for training a neural network architecture for speaker-independent embedding. Embodiments may include additional operations to those described in method 400, fewer operations than those described in method 400, or operations different from those described in method 400. Method 400 is implemented by a server that executes machine-readable software code of the neural network architecture, although various operations may be executed by one or more computing devices and / or processors.

[0082] In the multi-task learning method 400, instead of separately training and developing task-specific models (as in the signal task configuration of method 300 in FIG. 3), the server only trains two task-specific models (e.g., speech, non-speech) or three task-specific models (e.g., speech, non-speech, full audio signal) for multi-task learning. In the multi-task learning method 400, the neural network architecture includes a shared hidden layer (e.g., embedding extraction layer) that is shared by (and common to) multiple task-specific models such that the hidden layer generates feature vectors for given portions 403, 405 of the audio signal. Additionally, the input audio signal 401 (or its speech portion 403 and non-speech portion 405) is shared by the task-specific models. The neural network architecture includes a single final loss function that is a weighted sum of single-task losses. The shared hidden layer is followed by separate task-specific fully-connected (FC) layers and task-specific output layers. In the multi-task learning method 400, the server executes a single loss function across all task-specific models.

[0083] In step 402, the input layer of the neural network architecture captures an input audio signal 401 that includes a training audio signal in the training phase or an enrollment audio signal in the enrollment phase. The input layer performs various preprocessing operations (e.g., training audio signal, enrollment audio signal) on the input audio signal 401 before feeding the input audio signal 401 to various other layers of the neural network architecture. The preprocessing operations include, for example, applying a VAD operation, extracting various types of features (e.g., MFCC, metadata), and performing data conversion operations.

[0084] In some embodiments, the input layer performs various data augmentation operations during the training phase or the enrollment phase. The data augmentation operations may generate or obtain a specific training audio signal that includes a clean audio signal and noise samples. The server may receive or request a clean audio signal from one or more corpus databases. The clean audio signal may include audio signals from various types of audio sources having diverse speaker-independent characteristics. The clean audio signal may be stored in a non-transitory storage medium that is accessible to the server or received via a network or other data source. The data augmentation operations may further receive or generate a simulated audio signal from one or more databases based on the clean audio signal or the input audio signal 401 by applying various forms of data augmentation to the input audio signal 401 or the clean audio signal. Examples of data augmentation techniques are described in U.S. Patent Application No. 17 / 155,851, which is incorporated by reference in its entirety.

[0085] In step 404, the server applies a VAD to the input speech signal 401. The VAD's neural network layer detects the occurrence of speech and non-speech windows in the input speech signal 401. The VAD parses the input speech signal 401 into speech portions 403 and non-speech portions 405. The VAD includes a classification layer (e.g., a classifier) that the server trains separately or jointly with other layers of the neural network architecture. The VAD may directly output a binary result (e.g., speech, non-speech) for a portion of the input signal 401, or may generate an argument value (e.g., probability) for each portion of the input speech signal 401 that the server evaluates against a speech detection threshold to output a binary result for a given portion. From the input speech signal 401, the VAD generates a set of speech portions 403 and a set of non-speech portions 405 parsed by the VAD.

[0086] In step 406, the server extracts features from the set of speech portions 403, and in step 408, the server extracts corresponding features from the set of non-speech portions 405. The server performs various pre-processing operations on the portions 403, 405 of the input audio signal 401, for example to extract various types of features from the portions 403, 405 and to convert one or more extracted features from a time domain representation to a frequency domain representation by performing an SFT or FFT operation.

[0087] While the server and neural network architecture of method 400 are configured to extract features related to speaker-independent characteristics of the input speech signal 401, in some embodiments the server is further configured to extract features related to utterance-dependent speaker-dependent characteristics.

[0088] In step 410, the server applies the task-specific model to the utterance part 403 and jointly trains the neural network architecture for the utterance part 403. The task-specific models share a hidden layer for modeling and generating the joint feature vector based on the utterance part of a specific input audio signal 401. Each task-specific model is separate from the fully connected layer and output layer of other task-specific models and includes a fully connected layer and an output layer that independently affect a specific loss function. In the current step 410, the shared loss function evaluates the output related to the utterance part 403. For example, the loss function can be the sum, concatenation, or combination of other algorithms of several output layers that process the utterance part 403.

[0089] In particular, the server receives the input audio signal 401 (e.g., training audio signal, registration audio signal) together with one or more labels indicating specific aspects of the specific input audio signal 401, such as, among other aspects, the expected classification, expected features, expected feature vector, expected metadata, and the type or degree of degradation present in the input audio signal 401. The shared hidden layer (e.g., embedding extraction layer) generates a feature vector based on the features extracted from the utterance part 403. During the training phase (and, optionally, during the registration phase), the fully connected layer and output layer (e.g., classifier layer) of each task-specific model determine the predicted classification of the utterance part 403 based on the common feature vector generated by the shared hidden layer of the utterance part 403. The server executes a common loss function to determine the level of error according to the difference, for example, between one or more predicted utterance outputs (e.g., predicted feature vector, predicted classification) and one or more expected utterance outputs (e.g., expected feature vector, expected classification) indicated by the labels associated with the specific input audio signal 401. The shared loss function or other calculation layer of the neural network architecture adjusts one or more hyperparameters of the neural network architecture until the level of error meets the threshold error degree.

[0090] Similarly, in step 412, the server applies a shared hidden layer (e.g., an embedding extraction layer) and a task-specific model to the non-speech portion 405. The embedding extraction layer generates a non-speech feature vector based on the features extracted from the non-speech portion 405. During the training phase (and optionally, during the registration phase), the fully connected layer and the output layer (e.g., the classification layer) of each task-specific model generate a predicted non-speech output (e.g., a predicted feature vector, a predicted classification) of the non-speech portion 405 based on the non-speech feature vector generated by the shared hidden layer. The server executes a shared loss function to determine the level of error based on the difference between the predicted non-speech output (e.g., a predicted feature vector, a predicted classification) and the expected non-speech output (e.g., an expected feature vector, an expected classification) according to the label associated with the specific input audio signal 401. The shared loss function or other computational layers of the neural network architecture adjust one or more hyperparameters of the neural network architecture until the level of error meets a threshold error degree.

[0091] In step 414, the server extracts features from the entire input audio signal 401. The server performs various preprocessing operations on the input audio signal 401, for example, to extract various types of features and to convert specific extracted features from the time domain representation to the frequency domain representation by means of an SFT function or an FFT function.

[0092] In step 416, the server applies a hidden layer (e.g., an embedding extraction layer) to the features extracted in (step 414). The hidden layer is shared by specific task - specific models that evaluate the entire input audio signal 401 (e.g., an audio event neural network). The embedding extraction layer generates a full - audio feature vector based on the features extracted from the input audio signal 401. During the training phase (and optionally, during the registration phase), the fully - connected layer and the output layer (e.g., a classification layer) specific to each specific task - specific model generate the predicted full - audio output (e.g., the predicted classification, the predicted full - audio feature vector) according to the full - audio feature vector. The server executes a shared loss function to determine the level of error based on the difference between the predicted full - audio output (e.g., the predicted feature vector, the predicted classification) and the expected full - audio output (e.g., the expected feature vector, the expected classification) indicated by the label associated with the specific input audio signal 401. The shared loss function or other computational layers of the neural network architecture adjust one or more hyperparameters of the neural network architecture until the level of error meets a threshold error degree.

[0093] Speaker - independent embedding and extraction of DP vectors FIG. 5 shows the execution steps of a method 500 for applying a task - specific model of a neural network architecture and extracting speaker - independent embedding and DP vectors of an input audio signal. Embodiments may include additional operations, fewer operations than those described in method 500, or operations different from those described in method 500. Method 500 is implemented by a server that executes machine - readable software code of a neural network architecture, although various operations may be executed by one or more computing devices and / or processors.

[0094] In step 502, the server receives an input voice signal (e.g., a training voice signal, a registration voice signal, an inbound voice signal) from a specific sound source. The server may perform any number of preprocessing and / or data augmentation operations on the input voice signal before feeding the input voice signal into the neural network architecture.

[0095] In step 504, the server optionally applies VAD to the input voice signal. The neural network layer of VAD detects the occurrence of an active window or an inactive window of the input voice signal. VAD parses the input voice signal into an active part and an inactive part. VAD includes a classification layer (e.g., a classifier) that the server trains separately or jointly with other layers of the neural network architecture. VAD may directly output the binary result (e.g., active, inactive) of a part of the input signal, or generate a confidence value (e.g., probability) for each part of the input voice signal that the server evaluates against a voice activity detection threshold for outputting the binary result of a given part. VAD generates a set of active parts and a set of inactive parts parsed by VAD from the input voice signal.

[0096] In step 506, the server extracts various features from the input voice signal. The extracted features may include, among other data types, low-level spectral time features (e.g., MFCC) and communication metadata. These features are associated with the set of active parts, the set of inactive parts, and the entire voice signal.

[0097] In step 508, the server applies the neural network architecture to the input voice signal. In particular, the server applies the task-specific model to the features extracted from the active parts to generate a set of first speaker-independent feature vectors corresponding to each of the task-specific models. The server applies the task-specific model again to the features extracted from the inactive parts to generate a set of second speaker-independent feature vectors corresponding to each of the task-specific models. The server also applies the full voice task-specific model (e.g., a voice event neural network) to the entire input voice signal.

[0098] In a single-task configuration, each task-specific model includes a separate modeling layer for generating feature vectors. In a multi-task configuration, a shared hidden layer generates feature vectors and a shared loss layer tunes hyperparameters, but each task-specific model includes a separate set of post-modeling layers, such as a fully connected layer, a classification layer, and / or an output layer. In either configuration, the fully connected layer of each particular task-specific model statistically or algorithmically combines feature vectors to extract one or more corresponding speaker-independent embeddings. For example, a speech event neural network extracts a full speech embedding based on one or more full speech feature vectors, while a device neural network extracts an utterance embedding and then extracts a non-speech feature vector based on one or more utterance feature vectors and one or more non-speech feature vectors.

[0099] Then, after extracting speaker-independent embeddings from the input speech signal, the server extracts DP vectors for the speech source based on the extracted speaker-independent embeddings. In particular, the server algorithmically combines the extracted embeddings from each of the task-specific models to generate the DP vectors.

[0100] As shown in FIG. 5, the neural network architecture includes nine task-specific models that extract speaker-independent embeddings. The server applies an audio event neural network to the entire input audio signal to extract the corresponding speaker-independent embedding. The server also applies each of the remaining eight task-specific models to the speaking and non-speaking portions. The server extracts the following pair of speaker-independent embeddings output by each of the task-specific models: a first speaker-independent embedding of the speaking portion and a second speaker-independent embedding of the non-speaking portion. In this example, the server extracts eight of the first speaker-independent embeddings and eight of the second speaker-independent embeddings. The server then extracts the DP vector of the input audio signal by concatenating the speaker-independent embeddings. In the example of FIG. 5, the DP vector is generated by concatenating 17 speaker-independent embeddings output by nine task-specific models.

[0101] Additional embodiments Speaker-independent DP vectors complementary to voice biometrics A deep phone-printing system may be used for authentication operations of audio features and / or metadata associated with an audio signal. The DP vector or speaker-independent embedding may be evaluated as a complement to voice biometrics.

[0102] In a voice biometric system, in the training phase, VAD is used to preprocess the input audio signal to parse only the vocal part of the voice. A discriminable DNN model is trained on only the vocal part of the input audio signal using one or more speaker labels. In feature extraction, the audio signal is passed through VAD to extract one or more features of only the vocal part of the voice. The voice is taken as input to the DNN model to extract a speaker embedding vector (e.g., voiceprint) representing the speaker characteristics of the input audio signal.

[0103] In a DP system, in a training phase, VAD is used to process speech to parse both the voiced and unvoiced portions of an input speech signal. The voiced portion of the speech is used to train a set of DNN models using a set of metadata labels. Then, the unvoiced portion of the speech is used to train another set of DNN models using the same set of metadata labels. In some implementations, the metadata labels are independent of speaker labels. Because of these differences, the DP vectors generated by the DP system add complementary information regarding the speaker-independent aspect of speech to a voice biometric system. The DP vectors capture information that is complementary to the features extracted in the voice biometric system such that a neural network architecture can use the speaker embeddings (e.g., voiceprints) and the DP vectors in various voice biometric operations such as authentication. The fusion system described herein can be used in various downstream operations such as authenticating a sound source associated with an input speech signal or creating and enforcing a voice-based exclusion list.

[0104] FIG. 6 is a diagram showing the data flow between layers of a neural network architecture 600 that uses a DP vector as a complement to speaker-dependent embedding. The DP vector can be fused with the speaker-dependent embedding either at the "feature level" or at the "score level". In a feature-level fusion embodiment, the DP vector is concatenated with the speaker-dependent embedding (e.g., voiceprint) or otherwise algorithmically combined to generate a combined embedding. The concatenated combined feature vector is then used to train a machine learning classifier for a classification task and / or to use it in any number of downstream operations. The neural network architecture 600 is executed by a server during a training phase as well as an optional enrollment and deployment phase, but the neural network architecture 600 may be executed by any computing device including a processor capable of executing the operations of the neural network architecture 600 and by any number of such computing devices.

[0105] The audio input layer 602 receives one or more input audio signals (e.g., training audio signal, enrollment audio signal, inbound audio signal) and performs various preprocessing operations including applying VAD to the input audio signal and extracting various types of features. The VAD detects the speaking and non-speaking portions and outputs a voice signal with only speech and a voice signal with no speech only. The input layer 602 may also extract various types of features from the input audio signal, the voice signal with only speech, and / or the voice signal with no speech only.

[0106] The speaker-dependent embedding layer 604 is applied to the input audio signal to extract a speaker-dependent embedding (e.g., voiceprint). The speaker-dependent embedding layer 604 includes an embedding extraction layer for generating a feature vector according to the extracted features used to model the speaker-dependent aspect of the input audio signal. Examples of extracting such speaker-dependent embeddings are described in U.S. Patent Application Nos. 17 / 066,210 and 17 / 079,082, which are incorporated herein by reference.

[0107] Sequentially or simultaneously, the DP vector layer 606 is applied to the input audio signal. The DP vector layer includes task-specific models that extract various speaker-independent embeddings from the input audio signal. Next, the DP vector layer 606 extracts the DP vector of the input audio signal by concatenating (or otherwise combining) the various speaker-independent embeddings.

[0108] The neural network layer that defines the pooling classifier 608 is applied to the pooling embedding to output a classification or other output indicating a classification. The DP vector and the speaker-dependent embedding are concatenated or otherwise algorithmically combined to generate a pooling embedding. Then, the concatenated pooling embedding is used to train the pooling classifier for a classification task or any number of downstream operations that use a classification decision. During the training phase, the pooling classifier outputs the predicted outputs (e.g., predicted classifications, predicted feature vectors) of the training audio signals and uses these to determine the level of error according to the training labels indicating the expected outputs. Hyperparameters of the pooling classifier 608 and other layers of the neural network architecture 600 to minimize the level of error.

[0109] FIG. 7 is a diagram showing the data flow between the layers of a neural network architecture 700 that uses the DP vector as a complement to the speaker-dependent embedding. The neural network architecture 700 is executed by a server during the training phase as well as an optional enrollment and deployment phase, but the neural network architecture 700 may be executed by any computing device that includes a processor capable of executing the operations of the neural network architecture 700 and by any number of such computing devices.

[0110] In a score-level fusion embodiment, such as in Figure 7, one or more DP vectors are used to train a speaker-independent classifier 708, and speaker-dependent embeddings are used to train a speaker-dependent classifier 710. The two types of classifiers 708, 710 independently generate predicted classifications (e.g., predicted classification labels, predicted classification probabilities) for corresponding types of embeddings. A classification score fusion layer 712 algorithmically combines the predicted labels or predicted probabilities, for example, using an ensemble algorithm (e.g., a logistic regression model), to output a joint prediction.

[0111] The speech intake layer 702 receives one or more input speech signals (e.g., a training speech signal, an enrollment speech signal, an inbound speech signal) and performs various preprocessing operations, including applying a VAD to the input speech signals and extracting various types of features. The VAD detects speech and non-speech portions and outputs a speech-only speech signal and a no-speech-only speech signal. The intake layer 702 may also extract various types of features from the input speech signal, the speech-only speech signal, and / or the no-speech-only speech signal.

[0112] A speaker-dependent embedding layer 704 is applied to the input speech signal to extract a speaker-dependent embedding (e.g., a voiceprint). The speaker-dependent embedding layer 704 includes an embedding extraction layer for generating a feature vector according to the extracted features used to model speaker-dependent aspects of the input speech signal. Examples of extracting such speaker-dependent embeddings are described in U.S. Application Nos. 17 / 066,210 and 17 / 079,082, which are incorporated herein by reference.

[0113] Sequentially or simultaneously, a DP vector layer 706 is applied to the input speech signal. The DP vector layer includes task-specific models that extract various speaker-independent embeddings from the input speech signal. The DP vector layer 706 then extracts a DP vector for the input speech signal by concatenating (or otherwise combining) the various speaker-independent embeddings.

[0114] The neural network layer that defines the speaker-dependent classifier 708 is applied to the speaker-dependent embedding to output a speaker classification (e.g., genuine, spoofed) or other output indicating the speaker classification. During the training phase, the speaker-dependent classifier 708 generates predicted outputs (e.g., predicted classification, predicted feature vector, predicted DP vector) of the training speech signals and uses these to determine the level of error according to the training labels indicating the expected outputs. The hyperparameters of the speaker-dependent classifier 708 or other layers of the neural network architecture 700 are adjusted by a loss function or other function of the server to minimize the level of error of the speaker-dependent classifier 708.

[0115] The neural network layer that defines the speaker-independent classifier 710 is applied to the DP vector to output a speaker classification (e.g., genuine, spoofed) or other output indicating the speaker classification. During the training phase, the speaker-independent classifier 710 generates predicted outputs (e.g., predicted classification, predicted feature vector, predicted DP vector) of the training speech signals and uses these to determine the level of error according to the training labels indicating the expected outputs. The loss function or other function executed by the server adjusts the hyperparameters of the speaker-independent classifier 710 or other layers of the neural network architecture 700 to minimize the level of error of the speaker-independent classifier 710.

[0116] The server executes a classification score fusion operation 712 based on the output classification scores or decisions generated by the speaker-dependent classifier 708 and the speaker-independent classifier 710. In particular, each classifier 708, 710 independently generates its respective classification output (e.g., classification label, classification probability value). The classification score fusion operation 712 algorithmically combines the predicted outputs, for example, using an ensemble algorithm (e.g., logistic regression model), to output a consensus embedding.

[0117] Speaker-independent DP vectors of the exclusion list FIG. 8 is a diagram showing the data flow between layers of a neural network architecture 800 that uses DP vectors to authenticate sound sources according to an exclusion list and / or a permission list. The exclusion list can operate as a rejection list (sometimes called a "blacklist") and / or a permission list (sometimes called a "whitelist"). In the authentication operation, the exclusion list is used to exclude and / or permit specific sound sources. The neural network architecture 800 is executed by a server during a training phase and an optional registration and deployment phase, but the neural network architecture 800 may be executed by any computing device including a processor capable of executing the operations of the neural network architecture 800, and by any number of such computing devices.

[0118] In particular, the DP vector can be referenced to create and / or enforce the exclusion list. The DP vector is extracted from an input audio signal (e.g., a training audio signal, a registration audio signal, an inbound audio signal) for various operation phases. The DP vector is associated with a sound source that is part of the exclusion list and with audio associated with a normal population. The machine learning model is trained for the DP vector using labels. At test time, this model predicts new sample audio to determine whether the audio belongs to the exclusion list.

[0119] The audio input layer 802 receives one or more input audio signals (e.g., a training audio signal, a registration audio signal, an inbound audio signal) and performs various preprocessing operations including applying VAD to the input audio signal and extracting various types of features. The VAD detects the speaking part and the non-speaking part and outputs a voice signal with only speech and a voice signal with no speech only. The input layer 802 may also extract various types of features from the input audio signal, the voice signal with only speech, and / or the voice signal with no speech only.

[0120] The DP vector layer 804 is applied to the input audio signal. The DP vector layer includes task-specific models that extract various speaker-independent embeddings from the input audio signal. The DP vector layer 804 then extracts the DP vector of the input audio signal by concatenating (or otherwise combining) the various speaker-independent embeddings.

[0121] The exclusion list modeling layer 806 is applied to the DP vector. The exclusion list modeling layer 806 is trained to determine whether the sound source that emitted the input audio signal is on the exclusion list. The exclusion list modeling layer 806 determines a similarity score based on the relative distance (e.g., similarity, difference) between the DP vector of the input audio signal and the DP vectors of the sound sources in the exclusion list. During the training phase, the exclusion list modeling layer 806 generates the predicted outputs (e.g., predicted classifications, predicted similarity scores) of the training audio signals and uses these to determine the level of error according to the training labels indicating the expected outputs. The hyperparameters of the exclusion list modeling layer 806 or other layers of the neural network architecture 800 minimize the level of error of the exclusion list modeling layer 806.

[0122] As an example, the neural network architecture 800 uses an improper exclusion list applied to inbound sound sources. The training audio signals, X = {x1, x2, x3... x n} are associated with the corresponding training labels, Y = {y1, y2, y3,... y n}, where each label (y i ) indicates the classification {illegal, pure} of the corresponding training audio signal (x i ). The DP vectors (v i ) are extracted from the input audio signals (x i ) to obtain the corresponding DP vectors, V = {v1, v2, v3,... v nExtract {}. The modeling layer 806 is trained on the DP vectors using the labels. At test time, the trained modeling layer 806 determines a classification score based on a similarity score or likelihood score that the inbound audio signal is within a threshold distance of the DP vector of the inbound audio signal.

[0123] DP vectors for authentication Figure 9 is a diagram showing the data flow between the layers of a neural network architecture 900 that uses DP vectors to authenticate a device using a device identifier. The neural network architecture 900 is executed by a server during a training phase and an optional registration and deployment phase, but the neural network architecture 900 may be executed by any computing device including a processor capable of executing the operations of the neural network architecture 900, and by any number of such computing devices.

[0124] An authentication embodiment may register a sound source entity having a device ID or source ID, and the DP vectors used to register the sound source are stored in a database as stored DP vectors. At test time, each sound source associated with the device ID should be authenticated (or rejected) based on an algorithmic comparison between the inbound DP vector and the stored DP vector associated with the device ID of the inbound audio device.

[0125] The audio capture layer 902 receives one or more input audio signals (e.g., training audio signals, registration audio signals, inbound audio signals) and performs various preprocessing operations including applying VAD to the input audio signals and extracting various types of features. The VAD detects the talking and non-talking portions and outputs audio signals with only talking and audio signals with no talking only. The capture layer 902 may also extract various types of features from the input audio signals, audio signals with only talking, and / or audio signals with no talking only.

[0126] The voice input layer 902 includes identifying the input voice signal and the device identifier (device ID 903) associated with the sound source. The device ID 903 can be any identifier associated with the originating device or the sound source. For example, if the originating device is a phone, the device ID 903 can be the ANI or the phone number. If the originating device is an IoT device, the device ID 903 can be the MAC address, IP address, computer name, etc.

[0127] The voice input layer 902 also performs various preprocessing operations (e.g., feature extraction, VAD operation) on the input voice signal and outputs the preprocessed signal data 905. The preprocessed signal data 905 includes, for example, signals of speech only, signals of no speech only, and various types of extracted features.

[0128] The DP vector layer 904 is applied to the input voice signal. The DP vector layer includes a task-specific model that extracts various speaker-independent embeddings from substantially all of the preprocessed signal data 905 and the input voice signal. Then, the DP vector layer 904 extracts the DP vector of the input voice signal by concatenating (or otherwise combining) the various speaker-independent embeddings.

[0129] Additionally, the server uses the device ID 903 to query a database containing the stored DP vectors 906. The server identifies one or more stored DP vectors 906 in the database associated with the device ID 903.

[0130] The classifier layer 908 determines a similarity score based on the relative distance (e.g., similarity, difference) between the DP vectors extracted for the inbound voice signal and the stored DP vectors 906. In response to determining that the similarity score meets the authentication threshold score, the classifier layer 908 authenticates the originating device as a legitimate registered device. Any type of metric can be used to calculate the similarity between the stored vectors and the inbound DP vectors. Non-limiting examples of potential similarity calculations can include, among others, probabilistic linear discriminant analysis (PLDA) similarity, the reciprocal of the Euclidean distance, or the reciprocal of the Manhattan distance.

[0131] The server performs one or more authentication output operations 910 based on the similarity score generated by the classifier layer 908. The server can connect the originating device to a destination device (e.g., a provider server) or an agent phone, for example, in response to authenticating the originating device.

[0132] DP vectors for voice quality measurement DP vectors can be created by concatenating embeddings of different speaker-independent aspects of the voice signal. Speaker-independent aspects such as the microphone type used to capture the voice signal, the codec used to compress and decompress the voice signal, and the voice events (e.g., background noise present in the voice) provide useful information regarding the quality of the voice. The DP vectors can be used to represent voice quality measurements based on speaker-independent aspects and the voice quality measurements can be used in various downstream operations such as an authentication system.

[0133] For example, in an authentication system (e.g., a voice biometric authentication system), it is important to register speaker embeddings extracted from high-quality voice. If the voice signal contains high levels of noise, the quality of the speaker embedding can be adversely affected. Similarly, voice-related artifacts generated by the device and microphone can degrade the quality of the voice for registration embedding. Furthermore, high-energy voice events (e.g., overwhelming noise) can be present in the input voice signal such as a crying baby, barking dog, music, etc. These high-energy voice events are evaluated as voice parts by the VAD applied to the voice signal. This affects the quality of the speaker embedding extracted from the voice signal for registration purposes. If a poor-quality speaker embedding is registered and stored in the authentication system, the registered speaker embedding can adversely affect the performance of the entire authentication system. The quality of the voice can be characterized using a rule-based thresholding technique using DP vectors or by using a machine learning model as a binary classifier.

[0134] FIG. 10 is a diagram showing the data flow between layers of a neural network architecture 1000 that uses deep phone printing for dynamic registration in which a server determines whether to register based on voice quality. The neural network 1000 is executed by a server during a training phase as well as an optional registration phase and deployment phase, but the neural network architecture 1000 may be executed by any computing device having a processor capable of executing the operations of the neural network architecture 1000, and by any number of such computing devices.

[0135] The voice input layer 1002 receives one or more registered voice signals, applies VAD to the input voice signal, and extracts various types of features, performing various preprocessing operations. VAD detects the speaking part and the non-speaking part and outputs a voice signal with only speech and a voice signal with no speech only. The input layer 1002 may also extract various types of features from the v voice signal, the voice signal with only speech, and / or the voice signal with no speech only.

[0136] The DP vector layer 1006 is applied to the registered voice signal. The DP vector layer includes a task-specific model that extracts various speaker-independent embeddings from the registered voice signal. Then, the DP vector layer 1006 extracts the DP vector of the registered voice signal by concatenating (or otherwise combining) various speaker-independent embeddings. The DP vector represents the quality of the registered voice signal. The server applies a layer of a voice quality classifier, and the voice quality classifier generates a quality score and performs a binary classification determination as to whether to generate a registered speaker embedding based on the registered quality threshold using a specific registered voice signal.

[0137] In response to determining that the quality meets the registered quality threshold, the server performs various registration operations 1004 such as generating, storing, or updating the registered speaker embedding.

[0138] DP Vector for Replay Attack Detection FIG. 11 is a diagram showing the data flow of a neural network architecture 1100 for detecting replay attacks using a DP vector. The neural network architecture 1100 is executed by a server during a training phase as well as an optional enrollment and deployment phase, but the neural network 1100 may be executed by any computing device including a processor capable of executing the operations of the neural network architecture 1100, and by any number of such computing devices. The neural network architecture 1100 does not always need to execute the operations of the enrollment phase. Thus, in some embodiments, the neural network architecture 1100 includes a training phase and a deployment phase.

[0139] As a reliable solution for person authentication, an automatic speaker verification (ASV) system is commonly used. However, the ASV system is vulnerable to voice spoofing. Voice spoofing can be either local access (LA) or physical access (PA). The text-to-speech (TTS) function that generates synthetic speech and voice conversion is under the protection of LA. A replay attack is an example of PA spoofing. In the PA scenario, it is assumed that the speech data is captured by a microphone in a physical and reverberant space. A replay spoofing attack is assumed to be a recording of a genuine speech that is captured and then re-presented to the microphone of the ASV system using a replay device. The speech to be replayed is assumed to be first captured by a recording device before being replayed using a non-linear replay device.

[0140] The voice input layer 1102 receives one or more input voice signals (e.g., training voice signals, registration voice signals, inbound voice signals), applies VAD to the input voice signals, and executes various preprocessing operations including extracting various types of features. The VAD detects the speaking part and the non-speaking part, and outputs voice signals with only speech and voice signals with no speech only. The input layer 1102 may also extract various types of features from the input voice signal, the voice signal with only speech, and / or the voice signal with no speech only.

[0141] The DP vector layer 1104 is applied to the input voice signal. The DP vector layer includes a task-specific model that extracts various speaker-independent embeddings from the input voice signal. Then, the DP vector layer 1104 extracts the DP vector of the input voice signal by concatenating (or otherwise combining) various speaker-independent embeddings.

[0142] The device type and the microphone type are important speaker-independent aspects used by the deep fingerprinting system to detect replay attacks using the DP vector. For this reason, during the training phase, the DP vector is extracted from one or more corpora of training voice signals that also include simulated voice signals with various forms of degradation (e.g., additive noise, reverberation), which are used for training speaker embeddings. The DP vector layer 1104 is applied to the training voice signal including the simulated voice signal.

[0143] The replay attack detection classifier 1106 determines a detection score indicating the likelihood that the inbound voice signal is a replay attack. The binary classification model of the replay attack detection classifier 1106 is trained with training data and training labels. At test time, the replay attack detection classifier 1106 model outputs an indication of whether the neural network architecture 1100 detected a replay attack or a genuine voice signal in the inbound voice.

[0144] DP vectors for spoof detection In recent years, voice over IP (VoIP) services have risen due to their cost-effectiveness. VoIP services, such as Google Voice or Skype, use virtual phone numbers, also known as direct inward dialing (DID) or access numbers, which are phone numbers not directly associated with a telephone line. These phone numbers are also called gateway numbers. When a user makes a call through a VoIP service, the VoIP service may present a common gateway number or a number from a set of gateway numbers as a caller ID that is not necessarily unique to the device or caller. Some VoIP services allow loopholes, such as intentional caller ID / ANI spoofing, which allows callers to intentionally alter the information sent to the recipient's caller ID display to disguise their identity. There are several caller ID / ANI spoofing services available in the form of Android and iOS applications.

[0145] The deep phoneprinting system has a component for spoof service classification. Embeddings extracted from a DNN model trained for the ANI spoof service classification task can be used for spoof service recognition for new voices. These embeddings can also be used for spoof detection by considering all ANI spoof services as a spoofed class.

[0146] FIG. 12 is a diagram showing the data flow between layers of a neural network architecture 1200 that uses a DP vector to identify a spoofing service associated with an audio signal. The neural network architecture 1200 is executed by a server during a training phase as well as an optional enrollment and deployment phase, but the neural network architecture 1200 may be executed by any computing device that includes a processor capable of executing the operations of the neural network architecture 1200, and by any number of such computing devices. The neural network architecture 1200 does not always need to execute the operations of the enrollment phase. Thus, in some embodiments, the neural network architecture 1200 includes a training phase and a deployment phase.

[0147] The audio input layer 1202 receives one or more input audio signals (e.g., training audio signals, enrollment audio signals, inbound audio signals) and performs various preprocessing operations including applying VAD to the input audio signals and extracting various types of features. The VAD detects speech and non-speech portions and outputs speech-only audio signals and non-speech-only audio signals. The input layer 1202 may also extract various types of features from the input audio signals, speech-only audio signals, and / or non-speech-only audio signals.

[0148] The DP vector layer 1204 is applied to the input audio signals. The DP vector layer includes task-specific models that extract various speaker-independent embeddings from the input audio signals. The DP vector layer 1204 then extracts the DP vector of the input audio signal by concatenating (or otherwise combining) the various speaker-independent embeddings.

[0149] The spoof detection layer 1206 defines a binary classifier trained to determine a spoof detection likelihood score. The spoof detection layer 1206 is trained to determine whether a sound source is spoofing the input audio signal (spoof detection) or whether spoofing has not been detected. The spoof detection layer 1206 determines the likelihood score based on the relative distance (e.g., similarity, difference) between the DP vector of the input audio signal and the DP vector of the trained spoof detection classification or the stored spoof detection DP vector. During the training phase, the spoof detection layer 1206 generates the predicted outputs (e.g., predicted classification, predicted similarity score) of the training audio signals and uses these to determine the level of error according to the training labels indicating the expected outputs. A loss function or other server-executed process adjusts the hyperparameters of the spoof detection layer 1206 or other layers of the neural network architecture 1200 to minimize the level of error of the spoof detection layer 1206. The spoof detection layer 1206 determines whether a given inbound audio signal has a spoof detection likelihood score that meets a spoof detection threshold at test time.

[0150] When neural network architecture 1200 detects spoofing in an inbound voice signal, the neural network architecture applies a spoofing service recognition layer 1208. The spoofing service recognition layer 1208 is a multi-class classifier trained to determine likely spoofing services used to generate the spoofed inbound signal. The spoof service recognition layer 1208 determines a spoofing service likelihood score based on the relative distance (e.g., similarity, difference) between the DP vector of the input voice signal and the DP vectors of the trained spoof service recognition classification or the stored spoof service recognition DP vectors. During the training phase, the spoofing service recognition layer 1208 generates predicted outputs (e.g., predicted classifications, predicted similarity scores) of the training voice signals and uses these to determine the level of error according to training labels indicating the expected outputs. Hyperparameters of the spoofing service recognition layer 1208 or other layers of the neural network architecture 1200 minimize the level of error of the spoofing service recognition layer 1208. The spoofing service recognition layer 1208 determines the specific spoofing service applied to generate the spoofed inbound voice signal at test time when the spoofing service likelihood score meets the spoofing detection threshold.

[0151] FIG. 13 is a diagram showing the data flow between layers of a neural network architecture 1300 for identifying a spoofing service associated with an audio signal according to a label mismatch approach using a DP vector. The neural network architecture 1300 is executed by a server during a training phase as well as an optional enrollment and deployment phase, but the neural network architecture 1300 may be executed by any computing device including a processor capable of executing the operations of the neural network architecture 1300, and by any number of such computing devices. The neural network architecture 1300 does not always need to execute the operations of the enrollment phase. Thus, in some embodiments, the neural network architecture 1300 includes a training phase and a deployment phase.

[0152] The audio intake layer receives one or more input audio signals 1301 and performs various preprocessing operations including applying VAD to the input audio signals and extracting various types of features. The VAD detects speech and non-speech portions and outputs speech-only audio signals and non-speech-only audio signals. The intake layer may also extract various types of features from the input audio signals 1301, speech-only audio signals, and / or non-speech-only audio signals.

[0153] The DP vector layer 1302 is applied to the input audio signal 1301. The DP vector layer includes task-specific models that extract various speaker-independent embeddings from the input audio signal. The DP vector layer 1302 then extracts the DP vector of the input audio signal 1301 by concatenating (or otherwise combining) the various speaker-independent embeddings.

[0154] A deep phone printing system that executes a neural network architecture 1300 has components based on speaker-independent aspects such as, among other things, carrier, network type, and geography. Using DP vectors, one or more label classification models 1304, 1306, 1308 of speaker-independent embeddings can be trained for each of these speaker-independent aspects. At test time, the DP vectors extracted from the inbound audio signal are provided as input to the classification models 1304, 1306, 1308 to predict various types of labels such as, among other things, carrier, network type, and geography. For example, for an inbound audio signal received via a phone channel, using the phone number or ANI of the originating sound source, certain metadata such as, among other things, carrier, network type, or geography can be collected or determined. The classification models 1304, 1306, 1308 can detect a mismatch between the classification labels predicted based on the DP vectors and the metadata associated with the ANI. The output layers 1310, 1312, 1314 feed the corresponding classifications of the classification models 1304, 1306, 1308 as input to the spoof detection layer 1316.

[0155] In some implementations, the classification models 1304, 1306, 1308 can be the classifiers of the respective task-specific models or other types of neural network layers (e.g., fully connected layers). Similarly, in some implementations, the output layers 1310, 1312, 1314 can be the layers of the respective task-specific models.

[0156] The spoof detection layer 1316 defines a binary classifier trained to determine a spoof detection likelihood score based on classification scores or labels received from specific output layers 1310, 1312, 1314. The spoof detection layer 1316 is trained to determine whether a sound source is spoofing the input audio signal 1310 (spoofing detection), or whether spoofing has not been detected. The spoof detection layer 1316 determines a likelihood score based on the relative distance (e.g., similarity, difference) between the DP vector 1302 of the input audio signal 1301 and the DP vector of the trained spoof detection classification, score, cluster, or stored spoof detection DP vector. During the training phase, the spoof detection layer 1316 generates predicted outputs (e.g., predicted classifications, predicted similarity scores) of the training audio signals and uses these to determine the level of error according to the training labels indicating the expected outputs. A loss function (or other function executed by the server) adjusts the hyperparameters of the spoof detection layer 1316 or other layers of the neural network architecture 1300 to minimize the level of error of the spoof detection layer 1316. The spoof detection layer 1316 determines whether a given inbound audio signal 1301 has a spoof detection likelihood score that meets a spoof detection threshold at test time.

[0157] DP vector for speaker-independent feature classification FIG. 14 is a diagram showing the data flow between layers of a neural network architecture 1400 that uses a DP vector to determine the microphone type associated with an audio signal. The neural network architecture 1400 is executed by one or more computing devices such as a server using various other types of data such as various types of input audio signals (e.g., training, audio signals, enrollment audio signals, inbound audio signals) and training labels. The neural network architecture 1400 includes an audio capture layer, a DP vector layer 1404, and a classifier layer 1406.

[0158] The voice input layer 1402 captures the input voice signal and performs one or more preprocessing operations on the input voice signal. The DP vector layer 1404 extracts a plurality of speaker-independent embeddings for various types of speaker-independent characteristics and extracts the DP vector of the inbound voice signal based on the speaker-independent embeddings. The classifier 1406 generates a classification score (or other type of score) based on the DP vector of the inbound signal, as well as various types of stored data used for classification (e.g., stored embeddings, stored DP vectors), along with any number of additional outputs related to the classification result.

[0159] The DP vector can be used to recognize the type of microphone used to capture the input audio signal at the transmitting device. A microphone is a transducer that converts sound impulses into electrical signals. There are several types of microphones used for various purposes, such as dynamic microphones, condenser microphones, piezoelectric microphones, microelectromechanical system (MEMS) microphones, ribbon microphones, carbon microphones, etc. Microphones respond differently to sound based on direction. Microphones have different polar patterns, such as omnidirectional, bidirectional, and unidirectional, and have a proximity effect. The distance from the sound source to the microphone and the direction of the sound source significantly affect the response of the microphone. For telephone communication, during a voice call, the user can use different types of microphones, such as a conventional landline receiver, a speakerphone, a mobile phone microphone, a wired headset with a microphone, a wireless microphone, etc. For IoT devices (e.g., Alexa (registered trademark), Google Home (registered trademark)), a microphone array with direction detection, noise cancellation, and echo cancellation capabilities is used. The microphone array uses the beamforming principle to improve the quality of the captured audio. The neural network architecture 1400 can recognize or be trained to recognize the microphone type by evaluating the DP vector of the inbound audio signal against trained classifications, clusters, pre-stored embeddings, etc.

[0160] The DP vector can be used to recognize the type of device that captured the input voice signal. For telephone communication, there are several types of devices in use. There are several telephone device manufacturers that produce different types of telephones. For example, Apple manufactures the iPhone, and the iPhone has several models such as, to name a few, the iPhone 8, iPhone X, iPhone 11, etc., and Samsung manufactures the Samsung Galaxy series, etc.

[0161] Regarding IoT, for example, there are a variety of devices involved starting from voice-based assistant devices such as Amazon Echo with Alexa, Google Home, smartphones, smart refrigerators, smartwatches, smart fire alarms, smart door locks, smart bicycles, medical sensors, fitness trackers, smart security systems, etc. By recognizing the type of device on which the voice is captured, useful information regarding the voice is added.

[0162] The DP vector can be used to recognize the type of codec used to capture the input voice signal at the transmitting device. A voice codec is a device or computer program that can encode and decode a voice data stream or voice signal. The codec is used to compress the voice at one end before transmission or storage and then decompress the voice at the receiving end. There are several types of codecs used for telephone voice. For example, G.711, G.721, G729, Adaptive Multi-Rate (AMR), Enhanced Variable Rate Codec (EVRC) for CDMA networks, GSM used in GSM-based mobile phones, SILK used in Skype, Speex used in WhatsApp, VoIP apps, iLBC used in open-source VoIP apps, etc.

[0163] The DP vector can be used to recognize carriers (e.g., AT&T, Verizon Wireless, T-Mobile, Sprint) associated with the transmission of the input voice signal and / or associated with the transmitting device. A carrier is a component of a telecommunications system that transmits information such as voice signals and / or protocol metadata. Carrier networks distribute large amounts of data over long distances. Carrier systems typically use various forms of multiplexing to simultaneously transmit multiple communication channels over a shared medium.

[0164] The DP vector can be used to recognize the geography associated with the transmitting device. The DP vector can be used to recognize the geographical location of the sound source. The geographical location can be a broad classification such as a particular continent, country, state / province, city / county, etc., within the country or internationally.

[0165] The DP vector can be used to recognize the type of network associated with the transmission of the input voice signal and / or associated with the transmitting device. Generally, there are the following three classes of telephone networks: the Public Switched Telephone Network (PSTN), the cellular telephone network, and the Voice over Internet Protocol (VoIP) network. The PSTN is a traditional circuit-switched telephone communication system. Similar to the PSTN system, the cellular network has a circuit-switched core that is partly replaced by current IP links. These networks can deploy different technologies on the wireless interface, for example, but the core technology and protocol of the cellular network are the same as those of the PSTN network. Finally, the VoIP network operates over IP links and generally shares the same path as other Internet-based traffic.

[0166] The DP vector can be used to recognize voice events in an input voice signal. The voice signal encompasses several types of voice events. These voice events carry information regarding the daily environment and the physical events occurring within this environment. Recognizing such voice events and specific classes, and detecting the exact position of voice events in a voice stream, has downstream benefits provided by a deep phone printing system, such as multimedia search based on voice events, enabling context recognition in IoT devices such as mobile phones, cars, etc., and intelligent monitoring in security systems based on voice events. For example, during the training phase, a training voice signal can be received along with a label or other metadata indicator that indicates a specific voice event in the training voice signal. This voice event indicator is used to train a machine learning model for each voice event task according to a classifier, a fully connected layer, a regression algorithm, and / or a loss function that evaluates the distance between the predicted output of the training voice signal and the expected output indicated by the label.

[0167] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Those skilled in the art can implement the described functionality in various ways for each particular application, but such implementation decisions should not be construed as causing a departure from the scope of the present invention.

[0168] Embodiments implemented in computer software may be implemented in software, firmware, middleware, microcode, hardware description language, or any combination thereof. A code segment or machine-executable instruction may represent a procedure, function, subprogram, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. The information, arguments, parameters, data, etc. may be passed, transferred, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0169] The actual software code or dedicated control hardware used to implement these systems and methods is not limiting of the present invention. Thus, the operation and behavior of the systems and methods are described without reference to specific software code that can be designed such that software and control hardware implement the systems and methods based on the description herein.

[0170] When implemented in software, these functions may be stored as one or more instructions or codes on a non-transitory computer-readable or processor-readable storage medium. The steps of the methods or algorithms disclosed herein may be embodied in a processor-executable software module present on a computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable medium includes both a computer storage medium and a tangible storage medium that facilitate the transfer of a computer program from one place to another. A non-transitory processor-readable storage medium may be any available medium that can be accessed by a computer. By way of example and not limitation, such non-transitory processor-readable media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, or any other tangible storage medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer or a processor. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disk typically magnetically reproduces data, while disc optically reproduces data with a laser. Combinations of the above are also included within the scope of computer-readable media. Additionally, the operations of a method or algorithm may exist as one or any combination or set of codes and / or instructions on a non-transitory processor-readable medium and / or computer-readable medium incorporated in a computer program product.

[0171] The foregoing description of the disclosed embodiments is provided to enable a person of ordinary skill in the art to make or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments without departing from the spirit or scope of the invention. Accordingly, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.

[0172] Although various aspects and embodiments are disclosed, other aspects and embodiments are contemplated. The various aspects and embodiments disclosed are for purposes of illustration and not intended to be limiting, and the true scope and spirit are indicated by the following claims.

Claims

1. A computer-implemented method, comprising: applying, by a computer, a plurality of task-specific machine learning models to an inbound voice signal having one or more speaker-independent characteristics to extract a plurality of speaker-independent embeddings for the inbound voice signal; extracting, by the computer, a deep phone print (DP) vector of the inbound voice signal based on the plurality of speaker-independent embeddings extracted for the inbound voice signal; applying, by the computer, one or more post-modeling operations to the plurality of speaker-independent embeddings extracted for the inbound voice signal to generate one or more post-modeling outputs for the inbound voice signal.

2. extracting, by the computer, a plurality of features from a training voice signal, the plurality of features including at least one of spectral temporal features and metadata associated with the training voice signal; extracting, by the computer, the plurality of features from the inbound voice signal; the method of claim 1, further comprising.

3. wherein the metadata includes at least one of a microphone type used to capture the training voice signal, a device type from which the training voice signal was emitted, a codec type applied for compression and decompression of the training voice signal for transmission, a carrier on which the training voice signal was emitted, a spoofing service used to change a source identifier associated with the training voice signal, a geography associated with the training voice signal, a network type from which the training voice signal was emitted, and a voice event indicator associated with the training voice signal; wherein one or more classifications of the one or more post-modeling outputs include at least one of the microphone type, the device type, the codec type, the carrier, the spoofing service, the geography, the network type, and the voice event indicator; the method of claim 2.

4. The method of claim 1, wherein the one or more post-modeling operations include at least one of a classification operation and a regression operation.

5. The method according to claim 1, wherein the task-specific machine learning model includes at least one of a neural network and a Gaussian mixture model.

6. The method according to claim 1, further comprising training the plurality of task-specific machine learning models by applying each of the plurality of task-specific machine learning models to a plurality of training speech signals having the one or more speaker-independent characteristics by the computer.

7. Training the task-specific machine learning model comprises: extracting, by the computer, one or more first spectral time features of the speaking part of the training speech signal and one or more second spectral time features of the non-speaking part of the training speech signal; applying, by the computer, one or more modeling layers of the task-specific machine learning model to the one or more first spectral time features and the one or more second spectral time features to generate a first speaker-independent embedding and a second speaker-independent embedding; generating, by the computer, predicted output data of the training speech signal by applying one or more post-modeling layers of the task-specific machine learning model to the first speaker-independent embedding and the second speaker-independent embedding; adjusting, by the computer, one or more hyperparameters of the one or more modeling layers of the task-specific machine learning model by executing a loss function using the predicted output data and the expected output data indicated by one or more labels associated with the training speech signal, the method according to claim 6.

8. Training the task-specific machine learning model comprises: extracting, by the computer, one or more first spectral time features of the speaking part of the training speech signal and one or more second spectral time features of the non-speaking part of the training speech signal; applying, by the computer, one or more shared modeling layers of one or more task-specific machine learning models to the one or more first spectral time features and the one or more second spectral time features to generate a first speaker-independent embedding and a second speaker-independent embedding; The computer generates predicted output data of the training audio signal by applying one or more post-modeling layers of the task-specific machine learning model to the first speaker-independent embedding and to the second speaker-independent embedding. The method according to claim 6, further comprising: the computer adjusting one or more hyperparameters of the task-specific machine learning model by executing a shared loss function of the one or more task-specific machine learning models using the predicted output data and the expected output data indicated by one or more labels associated with the training audio signal. **Claim 9** Training the plurality of task-specific machine learning models comprises: The computer extracting one or more spectral temporal features of a substantial portion of the training audio signal. The computer applies a task-specific machine learning model to the one or more spectral temporal features of the training audio signal to generate a speaker-independent embedding. The computer generates predicted output data of the training audio signal by applying one or more post-modeling layers of the task-specific machine learning model to the speaker-independent embedding of the substantial portion of the training audio signal. The method according to claim 6, further comprising: the computer adjusting one or more hyperparameters of the task-specific machine learning model by executing a loss function using the predicted output data and the expected output data indicated by one or more labels associated with the training audio signal. **Claim 10** The method according to claim 1, wherein the task-specific machine learning model includes at least one of a convolutional neural network, a recurrent neural network, and a fully-connected neural network. **Claim 11** The computer performs a voice activity detection (VAD) operation on the training audio signal, thereby generating one or more speech portions of the training audio signal and one or more non-speech portions of the training audio signal. The computer further performs the VAD operation on the inbound voice signal, thereby generating the one or more speech portions of the inbound voice signal and the one or more non-speech portions of the inbound voice signal, the method according to claim 1.

12. extracting, by the computer, a plurality of registered speaker-independent embeddings and DP vectors of a registered sound source by applying the plurality of task-specific machine learning models to a plurality of registered voice signals associated with the registered sound source; the computer further extracting the DP vector of the registered sound source based on the plurality of registered speaker-independent embeddings; The method according to claim 1, wherein the computer generates the one or more post-modeling outputs using the DP vector of the registered sound source and the DP vector of the inbound voice signal.

13. The method according to claim 12, further comprising the computer determining an authentication classification of the inbound voice signal based on one or more similarity scores between the DP vector of the registered sound source and the DP vector of the inbound voice signal.

14. A system comprising: a database comprising a non-transitory memory configured to store a plurality of voice signals having one or more speaker-independent characteristics; a server comprising a processor, wherein the processor is applying a plurality of task-specific machine learning models to an inbound voice signal having the one or more speaker-independent characteristics to extract a plurality of speaker-independent embeddings for the inbound voice signal; extracting a deep phone print (DP) vector of the inbound voice signal based on the plurality of speaker-independent embeddings extracted for the inbound voice signal; applying one or more post-modeling operations to the plurality of speaker-independent embeddings extracted for the inbound voice signal to generate one or more post-modeling outputs for the inbound voice signal, a system.

15. The server Extracting a plurality of features from a training audio signal, wherein the plurality of features includes at least one of spectral time features and metadata associated with the training audio signal; The system according to claim 14, further configured to extract the plurality of features from the inbound audio signal. **Claim 16** The metadata includes at least one of a microphone type used to capture the training audio signal, a device type from which the training audio signal was emitted, a codec type applied for compression and decompression of the training audio signal for transmission, a carrier on which the training audio signal was emitted, a spoofing service used to change a source identifier associated with the training audio signal, a geography associated with the training audio signal, a network type from which the training audio signal was emitted, and an audio event indicator associated with the training audio signal; The system according to claim 15, wherein one or more classifications of the one or more post-modeling outputs includes at least one of the microphone type, the device type, the codec type, the carrier, the spoofing service, the geography, the network type, and the audio event indicator. **Claim 17** The system according to claim 14, wherein the one or more post-modeling operations includes at least one of a classification operation and a regression operation. **Claim 18** The system according to claim 14, wherein the task-specific machine learning model includes at least one of a neural network and a Gaussian mixture model. **Claim 19** The system according to claim 14, wherein the server is further configured to train the plurality of task-specific machine learning models by applying each of the plurality of task-specific machine learning models to a plurality of training audio signals having the one or more speaker-independent characteristics. **Claim 20** When training a task-specific machine learning model, the server extracts one or more first spectral time features of a speaking portion of a training audio signal and one or more second spectral time features of a non-speaking portion of the training audio signal; Apply one or more modeling layers of the task-specific machine learning model to the one or more first spectral time features and the one or more second spectral time features to generate a first speaker-independent embedding and a second speaker-independent embedding; Generate predicted output data of the training audio signal by applying one or more post-modeling layers of the task-specific machine learning model to the first speaker-independent embedding and to the second speaker-independent embedding; The system according to claim 19, further configured to adjust one or more hyperparameters of the one or more modeling layers of the task-specific machine learning model by executing a loss function using the predicted output data and the expected output data indicated by one or more labels associated with the training audio signal.

21. When training a task-specific machine learning model, the server Extract one or more first spectral time features of the speaking portion of the training audio signal and one or more second spectral time features of the non-speaking portion of the training audio signal; Apply one or more shared modeling layers of one or more task-specific machine learning models over a plurality of features to the one or more first spectral time features and the one or more second spectral time features to generate a first speaker-independent embedding and a second speaker-independent embedding; Generate predicted output data of the training audio signal by applying one or more post-modeling layers of the task-specific machine learning model to the first speaker-independent embedding and to the second speaker-independent embedding; The system according to claim 19, further configured to adjust one or more hyperparameters of the task-specific machine learning model by executing a shared loss function of the task-specific machine learning model using the predicted output data and the expected output data indicated by one or more labels associated with the training audio signal.

22. When training the plurality of task-specific machine learning models, the server One or more spectral time features of a substantial portion of the training audio signal, Applying a task-specific machine learning model to the one or more spectral time features of the training audio signal to generate speaker-independent embeddings; Generating predicted output data of the training audio signal by applying one or more post-modeling layers of the task-specific machine learning model to the speaker-independent embeddings of the corresponding portion of the training audio signal; Further configured to perform adjusting one or more hyperparameters of the task-specific machine learning model by executing a loss function using the predicted output data and the expected output data indicated by one or more labels associated with the training audio signal, the system according to claim 19.

23. The system according to claim 14, wherein the task-specific machine learning model includes at least one of a convolutional neural network, a recurrent neural network, and a fully-connected neural network.

24. The server is Performing a voice activity detection (VAD) operation on the training audio signal, thereby generating one or more speech portions of the training audio signal and one or more non-speech portions of the training audio signal; Further configured to perform the VAD operation on the inbound audio signal, thereby generating one or more speech portions of the inbound audio signal and one or more non-speech portions of the inbound audio signal, the system according to claim 14.

25. The server is Extracting a plurality of registered speaker-independent embeddings and DP vectors of registered sound sources by applying the plurality of task-specific machine learning models to a plurality of registered audio signals associated with the registered sound sources; Extracting the DP vector of the registered sound source based on the plurality of registered speaker-independent embeddings; and The server generates the one or more post-modeling outputs using the DP vector of the registered sound source and the DP vector of the inbound audio signal, the system according to claim 14.

26. The server is The system according to claim 25, further configured to determine an authentication classification of the inbound voice signal based on one or more similarity scores between the DP vector of the registered sound source and the DP vector of the inbound voice signal.

Citation Information

Patent Citations

  • Custom acoustic models

    JP2019211752A

  • Fraud detection

    US20150269941A1

  • Method and apparatus for detecting spoofing conditions

    US20180254046A1

  • Adaptive Convolutional Neural Knowledge Graph Learning System Leveraging Entity Descriptions

    US20190122111A1

  • Replay spoofing detection for automatic speaker verification system

    US20200053118A1