Self-supervised speech representation for fake audio detection
A self-supervised model and shallow classifier system effectively detect synthetic speech in audio data by using human-derived training data and a neural network to generate feature vectors, addressing the challenge of limited synthetic speech datasets and maintaining accuracy on user devices with limited resources.
Patent Information
- Application Number
- JP2023533746
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-02
- Filing Date
- 2021-11-11
- Publication Date
- 2025-11-12
- Estimated Expiration
- 2041-11-11
AI Technical Summary
Existing methods for speaker verification and authentication systems fail to accurately distinguish between human and synthetic speech, existing technologies have not addressed or effectively solve the technical problem of existing technologies have not addressed the technical problem of detecting synthetic speech in audio data based on a self-supervised model that extracts audio features from the audio data and a shallow classifier model that determines the probability of synthetic speech being present in the audio features and therefore in the audio data.
A self-supervised model trained exclusively on human-derived speech and a shallow classifier model trained on a smaller number of synthetic speech samples are used to detect synthetic speech in audio data, utilizing a neural network to generate audio feature vectors and determine the presence of synthetic speech through a score that satisfies a detection threshold.
The system effectively identifies synthetic speech even when interspersed with human-derived speech, maintaining high accuracy despite limited training data, and can be implemented on user devices with limited computational resources.
Smart Images

Figure 0007768989000001 
Figure 0007768989000002 
Figure 0007768989000003
Abstract
Description
[Technical Field]
[0001] This disclosure relates to self-supervised speech representations for the detection of fake or synthetic audio. [Background technology]
[0002] In a voice-enabled environment (e.g., at home, at work, at school, in an automobile, etc.), a user can speak queries or commands to a computer-based system, which then responds and replies to the query and / or performs a function based on the command. For example, a voice-enabled environment is implemented using a network of connected microphone devices distributed throughout various rooms or areas of the environment. As these environments become more prevalent and as voice recognition devices become more sophisticated, voice is increasingly used for critical functions including, for example, speaker identification and authentication. These functions greatly increase the need to ensure that voice is human-derived and not synthetic (i.e., digitally created or modified and played through a speaker). Summary of the Invention [Means for solving the problem]
[0003] One aspect of the present disclosure provides a method for classifying whether audio data contains synthetic speech. The method includes receiving, in data processing hardware, audio data characterizing speech within the audio data acquired by a user device. The method also includes generating, by the data processing hardware, a plurality of audio feature vectors, each representing audio features of a portion of the audio data, using a trained self-supervised model. The method also includes generating, by the data processing hardware, a score indicative of the presence of synthetic speech within the audio data based on the corresponding audio feature of each audio feature vector of the plurality of audio feature vectors, using a shallow discriminator model. The method also includes determining, by the data processing hardware, whether the score satisfies a synthetic speech detection threshold. The method also includes determining, by the data processing hardware, that speech within the audio data acquired by the user device likely contains synthetic speech when the score satisfies the synthetic speech detection threshold.
[0004] Implementations of the present disclosure may include one or more of the following optional features: In some implementations, the shallow classifier model includes an intelligent pooling layer. In some examples, the method further includes generating, by data processing hardware, a single final audio feature vector based on each audio feature vector of the plurality of audio feature vectors using the intelligent pooling layer of the shallow classifier model. Generating a score indicative of the presence of synthetic speech in the audio data may be based on the single final audio feature vector.
[0005] Optionally, the single final audio feature vector comprises an averaging of each audio feature vector of the multiple audio feature vectors. Alternatively, the single final audio feature vector comprises a sum of each audio feature vector of the multiple audio feature vectors. The shallow classifier model may include a fully connected layer configured to receive the single final audio feature vector as an input and to generate a score as an output.
[0006] In some implementations, the shallow classifier model includes one of a logistic regression model, a linear discriminant analysis model, or a random forest model. In some examples, the trained self-supervised model is trained on a first training dataset including only training samples of human-derived speech. The shallow classifier model can be trained on a second training dataset including training samples of synthetic speech. The second training dataset can be smaller than the first training dataset. Optionally, the data processing hardware resides on the user device. The trained self-supervised model can include a representation model derived from a larger trained self-supervised model.
[0007] Another aspect of the present disclosure provides a system for classifying whether audio data contains synthetic speech. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations. The operations include receiving audio data characterizing speech in the audio data acquired by a user device. The operations also include generating, using a trained self-supervised model, a plurality of audio feature vectors, each representing audio features of a portion of the audio data. The operations also include generating, using a shallow classifier model, a score indicative of the presence of synthetic speech in the audio data based on a corresponding audio feature of each audio feature vector of the plurality of audio feature vectors. The operations also include determining whether the score satisfies a synthetic speech detection threshold. The operations also include determining that the speech in the audio data acquired by the user device likely contains synthetic speech when the score satisfies the synthetic speech detection threshold.
[0008] This aspect may include one or more of the following optional features: In some implementations, the shallow classifier model includes an intelligent pooling layer. In some examples, the operations further include generating a single final audio feature vector based on each audio feature vector of the plurality of audio feature vectors using the intelligent pooling layer of the shallow classifier model. Generating a score indicative of the presence of synthetic speech in the audio data may be based on the single final audio feature vector.
[0009] Optionally, the single final audio feature vector comprises an average of each of the multiple audio feature vectors. Alternatively, the single final audio feature vector comprises a sum of each of the multiple audio feature vectors. The shallow classifier model may include a fully connected layer configured to receive the single final audio feature vector as an input and to generate a score as an output.
[0010] In some implementations, the shallow classifier model includes one of a logistic regression model, a linear discriminant analysis model, or a random forest model. In some examples, the trained self-supervised model is trained on a first training dataset including only training samples of human-derived speech. The shallow classifier model can be trained on a second training dataset including training samples of synthetic speech. The second training dataset can be smaller than the first training dataset. Optionally, the data processing hardware resides on the user device. The trained self-supervised model can include a representation model derived from a larger trained self-supervised model.
[0011] The details of one or more implementations of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will be apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a schematic diagram of an exemplary system for classifying audio data as synthetic speech. [Figure 2] FIG. 1 is a schematic diagram of exemplary components of an audio feature extractor and a synthetic speech detector. [Figure 3] FIG. 3 is a schematic diagram of the synthetic speech detector of FIG. 2; [Figure 4A]FIG. 3 is a schematic diagram of a training architecture for the audio feature extractor of FIG. [Figure 4B] FIG. 3 is a schematic diagram of a training architecture for the synthetic speech detector of FIG. 2. [Figure 5] FIG. 1 is a schematic diagram of an audio feature extractor that provides extracted audio features to multiple shallow classifier models. [Figure 6] 1 is a flowchart of an exemplary sequence of operations for classifying audio data as synthetic speech. [Figure 7] FIG. 1 is a schematic diagram of an exemplary computing device that can be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION
[0013] Like reference symbols in the various drawings indicate like elements.
[0014] As voice-enabled environments and devices become more commonplace and sophisticated, confidence in the use of audio as a reliable indicator of human-origin speech is becoming increasingly important. For example, voice biometrics are commonly used for speaker verification. Automatic speaker verification (ASV) authenticates individuals by performing analysis on their spoken utterances. However, with the emergence of synthetic media (e.g., "deepfakes"), it is crucial for these systems to accurately determine when a spoken utterance contains synthetic speech (i.e., computer-generated audio output that resembles a human voice). For example, state-of-the-art text-to-speech (TTS) and voice conversion (VC) systems can now closely mimic human speakers, providing a means to attack and deceive ASV systems.
[0015] In one example, an ASV system implementing a speaker verification model is used in conjunction with a hotword detection model to enable an authorized user to speak a predefined fixed phrase (e.g., a hotword, wake word, keyword, activation phrase, etc.) to activate and wake up a voice-enabled device and process subsequent verbal input from the user. In this example, the hotword detection model is configured to detect audio features in audio data that characterize the predefined fixed phrase, and the speaker verification model is configured to verify whether the audio features characterizing the predefined fixed phrase were spoken by the authorized user. Generally, the speaker verification model extracts a matched speaker embedding from the input audio features and compares the matched speaker embedding to a reference speaker embedding for the authorized user. In this case, the reference speaker embedding may be previously obtained by having a particular user speak the same predefined fixed phrase (e.g., during an enrollment process) and stored as part of a user profile for the authorized user. When the verification speaker embedding matches the reference speaker embedding, the detected hot word in the audio data is confirmed as having been spoken by an authorized user, thereby allowing the voice-enabled device to wake up and process subsequent speech spoken by the authorized user. Using the state-of-the-art TTS and VC systems described above, it is possible to generate synthetic speech representations of pre-defined fixed phrases in the voice of an authorized user, thereby fooling the speaker verification model into confirming that the synthetic speech representations were spoken by the authorized user.
[0016] Machine learning (ML) algorithms, such as neural networks, have largely driven the proliferation of ASV systems and other voice-enabled technologies. However, these algorithms traditionally require vast amounts of training samples, with the main bottleneck in training accurate models often being the lack of sufficiently large, high-quality datasets. For example, while large datasets containing human-derived speech are readily available, similar datasets containing synthetic speech are not. Therefore, training a model that can accurately identify synthetic speech without a traditional training set poses a significant challenge for the development of synthetic speech detection systems.
[0017] Implementations herein are directed to detecting synthetic speech in audio data based on a self-supervised model that extracts audio features from the audio data and a shallow classifier model that determines the probability of synthetic speech being present in the audio features and therefore in the audio data. The self-supervised model is trained exclusively on data containing human-derived speech rather than synthetic speech, and is therefore able to avoid bottlenecks resulting from a lack of sufficient synthetic speech samples. On the other hand, the shallow classifier model can maintain high accuracy even when trained on a smaller number of training samples containing synthetic speech (compared to the self-supervised model).
[0018] 1 , in some implementations, an exemplary system 100 includes a user device 102. The user device 102 may correspond to a computing device such as a mobile phone, a computer (laptop or desktop), a tablet, a smart speaker / display, a smart appliance, smart headphones, a wearable, a vehicle entertainment system, etc., and is equipped with data processing hardware 103 and memory hardware 105. The user device 102 includes or is in communication with one or more microphones 106 for capturing speech from an audio source 10. The audio source 10 can be a person emitting human-derived speech 119, or can be an audio device (e.g., a loudspeaker) that converts an electrical audio signal into corresponding speech 119. The loudspeaker can be part of or in communication with any type of computing or user device (e.g., a mobile phone, a computer, etc.).
[0019] The user device 102 includes an audio feature extractor 210 configured to extract audio features from audio data 120 that characterize speech captured by the user device 102. For example, the audio data 120 is captured by the user device 102 from streaming audio 118. In other examples, the user device 102 generates the audio data 120. In some implementations, the audio feature extractor 210 includes a trained neural network (e.g., a stored neural network such as a convolutional neural network) received over the network 104 from a remote system 110. The remote system 110 can be a single computer, multiple computers, or a distributed system (e.g., a cloud environment) with scalable / elastic computing resources 112 (e.g., data processing hardware) and / or storage resources 114 (e.g., memory hardware).
[0020] In some examples, the audio feature extractor 210 running on the user device 102 is a self-supervised model. That is, the audio feature extractor 210 is trained using self-supervised learning (also called "unsupervised learning"), where the labels are necessarily part of the training samples and do not involve separate external labels. More specifically, in self-supervised learning methods, the model looks for patterns in a dataset without any pre-existing labels (i.e., annotations) and with minimal human supervision.
[0021] In the illustrated example, audio source 10 emits utterance 119, including the sound "My name is Jane Smith." Audio feature extractor 210 receives audio data 120 characterizing the utterance 119 in streaming audio 118 and generates multiple audio feature vectors 212, 212a-n, from the audio data 120. Each audio feature vector 212 represents audio features (i.e., audio characteristics such as spectrograms (e.g., Mel-frequency spectrograms and Mel-frequency cepstral coefficients (MFCCs)) of a chunk or portion of audio data 120 (i.e., a portion of streaming audio 118 or utterance 119). For example, each audio feature vector represents features for a 960-millisecond portion of audio data 120. The portions may overlap. As an example, for five seconds of audio data 120, audio feature extractor 210 generates eight audio feature vectors 212 (each representing 960 milliseconds of audio data 120). The audio feature vectors 212 from the audio feature extractor 210 capture a number of acoustic characteristics of the audio data 120 based on self-supervised learning.
[0022] After generating the audio feature vectors 212, the audio feature extractor 210 sends the audio feature vectors 212 to a synthetic speech detector 220 that includes a shallow classifier model 222. As discussed in more detail below, the shallow classifier model 222 is a shallow neural network (i.e., with few or no hidden layers) that generates, based on each of the audio feature vectors 212, a score 224 ( FIG. 2 ) indicative of the presence of synthetic speech in the streaming audio 118, based on the corresponding audio features of each audio feature vector 212. The synthetic speech detector 220 determines whether the score 224 (e.g., a probability score) satisfies a synthetic speech detection threshold. When the score 224 satisfies the synthetic speech detection threshold, the synthetic speech detector 220 determines that the speech (i.e., the speech 119) in the streaming audio 118 captured by the user device 102 contains synthetic speech. The synthetic speech detector 220 can determine that the utterance 119 contains synthetic speech even when the majority of the utterance 119 contains human-derived speech (i.e., small portions of the synthetic speech are interspersed or interspersed with human-derived speech).
[0023] In some implementations, the synthetic speech detector 220 generates an indicator 150 to the user device 102 to indicate whether the streaming audio 118 includes synthetic speech based on whether the score 224 satisfies a synthetic speech detection threshold. For example, when the score 224 satisfies the synthetic speech detection threshold, the indicator 150 indicates that the utterance 119 includes synthetic speech. In response, the user device 102 can generate a notification 160 to a user of the user device 102. For example, the user device 102 executes a graphical user interface (GUI) 108 for display on a screen of the user device 102 that is in communication with the data processing hardware 103. The user device 102 can render the notification 160 within the GUI 108. In this case, the indicator 150 indicates that the streaming audio 118 included synthetic speech by rendering a message on the GUI 108 that reads, "Notification: Synthetic speech detected." The displaying notification 160 is merely exemplary, and the user device 102 may notify the user of the user device 102 in any other suitable manner. Additionally or alternatively, the synthetic speech detector 220 notifies other applications running on the user device 102. For example, an application running on the user device 102 may authenticate a user of the user device 102 to allow the user access to one or more restricted resources. The application may authenticate the user using biometric voice (e.g., via the utterance 119). The synthetic speech detector 220 may provide an indicator 150 to the application to alert the application that the utterance 119 included synthetic speech, so that the application may deny authentication to the user.In another scenario, when an utterance 119 contains a hot word that is detected by the user device 102 in the streaming audio 118 to trigger the user device 102 to wake up from a sleep state and begin processing subsequent audio, an indicator 150 generated by the synthetic speech detector 220 indicating that the hot word utterance 119 contains synthetic speech can suppress the wake-up process on the user device 102.
[0024] The user device 102 can forward the indicator 150 to the remote system 110 over the network 104. In some implementations, the remote system 110 executes the audio feature extractor 210 and / or the synthetic speech detector 220 instead of or in addition to the user device 102. For example, the user device 102 receives streaming audio 118 and forwards the audio data 120 (or some features of the audio data 120) to the remote system for processing. The remote system 110 may include significantly more computational resources than the user device 102. Additionally or alternatively, the remote system 110 may be more secure from potential adversaries. In this scenario, the remote system 110 can send the indicator 150 to the user device 102. In some examples, the remote server 110 performs multiple authentication operations using the audio data 120 and returns a value indicating whether the authentication was successful. In other implementations, audio source 10 transmits audio data 120 of streaming audio 118 directly to remote system 110 (e.g., over network 104) without any separate user device 102. For example, remote system 110 runs an application that uses voice biometrics. In this case, audio source 10 includes a device that transmits audio data 120 directly to remote system 110. For example, audio source 10 is a computer that generates synthesized speech and transmits the synthesized speech to remote system 110 (via audio data 120) without the synthesized speech being verbalized.
[0025] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.
[0026] Referring now to FIG. 2 , a schematic diagram 200 includes an audio feature extractor 210 running a deep neural network 250. The deep neural network 250 may include any number of hidden layers configured to receive audio data 120. In some implementations, the deep neural network 250 of the audio feature extractor 210 generates multiple audio feature vectors 212, 212a-n (i.e., embeddings) from the audio data 120. A shallow classifier model 222 receives the multiple audio feature vectors 212 simultaneously, sequentially, or concatenated with one another. The multiple audio feature vectors 212 may undergo some processing between the audio feature extractor 210 and the shallow classifier model 222. The shallow classifier model 222 may generate a score 224 based on the multiple audio feature vectors 212 generated / extracted by the deep neural network 250 of the audio feature extractor 210.
[0027] 3 , in some examples, the shallow classifier model 222 includes an intelligent pooling layer 310, 310P. The intelligent pooling layer 310P may receive multiple audio feature vectors 212 and generate a single final audio feature vector 212F based on each audio feature vector 212 received from the audio feature extractor 210. The shallow classifier model 222 may generate a score 224 indicative of the presence of synthetic speech in the streaming audio 118 based on the single final audio feature vector 212F. In some examples, the intelligent pooling layer 310P averages each audio feature vector 212 to generate the final audio feature vector 212F. In other examples, the intelligent pooling layer 310P sums each audio feature vector 212 to generate the final audio feature vector 212F. Finally, the intelligent pooling layer 310P refines the multiple audio feature vectors 212 in some manner into a final audio feature vector 212F that includes or emphasizes audio features that characterize the relationship between human-derived and synthetic speech. In some examples, the intelligent pooling layer 310P focuses the final audio feature vector 212F on a portion of the audio data 120 that is most likely to contain synthetic speech. For example, the audio data 120 includes a small or narrow portion that provides indicators (e.g., audio characteristics) that indicate that the utterance 119 contains synthetic audio, while the remainder of the audio data 120 provides few or no indicators that the utterance 119 contains synthetic speech. In this example, the intelligent pooling layer 310P emphasizes the audio feature vector 212 associated with that portion of the audio data 120 (or vice versa, de-emphasizes the other remaining audio feature vectors 212).
[0028] In some implementations, the shallow classifier model 222 includes only one other layer 310 in addition to the intelligent pooling layer 310P. For example, the shallow classifier model 222 includes a fully connected layer 310F configured to receive a single final audio feature vector 212F from the intelligent pooling layer 310P as input and generate a score 224 as output. Thus, in some examples, the shallow classifier model 222 is a shallow neural network including a single intelligent pooling layer 310P and only one other layer 310, such as a fully connected layer 310F. Each layer includes any number of neurons / nodes 332. The single fully connected layer 310F can map results to logits. In some examples, the shallow classifier model 222 includes one of a logistic regression model, a linear discriminant analysis model, or a random forest model.
[0029] Referring now to FIG. 4A , in some implementations, a training process 400, 400a trains the audio feature extractor 210 on a pool 402A of human-derived speech samples. These human-derived speech samples provide unlabeled audio extractor training samples 410A that train the untrained audio feature extractor 210. The pool 402A of human-derived speech can be quite large, resulting in a significant number of audio extractor training samples 410A. Thus, in some examples, the training process 400a trains the untrained audio feature extractor 210 on a large number of audio extractor training samples 410A that include only human-derived speech and do not include any synthetic speech. This is advantageous because large pools of synthetic speech are typically expensive and / or difficult to obtain. However, in some examples, the audio extractor training samples 410A include samples with human-derived speech and synthetic speech. Optionally, the audio feature extractor 210 includes a representation model derived from a larger, trained self-supervised model. In this scenario, the larger trained self-supervised model may be a very large model that is computationally expensive to run and is not well suited to the user device 102. However, because of the potential benefits (e.g., latency, privacy, bandwidth, etc.) of running the audio feature extractor 210 locally on the user device 102, the audio feature extractor 210 may be a representation model of the larger trained self-supervised model that reduces the model size and complexity without substantially sacrificing accuracy. This allows the model to run on the user device 102 despite limited computational power or memory capacity. The representation model improves performance by converting high-dimensional data (e.g., audio) to lower dimensions to train a smaller model and by using the representation model as pre-training.
[0030] 4B , in some examples, the shallow classifier model 222 is trained after the audio feature extractor 210 is trained via a training process 400, 400b. In this example, the trained audio feature extractor 210 receives audio data 120 from a pool of synthetic speech samples 402B. The trained audio feature extractor 210 generates audio feature vectors 212 corresponding to classifier training samples 410b based on the audio data 120 from the pool 402B. These classifier training samples 410b (i.e., the multiple audio feature vectors 212 generated by the trained audio feature extractor 210) train the shallow classifier model 222. When the shallow classifier model 222 may be trained using synthetic speech from the pool of synthetic speech 402B, the pool of synthetic speech 402B may be significantly smaller than the pool of human-derived speech 402A.
[0031] In some examples, the shallow classifier model 222 is trained exclusively on training samples 410b including synthetic speech, while in other examples, the shallow classifier model 222 is trained on a mixture of training samples 410b including synthetic speech and training samples 410b including only human-derived speech. The samples 410b including synthetic speech may include only synthetic speech (i.e., may not include human-derived speech). The samples 410b may include a mixture of synthetic and human-derived speech. For example, in the example of FIG. 1, the utterance 119 includes the speech "My name is Jane Smith." Possible training samples 410b from this utterance 119 include those in which the "My name is" portion of the utterance 119 is human-derived speech and those in which the "Jane Smith" portion of the utterance 119 is synthetic speech. The remote system 110 and / or user device 102 may perturb existing training samples 410b to generate additional training samples 410b. For example, the remote system replaces a portion of the human-derived speech with a synthetic speech, replaces a portion of the synthetic speech with a portion of the human-derived speech, replaces a portion of the synthetic speech with another portion of the synthetic speech, and replaces a portion of the human-derived speech with another portion of the human-derived speech.
[0032] In some implementations, the remote system 110 performs the training processes 400a, 400b to train the audio feature extractor 210 and the shallow classifier model 222 and then transmits the trained models 210, 222 to the user device 102. However, in other examples, the user device 102 performs the training processes 400a, 400b to train the audio feature extractor 210 and / or the shallow classifier model 222 on the user device 102. In some examples, the remote system 110 or the user device 102 fine-tunes the shallow classifier model 222 based on new or updated training samples 410b. For example, the user device 102 updates, fine-tunes, or partially re-trains the shallow classifier model 222 on the audio data 120 received from the audio source 10.
[0033] Referring now to the schematic diagram 500 of FIG. 5 , in some examples, the user device 102 and / or the remote system 110 utilize the same audio feature extractor 210 to provide audio feature vectors 212 to multiple shallow classifier models 222, 222a-n. In this manner, the audio feature extractor 210 serves as a “front-end” model, while the shallow classifier models 222 serve as “back-end” models. Each shallow classifier model 222 may be trained for a different purpose. For example, a first shallow classifier model 222a determines whether a voice is human or synthetic, while a second shallow classifier model 222b recognizes and / or classifies emotions in the streaming audio 118. That is, the self-supervised audio feature extractor 210 is well-suited for “non-semantic” tasks (i.e., aspects other than the meaning of human speech), which the shallow classifier models 222 can utilize for a variety of different purposes. Because the shallow classifier models 222 are potentially small in size and complexity, the user device can store and execute each of these models as needed to process the audio feature vectors 212 generated by the audio feature extractor 210.
[0034] 6 shows a flowchart of example operations of a method 600 for determining whether audio data 120 includes synthetic speech. At operation 602, the method 600 includes receiving, at data processing hardware 103, audio data 120 characterizing speech captured by the user device 102. At operation 604, the method 600 includes generating, by the data processing hardware 103, a plurality of audio feature vectors 212 using the trained self-supervised model 210 (i.e., the audio feature extractor 210), each audio feature vector 212 representing audio features of a portion of the audio data 120. At operation 606, the method 600 also includes generating, by the data processing hardware 103, a score 224 indicating the presence of synthetic speech in the audio data 120 based on the corresponding audio features of each audio feature vector 212 of the plurality of audio feature vectors 212 using the shallow classifier model 222. The method 600 includes, at operation 608, determining by the data processing hardware 103 whether the score 224 satisfies a synthetic voice detection threshold, and, at operation 610, determining by the data processing hardware 103 that the voice in the audio data 120 captured by the user device 102 includes synthetic voice when the score 224 satisfies the synthetic voice detection threshold.
[0035] 7 is a schematic diagram of an exemplary computing device 700 that can be used to implement the systems and methods described herein. Computing device 700 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The components shown, their connections and relationships, and their functionality are intended to be merely exemplary and are not intended to limit the implementation of the invention(s) described and / or claimed herein.
[0036] Computing device 700 includes a processor 710, memory 720, a storage device 730, a high-speed interface / controller 740 connecting to memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 connecting to a low-speed bus 770 and storage device 730. Components 710, 720, 730, 740, 750, and 760 are each interconnected using various buses and may be mounted on a common motherboard or in other suitable manners. Processor 710 may process instructions for execution within computing device 700, including instructions stored in memory 720 or on storage device 730, for displaying graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 780 coupled to high-speed interface 740. In other implementations, multiple processors and / or multiple buses may be used, along with multiple memories and multiple memory types, as appropriate. Additionally, multiple computing devices 700 can be connected together (eg, as a server bank, a group of blade servers, or a multi-processor system), with each device providing some portion of the required operations.
[0037] Memory 720 stores information non-transiently within computing device 700. Memory 720 may be a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. Non-transient memory 720 may be a physical device used to temporarily or permanently store programs (e.g., sequences of instructions) or data (e.g., program state information) for use by computing device 700. Examples of non-volatile memory include, but are not limited to, flash memory (e.g., typically used for firmware such as boot programs) and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase change memory (PCM), as well as disk or tape.
[0038] The storage device 730 is capable of providing mass storage for the computing device 700. In some implementations, the storage device 730 is a computer-readable medium. In various different implementations, the storage device 730 can be a floppy disk device, a hard disk device, an optical disk device, or an array of devices including a tape device, a flash memory or other similar solid-state memory device, or a device in a storage area network or other configuration. In further implementations, a computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as the memory 720, the storage device 730, or memory on the processor 710.
[0039] The high-speed controller 740 manages bandwidth-intensive operations of the computing device 700, while the low-speed controller 760 manages less bandwidth-intensive operations. Such allocation of roles is merely exemplary. In some implementations, the high-speed controller 740 is coupled to memory 720, to a display 780 (e.g., through a graphics processor or accelerator), and to a high-speed expansion port 750 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 760 is coupled to a storage device 730 and a low-speed expansion port 790. The low-speed expansion port 790, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices such as a keyboard, pointing device, scanner, etc., or to a networking device such as a switch or router, for example, through a network adapter.
[0040] Computing device 700 can be implemented in several different forms as shown in the figure. For example, computing device 700 can be implemented as a standard server 700a, or multiple times within a group of such servers 700a, or as a laptop computer 700b, or as part of a rack server system 700c.
[0041] Various implementations of the systems and techniques described herein may be realized as digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation as one or more computer programs executable and / or interpretable on a programmable system that includes at least one programmable processor, which may be special-purpose or general-purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0042] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a procedural and / or object-oriented high-level programming language and / or in assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0043] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably coupled to receive data from or transfer data to them, or both. However, a computer need not have such devices. Suitable computer-readable media for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0044] To enable user interaction, one or more aspects of the present disclosure can be implemented on a computer having a display device, such as a CRT (cathode ray tube), LCD (liquid crystal display) monitor, or touch screen, for displaying information to the user, and optionally a keyboard and pointing device, such as a mouse or trackball, by which the user can input to the computer. Other types of devices can also be used to enable user interaction; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback, and input from the user can be received in any form, including acoustic, speech, or tactile input. Additionally, the computer can interact with the user by sending documents to and receiving documents from a device being used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.
[0045] Although several implementations have been described above, it will be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Accordingly, other implementations are within the scope of the following claims. [Explanation of symbols]
[0046] 10 Audio Sources 100 systems 102 User Devices 103 Data Processing Hardware 104 Network 105 Memory Hardware 106 Microphone 108 Graphical User Interface (GUI) 110 Remote System, Remote Server 112 computing resources 114 Memory Resources 118 Streaming Audio 119 utterances 120 Audio Data 150 signs 160 notifications 200 Schematic 210 Self-supervised audio feature extractor, self-supervised model 212 audio feature vectors 212a~n Audio feature vector 212F Final audio feature vector 220 Synthetic Speech Detector 222 Shallow Classifier Model 222a~n Shallow classifier model 222a First shallow classifier model 222b Second shallow classifier model 224 score 250 Deep Neural Networks 300 Schematic 310 Intelligent Pooling Layer, Other Layers 310F fully connected layer 310P Intelligent Pooling Layer 332 neurons / nodes 400 training processes 400a Training Process 400b Training Process 402A Human-derived audio pool 402B Synthetic Voice Pool 410A Unlabeled Audio Extractor Training Samples 410b Classifier training samples 500 Schematic 600 ways 700 computing devices 700a Standard Server 700b laptop computer 700c Rack Server System 710 Processors, Components 720 Memory, Components 730 Storage Devices, Components 740 High-Speed Interface / Controller, Components 750 High-Speed Expansion Port, Component 760 Low-Speed Interfaces / Controllers, Components 770 Slow Bus 780 Display 790 Low-Speed Expansion Port
Claims
1. receiving, in data processing hardware (103), audio data (120) characterizing speech captured by the user device (102); generating, by the data processing hardware (103), a plurality of audio feature vectors (212) using the trained self-supervised model (210), each of which represents audio features of a different portion of the audio data (120); generating, by the data processing hardware (103), a score (224) indicative of the presence of synthetic speech in the audio data (120) based on the audio features of each audio feature vector (212) of the plurality of audio feature vectors (212) using a shallow classifier model (222), wherein the shallow classifier model (222) is a shallow neural network including only an intelligent pooling layer (310) and a fully connected layer, the intelligent pooling layer (310) configured to generate a single final audio feature vector (212) based on each audio feature vector (212) of the plurality of audio feature vectors (212), and the fully connected layer configured to receive the single final audio feature vector (212) as an input and generate the score (224) as an output; determining, by said data processing hardware (103), whether said score (224) satisfies a synthetic speech detection threshold; determining, by the data processing hardware (103), that the speech in the audio data (120) acquired by the user device (102) includes synthetic speech when the score (224) satisfies the synthetic speech detection threshold; Including, The method (600), wherein the shallow classifier model (222) is trained on a plurality of training samples, each of the plurality of training samples including a synthetic speech portion and a human-derived speech portion.
2. 10. The method of claim 1, wherein generating the score indicative of the presence of the synthetic speech in the audio data is based on the single final audio feature vector.
3. 3. The method of claim 2, wherein the single final audio feature vector comprises an average of each audio feature vector of the plurality of audio feature vectors.
4. 3. The method of claim 2, wherein the single final audio feature vector comprises a sum of each audio feature vector of the plurality of audio feature vectors.
5. 5. The method of claim 1, wherein the shallow classifier model comprises one of a logistic regression model, a linear discriminant analysis model, or a random forest model.
6. 6. The method (600) of any one of claims 1 to 5, wherein the trained self-supervised model (210) is trained on a first training data set that includes only training samples (410) of human-derived speech.
7. The method (600) of any one of claims 1 to 6, wherein the data processing hardware (103) resides on the user device (102).
8. data processing hardware (103); memory hardware (105) in communication with the data processing hardware (103), storing instructions that, when executed on the data processing hardware (103), cause the data processing hardware (103) to: receiving audio data (120) characterizing speech within audio data (120) captured by a user device (102); generating a plurality of audio feature vectors (212) using the trained self-supervised model (210), each of which represents audio features of a different portion of the audio data (120); generating a score (224) indicating the presence of synthetic speech in the audio data (120) based on the audio features of each audio feature vector (212) of the plurality of audio feature vectors (212) using a shallow classifier model (222), wherein the shallow classifier model (222) is a shallow neural network including only an intelligent pooling layer (310) and a fully connected layer, the intelligent pooling layer (310) being configured to generate a single final audio feature vector (212) based on each audio feature vector (212) of the plurality of audio feature vectors (212), and the fully connected layer being configured to receive the single final audio feature vector (212) as an input and generate the score (224) as an output; determining whether the score (224) satisfies a synthetic speech detection threshold; and determining that the speech in the audio data captured by the user device includes synthetic speech when the score satisfies the synthetic speech detection threshold; and memory hardware (105) that performs operations including: Equipped with The system, wherein the shallow classifier model (222) is trained on a plurality of training samples, each of the plurality of training samples including a synthetic speech portion and a human-derived speech portion.
9. 9. The system of claim 8, wherein generating the score (224) indicative of the presence of the synthetic speech in the audio data (120) is based on the single final audio feature vector (212).
10. 10. The system of claim 9, wherein the single final audio feature vector (212) comprises an average of each audio feature vector (212) of the plurality of audio feature vectors (212).
11. 10. The system of claim 9, wherein the single final audio feature vector (212) comprises a sum of each audio feature vector (212) of the plurality of audio feature vectors (212).
12. 12. The system of claim 8, wherein the shallow classifier model (222) comprises one of a logistic regression model, a linear discriminant analysis model, or a random forest model.
13. 13. The system of claim 8, wherein the trained self-supervised model is trained on a first training data set that includes only training samples (410) of human-derived speech.
14. 14. The system of claim 8, wherein the data processing hardware (103) resides on the user device (102).
Citation Information
Patent Citations
Method and apparatus for detecting spoofing conditions
US20180254046A1
Systems and methods for end-to-end architectures for voice spoofing detection
US20200322377A1