Audiovisual deepfake detection

The machine learning architecture addresses the limitations of conventional deepfake detection by integrating audio and visual analysis, enhancing detection accuracy and authentication through combined scoring and evaluation.

JP7865955B2Active Publication Date: 2026-05-26PINDROP SECURITY INC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
PINDROP SECURITY INC
Filing Date
2021-10-15
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Conventional deepfake detection systems are limited to evaluating either audio or visual data separately, requiring additional computational resources and often failing to detect deepfakes in integrated audiovisual data.

Method used

A machine learning architecture that integrates audio and visual deepfake detection by employing layers for speaker recognition, face recognition, and lip-sync estimation, generating and combining scores to determine the likelihood of deepfake content in audiovisual data.

Benefits of technology

Enhances the accuracy of deepfake detection by evaluating both audio and visual data simultaneously, improving the system's ability to authenticate identities and detect deepfakes in integrated audiovisual data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007865955000001
    Figure 0007865955000001
  • Figure 0007865955000002
    Figure 0007865955000002
  • Figure 0007865955000003
    Figure 0007865955000003
Patent Text Reader

Abstract

In an embodiment, a machine learning architecture is implemented for biometric-based identity recognition (e.g., speaker recognition, face recognition) and deepfake detection (e.g., speaker deepfake detection, face deepfake detection). The machine learning architecture includes layers that define multiple score components, including sub-architectures for speaker deepfake detection, speaker recognition, face deepfake detection, face recognition, and a lip-sync estimation engine. The machine learning architecture extracts and analyzes various types of low-level features from both audio and visual data, combines various scores, and uses the scores to determine the likelihood that the audiovisual data contains deepfake content and the likelihood that the claimed identities of people in the video match the predicted or enrolled identities of people. This enables the machine learning architecture to perform integrated identity recognition and authentication and deepfake detection for both audio and visual data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims priority to U.S. Provisional Application No. 63 / 092,956, filed October 16, 2020, which is incorporated herein by reference in its entirety.

[0002] This application generally relates to U.S. Application No. 15 / 262,748, issued as U.S. Patent No. 9,824,692, entitled "End - to - End Speaker Recognition Using Deep Neural Network", filed September 12, 2016, which is incorporated herein by reference in its entirety.

[0003] This application generally relates to U.S. Application No. 17 / 155,851, entitled "Robust spoofing detection system using deep residual neural networks", filed August 21, 2020, which is incorporated herein by reference in its entirety.

[0004] This application generally relates to U.S. Application No. 15 / 610,378, issued as U.S. Patent No. 10,141,009, entitled "System and Method for Cluster - Based Audio Event Detection", filed May 31, 2017, which is incorporated herein by reference in its entirety.

[0005] This application generally relates to systems and methods for managing, training, and deploying machine - learning architectures for audio processing.

Background Art

[0006] Deepfakes of manipulated audiovisual data are becoming increasingly common and sophisticated. This allows for the dissemination of personal videos through social media websites and video-sharing platforms. Fraudsters, pranksters, or other malicious actors may use deepfakes to damage a person's reputation or disrupt interpersonal speech by publishing false and / or misleading information about the person shown in the fake video. Furthermore, communication and computing systems that rely on audiovisual data to authenticate users or monitor user behavior face additional problems. Deepfakes can be used by fraudsters or other malicious actors to impersonate the identities of certain authorized access system functions.

[0007] "Deepfakes" are manipulated video, audio, or other digital content generated by artificial intelligence algorithms that can so skillfully recreate human images and voices. These algorithms often produce audio and / or visual content that appears authentic. Recent improvements in deepfake algorithms have made deepfake videos and audio highly sophisticated, sometimes making them almost indistinguishable from real people. These fake videos and audio pose a significant threat to social media platforms, as deepfakes can be used to manipulate factual narratives, make fake news seem real, or damage people's reputations. Improvements in deepfake detection and biometric authentication systems will be beneficial in a variety of situations. [Overview of the Initiative] [Means for solving the problem]

[0008] Many conventional deepfake detection systems focus on detecting deepfake content contained in speech or facial images. These deepfake detection systems can only evaluate and protect one form of data at a time (e.g., audio or visual), requiring additional computational resources to evaluate audio data separately from visual data, and may fail to detect deepfakes. A means is needed to evaluate audio data, visual data, and / or audiovisual data in a system that integrates one or more machine learning architectures.

[0009] Disclosed herein are systems and methods that can address the aforementioned shortcomings and may provide any number of additional or alternative benefits and advantages. Embodiments include a computing device that runs software routines of one or more machine learning architectures. The machine learning architectures perform integrated evaluation operations for audio and visual deepfake detection, evaluating and protecting audio data, visual data, and audiovisual data. Furthermore, this combination can improve the overall accuracy of the deepfake detection system.

[0010] Embodiments disclosed herein include systems and methods for implementing a machine learning architecture for biometric-based identity recognition (e.g., speaker recognition, face recognition) and deepfake detection (e.g., speaker deepfake detection, face deepfake detection). The machine learning architecture includes layers defining multiple scoring components, including sub-architectures for speaker deepfake detection (generation of speaker deepfake scores), speaker recognition (generation of speaker recognition similarity scores), face deepfake detection (generation of face deepfake scores), face recognition (generation of face recognition similarity scores), and a lip-sync estimation engine (generation of lip-sync estimation scores). The machine learning architecture extracts and analyzes various types of low-level features from both audio and visual data of a given video (audiovisual data sample), combines various scores generated by the scoring components, and uses the various scores to determine the likelihood that the audiovisual data contains deepfake content, the likelihood that the claimed identity of a person in the video matches the identity of a predicted or registered person. This enables the machine learning architecture to perform identity recognition and authentication, and deepfake detection in an integrated manner for both audio and visual data.

[0011] In one embodiment, the computer implementation method includes the steps of: acquiring an audiovisual data sample including audiovisual data using a computer; applying a machine learning architecture to the audiovisual data using the computer to generate a similarity score using biometric information embedding extracted from the audiovisual data and generating a deepfake score using a forged print extracted from the audiovisual data; and using the similarity score and the deepfake score, generating a final output score using the computer that indicates the likelihood that the audiovisual data is authentic.

[0012] In another embodiment, the computer comprises a processor configured to acquire an audiovisual data sample including audiovisual data, apply a machine learning architecture to the audiovisual data, generate a similarity score using biometric embeddings extracted from the audiovisual data, generate a deepfake score using a fake print extracted from the audiovisual data, and use the similarity score and the deepfake score to generate a final output score that indicates the likelihood by the computer that the audiovisual data is authentic.

[0013] Both the above-mentioned overview and the following embodiments for carrying out the invention are examples and descriptions and are intended to provide a further explanation of the claimed invention. [Brief explanation of the drawing]

[0014] This disclosure can be better understood by referring to the following figures. The elements in the figures are not necessarily to scale, and instead the focus is on illustrating the principles of this disclosure. In the figures, reference numbers indicate the corresponding parts across different figures. [Figure 1] Figure 1 shows the components of a system for receiving and analyzing audiovisual data. [Figure 2] Figure 2 shows the data flow between the components of a system that performs identity recognition and deepfake detection calculations. [Figure 3] Figure 3 shows the data flow between the components of a system that performs registration calculations to build a registered audiovisual profile for a specific person. [Figure 4] Figure 4 shows the execution steps of a method for implementing one or more machine learning architectures for deepfake detection and identity recognition. [Figure 5] Figure 5 shows the data flow of system components for implementing a machine learning architecture for deepfake detection and identity recognition, following score fusion calculations of score levels applied to various biometric measurements. [Figure 6] Figure 6 shows the data flow of system components for implementing a machine learning architecture for deepfake detection and identity recognition, according to embedded-level score fusion calculations applied to various biometric measurements. [Figure 7] Figure 7 shows the data flow of system components for implementing a machine learning architecture for deepfake detection and identity recognition, according to feature-level score fusion operations applied to various biometric measurements. [Figure 8] Figure 8 shows the execution steps of a method for implementing one or more machine learning architectures for deepfake detection and identity recognition. [Modes for carrying out the invention]

[0015] Here, exemplary embodiments shown in the drawings are used to describe them, and the specification uses specific terminology to describe them as well. Nevertheless, it will be understood that no claims or limitations of this disclosure are intended therein. Modifications and further alterations of the features of the invention shown herein, as well as additional applications of the principles of the subject matter shown herein, which a person skilled in the art and possessing this disclosure might conceive of, should be considered within the scope of the subject matter disclosed herein.

[0016] Many conventional deepfake detection systems focus on detecting deepfake content contained in either speech or facial images. While these systems are effective, data streams and computer files often contain audiovisual data with both audio and visual components, and they can only detect deepfakes (or disguises) in one type of data, such as audio data or image data. Therefore, conventional approaches are often insufficient or inefficient. Embodiments disclosed herein include a computing device running software for one or more machine learning architectures, which perform integrated analytical calculations for audio and visual deepfake detection to evaluate audio data, visual data, and audiovisual data.

[0017] Embodiments disclosed herein include systems and methods for implementing a machine learning architecture for biometric-based identity recognition (e.g., speaker recognition, face recognition) and deepfake detection (e.g., speaker deepfake detection, face deepfake detection). The machine learning architecture includes layers defining multiple scoring components, including sub-architectures for speaker deepfake detection (generation of speaker deepfake scores), speaker recognition (generation of speaker recognition similarity scores), face deepfake detection (generation of face deepfake scores), face recognition (generation of face recognition similarity scores), and a lip-sync estimation engine (generation of lip-sync estimation scores). The machine learning architecture extracts and analyzes various types of low-level features from both audio and visual data of a given video (audiovisual data sample), combines various scores generated by the scoring components, and uses the various scores to determine the likelihood that the audiovisual data contains deepfake content, the likelihood that the claimed identity of a person in the video matches the identity of a predicted or registered person. This enables the machine learning architecture to perform identity recognition and authentication, and deepfake detection in an integrated manner for both audio and visual data.

[0018] Figure 1 shows the components of a system 100 for receiving and analyzing audiovisual data. System 100 comprises an analysis system 101 and an end-user device 114. The analysis system 101 includes an analysis server 102, an analysis database 104, and a management device 103. Embodiments may include additional or alternative components, or certain components may be omitted from the components of Figure 1, and still remain within the scope of the disclosure. For example, it may be common to include multiple analysis servers 102. Embodiments may include any number of devices capable of performing the various features and tasks described herein, or may otherwise be implemented in other ways. For example, Figure 1 shows an analysis server 102 as a computing device separate from the analysis database 104. In some embodiments, the analysis database 104 includes an integrated analysis server 102. In computation, the analysis server 102 receives and processes audiovisual data from the end-user device 114, recognizes the voice and face of a speaker in a video, and / or detects whether the video contains a deepfake of the speaker's voice or face image. The analysis server 102 outputs a score or display indicating whether the audiovisual input is likely to contain either genuine or falsified audiovisual data.

[0019] System 100 comprises various hardware and software components of one or more public or private networks 108 that interconnect various components of System 100. Non-limiting embodiments of such networks 108 may include local area networks (LANs), wireless local area networks (WLANs), metropolitan area networks (MANs), wide area networks (WANs), and the Internet. Communication over networks 108 may be carried out according to various communication protocols such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE Communication Protocol. Similarly, end-user devices 114 may communicate with the analysis system 101 or other customer-facing systems via telephony and telecommunications protocols, hardware, and software that can host, transmit, and exchange audiovisual data (e.g., computer files, data streams). Non-limiting embodiments of telecommunications and / or computing networking hardware may include switches and trunks, among other additional or alternative hardware used to host, route, or manage data communications, circuits, and signals over the Internet or other device communication media.

[0020] The analysis system 101 represents a computing network infrastructure that includes physically and logically related software and electronic devices managed or operated by a corporate organization that hosts a particular service (e.g., video conferencing software). The devices of the network system infrastructure 101 can provide the intended services of a particular corporate organization and communicate via one or more internal networks. In some embodiments, the analysis system 101 operates in place of an intermediate computing network infrastructure of a third-party customer support company (e.g., a corporation, government agency, university). In such embodiments, the third-party infrastructure includes computing devices (e.g., servers) that capture, store, and transfer visual data to the analysis system 101. The analysis server 102 hosts a cloud-based service or communicates with a server that hosts a cloud-based service.

[0021] The end-user device 114 can be any communication device or computing device that an end user operates to send visual data to a specific destination (e.g., the analysis system 101, the customer support system). The end-user device 114 includes a processor and software for sending visual data to the analysis system 101 via one or more networks 108. In some cases, the end-user device 114 includes software and hardware for generating visual data, including a camera and a microphone. Non-limiting examples of the end-user device 114 can include a mobile device 114a (e.g., a smartphone, a tablet) and an end-user computer 114b (e.g., a laptop, a desktop, a server). For example, the end-user device 114 can be an end-user computer 114b that runs video conferencing software that captures and transmits visuals to a central host server that functions as or communicates with the analysis server 102.

[0022] The analysis server 102 of the call analysis system 101 includes one or more processors and software, and can be any computing device capable of executing various processes and tasks described herein. The analysis server 102 receives and processes the audiovisual data transmitted from the end-user device 114 during training or received from the analysis database 104. The analysis server 102 may host or communicate with an analysis database 104 that includes various types of information that the analysis server 102 refers to or queries when executing the layers of the machine learning architecture. The analysis database 104 may store, for example, registered audiovisual profiles of registered people (e.g., registered users, celebrities), trained models of the machine learning architecture, and other types of information. Although FIG. 1 shows only a single analysis server 102, the analysis server 102 may include any number of computing devices. In some cases, the computing devices of the analysis server 102 may be able to execute all or some of the processes and advantages of the analysis server 102. The analysis server 102 may comprise computing devices operating in a distributed or cloud computing configuration and / or in a virtual machine configuration. Also, in some embodiments, it should be understood that the functions of the analysis server 102 may be partially or fully executed by various computing devices of the analysis system 101 or other computing infrastructure.

[0023] The analysis server 102 receives audiovisual data in a data stream, which may include discrete computer files or continuous streams containing the audiovisual data. In some cases, the analysis server 102 receives audiovisual data from end-user devices 114 or third-party devices (e.g., web servers, third-party servers for computing services). For example, an end-user device 114 transmits a multimedia computer file containing audiovisual data (e.g., an MP4 file, a MOV file) or a hyperlink to a third-party server hosting audiovisual data (e.g., a YouTube® server). In some cases, the analysis server 102 receives audiovisual data such as that generated and distributed by a server hosting audiovisual communication software. For example, two or more end-user devices 114 run communication software (e.g., Skype®, MS Teams®, Zoom®) that establishes communication event sessions between the end-user devices 114, either directly or indirectly through a server that establishes and hosts the communication event session. Communication software running on the end-user device 114 and the server captures, stores, and distributes audiovisual data between the communication event session and the end-user device 114 communicating with it.

[0024] In some embodiments, the analysis server 102 or a third-party server in the third-party infrastructure hosts a cloud-based, audiovisual communication service. Such software performs, among other potential software operations, the process of managing device communication queues for end-user devices 114 involved in a specific device communication event session (e.g., a conference call, a video call), and / or the process of routing data packets for device communication, including audiovisual data, between end-user devices 114 over one or more networks 108. Audiovisual communication software, run by a specific server (e.g., the analysis server 102, a third-party server), can ingest, query, or generate various types of information about end-user devices 114 and / or end-users. When run by a third-party server, the third-party server transmits audiovisual data and other types of information to the analysis server 102.

[0025] The analysis server 102 runs analysis software for processing audiovisual data samples (e.g., computer files, machine-readable data streams). The input audiovisual data includes audio data representing the audio signal from a speaker and visual image data including a facial image of a specific person. The processing software of the analysis server 102 includes machine learning software routines organized as various types of machine learning architectures or models, such as Gaussian mixture models (GMMs) and neural networks (e.g., convolutional neural networks (CNNs), deep neural networks (DNNs)). A machine learning architecture comprises functions or layers that perform various processing operations as described herein. For example, the analysis software includes one or more machine learning architectures for speaker and face identification and speaker and face spoofing detection. The layers and operations of a machine learning architecture define the components of the machine learning architecture, which may be separate architectures or sub-architectures. Components may include various types of machine learning techniques and functions, such as neural network architectures and Gaussian mixture models (GMMs).

[0026] The machine learning architecture operates in several operational phases, including a training phase, an optional enrollment phase, and a deployment phase (sometimes called a "test" phase or simply "test"). The input audiovisual data processed by the analysis server 102 may include training audiovisual data, training audio signals, training visual data, enrollment audiovisual data, enrollment audio signals, enrollment visual data, and inbound audiovisual data received and processed during the deployment phase. The analysis server 102 applies the machine learning architecture to each type of input audiovisual data during the corresponding operational phase.

[0027] The analysis server 102 or other computing devices of the system 100 (e.g., the call center server 111) can perform various preprocessing and / or data augmentation operations on the input audiovisual data. Non-limiting embodiments of preprocessing operations include extracting low-level features from audio signals or image data, analyzing and segmenting audio signals or image data into frames and segments, and, among other potential preprocessing operations, performing one or more transform functions such as a Short-Time Fourier Transform (SFT) or a Fast Fourier Transform (FFT). The analysis server 102 may perform preprocessing or data augmentation operations before feeding the input audiovisual data to the input layer of the machine learning architecture, or the analysis server 102 may perform these operations as part of the execution of the machine learning architecture, where the input layer (or other layers) of the machine learning architecture performs these operations. For example, the machine learning architecture may include an in-network preprocessing or data augmentation layer that performs specific preprocessing or data augmentation operations on the input audiovisual data.

[0028] The machine learning architecture includes layers that define the components of an audiovisual deepfake detection architecture. These layers define engines for scoring aspects of audiovisual data, including voice spoofing detection (or "voice deepfake" detection), speaker recognition, face spoofing detection (or "face deepfake detection"), face recognition, and lip-sync estimation. Components of the machine learning architecture include, for example, a speaker recognition engine, a speaker deepfake engine, a face recognition engine, a face deepfake engine, and, in some embodiments, a lip-sync estimation engine. The machine learning architecture applies these scoring engines to various types of audiovisual data to analyze both the audio signals and image data of the audiovisual data, generate specific types of scores, combine the scores, and determine whether the audiovisual data contains deepfake components.

[0029] speaker engine The machine learning architecture includes a layer that defines one or more speaker embedding engines (sometimes called "speaker biometric engines" or "speaker engines"), including speaker recognition engines and speaker deepfake detection engines.

[0030] The speaker recognition engine extracts a set of speech features from speech data or segments of speech data. These features may include spectral-temporal features, such as Mel-frequency septal coefficients (MFCCs) and linear filter banks (LFBs). The analysis server 102 applies the speaker recognition engine to the speech features and extracts embeddings as feature vectors representing the set of speaker features. During the enrollment phase, the analysis server 102 takes in one or more registered audiovisual data samples to generate one or more corresponding registered speech embeddings. The machine learning architecture algorithmically combines (e.g., averages) the registered speech embeddings to generate a registered voiceprint of the registrant audiovisual profile, which the analysis server 102 stores in the analysis database 104.

[0031] During the deployment phase, the analysis server 102 takes in inbound audiovisual data samples to extract inbound voice embeds as inbound voiceprints. In some cases, the analysis server 102 also receives identity claims of the person associated with the inbound voiceprint. For speaker recognition, the machine learning architecture generates speaker similarity scores that represent the likelihood of similarity between the speaker of the registered voiceprint and the speaker of the inbound voiceprint. The analysis server 102 outputs one or more speaker recognition similarity scores.

[0032] Speaker embedded representations can be generated, for example, by implementing a GMM-based system or a neural network architecture (e.g., a deep neural network, a convolutional neural network). Exemplary embodiments of speaker recognition engines are described in U.S. Patents 9,824,692 and 10,141,009, and U.S. Application No. 17 / 155,851, each of which is incorporated in whole by reference.

[0033] The analysis server 102 applies a speaker deepfake engine to the speech features to extract speaker spoofing embeddings as feature vectors representing a set of features for the spoofed speech signal artifacts. During any enrollment phase, the analysis server 102 takes in one or more registered audiovisual data samples to generate one or more corresponding registered spoofing print embeddings. The machine learning architecture algorithmically combines (e.g., averages) the registered spoofing print embeddings to generate registered spoofing prints for registrant audiovisual profiles or other people, which the analysis server 102 stores in the analysis database 104.

[0034] During the deployment phase, the analysis server 102 takes in inbound audiovisual data samples to extract inbound impersonation embeds as inbound impersonation prints. In some cases, the analysis server 102 also receives identity claims of individuals associated with the inbound impersonation prints. For speaker deepfake detection, the machine learning architecture generates an impersonation print similarity score that represents the likelihood that the inbound audiovisual data sample contains a speaker deepfake, based on the similarity between one or more pre-configured or registered impersonation prints and the speaker in the inbound impersonation print. The analysis server 102 outputs a similarity or detection score for one or more speaker-deepfakes.

[0035] The audio deepfake detection engine may implement, for example, a neural network architecture or a GMM-based architecture. Exemplary embodiments of speaker-deepfake detection engines are described in U.S. Patent No. 9,824,692 and U.S. Patent Application No. 17 / 155,851, each of which is incorporated in whole by reference.

[0036] Face engine The machine learning architecture includes a layer that defines one or more face embedding engines (sometimes called "face biometric engines" or "face engines"), including face recognition engines and face deepfake detection engines.

[0037] The face recognition engine extracts embedded representations of faces contained within frames of audiovisual data. The face recognition engine extracts a set of image features from image data or segments of image data. Features may include low-level image features, such as (e.g., pixel vectors, linear binary patterns (LBP), discrete cosine transform (DCT)). The analysis server 102 applies the face recognition engine to the image features to extract face embeddings as feature vectors representing a set of features of a person's face. During the enrollment phase, the analysis server 102 takes in one or more registered audiovisual data samples to generate one or more corresponding registered face embeddings. The machine learning architecture algorithmically combines (e.g., averages) the registered face print embeddings to generate registered face prints for the registered audiovisual profile, which the analysis server 102 stores in the analysis database 104.

[0038] During the deployment phase, the analysis server 102 takes in inbound audiovisual data samples and extracts inbound face embeddings as inbound faceprints. In some cases, the analysis server 102 also receives identity claims of the people associated with the inbound faceprints. For face recognition, the machine learning architecture generates a face similarity score that represents the likelihood of similarity between the faces in the registered faceprints and the faces in the inbound faceprints. The analysis server 102 outputs one or more face recognition similarity scores.

[0039] A face recognition engine can implement a neural network architecture such as vggface (e.g., a deep neural network). Exemplary embodiments of face recognition engines can be found in Cao, et al, “Vggface2: A Dataset for Recognizing Faces across Pose and Age,” IEEE, 13th IEEE International Conference on Automatic Face & Gesture Recognition, pp. 67-74 (2018), which are incorporated in their entirety by reference.

[0040] The analysis server 102 applies a facial deepfake engine to the image features and extracts fake face embeddings as feature vectors representing a set of artifact features in the fake face image. During any enrollment phase, the analysis server 102 takes in one or more enrollment audiovisual data samples to generate one or more corresponding enrollment face fake print embeddings. The machine learning architecture algorithmically combines (e.g., averages) the enrolled face fake print embeddings to generate registered face fake prints for registrant audiovisual profiles or other people, and the analysis server 102 stores them in the analysis database 104.

[0041] During the deployment phase, the analysis server 102 takes in inbound audiovisual data samples and extracts inbound face spoofing embeds as inbound faceprints. In some cases, the analysis server 102 also receives identity claims of the person associated with the inbound faceprint. For face deepfake detection, the machine learning architecture generates a faceprint similarity score that represents the likelihood that the inbound audiovisual data sample contains a face deepfake, based on the similarity between one or more pre-configured or registered faceprints and the faces in the inbound faceprints. The analysis server 102 outputs one or more face-deepfake similarity or detection scores.

[0042] A face deepfake detection engine can implement, for example, a neural network architecture such as residual networks, Xception networks, and EffecientNets, or a GMM-based architecture.

[0043] Lip-sync estimation engine The machine learning architecture includes a layer that defines a lip-sync estimation engine to determine whether the mismatch between the speaker's voice signal and the speaker's facial gestures exceeds a synchronization threshold. The machine learning architecture applies the lip-sync estimation engine to audiovisual data, or to both audio and image data. The lip-sync estimation engine analyzes the synchronization between the speaker's speech voice and the mouth and facial gestures of a specific speaker as captured in the audiovisual video data. The lip-sync estimation engine generates a lip-sync score that indicates the synchronization quality between the speaker's mouth and voice, thereby indicating the likelihood that the voice seen and heard in the video was emitted by the speaker.

[0044] In some implementations, the lip-sync estimation engine implements signal processing techniques; a non-limiting example can be found in F. Pitie, et al., “Assessment of Audio / Video Synchronisation in Streaming Media,” IEEE, 2014 Sixth International Workshop on Quality of Multimedia Experience (QoMEX), pp. 171-176 (2014). In some implementations, the lip-sync estimation engine implements deep learning algorithms or neural network architectures. A non-limiting example of a deep learning approach can be found in JSChung, et al., “Out of Time: Automated Lip Sync in the Wild,” ACCV, Workshop on Multi-view Lip-Reading (2016), which is incorporated in its entirety by reference.

[0045] The lip-sync estimation engine analyzes facial gestures or mouth / lip image data moving from snapshot to snapshot, and phonemes from segments of audio data. For example, the lip-sync estimation engine can extract lip-sync embeddings as feature vectors representing low-level features of facial gestures, speech phonemes, and associated timing data extracted from audiovisual data. The lip-sync estimation engine focuses on mouth movements by extracting features from rectangular regions around the mouth. The lip-sync estimation engine uses associated pixels or image maps to create lip motion estimates or visual descriptors and combines the motion estimates with audio features detected from segments of audio data or audiovisual data. The lip-sync estimation engine determines the timing lag between the speech phonemes and the lip movements and / or facial gestures. The lip-sync estimation engine generates a lip-sync score that indicates the quality or likelihood of synchronization between the audio and video of a video. In some cases, the lip-sync estimation engine's binary classifier determines whether the audio and visual aspects of a segment or video are synchronized or not based on whether the lip-sync score meets a pre-set synchronization score.

[0046] Biometric information scores, score fusion, and classifiers The machine learning architecture includes one or more layers for scoring operations and / or score fusion operations. As previously mentioned, the machine learning architecture generates various biometric similarity scores for a specific end-user's claimed identity using inbound embeddings (e.g., inbound voiceprint, inbound faceprint, inbound speaker spoof print, inbound face spoof print) extracted from inbound audiovisual data, to be compared against registered embeddings (e.g., registered voiceprint, registered faceprint, pre-configured speaker spoof print, pre-configured face spoof print) such as one or more registered identities' registered audiovisual profiles, which may include the audiovisual profile of the claimed identity. Given the inbound faceprint and inbound voiceprint of the claimed identity, and the registered faceprint and registered voiceprint of the claimed identity, the biometric scorer calculates the mathematical similarity between the corresponding embeddings. The similarity score may be, for example, cosine similarity, the output of a layer defining stochastic linear discriminant analysis (PLDA), the output of a layer defining a support vector machine (SVM), or the output of a layer defining an artificial neural network (ANN).

[0047] The machine learning architecture can perform one or more fusion operations to generate a final output score for specific audiovisual data. In some embodiments, the fusion operation includes score fusion, which algorithmically combines previously generated speaker deepfake detection scores, face deepfake detection scores, and biometric similarity scores. The score fusion operation generates, for example, a final output score, a final audiovisual deepfake score, and / or a final auditory recognition score. The layer for the score fusion operation may be, for example, a single rule-based model or a linear machine learning model (e.g., logistic regression). The machine learning architecture may further include one or more classifier models applied to one or more final scores to classify audiovisual data between "authentic" classification and "fake" classification (sometimes called "deepfake" classification). In the training phase, the analysis server 102 trains the classifier models to perform "authentic" and "fake" classifications according to the labeled data of the training audiovisual data.

[0048] In some embodiments, the fusion operation includes embedding-level fusion (sometimes called “intermediate-level” fusion), which algorithmically combines various embeddings. The speaker engine, face engine, and / or lip-sync estimation engine of the machine learning architecture extracts and concatenates embeddings extracted from audiovisual data to compute one or more scores. For example, the machine learning architecture extracts joint embeddings (e.g., joint inbound embeddings, joint registered embeddings) and a layer of machine learning classifiers trained to classify the joint embeddings into “true” and “fake” classifications to generate one or more scores. The classifier layer may implement linear discriminant analysis (LDA), stochastic linear discriminant analysis (PLDA), support vector machines (SVM), or artificial neural networks (ANN), etc.

[0049] In some embodiments, the fusion operation includes "feature-level" fusion. The analysis server 102 extracts spectral time-series features or other features from segments of audio data (e.g., Mel frequency septal coefficients (MFCCs), linear filter banks (LFBs)) and visual data (e.g., pixel vectors, linear binary patterns (LBPs), discrete cosine transforms (DCTs)). The machine learning architecture then concatenates the features to determine one or more similarity scores. The machine learning architecture includes layers of machine learning classifiers trained to classify between "authentic" and "fake" classifications according to embeddings extracted from the concatenated joint features.

[0050] Area of ​​interest suggestion engine In some embodiments, the machine learning architecture includes a layer that defines a region of interest (ROI) suggestion engine. If the analysis server 102 determines that an audiovisual data sample is not authentic, the analysis server 102 then applies the ROI suggestion engine. The ROI suggestion engine refers to deepfake detection scores to identify a set of one or more trouble segments that are likely to contain speaker deepfake content and / or facial deepfake content. The ROI suggestion engine generates a notification to be displayed on the end-user device 114 or the management device 103. The notification indicates to the end-user or management user the set of one or more trouble segments. In some implementations, to identify trouble segments, the ROI suggestion engine compares one or more segment-level deepfake scores to one or more corresponding pre-configured fake segment thresholds. For example, the ROI suggestion engine determines that a particular segment is likely to contain speaker deepfake content if the speaker deepfake detection score for that segment does not meet the speaker fake segment threshold. In some implementations, the ROI suggestion engine may perform additional or alternative operations (e.g., score smoothing) to detect trouble segments.

[0051] The analysis database 104 or other databases of the system 100 contain any number of corpora of training audiovisual data samples, training audio signals, or training image data, and can access the analysis server 102 via one or more networks 108. In some embodiments, the analysis server 102 employs supervised training to train various layers of a machine learning architecture, and the analysis database 104 contains labels associated with training audiovisual data samples that indicate expected features, embeddings, or classifications for specific training audiovisual data. The analysis server 102 adjusts the weights or hyperparameters of the machine learning architecture according to one or more loss layers during training. The loss layer outputs an error level representing the distance between the expected output indicated by the label (e.g., expected feature, expected embedding, expected classification) and the corresponding predicted output generated by the machine learning architecture (e.g., predicted feature, predicted embedding, predicted classification). In response to the determination that the error level meets the training audiovisual data error threshold, the analysis server 102 modifies the hyperparameters or weights and stores them in the analysis database 104.

[0052] The analysis database 104 can further store any number of registration embeds for audiovisual profiles. The analysis server 102 can generate audiovisual profiles for specific registrants-users of a particular service. In some cases, the analysis server 102 generates audiovisual profiles for celebrities or other well-known individuals.

[0053] The management device 103 or other computing device of system 100 includes a graphical user interface that executes software programming and allows personnel of the analysis system 101 to perform various management tasks, such as configuring the analysis server 102 or user-facilitated analysis calculations performed by the analysis server 102. The management device 103 may be any computing device equipped with a processor and software capable of performing the various tasks and processes described herein. Non-limiting embodiments of the management device 103 may include a server, personal computer, laptop computer, or tablet computer. In calculations, the management user uses the management device 103 to configure calculations for various components of system 100 and to query or instruct such components.

[0054] In some cases, the analysis server 102 or other servers in system 100 transmit the output results generated by the machine learning architecture to the management device 103. The graphical user interface of the management device 103 or other computing device displays some or all of the output result data, such as notifications indicating that the audiovisual data of a particular communication event session contains genuine or fake data, or one or more scores generated by components of the machine learning architecture.

[0055] Components of a system that performs deepfake detection Figure 2 shows the data flow between the components of system 200 that performs human recognition and deepfake detection calculations. Server 202 (or other computing device) applies one or more layers and calculations of machine learning architecture 203 to registered audiovisual profiles on audiovisual data 206 associated with a person's target identity and claimed identity, indicated by one or more inputs from an end-user device. Server 202 determines whether the claimed identity is present in the audiovisual data 206 by performing speaker recognition and facial recognition calculations. Server 202 further determines whether the claimed identity is genuine or fake by performing deepfake detection and / or lip-sync estimation calculations of machine learning architecture 203.

[0056] System 200 includes one or more layers of machine learning architecture 203 and a server 202 containing software configured to perform operations. System 200 further includes a database 204 configured to store one or more registered profiles. In operations, server 202 receives audiovisual data 206 as a media data file or data stream, the audiovisual data 206 including a specific audiovisual media format (e.g., MP4, MOV). The audiovisual data 206 includes audio data 208 containing an audio signal of a speaker's voice and image data 210 containing a video or one or more images of a person. Server 202 further receives end-user input or other data indicating the claimed identity of a specific person who is supposedly spoken in the audiovisual data 206.

[0057] In the training phase, server 202 receives training audiovisual data 206 and applies the machine learning architecture 203 to the training audiovisual data 206 to train the machine learning architecture 203. In the enrollment phase, server 202 receives enrollment audiovisual data 206 and applies the machine learning architecture 203 to the enrollment audiovisual data 206 to develop the machine learning architecture 203 for specific people (e.g., registered users of the service, celebrities). Server 202 generates enrolled profiles that include biometric feature embeddings representing registrant-person aspects, such as enrolled voiceprints and enrolled faceprints. Server 202 stores the profile data in database 204, which it references during the deployment phase.

[0058] During the deployment phase, server 202 receives inbound audiovisual data 206 and applies a machine learning architecture 203 to the inbound audiovisual data 206 to determine whether the inbound audiovisual data 206 is likely to be either a genuine video of a person or a deepfake video of a person. Server 202 generates an inbound profile that includes embedded biometric features representing the characteristics of one or more people (e.g., a speaker's voice, a person's face), such as an inbound voiceprint and an inbound faceprint. In some implementations, server 202 generates one or more scores indicating the similarity between the registered profiles of registered people (e.g., registered voiceprint, registered faceprint) and the inbound profile (e.g., inbound voiceprint, inbound faceprint).

[0059] Server 202 may analyze the audiovisual data 206 into segments of audio data 208 and image data 210. Server 202 can convert the data format of these analyzed segments into different data formats. For example, Server 202 may analyze the audio data 208 of registered audiovisual data 206 into a set of one or more one-second audio segments, and analyze the image data 210 of registered audiovisual data 306 into snapshot images of each second of audiovisual data 206. In this embodiment, the server generates the set of audio data 208 in an audio format (e.g., mp3, wav) and the set of image data 210 in an image format (e.g., jpg, gif).

[0060] Server 202 applies a machine learning architecture 203 to audiovisual data 206 to generate one or more scores. Server 202 refers to the scores and determines the likelihood that the audiovisual data 206 contains genuine video of a person or a deepfake of a person. The machine learning architecture 203 includes layers that define various components, including a speaker recognition engine, a speaker deepfake engine, a face recognition engine, a face deepfake engine, and, in some embodiments, a lip-sync estimation engine. In the computation, Server 202 extracts a set of audio features from audio data 208 and a set of visual features from image data 210. Components of the machine learning architecture 203 extract embeddings, each embedding containing a vector representing a specific set of features extracted from a particular segment of the audio data 208 or image data 210.

[0061] The machine learning architecture 203 includes a speaker engine, which includes a speaker recognition engine and a speaker deepfake detection engine. The speaker recognition engine of the machine learning architecture 203 extracts speaker voiceprints (e.g., training voiceprint, registered voiceprint, inbound voiceprint) based on speaker recognition features and embeddings extracted from audio data 208. The speaker deepfake detection engine of the machine learning architecture 203 extracts speaker spoof prints (e.g., training speaker spoof print, registered speaker spoof print, inbound speaker spoof print) based on speaker deepfake detection features and embeddings extracted from audio data 208. The speaker engine outputs one or more similarity scores for speaker recognition and speaker deepfake detection.

[0062] The machine learning architecture 203 includes a face engine, which includes a face recognition engine and a face deepfake detection engine. The face recognition engine of the machine learning architecture 203 extracts face prints (e.g., training face print, registered face print, inbound face print) based on face recognition features and embeddings extracted from image data 210. The face deepfake engine of the machine learning architecture 203 extracts face spoof prints (e.g., training face spoof print, registered face spoof print, inbound face spoof print) based on face deepfake detection features and embeddings extracted from image data 210. The face engine outputs one or more similarity scores for face recognition and face deepfake detection.

[0063] The lip-sync estimation engine of machine learning architecture 203 generates a lip-sync score. The lip-sync estimation engine analyzes image data of facial gestures or mouth / lips moving from snapshot to snapshot, and phonemes from segments of audiovisual data 206. For example, the lip-sync estimation engine can extract lip-sync embeddings as feature vectors representing low-level features of facial gestures, speech phonemes, and associated timing data extracted from audiovisual data 206. The lip-sync estimation engine focuses on mouth movements by extracting features of a rectangular region around the mouth in image data 210 or audiovisual data 206. The lip-sync estimation engine uses associated pixels or image maps to create lip movement estimates or visual descriptors and combines the movement estimates with speech features detected from segments of audio data 208 or audiovisual data 206. The lip-sync estimation engine determines the timing lag between the speech phonemes and the lip movements and / or facial gestures. The lip-sync estimation engine generates a lip-sync score that indicates the quality or likelihood of synchronization between the audio and video aspects of the video.

[0064] The machine learning architecture 203 includes one or more layers for scoring operations and / or score fusion operations. The score fusion layer outputs one or more final output scores indicating the likelihood of identity recognition (e.g., speaker recognition, face recognition), the likelihood of deepfake detection (e.g., speaker deepfake, face deepfake), and the likelihood of lip-sync quality. The machine learning architecture 203 includes one or more classification layers trained to classify audiovisual data 206 as either genuine or fake.

[0065] Figure 3 shows the data flow between components of system 300 that perform registration calculations to build registered audiovisual profiles. System 300 includes a server 302 and a database 304. Server 302 receives registered audiovisual data 306 as a media data file or data stream, and the registered audiovisual data 306 includes a specific audiovisual media format (e.g., mp4, mov). The machine learning architecture includes layers that define a speaker engine and an image engine. Server 302 applies the components of the machine learning architecture to the registered audiovisual data 306 to generate a registered profile of a specific person, and the registered profile includes a registered voice print 312 and a registered face print 314. Server 302 applies the speaker engine to registered audio data 308 to generate a registered voice print 312 and applies the image engine to registered image data 310 to generate a registered face print 314.

[0066] In some cases, the server 302 may receive registered audio data 308 and / or registered image data 310 that are different from the registered audiovisual data 306. In these cases, the registered audio data 308 includes a specific audio format (e.g., mp3, wav) or image format (e.g., jpg, gif). The server 302 can analyze the registered audiovisual data 306 into segments of registered audio data 308 and registered image data 310. The server 302 can convert the data formats of these analyzed segments into different data formats. For example, the server 302 may analyze the audio data of the registered audiovisual data 306 into a set of one or more one-second audio segments, and analyze the image data of the registered audiovisual data 306 into snapshot images of each second of the registered audiovisual data 306. In this embodiment, the server generates a set of registered audio data 308 in audio format and a set of registered image data 310 in image format.

[0067] Server 302 extracts a set of features from registered voice data 308 and registered image data 310, and extracts a set of features from registered image data 310. The speaker engine extracts speaker embeddings as vectors representing the features of a specific segment of the registered voice data 308. The speaker engine algorithmically combines the speaker embeddings (e.g., mean values) to extract the registered voiceprint 312. Similarly, the image engine extracts image embeddings as vectors representing the features of a specific registered image data 310. The image engine algorithmically combines the image embeddings to extract the registered faceprint 314. Server 302 stores the registered voiceprint 312 and registered faceprint 314 in database 304, which Server 302 later references during the deployment phase.

[0068] Examples of process calculations Figure 4 illustrates the steps of implementing Method 400 for implementing one or more machine learning architectures for deepfake detection (e.g., speaker impersonation, facial impersonation) and identity recognition (e.g., speaker recognition, facial recognition) using various biometric information. Embodiments may include additional, fewer, or different operations than those described in Method 400. A server performs the steps of Method 400 by executing machine-readable software code that includes one or more machine learning architectures, but it should be understood that any number of computing devices and / or processors may perform the various operations of Method 400.

[0069] In step 402, the server acquires training audiovisual data during the training phase, including training image data and training audio data for specific individuals. During the training phase, the server receives training audiovisual data (e.g., training audiovisual data samples) or generates various simulated audiovisual data samples, which may include degraded or mixed copies of training audiovisual data, training image data, or training audio data.

[0070] Servers or layers of a machine learning architecture can perform various preprocessing operations on input audiovisual data (e.g., training audiovisual data, registered audiovisual data, inbound audiovisual data), including audio data (e.g., speaker audio signals) and visual data (e.g., facial images). These preprocessing operations may include, for example, extracting low-level features from speaker audio signals or visual image data and transforming these features into various alternative representations of the features by performing Short-Time Fourier Transform (SFT), Fast Fourier Transform (FFT), or other transformation operations (e.g., transforming audio data from a time-domain representation to a frequency-domain representation). Preprocessing operations may also include parsing the audio signals or visual data into frames or subframes and performing various normalization or scaling operations. Selectively, the server performs any number of preprocessing operations before feeding the audiovisual data to the layers of the machine learning architecture. The server may perform various preprocessing operations in one or more operational phases, but the specific preprocessing operations performed may differ between operational phases. The server can perform various preprocessing operations independently of the machine learning architecture, or as an inner layer of the machine learning architecture's network.

[0071] A server or layer in a machine learning architecture can perform various augmentation operations on audiovisual data for training or registration purposes. These augmentation operations generate various types of distortion or degradation in the input audio signal, which is then consumed by, for example, a convolution operation that generates feature vectors. The server can perform these augmentation operations as separate operations from the neural network architecture or as an augmentation layer within the network. The server can perform various augmentation operations in one or more operational phases, but the specific augmentation operations performed may differ between operational phases.

[0072] In step 404, the server trains the machine learning architecture by applying layers of the machine learning architecture onto training audiovisual data. The server applies layers of the machine learning architecture according to the operational layers of a particular component of the machine learning architecture to generate predicted outputs. The loss layer or another function of the machine learning architecture determines the level of error (e.g., one or more similarities, distances) between the predicted output and the labels or other data that represent the expected output. The loss layer or another aspect of the machine learning architecture tunes the hyperparameters until the error level of the predicted output (e.g., predicted embedding, predicted score, predicted classification) satisfies a threshold level or error with respect to the expected output (e.g., expected embedding, expected score, expected classification). The server then stores the hyperparameters, weights, or other terms of a particular machine learning architecture in a database, thereby "fixing" a particular component of the machine learning architecture and one or more models.

[0073] In step 406, the server places the neural network into any enrollment operational phase and acquires enrollment audiovisual data to generate enrollment embeddings for the enrolled profile. The server applies layers of the machine learning architecture to the enrollment audiovisual data to generate enrollment embeddings for the enrollment audiovisual profile of a particular person. The server receives a sample of the enroller's enrollment audiovisual data and applies the machine learning architecture to generate various enrollment feature vectors, including, for example, speaker spoof print, registrant voice print, face spoof print, and registrant face print. The server can enable and / or disable specific layers of the machine learning architecture during the enrollment phase. For example, the server typically enables and applies each layer during the enrollment phase, but in some implementations, the server can disable specific classification layers.

[0074] When extracting specific embeddings of a registrant (e.g., voiceprint, faceprint, or disguise print(s)), the machine learning architecture generates a set of registrant embeddings as a feature vector based on corresponding type features associated with the specific type of embedding. The machine learning architecture then algorithmically combines the corresponding type of embeddings to generate a voiceprint, faceprint, or speaker / face disguise print. The server stores each registrant embedding in a non-transient storage medium in the database.

[0075] In step 408, the server deploys the neural network architecture during the deployment phase and receives inbound audiovisual data. The server analyzes the inbound audiovisual data into segments and extracts low-level features from the segments. Next, the server extracts various types of embeddings (e.g., inbound voiceprint, inbound faceprint, inbound disguise print(s)) associated with a specific person from the inbound audiovisual data. In some cases, the server receives data input containing identity claims that identify a specific person.

[0076] In step 410, the server determines whether the inbound audiovisual data is authentic by applying a machine learning architecture to the features of the inbound audiovisual data. The machine learning architecture generates one or more similarity scores based on the similarity or difference between the inbound embedding and the corresponding registered embedding, which may be a registered embedding associated with the person of the identity claim.

[0077] In one embodiment, the machine learning architecture extracts inbound voiceprints and outputs a similarity score indicating the similarity between the inbound voiceprint and the registered voiceprint for speaker recognition. Similarly, for face recognition, the machine learning architecture extracts inbound faceprints and registered prints and outputs a similarity score indicating the distance between the inbound faceprint and the registered faceprint. A larger distance indicates lower similarity, which may indicate that the speaker or face in the inbound audiovisual data is less likely to match the voice or face of the registered speaker. In this embodiment, if the similarity score meets the speaker or face recognition threshold, the server identifies (or recognizes) the match with the speaker or face as the registered speaker.

[0078] In another embodiment, the neural network architecture extracts inbound face spoofs and inbound speaker spoofs and outputs a similarity score indicating the similarity between the inbound speaker / face spoof print and the corresponding registered speaker / face spoof print. A larger distance may indicate a lower similarity between the inbound speaker / face spoof print and the registered speaker / face spoof print, thus suggesting a lower probability that the inbound audiovisual data is spoofed. In this embodiment, the server determines that the speaker or face in the inbound audiovisual data is a deepfake if the similarity score meets the deepfake detection threshold.

[0079] In some embodiments, the machine learning architecture includes one or more fusion operations that generate a combined similarity score using a speaker / face similarity score (based on a comparison of voice prints) and a corresponding speaker / face deepfake detection score (based on a comparison of fake prints). The server generates the combined similarity score by summing or algorithmically combining the speaker / face similarity score and the corresponding speaker / face deepfake detection score. The server then determines whether the combined similarity score meets a confirmation or authentication threshold score. As discussed herein, the machine learning architecture may implement additional or alternative score fusion operations to determine various similarity scores and classifications.

[0080] Figure 5 shows the data flow of components of system 500 for implementing one or more machine learning architectures for deepfake detection (e.g., speaker deepfake disguise, face deepfake disguise) and biometric recognition (e.g., speaker recognition, face recognition) according to score-level score fusion calculations 524. A server or other computing device runs software for one or more machine learning architectures 507 configured to perform various calculations in system 500.

[0081] The machine learning architecture 507 receives audiovisual data 502 in the form of a computer file or data stream containing a video clip. The audiovisual data 502 includes audio data 504 containing a speaker's audio signal and image data 506 containing an image of a person's face. The machine learning architecture 507 includes a speaker engine 508 that takes in the audio data 504, a face engine 512 that takes in the image data 506, and a lip-sync estimation engine 510 that takes in the audiovisual data 502 and / or both the audio data 504 and the image data 506. The server parses the audio data 504, the image data 506, and / or the audiovisual data 502 into segments or frames of a predetermined size (e.g., length, snapshot, data size). The server then extracts various types of low-level features from the corresponding portions of the audiovisual data 502, the audio data 504, and / or the image data 506. The server applies a machine learning architecture 507 to features extracted from audiovisual data 502, audio data 504, and / or image data 506 to generate biometric similarity scores (e.g., speaker similarity score 514, face similarity score 520).

[0082] The speaker engine 508 of the machine learning architecture 507 extracts speaker recognition embeddings (voice prints) for specific features of the audio data 504 and extracts audio deepfake embeddings for specific features of the audio data 504. The speaker engine 508 refers to registered voice prints in the database and determines the similarity between the input voice print and the registered voice print to generate a speaker biometric similarity score 514. The speaker engine 508 refers to one or more pre-configured speaker disguise prints in the database and determines the similarity between the audio deepfake embedding and the pre-configured speaker disguise print to generate a speaker deepfake score 516 for audio deepfake detection. The speaker engine 508 outputs the speaker similarity score 514 and the speaker deepfake score 516 to the score fusion operation 524.

[0083] The face engine 512 of the machine learning architecture 507 extracts face recognition embeddings (face prints) for specific features of the image data 506 and extracts face deepfake embeddings for specific features of the image data 506. The face engine 512 refers to registered face prints in the database and determines the similarity between the input face print and the registered voice print to generate a face biometric similarity score 520. The face engine 512 refers to one or more pre-configured face spoof prints in the database and determines the similarity between the face deepfake embedding and the pre-configured face spoof print to generate a face deepfake score 522 for face deepfake detection. The face engine 512 outputs the face similarity score 520 and the face deepfake score 522 to the score fusion operation 524.

[0084] For segments of audiovisual data 502, audio data 504, and / or image data 506, the lip-sync estimation engine 510 outputs a lip-sync score 518. For a specific segment of the video in the audiovisual data 502, the lip-sync estimation engine 510 can extract features of lip / mouth gestures, phonemes, and / or timing data, and extract feature vector embeddings that represent estimated lip-sync features for a given segment. The lip-sync score 518 indicates the likelihood that both the speech and lip movements are synchronized or not synchronized to a predetermined degree. The lip-sync estimation engine 510 outputs the lip-sync score 518 to a score fusion function 524.

[0085] The score fusion function 524 of method 500 algorithmically combines scores 514, 516, 518, 520, and 522 generated using audiovisual data 502, audio data 504, and image data 506 to output a final audiovisual score 526. The machine learning architecture 507 determines whether the audiovisual data 502 is genuine or fake if the final output score 526 meets a certain threshold score. In some cases, the machine learning architecture 507 includes a classifier layer trained to classify the audiovisual data 502 as genuine or fake based on the final output score 526 when represented as a vector.

[0086] Figure 6 shows the data flow of components of system 600 for implementing one or more machine learning architectures for deepfake detection (e.g., speaker spoofing, face spoofing) and human recognition (e.g., speaker recognition, face recognition) according to embedded-level score fusion calculations 624. A server or other computing device runs software for one or more machine learning architectures 607 configured to perform various calculations in system 600.

[0087] The machine learning architecture 607 receives audiovisual data 602 in the form of a computer file or data stream containing video clips. The audiovisual data 602 includes audio data 604 containing a speaker's audio signal and image data 606 containing an image of a person's face. The machine learning architecture 607 includes a speaker engine 608 that takes in the audio data 604, a face engine 612 that takes in the image data 606, and a lip-sync estimation engine 610 that takes in the audiovisual data 602 and / or both the audio data 604 and the image data 606. The server parses the audio data 604, the image data 606, and / or the audiovisual data 602 into segments or frames of a predetermined size (e.g., length, snapshot, data size). The server then extracts various types of low-level features from the corresponding portions of the audiovisual data 602, the audio data 604, and / or the image data 606. The server applies a machine learning architecture 607 to features extracted from audiovisual data 602, audio data 604, and / or image data 606, and extracts various types of embeddings 614, 616, 618, 620, and 622 using the corresponding types of features.

[0088] The speaker engine 608 of the machine learning architecture 607 extracts speaker recognition embeddings 614 (voice prints) for specific features of the audio data 604 and voice spoofing print embeddings 616 for specific features of the audio data 604. The speaker engine 608 outputs the speaker voice prints 614 and speaker spoofing prints 616 to the score fusion operation 624.

[0089] The face engine 612 of the machine learning architecture 607 extracts face recognition embeddings 620 (face prints) for specific features of the image data 606, and extracts face spoof print embeddings 622 for specific features of the image data 606. The face engine 612 outputs the face print embeddings 620 and face spoof prints 622 to the score fusion operation 624.

[0090] For segments of audiovisual data 602, audio data 604, and / or image data 606, the lip-sync estimation engine 610 outputs a lip-sync score 618. For a specific segment of the video in the audiovisual data 602, the lip-sync estimation engine 610 can extract features of lip / mouth gestures, phonemes, and / or timing data, and extract feature vectors as lip-sync embeddings 618 that represent estimated lip-sync features for a given segment. The lip-sync estimation engine 610 outputs the lip-sync embeddings 618 to the score fusion function 624.

[0091] The score fusion function 624 of method 600 algorithmically combines (e.g., concatenates) the embeddings 614, 616, 618, 620, and 622 to generate a joint embedding using audiovisual data 602, audio data 604, and image data 606. The score fusion function 624 or other functions of the machine learning architecture 607 determine a joint similarity score (shown as the final score 626) based on the distance or similarity between the joint embedding of the audiovisual data 602 and the registered joint embedding stored in the database.

[0092] The machine learning architecture 607 determines whether the audiovisual data 602 is genuine or fake based on whether the final output score 626 meets a pre-set threshold score. In some cases, the machine learning architecture 607 includes a classifier layer trained to classify the audiovisual data 602 as genuine or fake based on the final output score 626, which is represented as a vector.

[0093] Figure 7 shows the data flow of components of system 700 for implementing one or more machine learning architectures for deepfake detection (e.g., speaker spoofing, face spoofing) and human recognition (e.g., speaker recognition, face recognition) according to feature level score fusion calculation 724. A server or other computing device runs software for one or more machine learning architectures 707 configured to perform various calculations in system 700.

[0094] The machine learning architecture 707 receives audiovisual data 702 in the form of a computer file or data stream containing video clips. The audiovisual data 702 includes audio data 704 containing a speaker's audio signal and image data 706 containing an image of a person's face. The machine learning architecture 707 includes a speaker engine 708 that takes in the audio data 704, a face engine 712 that takes in the image data 706, and a lip-sync estimation engine 710 that takes in the audiovisual data 702 and / or both the audio data 704 and the image data 706. The server parses the audio data 704, the image data 706, and / or the audiovisual data 702 into segments or frames of a predetermined size (e.g., length, snapshot, data size). The server then extracts various types of low-level features from the corresponding portions of the audiovisual data 702, the audio data 704, and / or the image data 706. The server applies the machine learning architecture 707 and the feature level score fusion function 724 to various types of features 714, 716, 718, 720, and 722, and extracts one or more joint embeddings that the machine learning architecture 707 compares with one or more corresponding registered joint embeddings stored in the database.

[0095] The speaker engine 708 of the machine learning architecture 607 extracts specific low-level speaker recognition features 714 and voice spoofing print features 716 from the audio data 704. The speaker engine 708 concatenates the speaker voice print features 714 and the speaker spoofing print features 716 and outputs them to the score fusion operation 724.

[0096] The face engine 712 of the machine learning architecture 707 extracts specific low-level face recognition features 720 and face camouflage print features 722 from the image data 706. The face engine 712 concatenates the face camouflage print features 720 and face camouflage print features 722 and outputs this to the score fusion operation 724.

[0097] For segments of audiovisual data 702, audio data 704, and / or image data 706, the lip-sync estimation engine 710 extracts low-level lip-sync features 718 of lip / mouth gestures, phonemes, and / or timing data for specific segments of the video in the audiovisual data 702. The lip-sync estimation engine 710 outputs the lip-sync features 718 to a score fusion function 724.

[0098] The method 700 of the score fusion function 724 algorithmically combines (e.g., concatenates) various types of features 714, 716, 718, 720, and 722 to extract joint embeddings using audiovisual data 702, audio data 704, and image data 706. The score fusion function 724 determines a joint similarity score (shown as the final score 726) based on the similarity between the joint embeddings and registered joint embeddings in the database. The machine learning architecture 707 determines whether the audiovisual data 702 is genuine or spoofed based on whether the final output score 726 meets a pre-set threshold score. In some cases, the machine learning architecture 707 includes a classifier layer trained to classify the audiovisual data 672 as genuine or spoofed based on the final output score 726, which is represented as a vector.

[0099] Figure 8 shows the execution steps of Method 800 for implementing one or more machine learning architectures for deepfake detection (e.g., speaker spoofing, face spoofing) and human recognition (e.g., speaker recognition, face recognition) according to an embodiment. The machine learning architecture includes a layer that analyzes clear audio and visual biometric embeddings (steps 806-814) and a layer that analyzes audiovisual embeddings as lip-sync estimates (step 805). In Method 800, the server generates a segment-level fused score for each segment using the audio and visual embeddings (step 814), and generates a recording-level fused score of segment-level scores and lip-sync estimate scores for most or all of the audiovisual data (e.g., video clips) (step 816). Embodiments may perform score fusion operations on various levels of data (e.g., feature level, embedding level) and various levels of data quantities (e.g., full recording, segment).

[0100] In step 802, the server acquires audiovisual data. During the training or enrollment phase, the server may receive training or enrollment audiovisual data samples from end-user devices, a database containing one or more corpora of training or enrollment audiovisual data, or a third-party data source hosting the training or enrollment audiovisual data. In some cases, the server applies data augmentation operations to the training audiovisual data to generate simulated audiovisual data for additional training audiovisual data. During the deployment phase, the server receives inbound audiovisual data samples from end-user devices or a third-party server hosting software services that generate inbound audiovisual data.

[0101] In step 804, the server analyzes the audiovisual data into segments or frames. The server applies a machine learning architecture to the segments for biometric information embedding (steps 806-814) and applies the machine learning architecture to some or all of the audiovisual data for lip-sync estimation (step 805). The server extracts various types of low-level features from the segments of audiovisual data.

[0102] In step 805, for some or all of the segments, the server extracts lip-sync embeddings using the features of the specific segment. The server then applies a lip-sync estimation engine to the lip-sync embeddings of the segments to determine the lip-sync score.

[0103] In step 806, for some or all of the segments, the server extracts biometric embeddings (e.g., voice prints, face prints) using the features of the specific segments. In step 808, the server generates a speaker recognition similarity score based on the similarity between the speaker voice print and the registered speaker voice print. The server further generates a face recognition similarity score based on the similarity between the face print and the registered face print.

[0104] In any step 810, the server determines whether both the speaker similarity score and the face similarity score satisfy one or more corresponding recognition threshold scores. The server compares the speaker similarity score with the corresponding speaker recognition score and the face similarity score with the corresponding face recognition score. If the server determines that one or more biometric similarity scores cannot satisfy the corresponding recognition threshold, method 800 proceeds to step 812. Alternatively, if the server determines that one or more biometric similarity scores satisfy the corresponding recognition threshold, method 800 proceeds to step 814.

[0105] Additionally or alternatively, in some embodiments, the server fuses the types of biometric information embeddings to generate a joint biometric information embedding (step 806) and generates a joint similarity score (step 808). The server then determines whether the joint similarity score satisfies the joint recognition score by comparing the inbound joint embedding with the registered joint embedding.

[0106] The decision-making step 810 is optional. The server does not need to determine whether one or more similarity scores meet the corresponding recognition threshold. In some embodiments, the server applies a deepfake detection function (step 812) in each embodiment, thereby omitting the optional step 810.

[0107] If, in step 812, the server determines that one or more biometric similarity scores do not meet the corresponding recognition threshold (step 810), the server then applies layers of machine learning architecture for speaker deepfake detection and face deepfake detection. The machine learning architecture extracts deepfake detection embeddings (e.g., speaker spoof prints, face spoof prints) using low-level features extracted for each specific segment. The server generates speaker deepfake detection scores based on the distance or similarity between the speaker spoof print and one or more registered speaker spoof prints. The server further generates face deepfake detection scores based on the distance or similarity between the face spoof print and one or more registered face spoof prints.

[0108] In step 814, the server applies a score fusion operation to generate a score-level fused score using the scores generated for each segment (e.g., face recognition similarity score, speaker recognition similarity score, speaker-deepfake detection score, face-deepfake detection score), thereby generating a segment-level score.

[0109] In step 816, the server applies a score fusion operation to generate a final fused score using the segment-level scores and the lip-sync estimation score generated by the server (step 805), thereby generating a recording-level score. In the current embodiment, one or more segment-level scores represent, for example, the final biometric evaluation score, the likelihood of deepfake, and the likelihood of speaker / face recognition. The lip-sync estimation score can be applied as a confidence adjustment or confidence check to determine whether the entire video contains genuine or deepfake content. The recording-level score may be calculated as a mean or median calculation, such as the mean of the Top-N scores (e.g., N=10), or as a heuristic calculation.

[0110] In step 818, the server generates a recording score for the audiovisual data. The server compares the recording score to a genuine video threshold to determine whether the inbound audiovisual data contains genuine or spoofed data. The server generates a notification based on the final output, which may include other potential information such as whether the audiovisual data is genuine or spoofed, or one or more scores generated by a machine learning architecture. The server is configured to generate notifications according to any number of protocols and machine-readable software code and to display them on the server or the user interface of an end-user device.

[0111] In some embodiments, the server runs a layer of a machine learning architecture's region of interest (ROI) suggestion engine when the audiovisual data does not meet the authentic video threshold. The ROI suggestion engine references segment-level audiovisual deepfake scores and identifies a set of one or more trouble segments that are likely to contain speaker and / or facial deepfake content. The ROI suggestion engine may generate a notification to be displayed on the end-user device. The notification shows the user the set of one or more trouble segments. In some implementations, to identify trouble segments, the ROI suggestion engine compares one or more segment-level deepfake scores to one or more corresponding pre-configured fake segment thresholds. For example, the ROI suggestion engine might determine that a particular segment is likely to contain speaker deepfake content if its speaker deepfake detection score does not meet the speaker fake segment threshold. The ROI suggestion engine may perform additional or alternative operations (e.g., score smoothing) to detect trouble segments.

[0112] Additional exemplary embodiments Detection of malicious deepfake videos on the internet In some embodiments, a website or cloud-based server, such as a social media site or forum website for exchanging video clips, includes one or more servers running the machine learning architecture described herein. The host infrastructure includes a web server, an analytics server, and a database for continuously storing registered voiceprints and faceprints for identity purposes.

[0113] In the first embodiment, a machine learning architecture detects deepfake videos posted to and hosted on a social media platform. The end user provides a sample of registered audiovisual data to the analysis server, which extracts registered voiceprints and registered faceprints. The host system can further generate registered voiceprints and registered faceprints of celebrities, as well as registered speaker impersonation prints and face impersonation prints. During deployment, the analysis server applies registered voiceprints, registered faceprints of celebrities, registered speaker impersonation prints, and registered face impersonation prints to any audiovisual data file or data stream on the social media platform. The analysis server determines whether the inbound audiovisual data contains deepfake content of a particular celebrity. If the analysis server detects deepfake content, it identifies the trouble segment containing the deepfake content and generates recommendations indicating the trouble segment.

[0114] In a second embodiment, the machine learning architecture detects malicious deepfake adult video content (typically of celebrities obtained without their consent) on internet forums. End users provide registered audiovisual data samples to the analysis server, which extracts registered voiceprints and registered faceprints. The host system can further generate registered voiceprints and registered faceprints of celebrities, as well as registered speaker and faceprints. During deployment, for any video on the internet forum, the analysis server downloads and analyzes samples of audiovisual data layer by layer for speaker and face deepfake detection. The analysis server determines whether the downloaded audiovisual data contains deepfake content of a particular celebrity. If the analysis server detects deepfake content, it identifies the trouble segment containing the deepfake content and generates recommendations indicating the trouble segment.

[0115] In the third embodiment, a host server and an analysis server host a celebrity reputation service on social media platforms such as Twitter and Facebook. The analysis server generates registered voiceprints and registered faceprints for celebrity-users who have purchased this add-on service. For audiovisual data samples posted and hosted on social media platforms, the analysis server can detect whether the audiovisual data contains deepfake content.

[0116] Various exemplary logic blocks, modules, circuits, and algorithmic processes described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this hardware and software compatibility, various exemplary components, blocks, modules, circuits, and processes are generally described above with respect to their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. A person skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as causing a departure from the claims of the present invention.

[0117] Embodiments implemented in computer software may be implemented in software, firmware, middleware, microcode, hardware description languages, or any combination thereof. A code segment or machine-executable instruction may represent a procedure, function, subprogram, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, attributes, or memory contents. Information, arguments, attributes, data, etc., may be passed, transferred, or transmitted via any suitable means, such as memory sharing, message passing, token passing, or network transmission.

[0118] Actual software code or specialized control hardware used to implement these systems and methods does not limit the invention. Therefore, while the calculations and behaviors of the systems and methods have been described without reference to specific software code, it should be understood that software and control hardware may be designed to implement the systems and methods based on the descriptions herein.

[0119] When implemented in software, a function may be stored as one or more instructions or codes on a non-temporary computer-readable or processor-readable storage medium. The steps of the methods or algorithms disclosed herein may be embodied in processor-executable software modules that may reside on a computer-readable or processor-readable storage medium. Non-temporary computer-readable or processor-readable media include both computer storage media and tangible storage media that facilitate the transfer of computer programs from one location to another. Non-temporary processor-readable storage media may be any available medium that can be accessed by a computer. Examples, but not limited to, include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, or any other tangible storage media that can be used to store desired program code in the form of instructions or data structures and can be accessed by a computer or processor. Disk and disc, as used herein, include compact discs (CDs), laser discs, optical discs, digital general-purpose discs (DVDs), floppy discs, and Blu-ray discs, where a disk typically reproduces data magnetically, and a disc reproduces data optically using a laser. The above combinations should also be included within the scope of computer-readable media. Furthermore, the calculations of a method or algorithm may exist as one or any combination or set of codes and / or instructions on non-temporary processor-readable media and / or computer-readable media that can be incorporated into a computer program product.

[0120] The above description of the disclosed embodiments is provided to enable those skilled in the art to create or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments without departing from the spirit or scope of the subject matter disclosed herein. Thus, this disclosure is not intended to be limited to the embodiments shown herein, but should be given the broadest scope that is consistent with the following claims and the principles and novel features disclosed herein.

[0121] While various aspects and embodiments are disclosed, other aspects and embodiments are contemplated. The various aspects and embodiments disclosed are for illustrative purposes only and are not intended to limit the scope, and the true scope and spirit are shown by the following claims.

Claims

1. A computer implementation method, A process of acquiring audiovisual data samples, including audiovisual data, by computer, The process of applying a machine learning architecture to the audiovisual data using the computer, generating a similarity score using biometric information embedding extracted from the audiovisual data, generating a lip-sync score using one or more lip-sync embedding extracted from the audiovisual data, and generating a deepfake score using speaker disguise print embedding and face disguise print embedding extracted from the audiovisual data, A step of causing the computer to generate a final output score indicating the likelihood that the audiovisual data is authentic, using the similarity score, the lip-sync score, and the deepfake score; Computer implementation methods, including those mentioned above.

2. The method according to claim 1, further comprising the step of having the computer identify the audiovisual data sample as a genuine data sample in response to the determination that the final output score meets a threshold.

3. The method according to claim 1, further comprising the step of the computer identifying deepfake content in the audiovisual data in response to the determination that the deepfake score satisfies a deepfake detection threshold.

4. The method according to claim 1, wherein the biometric information implantation includes at least one of a voice print and a face print.

5. The computer, The process involves applying the speaker embedding extraction engine of the machine learning architecture to the audio signal of the audiovisual data to extract the voice print of the audiovisual data sample, The process of extracting the speaker disguise print embedding from the audiovisual data by applying the audio disguise print embedding extraction engine of the machine learning architecture to the audio signal of the audiovisual data using the computer, The method according to claim 1, further comprising:

6. The steps of extracting face prints from audiovisual data by applying the face print embedding extraction engine of the machine learning architecture to the visual media of the audiovisual data using the computer, The computer performs the following steps: extracting face-fake prints from the audiovisual data by applying the visual disguise print embedding engine of the machine learning architecture to the visual media of the audiovisual data; The method according to claim 1, further comprising:

7. The method according to claim 1, further comprising the step of using the computer to extract features for embedding a speaker voiceprint of the biometric information, wherein the features are extracted from the audio data of the audiovisual data.

8. The method according to claim 1, further comprising the step of using the computer to extract features for one or more biometric information embeddings for faceprint embedding, wherein the features are extracted from image data of the audiovisual data.

9. The method according to claim 1, further comprising the step of analyzing the audiovisual data sample into a plurality of segments having a predetermined length using the computer, wherein the computer generates one or more biometric information embeds and at least one disguised print for each segment.

10. The method according to claim 1, further comprising the step of generating the lip-sync score by applying the lip-sync estimation engine of the machine learning architecture to the audiovisual data using the computer, wherein the computer uses the lip-sync score to generate the final output score.

11. The method according to claim 1, wherein the audiovisual data includes audio data, image data, or both audio data and image data.

12. It is a system, Computers Obtain an audiovisual data sample that includes audiovisual data, Applying a machine learning architecture to the audiovisual data, generating a similarity score using biometric data embedding extracted from the audiovisual data, generating a lip-sync score using one or more lip-sync embeddings extracted from the audiovisual data, generating a deepfake score using speaker camouflage print embeddings and face camouflage print embeddings extracted from the audiovisual data, and Using the similarity score, the lip-sync score, and the deepfake score, a final output score is generated that indicates the likelihood that the audiovisual data is authentic. A system having a processor configured in such a way.

13. The system according to claim 12, wherein the computer is further configured to identify the audiovisual data sample as a genuine data sample in response to determining that the final output score meets a threshold.

14. The system according to claim 12, wherein the computer is further configured to identify deepfake content in the audiovisual data in response to determining that the deepfake score meets a deepfake detection threshold.

15. The system according to claim 12, wherein the biometric information implantation includes at least one of a voice print and a face print.

16. The aforementioned computer further, By applying the speaker embedding extraction engine of the machine learning architecture to the audio signal of the audiovisual data, the voice print of the audiovisual data sample is extracted, and The speaker disguise print of the audiovisual data is extracted by applying the audio disguise print embedding extraction engine of the machine learning architecture to the audio signal of the audiovisual data. The system according to claim 12, configured as follows.

17. The aforementioned computer further, By applying the faceprint embedding extraction engine of the machine learning architecture to the visual media of the audiovisual data, faceprints are extracted from the audiovisual data, and By applying the visual camouflage print embedding extraction engine of the machine learning architecture to the visual media of the audiovisual data, face camouflage prints are extracted from the audiovisual data. The system according to claim 12, configured as follows.

18. The system according to claim 12, wherein the computer is further configured to extract features for the speaker voiceprint embedding of the biometric information embedding, the features being extracted from the audio data of the audiovisual data.

19. The system according to claim 12, wherein the computer is further configured to extract features for the biometric information embedding of a faceprint, the features being extracted from the audiovisual data image data.

20. The system according to claim 12, wherein the computer is further configured to analyze the audiovisual data sample into a plurality of segments having a predetermined length, and the computer generates the biometric information embedding and at least one disguised print for each segment.

21. The system according to claim 12, wherein the computer is further configured to generate the lip-sync score by applying the lip-sync estimation engine of the machine learning architecture to the audiovisual data, and the computer uses the lip-sync score to generate a final output score.

22. The system according to claim 12, wherein the audiovisual data includes audio data, image data, or both audio data and image data.