Multimodal neural lip synchronization for automatic dubbing

WO2025186768A8PCT designated stage Publication Date: 2025-10-02SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/052447
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-07
Filing Date
2025-03-06
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing audio-driven talking head generation techniques, particularly person-generic ATHG, struggle to maintain accurate lip-speech synchronization and visual quality without requiring speaker-specific training data, often overlooking the dynamic relationship between audio cues and facial movements, especially lip and jaw regions.

Method used

A system employing multimodal neural lip synchronization for automatic dubbing uses temporal shuffling augmentation, generation-based and representation-based temporal contrastive learning, and part-level audio-visual alignment to enhance lip-speech synchronization, focusing on the dynamic relationship between audio cues and facial movements, especially lip and jaw regions, without needing explicit labels.

Benefits of technology

The system generates realistic and temporally synchronized talking head videos with improved lip-speech synchronization and visual quality, enhancing adaptability to diverse speaking rates and preventing overfitting, thus producing high-fidelity talking face videos applicable to any speaker.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025052447_02102025_PF_FP_ABST
    Figure IB2025052447_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The system acquires an audio-visual input including a first set of video frames and a set of audio spectrograms associated with a human speaker. The system generates a second set of video frames from the input and computes a first temporal contrastive loss based on the second set of video frames. The system extracts a set of visual embeddings for the first set of video frames, and a set of audio embeddings based on the set of audio spectrograms. The system computes a second temporal contrastive loss based on the set of visual embeddings and the set of audio embeddings. The system further computes an alignment loss between the set of audio spectrograms and lower-half face of the human speaker in the first set of video frames. The system trains the generative neural network based on the first temporal contrastive loss, the second temporal contrastive loss, and the alignment loss.
Need to check novelty before this filing date? Find Prior Art

Description

DOCKET NO. SYP354694WO02 MULTIMODAL NEURAL LIP SYNCHRONIZATION FOR AUTOMATIC DUBBING CROSS-REFERENCE TO RELATED APPLICATIONS / INCORPORATION BY REFERENCE

[0001] This Application also makes reference to Indian Provisional Application Ser. No. 202411016441 which was filed on March 07, 2024. The above stated Patent Application is hereby incorporated herein by reference in its entirety. FIELD

[0002] Various embodiments of the disclosure relate to generative neural networks for audio-visual data. More specifically, various embodiments of the disclosure relate to a multimodal neural lip synchronization for automatic dubbing. BACKGROUND

[0003] Advancements in the field of neural networks have led to the development of various techniques for voice conversion and speech processing. Audio-driven Talking Head Generation (ATHG) is an emerging field of research focused on creating realistic and temporally synchronized talking human heads that move in accordance with input speech. ATHG focuses on speech-to-lip generation, i.e., it aims to reconstruct facial movements, particularly the movements of the lips, to match coherent input speech. ATHG holds substantial importance in numerous applications, including film dubbing, video editing, face animation, and digital assistants. New algorithms and approaches are being developed to create talking faces that closely mimic natural speech patterns and expressions, thus advancing ATHG significantly.

[0004] Person-specific ATHG techniques may generate photorealistic talking face videos that closely resemble a target speaker. However, such ATHG techniques may require access to a substantial amount of training data consisting of videos of the target speaker. On the other hand, person-generic ATHG techniques may aim to generate talking face videos without the need for speaker-specific training data. These techniques tackleDOCKET NO. SYP354694WO02 the more challenging task of generating realistic talking faces that may be applicable to any speaker. Apart from preserving the identity of the target speaker, person-generic ATHG techniques must also address key issues, including ensuring temporal synchronization between input speech and synthesized video streams of the target speaker, and preserving the visual quality of the synthesized video streams while maintaining proper lip-speech synchronization. However, despite significant progress, these techniques have tended to overlook a crucial aspect: the content associated with lip movements and its impact on the visual interpretability of the spoken words. This may involve focusing on the dynamic and subtle relationship between audio cues and facial movements, especially the movements of the lips and jaw regions.

[0005] Further limitations and disadvantages of conventional and traditional approaches will become apparent to one of skill in the art, through comparison of described systems with some aspects of the present disclosure, as set forth in the remainder of the present application and with reference to the drawings. SUMMARY

[0006] A system and method for multimodal neural lip synchronization for automatic dubbing is provided substantially as shown in, and / or described in connection with, at least one of the figures, as set forth more completely in the claims.

[0007] These and other features and advantages of the present disclosure may be appreciated from a review of the following detailed description of the present disclosure, along with the accompanying figures in which like reference numerals refer to like parts throughout. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] FIG.1 is a block diagram that illustrates an exemplary network environment for multimodal neural lip synchronization for automatic dubbing, in accordance with an embodiment of the disclosure.DOCKET NO. SYP354694WO02

[0009] FIG. 2 is a block diagram that illustrates an exemplary system of FIG. 1, in accordance with an embodiment of the disclosure.

[0010] FIG. 3A is a diagram that illustrates an exemplary processing pipeline for temporal shuffling augmentation, in accordance with an embodiment of the disclosure.

[0011] FIG. 3B is a diagram that illustrates an exemplary processing pipeline for multimodal neural lip synchronization for automatic dubbing, in accordance with an embodiment of the disclosure.

[0012] FIG.4 is a diagram that illustrates generation of an exemplary synthetic talking head through trained generative neural network, in accordance with an embodiment of the disclosure.

[0013] FIG. 5 is a flowchart that illustrates operations of an exemplary method for multimodal neural lip synchronization for automatic dubbing, in accordance with an embodiment of the disclosure. DETAILED DESCRIPTION

[0014] The following described implementation may be found in a system and a method for multimodal neural lip synchronization for automatic dubbing. Exemplary aspects of the disclosure may provide a system, which may include circuitry that may acquire an audio- visual input associated with a human speaker, wherein the audio-visual input may include a first set of video frames and a set of audio spectrograms associated with the first set of video frames. The circuitry may receive a dataset, which may a video clip and an audio clip, and may further partition, based on a time window, the dataset into sets of audio- visual pairs associated with the human speaker. Each set of audio-visual pairs of the sets of audio-visual pairs may include frames in a form of a plurality of video frames and a plurality of audio spectrograms associated with the plurality of video frames, where the plurality of video frames may include at least one video frame of the first set of video frames, and the plurality of audio spectrograms may include at least one audioDOCKET NO. SYP354694WO02 spectrogram of the set of audio spectrograms. Next, the circuitry may generate a second set of video frames based on application of a generative neural network on the acquired audio-visual input, where the second set of video frames may include a talking head of the human speaker at different time-instants which corresponds to that of the first set of video frames. Next, the circuitry may compute a first temporal contrastive loss based on the second set of video frames. Next, the circuitry may extract, from the generative neural network, a set of visual embeddings for the first set of video frames. The circuitry may also extract a set of audio embeddings based on the set of audio spectrograms. Next, the circuitry may compute a second temporal contrastive loss based on the extracted set of visual embeddings and the extracted set of audio embeddings. Next, the circuitry may compute an alignment loss between the set of audio spectrograms and a lower-half face of the human speaker in the first set of video frames. Thereafter, the circuitry may train the generative neural network based on the first temporal contrastive loss, the second temporal contrastive loss, and the alignment loss.

[0015] The circuitry may compute a loss based on a weighted sum of the computed first temporal contrastive loss, the computed second temporal contrastive loss, and the computed alignment loss. The generative neural network may be trained for a number of iterations until a computed loss is below a threshold loss.

[0016] Once the neural language model is trained, the system may utilize it for generating talking head of a human speaker that is aligned with the spoken words of corresponding audio data. The circuitry may receive, via a user interface of a user device, an input video clip associated with the human speaker. The input video clip may include initial video data and audio data associated with the initial video data, and the initial video data may include a talking head of a second human speaker that is misaligned with spoken words of the audio data. Further, the circuitry may prepare an input for the trained generative neural network based on the input video clip. Next, the circuitry may generateDOCKET NO. SYP354694WO02 output video data based on application of the trained generative neural network on the prepared input, where the output video data may include a synthetic talking head of the second human speaker that is aligned with the spoken words of the audio data.

[0017] In general, the ATHG may create realistic and temporally synchronized talking human heads that move in accordance with input speech. The ATHG may be classified as person-specific ATHG and person-generic ATHG. The person-specific ATHG techniques may focus on a target speaker and may generate photo-realistic talking face videos that closely resemble the target speaker. However, to generate such videos, the person specific ATHG techniques may require access to a substantial amount of training data consisting of multiple videos of the target speaker. In real-world scenarios, obtaining such training data for every individual may be challenging or impractical. The person specific ATHG techniques may often involve re-training or fine-tuning the model with the videos of the target speaker to achieve accurate results.

[0018] On the other hand, the person-generic ATHG techniques aim to generate talking face videos without the need for speaker-specific training data. The person generic ATHG techniques tackle the more challenging task of generating realistic talking faces that can be applied to any speaker. However, apart from preserving identity of the target speaker, the person-generic ATHG techniques are also required to handle key issues including - ensuring temporal synchronization between input speech and synthesized video streams of the target speaker and preserving visual quality of the synthesized video streams while also maintaining proper lip speech synchronization. Various studies have emphasized the significance of achieving accurate lip-speech synchronization as well as maintaining high visual quality in corresponding output. However, despite significant progress, these studies have tended to overlook a crucial aspect, i.e., content associated with lip movements and its impact on visual interpretability of the spoken words by focusing on the dynamic and subtle relationship between audio cues and facial movements, especially movements ofDOCKET NO. SYP354694WO02 the lips and jaw regions.

[0019] The proposed system may solve the issue of audio-driven talking head generation, also known as speech-to-lip generation, that aims to reconstruct facial movements, particularly the movements of the lips, to match coherent speech input. The system may focus on the dynamic and subtle relationship between audio cues and facial movements, especially movements of the lips and jaw regions. The system may perform contrastive learning in a unimodal manner on generated images and in a cross-modal manner on internal audio-visual feature embeddings generated from the generative neural network that helps in training the generative neural network in meaningful temporal representations without any need for explicit labels.

[0020] The proposed system may introduce temporal shuffling augmentation for the ATHG task, enhancing the robustness of the model by introducing variations in the temporal structure of the training data. To capture meaningful temporal relationships and patterns in audio-visual data, the system may introduce two temporal contrastive learning- based approaches: generation-based temporal contrastive learning and representation- based temporal contrastive learning. Additionally, the system may incorporate a part-level audio-visual alignment module to synchronize the audio signals with the corresponding lower half of the generated faces at a fine-grained level.

[0021] The system may include four key components: a) Audio-Visual Temporal Shuffling Augmentation, b) Generation-based Temporal Contrastive Learning, c) Representation-based Temporal Contrastive Learning to improve lip-speech synchronization by attracting audio embeddings and respective time-aligned visual context features while repelling the audio embeddings from different frames, and finally d) Part- Level Audio-Visual Alignment to improve synchronization at a more fine-grained level. Training with shuffled sequences may reduce sensitivity to input order, enhancing adaptability to diverse scenarios. Such temporal shuffling augmentation may aim toDOCKET NO. SYP354694WO02 improve the system’s handling of variations in speaking rates, pacing, and overall temporal dynamics found in real-world speech. The augmentation may also prevent overfitting to specific temporal patterns in training data, enhancing generalization on unseen data. Regardless of temporal shuffling, the generated frames may be reverted to the original sequence to obtain the final output.

[0022] FIG.1 is a block diagram that illustrates an exemplary network environment for multimodal neural lip synchronization for automatic dubbing, in accordance with an embodiment of the disclosure. With reference to FIG. 1, there is shown a network environment 100. The network environment 100 may include a system 102, a generative neural network 104, a database 108, a communication network 110, a server 106, and a user device 112.

[0023] The system 102 may include suitable logic, circuitry, interfaces, and / or code that may be configured to take a sequence of target frames, randomly chosen reference frames for a target speaker, and an aligned audio clip corresponding to the target visual frames as input. The audio clip may be transformed into a Mel-spectrogram. The system 102 may generate high-quality talking head faces where target face speaks with high synchronization, irrespective of the facial alignment in the reference frames.

[0024] The system 102 may implement a training pipeline for the generative neural network 104. As part of the pipeline, the system 102 may acquire an audio-visual input 114 associated with a human speaker. The audio-visual input 114 may include a first set of video frames 114-1 and a set of audio spectrograms 114-2 associated with the first set of video frames 114-2. The system 102 may further train the generative neural network 104 based on the acquired audio-visual input 114 for automatic dubbing with lip synchronization. Once trained, the system 102 may deploy the generative neural network 104 for inference. Examples of the system 102 may include, but are not limited to, a digital media player (DMP), a micro-console, a TV tuner, a digital media streamer, a mediaDOCKET NO. SYP354694WO02 extender / regulator, a digital media hub, a computer workstation, a mainframe computer, a handheld computer, a smart appliance, a plug-in device, and / or any other computing device with content streaming functionality.

[0025] The system 102 may store the generative neural network 104 or may be remotely connected to another system (such as the server 106) that hosts the generative neural network 104. When hosted on another system, the system 102 may send instructions to control training or inference of the generative neural network 104 via remote calls (e.g., API calls).

[0026] The generative neural network 104 be a deep learning-based person-generic lip- sync neural network for Audio-Driven Talking Head Generation (ATHG). The generative neural network 104 may be configured to generate realistic lip movements synchronized with arbitrary audio input, regardless of the specific identity of the person in the video. In terms of architecture, the person-generic lip-synch neural network for Audio-Driven Talking Head Generation (ATHG) may involve several components. For example, the person- generic lip-synch neural network may start with an audio encoder that may extract features from the input audio (or respective spectrograms), and a visual encoder that may process video frames to capture facial features. These features from the audio encoder and the visual encoder may be then fused in a feature fusion module to obtain fused features. The core of the person-generic lip-synch neural network may be referred to as a lip-sync generator which may use the fused features to produce synchronized lip movements, often employing generative models like GANs or diffusion models. To ensure smooth transitions between the frames, a temporal consistency module may be used. In GAN-based models, a discriminator may help to improve the realism of the generated frames. The person- generic lip-synch neural network may be trained using various loss functions to measure lip synchronization accuracy, visual quality, and temporal consistency, ensuring the generated talking head videos are both realistic and synchronized with the audio input.DOCKET NO. SYP354694WO02

[0027] In an embodiment, the generative neural network 104 may be referred to as an artificial deep neural network, which is a computational network or a system of artificial neurons, arranged in a plurality of layers, as nodes. As an artificial deep neural network, the plurality of layers of the generative neural network 104 may include an input layer, one or more hidden layers, and an output layer. Each layer of the plurality of layers may include one or more nodes (or artificial neurons, for example). Outputs of all nodes in the input layer may be coupled to at least one node of hidden layer(s). Similarly, inputs of each hidden layer may be coupled to outputs of at least one node in other layers of the generative neural network 104. Outputs of each hidden layer may be coupled to inputs of at least one node in other layers of the generative neural network 104. Node(s) in the final layer may receive inputs from at least one hidden layer to output a result. The number of layers and the number of nodes in each layer may be determined from hyper-parameters of the generative neural network 104. Such hyper-parameters may be set before or after training the generative neural network 104 on a training dataset.

[0028] Each node of the generative neural network 104 may correspond to a mathematical function (e.g., a sigmoid function or a rectified linear unit) with a set of parameters, tunable during training of the generative neural network 104. The set of parameters may include, for example, a weight parameter, a regularization parameter, and the like. Each node may use the mathematical function to compute an output based on one or more inputs from nodes in other layer(s) (e.g., previous layer(s)) of the generative neural network 104. All or some of the nodes of the generative neural network 104 may correspond to the same or a different mathematical function.

[0029] In an embodiment, the generative neural network 104 may include electronic data, which may be implemented as, for example, a software component of an application executable on the system 102. The generative neural network 104 may rely on libraries, external scripts, or other logic / instructions for execution by a processing device. TheDOCKET NO. SYP354694WO02 generative neural network 104 may include code and parameter file(s), that when executed on a computing system, such as the system 102, may enable the computing system to perform one or more operations for automatic dubbing with lip synchronization. Additionally, or alternatively, the generative neural network 104 may be implemented using hardware including, but not limited to, a processor, a microprocessor (e.g., to perform or control performance of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, the generative neural network 104 may be implemented using a combination of hardware and software.

[0030] The server 106 may be implemented as a cloud server that may be configured to receive the audio-visual input 114 or a dataset 116, which may include a video clip and an audio clip, where the audio clip is aligned with the video clip. The server 106 may also host the generative neural network 104 and / or an API to access the generation neural network 104. Additionally, or alternatively, the server 106 may be implemented using on- premises hosting (local servers), colocation hosting (third-party data centers), bare metal servers (dedicated servers), edge computing (local data processing), fog computing (decentralized data processing), mesh computing (distributed computing), hybrid cloud (combination of on-premises and cloud), or multi-cloud (multiple cloud providers).

[0031] The server 106 may execute operations through web applications, cloud applications, HTTP requests, repository operations, file transfer, and the like. Other example implementations of the server 106 may include, but are not limited to, a database server, a file server, a web server, a media server, an application server, a mainframe server, a machine learning server, or a cloud computing server.

[0032] In at least one embodiment, the server 106 may be implemented as a plurality of distributed cloud-based resources by use of several technologies that are well known to those ordinarily skilled in the art. A person with ordinary skill in the art will understand thatDOCKET NO. SYP354694WO02 the scope of the disclosure may not be limited to the implementation of the server 106 and the system 102, as two separate entities. In certain embodiments, the functionalities of the server 106 can be incorporated in its entirety or at least partially in the system 102 without a departure from the scope of the disclosure. In certain embodiments, the server 106 may host the database 108. Alternatively, the server 106 may be separate from the database 108 and may be communicatively coupled to the database 108.

[0033] The database 108 may include suitable logic, interfaces, and / or code that may be configured to store the dataset 116 including references to an audio-visual input 114 of the human speaker. The audio-visual input 114 may include the first set of video frames 114-1 and the set of audio spectrograms 114-2. The database 108 may store references to multiple datasets similar to the dataset 116. The database 108 may be derived from data off a relational or non-relational database, or a set of comma-separated values (csv) files in conventional or big-data storage. The database 108 may be stored or cached on a device, such as a server (e.g., the server 106) or the system 102. The device storing the database 108 may be configured to receive audio related commands or instructions from the system 102 or the server 106. In response, the device of the database 108 may be configured to retrieve and provide response of the query to the system 102 or the server 106, based on the received query.

[0034] In some embodiments, the database 108 may be hosted on a plurality of servers stored at the same or different locations. The operations of the database 108 may be executed using hardware including a processor, a microprocessor (e.g., to perform or control performance of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). In some other instances, the database 108 may be implemented using software.

[0035] The communication network 110 may include a communication medium through which the system 102 and the server 106 may communicate with one another. TheDOCKET NO. SYP354694WO02 communication network 110 may be one of a wired connection or a wireless connection. Examples of the communication network 110 may include, but are not limited to, the Internet, a cloud network, Cellular or Wireless Mobile Network (such as Long-Term Evolution and 5thGeneration (5G) New Radio (NR)), satellite communication system (using, for example, low earth orbit satellites), a Wireless Fidelity (Wi-Fi) network, a Personal Area Network (PAN), a Local Area Network (LAN), or a Metropolitan Area Network (MAN). Various devices in the network environment 100 may be configured to connect to the communication network 110 in accordance with various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of a Transmission Control Protocol and Internet Protocol (TIP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Zig Bee, EDGE, IEEE 802.11, light fidelity (Li-Fi), 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device to device communication, cellular communication protocols, and Bluetooth (BT) communication protocols.

[0036] The user device 112 may include a user-interface through which a user may interact with the system 102, send queries, feed commands and instructions, provide inputs similar to samples in the dataset 116 to train, finetune, or obtain inferences from the generative neural network 104. The user device 112 may also provide an input video clip associated with the human speaker for inference. The user may be a human speaker, a character associated with the inputs, or some other authorized person associated with the network environment 100. The user device 112 may be fixed at a place or may be portable. Examples of the user device 112 may include, but not limited to, a smartphone, a touchpad, a GUI interface, a personal computer, a wearable device such as a smartwatch or an eXtended-Reality (XR) device, a microphone, or a display device.

[0037] In operation, the system 102 may be configured to acquire audio-visual inputDOCKET NO. SYP354694WO02 114. The system 102 may acquire the audio-visual input 114 from the dataset 116 including a video clip and an audio clip of a human speaker. The audio clip may be aligned with the video clip. In an exemplary embodiment, the system 102 may partition the dataset 116 into sets of audio-visual pairs 322 (shown in FIG. 3A) associated with the human speaker, based on a time window, such that each set of audio-visual pairs of the sets of audio-visual pairs 322 includes frames in a form of a plurality of video frames 324 (shown in FIG.3A) and a plurality of audio spectrograms 326 (shown in FIG.3B) associated with the plurality of video frames 324. Further, the plurality of video frames 324 may include at least one video frame of the first set of video frames 114-1, and the plurality of audio spectrograms 326 may include at least one audio spectrogram of the set of audio spectrograms 114-2. The plurality of video frames 324 may capture facial expressions, jaw movement, lip movement, and lip shapes of the human speaker. The set of audio spectrograms 114-2 may capture voice, speech / dialogues, tone, pitch, of the human speaker, and background noise.

[0038] The system 102 may apply a shuffling operation on the frames in each set of audio-visual pairs of the sets of audio-visual pairs 322 to obtain sets of shuffled audio- visual pairs. For instance, the shuffling operation may include a random permutation function and may involve rearranging the frames in each set of audio-visual pairs of the sets of audio-visual pairs in a random or pseudo-random order. The system 102 may further apply a sampling operation on each set of shuffled audio-visual pairs of the sets of shuffled audio-visual pairs. For instance, the sampling operation may involve selecting a subset of audio-visual pairs from each set of shuffled audio-visual pairs. The system 102 may further acquire, based on the application of the sampling operation, the audio-visual input 114, which may include the first set of video frames 114-1 and the set of audio spectrograms 114-2 associated with the first set of video frames 114-1. Details related to the acquisition of the audio-visual input 114 are further provided, for example, in FIG.3A.DOCKET NO. SYP354694WO02

[0039] After the acquisition, the system 102 may be configured to generate a second set of video frames 118. For instance, the system 102 may apply the generative neural network 104 on the acquired audio-visual input 114 to generate the second set of video frames 118. The second set of video frames 118 may include a talking head of the human speaker at different time-instants, which may correspond to that of the first set of video frames 114-1. The system 102 may generate the second set of video frames 118, such that the talking head of the human speaker is temporally aligned with the set of audio spectrograms 114-2. For instance, the shape of lips or movement of lips of the talking head may be synchronized with time stamps of speech / dialogues / phonemes associated with the set of audio spectrograms 114-2. Details related to the generation of second set of video frames are further provided, for example, in FIG.3B.

[0040] After the generation of the second set of video frames 118, the system 102 may compute a first temporal contrastive loss 338 (shown in FIG.3B). The system 102 may compute the first temporal contrastive loss 338 based on the second set of video frames 118. For instance, the system 102 may process each video frame of the second set of video frames 118 and compute similarities between consecutive video frames of the second set of video frames 118. The system 102 may further compute feature embeddings based the similarities. Thereafter, the system 102 may compute the cross-entropy losses based on the feature embeddings and corresponding ground truth embeddings. The system 102 may further compute the first temporal contrastive loss 338 based on the computed cross-entropy losses. Details related to the computation of the first temporal contrastive loss are further provided, for example, in FIG.3B.

[0041] After the acquisition, the system 102 may extract from the generative neural network 104, a set of visual embeddings for the first set of video frames 114-1. The set of visual embeddings may be compact, numerical representations of visual data associated with the first set of video frames 114-1. The set of visual embeddings may capture theDOCKET NO. SYP354694WO02 essential features and patterns of the visual data in a lower-dimensional vector space. By transforming complex visual data into manageable vectors, the set of visual embeddings may enable efficient processing and retrieval of the visual data. Details related to the extraction of the set of visual embeddings are further provided, for example, in FIG. FIG. 3B.

[0042] The system 102 may further extract a set of audio embeddings based on the set of audio spectrograms 114-2. The set of audio embeddings may represent compact and numerical representations of audio data to fixed-size vectors associated with the set of audio spectrograms 114-2. The set of audio embeddings may capture essential features and characteristics of the sound and may be generated using neural networks. Details related to the extraction of the set of audio embeddings are further provided, for example, in FIG.3A and FIG.3B.

[0043] After the extraction, the system 102 may compute a second temporal contrastive loss 346 (as shown in FIG. 3B). The system 102 may compute the second temporal contrastive loss 346 based on the extracted set of visual embeddings and the extracted set of audio embeddings. In an exemplary embodiment, the system 102 may form audio- visual embedding pairs by pairing a visual embedding of the set of visual embeddings and an audio embedding of the set of audio embeddings. Thereafter, the system 102 may compute similarity values based on vector similarity between each embedding pair of the audio-visual embedding pairs. Upon computation of the similarity values, the system 102 may further compute the second temporal contrastive loss 346 based on the computed similarity values. The second temporal contrastive loss 346 may be referred to as representation loss. Details related to the computation of the second temporal contrastive loss are further provided, for example, in FIG.3B.

[0044] The system 102 may also compute an alignment loss 354 (shown in FIG.3B) between the set of audio spectrograms 114-2 and a lower-half face of the human speakerDOCKET NO. SYP354694WO02 in the first set of video frames 114-1. The alignment loss 354 may be computed based on misalignment between the set of audio spectrograms 114-2 and the lower-half face of the human speaker in the first set of video frames 114-1 while both the set of audio spectrograms 114-2 and the first set of video frames 114-1 are played simultaneously. Details related to the computation of the alignment loss are further provided, for example, in FIG.3B.

[0045] The system 102 may further train the generative neural network 104. The system 102 may train the generative neural network 104 based on the first temporal contrastive loss, the second temporal contrastive loss, and the alignment loss. For instance, the system 102 may compute a loss based on a weighted sum of the computed first temporal contrastive loss 338, the computed second temporal contrastive loss 346, and the computed alignment loss 354. Further, the generative neural network 104 may be trained for a number of iterations until the computed loss is below a threshold loss. Details related to the training of the generative neural network are further provided, for example, in FIG. 3B.

[0046] The system 102 may be integrated with any existing baseline lip-sync networks (not shown) during training to improve lip-sync capability of said baseline lip-sync networks. The system 102 may train the generative neural network 104 based on contrastive learning using generation-based temporal contrastive learning and representation-based temporal contrastive learning. The system 102 may also train the generative neural network 104 based on part-level alignment, in order to capture fine- grained correlation between audio embeddings and visual embeddings of lower-half face of the human speaker. The system 102 may generate more realistic, lip- synced, and high- fidelity talking face videos compared to the existing baseline lip-sync networks.

[0047] FIG. 2 is a block diagram that illustrates an exemplary system of FIG. 1, in accordance with an embodiment of the disclosure. FIG.2 is explained in conjunction withDOCKET NO. SYP354694WO02 elements from FIG.1. With reference to FIG.2, there is shown a block diagram 200 of the exemplary system 102. The system 102 may include circuitry 202, a memory 204, a network interface 206, and an input / output (I / O) device 208. The I / O device 208 may include a display device 208-1. The memory 204 may include a dataset (say, the dataset 116 of FIG.1) and a generative neural network (say, the generative neural network 104 of FIG.1). The network interface 206 may connect the system 102 with the server 106, via the communication network 110.

[0048] The circuitry 202 may include suitable logic, circuitry, and / or interfaces that may be configured to execute program instructions associated with different operations to be executed by the system 102. The operations may include reception of the dataset 116, application of the generative neural network 104, and training of the generative neural network 104. The circuitry 202 may include one or more processing units, which may be implemented as a separate processor. In an embodiment, the one or more processing units may be implemented as an integrated processor or a cluster of processors that perform the functions of the one or more specialized processing units, collectively. The circuitry 202 may be implemented based on a number of processor technologies known in the art. Examples of implementations of the circuitry 202 may be an X86-based processor, a Graphics Processing Unit (GPU), a Reduced Instruction Set Computing (RISC) processor, an Application-Specific Integrated Circuit (ASIC) processor, a Complex Instruction Set Computing (CISC) processor, a microcontroller, a central processing unit (CPU), and / or other control circuits.

[0049] The memory 204 may include suitable logic, circuitry, interfaces, and / or code that may be configured to store one or more instructions to be executed by the circuitry 202. The one or more instructions stored in the memory 204 may be configured to execute the different operations of the circuitry 202 (and / or the system 102). The memory 204 may be further configured to store the dataset 116 and the generative neural network 104.DOCKET NO. SYP354694WO02 Examples of implementation of the memory 204 may include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Hard Disk Drive (HDD), a Solid-State Drive (SSD), a CPU cache, and / or a Secure Digital (SD) card.

[0050] The network interface 206 may include suitable logic, circuitry, interfaces, and / or code that may be configured to facilitate communication between the system 102 and the server 106, via the communication network 110. The network interface 206 may be implemented by use of various known technologies to support wired or wireless communication of the system 102 with the communication network 110. The network interface 206 may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, or local buffer circuitry.

[0051] The network interface 206 may be configured to communicate via wireless communication with networks, such as the Internet, an Intranet, a wireless network, a cellular telephone network, a wireless local area network (LAN), or a metropolitan area network (MAN). The wireless communication may be configured to use one or more of a plurality of communication standards, protocols and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), wideband code division multiple access (W-CDMA), Long Term Evolution (LTE), 5thGeneration (5G) New Radio (NR), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (such as IEEE 802.11a, IEEE 802.11b, IEEE 802.11g or IEEE 802.11n), voice over Internet Protocol (VoIP), light fidelity (Li-Fi), Worldwide Interoperability for Microwave Access (Wi-MAX), a protocol for email, instant messaging, and a Short Message Service (SMS).

[0052] The I / O device 208 may include suitable logic, circuitry, interfaces, and / or codeDOCKET NO. SYP354694WO02 that may be configured to receive an input and provide an output based on the received input. For example, the I / O device 208 may receive the dataset 116 or acquire the audio- visual input 114. The I / O device 208 may be further configured to render output video on user interface of a user device (say, the user device 112 of FIG.1). The I / O device 208 may include the display device 208-1. Examples of the I / O device 208 may include, but are not limited to, a display (e.g., a touch screen), a keyboard, a mouse, a joystick, a microphone, or a speaker. Examples of the I / O device 208 may further include braille I / O devices, such as, braille keyboards and braille readers.

[0053] The display device 208-1 may include suitable logic, circuitry, and interfaces that may be configured to display or render the health condition associated with the user. The display device 208-1 may be a touch screen which may enable a user to provide a user- input via the display device 208-1. The touch screen may be at least one of a resistive touch screen, a capacitive touch screen, or a thermal touch screen. The display device 208-1 may be realized through several known technologies such as, but not limited to, at least one of a Liquid Crystal Display (LCD) display, a Light Emitting Diode (LED) display, a plasma display, or an Organic LED (OLED) display technology, or other display devices. In accordance with an embodiment, the display device 208-1 may refer to a display screen of a head mounted device (HMD), a smart-glass device, a see-through display, a projection-based display, an electro-chromic display, or a transparent display. Various operations of the circuitry 202 for training the generative neural network 104, are described further, for example, in FIG.3A and FIG.3B.

[0054] FIG. 3A is a diagram that illustrates an exemplary processing pipeline for temporal shuffling augmentation, in accordance with an embodiment of the disclosure. FIG.3A is explained in conjunction with elements from FIG.1 and FIG.2. With reference to FIG.3A, there is shown an exemplary processing pipeline 300 that illustrates exemplary operations 302 and 304 for temporal shuffling augmentation. The exemplary operationsDOCKET NO. SYP354694WO02 302 and 304 may be executed by any computing system, for example, by the system 102 of FIG.1 or by the circuitry 202 of FIG.2.

[0055] At 302, an operation for dataset reception may be executed. The circuitry 202 may receive a dataset 116 including a video clip and an audio clip, where the audio clip is aligned with the video clip. The dataset 116 may be received from a user or may be received from an external data source. The circuitry 202 may similarly receive many such datasets.

[0056] The circuitry 202 may further partition the dataset 116 into sets of audio-visual pairs 322 associated with the human speaker, based on a time window. Each set of audio- visual pairs of the sets of audio-visual pairs 322 may include frames in a form of a plurality of video frames 324 and a plurality of audio spectrograms 326 associated with the plurality of video frames 324. The plurality of video frames 324 may include at least one video frame of the first set of video frames 114-1. The plurality of audio spectrograms 326 include at least one audio spectrogram of the set of audio spectrograms 114-2. For instance, consider each set of audio-visual pairs of the sets of audio-visual pairs 322 in equation 1, as follows: ∧= {I, A} (1)where, I = {Ii}, andA = {Ai}, for i ^ (0, 1…n) represent the corresponding plurality of video frames 324 andthe plurality of audio spectrograms 326. Consider the dataset 116 be of size nb. The dataset 116 may be partitioned into three sets of the audio-visual pairs 322 (∧nb) for three different timeslots, say past (∧t−), present (∧t),and future (∧t+) timeslots, where, ∧t− = {It−, At−} │ It− = {It−i}, IAt− = {At−i} for i ^ (0…nb / 3), ∧t = {It, At} │ It = {Iti}, IAt = {Ati} for i ^, and ∧t+ = {It+, At+}│ It+ = {It+i}, IAt+ = {At+i} for i ^ (2nb / 3+1,… nb)322 (∧b) may include first set of audio-DOCKET NO. SYP354694WO02 visual pairs 322-1, second set of audio-visual pairs 322-2, and third set of audio-visual pairs 322-3. Further, consider that the plurality of video frames 324 associated with each of the sets of the audio-visual pairs 322 may include three video frames. The first set of audio-visual pairs 322-1 may pertain to the past timeslot (∧t−) and may include video frames 324-1, 324-2 and 324-3 and audio spectrograms 326-1, 326-2, and 326-3 associated with the video frames 324-1, 324-2, and 324-3. The second set of audio-visual pairs 322-2 may pertain to the present timeslot (∧t), and may include video frames 324-4, 324-5, and 324-6, and audio spectrograms 326-4, 326-5, and 326-6 associated with the video frames 324-4, 324-5, and 324-6. The third set of audio-visual pairs 322-3 may pertain to the future timeslot (∧t+), and may include video frames 324-7, 324-8, and 324-9, and audio spectrograms 326-7, 326-8, and 326-9 associated with the video frames 324-7, 324- 8, and 324-9.

[0058] At 304, an operation for frames shuffling may be executed. The circuitry 202 may apply a shuffling operation on the frames in each set of audio-visual pairs of the sets of audio-visual pairs to obtain sets of shuffled audio-visual pairs. The shuffling operation may include a random permutation function π(t). The shuffled sets of the audio-visual pairs 322(∧ ′nb) may be defined using equation 2, as follows:∧ ′nb = (∧ ′t−, ∧ ′t, ∧ ′t+) = π((∧t−), (∧t), (∧t+)) (2)where π represents shuffling operation,∧ ′t− represents shuffles ∧t−,∧ ′t represents shuffles ∧t , and∧ ′t+ represents shuffles ∧t+ .

[0059] In an embodiment, the video frames 324-1, 324-2, and 324-3, and the corresponding audio spectrograms 326-1, 326-2, and 326-3 may be shuffled within thefirst set of audio-visual pairs 322-1, and correspondingly ∧ ′t− may be obtained. Similarly,the video frames 324-4, 324-5, and 324-6, and the corresponding audio spectrogramsDOCKET NO. SYP354694WO02 326-4, 326-5, and 326-6 may be shuffled within the second set of audio-visual pairs 322-2, and correspondingly ∧ ′t may be obtained. Similarly, the video frames 324-7, 324-8, and324-9, and the corresponding audio spectrograms 326-7, 326-8, and 326-9 may beshuffled within the third set of audio-visual pairs 322-3, and correspondingly ∧ ′t+ may beobtained.

[0060] It should be further noted that after shuffling, the video frames 324-1, 324-2, and324-3 and the audio spectrograms 326-1, 326-2, and 326-3 of the ∧ ′t− may or may not bealigned with one another. It should be further noted that after shuffling, the video frames324-4, 324-5, and 324-6 and the audio spectrograms 326-4, 326-5, and 326-6 of the ∧ ′tmay or may not be aligned with one another. It should be further noted that after shuffling, the video frames 324-7, 324-8, and 324-9 and the audio spectrograms 326-7, 326-8, and326-9 of the ∧ ′t+ may or may not be aligned with one another.

[0061] These sets of shuffled audio-visual pairs may result in data augmentation, which may enhance ability of the system 102 to generate coherent facial expressions and lip movements regardless of sequence of the corresponding frames. The data augmentation may improve handling of variations in speaking rates, pacing, and overall temporal dynamics found in real-world speech. Additionally, the data augmentation may prevent, during training, overfitting of the generative neural network 104 to specific temporal patterns in training data, enhancing generalization on unseen data.

[0062] At 306, an operation for sampling may be executed. The circuitry 202 may apply a sampling operation to sample and select at least one video frame and corresponding audio spectrogram from each of the shuffled sets of audio-visual pairs 322. For instance, the circuitry 202 may sample and select video frame 324-2 and audio spectrogram 326-2 from the shuffled first set of audio-visual pairs 322-1. The circuitry 202 may sample and select video frame 324-5 and audio spectrogram 326-5 from the shuffled second set of audio-visual pairs 322-2. Similarly, the circuitry 202 may sample and select video frameDOCKET NO. SYP354694WO02 324-8 and audio spectrogram 326-8 from the shuffled sets of audio-visual pairs 322-3. Herein, the selected video frames 324-2, 324-5, and 324-8 may correspond to the first set of video frames 114-1 and the selected audio spectrograms 326-2, 326-5, and 326-8 may correspond to the set of audio spectrograms 114-2, forming the audio-visual input 114.

[0063] In an embodiment, the sampling operation may include a random selection of a video frame of the first set of video frames 114-1 and a corresponding audio spectrogram of the set of audio spectrograms 114-2 from a set of shuffled audio-visual pairs of the set of shuffled audio-visual pairs.

[0064] FIG. 3B is a diagram that illustrates an exemplary processing pipeline for multimodal neural lip synchronization for automatic dubbing, in accordance with an embodiment of the disclosure. With reference to FIG.3B, there is shown an exemplary processing pipeline 350 that illustrates exemplary operations 306 to 314 for multimodal neural lip synchronization for automatic dubbing. The exemplary operations from 306 to 314 may be executed by any computing system, for example, by the system 102 of FIG. 1 or by the circuitry 202 of FIG.2.

[0065] To capture meaningful temporal relationships within the augmented data, the circuitry 202 may execute operations for contrastive learning, as described herein. In general, contrastive learning may be utilized across various fields for executing distinct tasks, for instance, for text-image retrieval, image classification, and acoustic event detection. Contrastive learning may aim to minimize the distance between samples of same class (referred to as aligned pairs or positive pairs, herein) while maximizing distance between samples from different classes (also referred to as misaligned pairs or negative pairs). By mapping features into a common representation space, positive pairs are attracted while negative pairs are repelled. Selecting appropriate positive and negative pairs is crucial for guiding the generative neural network 104 in learning contrastive representations and enabling self-supervised learning without labeled data. Specifically inDOCKET NO. SYP354694WO02 the domain of auto-driven talking face generation (ATHG), contrastive learning may play a pivotal role in aiding lip-speech synchronization. The circuitry 202 may execute two temporal contrastive learning approaches, namely – generation-based temporal contrastive learning and representation-based temporal contrastive learning. These approaches may seek to optimize the representation / generation of the talking head in accordance with lip-speech and to improve the generative neural network's robustness against temporal misalignment. The performance of the generative neural network 104 may be improved by maximizing agreement in aligned pairs and minimizing agreement in misaligned pairs of generated images and their corresponding feature representations.

[0066] At 308, an operation for audio-visual input acquisition may be executed. The circuitry 202 may acquire the audio-visual input 114 from each set of shuffled audio-visual pairs of the sets of shuffled audio-visual pairs (obtained at 304). The acquired audio-visual input 114 may be associated with a human speaker and may include first set of video frames 114-1 and set of audio spectrograms 114-2 associated with the first set of video frames 114-1.

[0067] The acquired audio-visual input 114 may be fed to the generative neural network 104. Regardless of temporal shuffling, the generative neural network 104 may revert the generated sets of video frames to original sequence to obtain the final output. The acquired audio-visual input 114 may be fed to a facial feature extractor 328 of the generative neural network 104, which may be configured to extract facial features from the plurality of video frames 324 associated with the acquired audio-visual input 114.

[0068] At 310, generation-based temporal contrastive learning may be executed. The circuitry 202 may generate a second set of video frames 118 based on application of the generative neural network 104 on the acquired audio-visual input 114. The second set of video frames 118 may include a talking head of the human speaker at different time- instants which corresponds to that of the first set of video frames 114-1. The talking headDOCKET NO. SYP354694WO02 of the human speaker in the second set of video frames 118 may be accurately aligned with the set of audio spectrograms 114-2. For instance, the second set of video frames 118 may include a first video frame 118-1, a second video frame 118-2, and a third video frame 118-3, where the first video frame 118-1, the second video frame 118-2, and the third video frame 118-3 may be generated in a temporal sequence of a speech signal represented in the set of audio spectrograms 114-2.

[0069] The circuitry 202 may further compute a first similarity 330-1 between the first video frame 118-1 and the second video frame 118-2, a second similarity 330-2 between the second video frame 118-2 and the third video frame 118-3. The first similarity 330-1 and the second similarity 330-2 may correspond to cosine similarity (CS). Specifically, the generative neural network 104 gains the knowledge that if ^�^^^^^^^is the current or presentgenerated video frame, ^^^^ℎ^^^^^^^^ ^^�^^^^^^− should come before ^�^^^^^^^ in the entire second set of videoframes 118, indicatingvideo frame, and ^^�^^^^^^+ should come after ^�^^^^^^^ indicating thefuture video frame. This knowledge may play a crucial role in generating lip-synced videoswith high fidelity. In an instance, cosine similarity between ^^�^^^^^^− and ^�^^^^^^^ may be calculated asEsim1 or sim(^^�^^^^^^− , ^�^^^^^^^) and the cosine similarity between ^�^^^^^^^ and ^^�^^^^^^+ may be calculated asEsim2 or sim(^�^^^^^^^ , ^^�^^^^^^+ ). The cosine similarity may be defined for two variable x and y usingequation 3, as follows – sim(x.y x,y)=(3) ‖x‖.‖y‖

[0070] The circuitry 202 may compute a first feature embedding (Esim1) 334-1 based on application of a first fully connected neural network layer 332-1 on the first similarity 330-1. The circuitry 202 may compute a second feature embedding (Esim2) 334-2 based on application of a second fully connected neural network layer 332-2 on the second similarity 330-2. In an instance, the fully connected neural network layers 332-1 and 332- 2 may correspond to multilayer perceptron (MLP) and may act as feedforward layers forDOCKET NO. SYP354694WO02 facilitating training of the generative neural network 104. In an embodiment, similarity matrices may be processed through identical fully connected layers F1() and F2() in order to obtain the feature embeddings, as given by following equations 4 and 5: Esim1 = Esim(I� Dt− , I�t) = F1(sim(I�t− , I�t)) ^ R *1(4) Esim2 =Esim(I�t , I�t+ ) = F2(sim(I�t , I�t+ )) ^ RD*1(5) where D may represent embedding dimension.

[0071] The circuitry 202 may compute a first cross-entropy loss (loss1) 336-1 between the first feature embedding 334-1 and a first ground truth embedding. Similarly, the circuitry 202 may compute a second cross-entropy loss (loss2) 336-2 between the second feature embedding 334-2 and a second ground truth embedding. Further, the circuitry 202 may apply a first contrastive loss calculating model to compute a first temporal contrastive loss (Lgen-TCL) 338 based on the first cross-entropy loss 336-1 and the second cross-entropy loss 336-2. The computation of the contrastive loss (Lgen-TCL) 338 may be represented using equation 6, as follows: CE�Esim�I�t− , I�t�, O(It− , It)�│inputs ^ {∧′t−,∧ ′t}(6) Lgen−TCL =�^ {∧′ ,∧ ′ }t t+1where, CE denotes the standard cross entropy loss, stated as follows – CE �Esim�I�t− , I�t�, O(It− , It)� = −O(It− , It). log (Esim�I�t− , I�t�), andCE�Esim�I�t , I�t+�, O(It , It+)� = −O(It , It+). log (Esim�I�t , I�t+�)where terms O(It− , It) and O(It , It+) represent constructed ground truths, which may bedefined as follows – −1D OI … −1−andDOCKET NO. SYP354694WO02 The ground truths may be constructed in a manner to ensure that distance betweenI�t− (120°) and I�t+ (60°) with respect to I�t(0°) is maximized in Euclidean space. Byconstructing O(I , I ) and O(I , I ) with constant values o−1+1t− t t t+ f2 and 2 respectively, a clear geometric relationship may be established. This relationship may facilitate the desired separation of the generated images in orientation space. The main objective of the generation-based temporal contrastive learning is to make the generative neural network 104 aware of the temporal consistency among the generated second set of video frames 118.

[0072] At 312, representation-based temporal contrastive learning may be executed. The circuitry 202 may extract, via the facial feature extractor 328, a set of visual embeddings for the first set of video frames 114-1. For instance, the set of visual embeddings may include visual embedding 340-1, visual embedding 340-2, and visual embedding 340-3. The circuitry 202 may also extract a set of audio embeddings based on the set of audio spectrograms 114-2 (as described at 306 of FIG.3A). The set of audio embeddings may include audio embedding 344-1, audio embedding 344-2, and audio embedding 344-3. The set of audio embeddings may be extracted through an audio feature extractor 342, from the shuffled audio spectrograms. Thereafter, the circuitry 202 may determine audio-visual embedding pairs between the extracted set of visual embeddings and the extracted set of audio embeddings, based on the audio-visual input 114.

[0073] In one embodiment, the circuitry 202 may select an anchor visual embedding from the extracted set of visual embeddings and may further separate the audio-visual embedding pairs into first positive pairs and first negative pairs based on the anchor visual embedding. For instance, the circuitry 202 may select visual embedding 340-1 as the anchor visual embedding, and may separate corresponding audio-visual embedding pair,DOCKET NO. SYP354694WO02 i.e., (340-1, 344-1) as the first positive pair. In another instance, the circuitry 202 may select visual embedding 340-3 as the anchor visual embedding, and may separate corresponding audio-visual embedding pair, i.e., (340-3, 344-3) as the first negative pair.

[0074] In another embodiment, the circuitry 202 may select an anchor audio embedding from the extracted set of audio embeddings and may further separate the audio-visual embedding pairs into second positive pairs and second negative pairs based on the anchor audio embedding. For instance, the circuitry 202 may select audio embedding 344-1 as the anchor audio embedding, and may separate corresponding audio-visual embedding pair, i.e., (340-1, 344-1) as the first positive pair. In another instance, the circuitry 202 may select audio embedding 344-3 as the anchor audio embedding, and may separate corresponding audio-visual embedding pair, i.e., (340-3, 344-3) as the first negative pair.

[0075] According to some embodiments, during execution, the circuitry 202 may be configured to separate the audio-visual embedding pairs into the first positive pairs and the first negative pairs only. According to other embodiments, during execution, the circuitry 202 may be configured to separate the audio-visual embedding pairs into the second positive pairs and the second negative pairs only. According to some other embodiments, during execution, the circuitry 202 may be configured to separate the audio- visual embedding pairs into the first positive pairs and the first negative pairs (based on selection of an anchor visual embedding for some of the audio-visual embedding pairs) as well as the second positive pairs and the second negative pairs (based on selection of an anchor audio embedding for rest of the audio-visual embedding pairs).

[0076] Further, the circuitry 202 may compute similarity values based on a vector similarity between each pair of the first positive pairs, the second positive pairs, the first negative pairs, and the second negative pairs. For instance, the circuitry 202 may compute cosine similarity and correspondingly integrate representation-based temporal contrastive learning to quantify relationship between the audio-visual embedding pairs. Furthermore,DOCKET NO. SYP354694WO02 the circuitry 202 may compute second temporal contrastive loss 346 based on the computed similarity values. The second temporal contrastive loss 346 may be referred to as representation loss (Lrep−TCL), which may be formulated using equation 7, as follows: L exp(sim(Ev,Ea) rep−TCL = − log) exp�sim(Ev,Ea)�+exp�sim(Ev,Ea′)�+exp (sim(Ev′,Ea))(7)to the anchor visual embedding, Ea pertains to the anchor audio embedding,(Ev, Ea) denotes the first positive pairs and the second positive pairs,(Ev,Ea′)denotes the first negative pairs for any shifted audio embedding Ea′,(Ev′, Ea) denotes the second negative pairs for any shifted visual embedding Ev′, andsim(Ev, Ea), sim(Ev,Ea′), sim(Ev′, Ea) are cosine similarities.

[0077] At 314, part-level audio-visual alignment may be executed. In the audio-visual input 114, different parts of face of the human speaker may exhibit distinct temporal dynamics that correspond to specific aspects of the corresponding set of audio spectrograms 114-2. The circuitry 202 may create a correlation matrix 348 based on the set of audio spectrograms 114-2 and visual embeddings for the lower-half face of the human speaker in the first set of video frames 114-1. The lower-half face of the human speaker may include lips and jaw line, for instance.

[0078] In operation, visual embeddings of the lower-half face (i.e. Evlower^ Rd*1) may be extracted from a lip-sync network M() via a fully connected layer, and the corresponding audio embeddings (i.e. Ea^ Rd*1) may be extracted from the audio feature extractor 342. Here, ‘d’ may represent embedding dimension. Further, the correlation matrix 348 (Mav-corr^ Rd*d) may be computed using equation 8, as follows – Mav-corr= Evlower *EaT(8) Each cell in the Mav-corr may include values representing disentangled audio and visual features from the corresponding rows and columns of the correlation matrix 348. ForDOCKET NO. SYP354694WO02 instance, the correlation matrix 348 may be depicted as follows – 0.68 0.13 0.21 … 0.11 0.11 0.03 0.56 … 0.74 0.36 0.80 0.01 … 0.22 … … … … … 0.22 0.08 0.78 … 0.02 where columns may represent lower-half facial features and rows may represent audio features.

[0079] The circuitry 202 may compute a cross-entropy loss between the correlation matrix 348 (Mav-corr) and a ground truth matrix (MDSP^ Rd*d). The computed cross-entropy loss may be considered as alignment loss (Lav-align) 354. For instance, the alignment loss (Lav-align) 354 may be computed by applying standard cross-entropy function CE() on the correlation matrix 348 (Mav-corr) and the ground truth matrix (MDSP), which may be represented by equation 9, as follows: Lav-align= CE(Mav-corr, MDSP) (9) where, CE(Mav-corr, MDSP) may be formulated in equation 10, as follows: CE(Mav-corr, MDSP) = - MDSP. log(Mav-corr) (10)

[0080] In a preferred embodiment, the ground truth matrix (MDSP) may be a doubly stochastic permutation matrix 352, and may be illustrated using equation 11, as follows: d d (11)

[0081] permutation matrix 352 is to ensure that if a part ‘m’ of the set of audio embeddings highly correlates to a particular part ‘n’ of the lower-half face, then corresponding cell (m, n) in the Mav-corrshould have a higher value. Conversely, remaining cells in the same row ‘m’ and column ‘n’ should have low values,DOCKET NO. SYP354694WO02 indicating that other parts of the set of audio embeddings are not correlated well with the lower-half face. For instance, the doubly stochastic permutation matrix 352 may be depicted as follows: 1 0 0 … 0 0 0 0 … 1 0 1 0 … 0 … … … … … 0 0 0 … 1

[0082] The generative neural network 104 may optimize value of the created correlation matrix 348 during training, hence ensuring synchronization between the temporal dynamics of the lower-half face of the human speaker and corresponding audio embeddings. Further, the alignment loss (Lav-align) 354 may ensure proper disentanglement between the set of audio spectrograms 114-2 and features (lip shape, lip movement, jaw movement, etc.) of the lower-half face of the human speaker in the first set of video frames 114-1.Therefore, through the part-level alignment, the circuitry 202 may capture fine- grained temporal relationships between the lower-half face of the human speaker and the corresponding audio embeddings.

[0083] The circuitry 202 may further compute a loss based on a weighted sum of the computed first temporal contrastive loss, the computed second temporal contrastive loss, and the computed alignment loss. Thereafter, the generative neural network 104 may be trained for a number of iterations until the computed loss is below a threshold loss.

[0084] Further, the generative neural network 104 may be trained for disentanglement between the first set of video frames 114-1 and the set of audio spectrograms 114-2 of the audio-visual input 114 in a temporal domain. For training the generative neural networkDOCKET NO. SYP354694WO02 104 to ensure proper disentanglement, both the generation-based temporal contrastive learning and the representation-based temporal contrastive learning may be executed in a self-supervised manner.

[0085] To further enhance lip synchronization at a finer level of detail, the part-level audio-visual alignment may be executed. Overall, the generation-based temporal contrastive learning, the representation-based temporal contrastive learning, and the part- level audio-visual alignment may be integrated altogether into the training process of the generative neural network 104, enhancing capacity of the generative neural network 104 for lip synchronization by focusing on both the content of lip movements and synchronization with speech input / audio data.

[0086] FIG.4 is a diagram that illustrates an exemplary diagram illustrating generation of a synthetic talking head through trained generative neural network, in accordance with an embodiment of the disclosure. With reference to FIG.4, there is shown an exemplary diagram 400 illustrating generation of a synthetic talking head through trained generative neural network. The operations depicted in the diagram 400 may be executed by any computing system, for example, by the system 102 of FIG.1 or by the circuitry 202 of FIG. 2.

[0087] Once the generative neural network 104 is trained and starts performing satisfactorily, the generative neural network 104 may be deployed in any Natural Language Processing (NLP) application for inference, as described herein. In operation, the system 102 may communicate with the user device 112 and may receive an input video clip 402 associated with a second human speaker. The input video clip 402 may include initial video data 402-1 and audio data 402-2 associated with the initial video data 402-1. The initial video data 402-1 may include a talking head of the second human speaker, which may be misaligned with spoken words of the audio data 402-2. For instance, the initial video data 402-1 may include the talking head of the second humanDOCKET NO. SYP354694WO02 speaker as reference, and the audio data 402-2 may correspond to a sentence: “It took some great efforts”. The system 102 may prepare an input for the trained generative neural network 104 based on the input video clip 402.

[0088] The trained generative neural network 104 may further produce output video data 404, which may include a synthetic talking head of the second human speaker, where the synthetic talking head of the second human speaker may be aligned with the spoken words of the audio data. For instance, lip shape / movement of the lips of the synthetic talking head of the second human speaker is aligned with the audio spectrogram of the term “great”, as can be observed in a frame of the produced output video data 404. Hence, the trained generative neural network 104 may generate the synthetic talking head of the second human speaker with high-fidelity lip movements, enhancing reading interpretability and reducing speech perception ambiguity.

[0089] FIG. 5 is a flowchart that illustrates operations of an exemplary method for multimodal neural lip synchronization for automatic dubbing, in accordance with an embodiment of the disclosure. FIG.5 is described in conjunction with elements from FIG. 1, FIG.2, FIG.3A, FIG.3B, and FIG.4. With reference to FIG.5, there is shown a flowchart 500. The flowchart 500 may include operations from 502 to 518 and may be implemented by the system 102 of FIG.1 or by the circuitry 202 of FIG.2. The flowchart 500 may start at 502 and proceed to 504.

[0090] At 504, an audio-visual input may be acquired. The circuitry 202 may be configured to acquire the audio-visual input 114 associated with a human speaker. The audio-visual input 114 may include the first set of video frames 114-1 and the set of audio spectrograms 114-2 associated with the first set of video frames 114-1. In an operation, the system 102 may receive a dataset 116 comprising a video clip and an audio clip. Next, the system 102 may partition the dataset 116 into sets of audio-visual pairs associated with the human speaker, based on a time window. Each set of audio-visual pairs of theDOCKET NO. SYP354694WO02 sets of audio-visual pairs may include frames in a form of a plurality of video frames and a plurality of audio spectrograms associated with the plurality of video frames. The plurality of video frames may include at least one video frame of the first set of video frames. The plurality of audio spectrograms includes at least one audio spectrogram of the set of audio spectrograms. For, instance, the dataset 116 may be partitioned into three sets of the audio-visual pairs – first audio-visual pair, second audio-visual pair, and third audio-visual pair, and further each audio-visual pair of the three sets of the audio-visual pairs may include three video frames and three audio spectrograms. In one embodiment, all the three video frames and the three audio spectrograms may together form a part of any one of the first audio-visual pair, the second audio-visual pair, or the third audio-visual pair. In other embodiment, each of the three video frames and the three audio spectrograms may be segregated to form a part of each one of the first audio-visual pair, the second audio-visual pair, and the third audio-visual pair. In another embodiment, two of the three video frames and the corresponding audio spectrograms may together form a part of any one of the first audio-visual pair, the second audio-visual pair, or the third audio-visual pair, whereas third video frame and third audio spectrogram of the three video frames and the corresponding audio spectrograms may be a part of another audio-visual pair. Details related to the acquisition of the audio-visual input are further described, for example, in FIG.3A.

[0091] At 506, second set of video frames may be generated. The circuitry 202 may be configured to generate the second set of video frames 118 based on application of the generative neural network 104 on the acquired audio-visual input 114. The second set of video frames 118 may include a talking head of the human speaker at different time- instants which corresponds to that of the first set of video frames 114-1. The second set of video frames 118 may include a first video frame, a second video frame, and a third video frame, where the first video frame, the second video frame, and the third video frame may be generated in a temporal sequence of a speech signal represented in the set ofDOCKET NO. SYP354694WO02 audio spectrograms 114-2. Details related to the generation of the second set of video frames are further described, for example, in FIG.3B.

[0092] At 508, first temporal contrastive loss may be computed. The circuitry 202 may be configured to compute a first similarity between the first video frame and the second video frame. The circuitry 202 may also compute a second similarity between the second video frame and the third video frame. Next, the circuitry 202 may compute a first feature embedding based on application of a first fully connected neural network layer on the first similarity. The circuitry 202 may similarly compute a second feature embedding based on application of a second fully connected neural network layer on the second similarity. Next, the circuitry 202 may compute a first cross-entropy loss between the first feature embedding and a first ground truth embedding. The circuitry 202 may also compute a second cross-entropy loss between the second feature embedding and a second ground truth embedding. Further, the circuitry 202 may compute the first temporal contrastive loss 338 based on the first cross-entropy loss and the second cross-entropy loss. Details related to the computation of the first temporal contrastive loss are further described, for example, in FIG.3B.

[0093] At 510, a set of visual embeddings may be extracted. The circuitry 202 may be configured to extract, via the facial feature extractor, a set of visual embeddings for the first set of video frames 114-1 from the generative neural network 104. Details related to the extraction of the set of visual embeddings are further described, for example, in FIG. 3B.

[0094] At 512, a set of audio embeddings may be extracted. The circuitry 202 may be configured to extract a set of audio embeddings based on the set of audio spectrograms 114-2. Details related to the extraction of the set of audio embeddings are further described, for example, in FIG.3B.

[0095] At 514, second temporal contrastive loss may be computed. The circuitry 202DOCKET NO. SYP354694WO02 may be configured to compute the second temporal contrastive loss 346 based on the extracted set of visual embeddings and the extracted set of audio embeddings. Details related to the computation of the second temporal contrastive loss are further described, for example, in FIG.3B.

[0096] At 516, an alignment loss may be computed. The circuitry 202 may be configured to compute the alignment loss between the set of audio spectrograms 114-2 and the lower- half face of the human speaker in the first set of video frames 114-1. Details related to the computation of the alignment loss are further described, for example, in FIG.3B.

[0097] At 518, the generative neural network 104 may be trained. The circuitry 202 may be configured to train the generative neural network based on the first temporal contrastive loss, the second temporal contrastive loss, and the alignment loss. The circuitry 202 may compute a loss based on a weighted sum of the computed first temporal contrastive loss, the computed second temporal contrastive loss, and the computed alignment loss. Further, the generative neural network 104 may be trained for a number of iterations until the computed loss is below a threshold loss. Details related to the training of the generative neural network are further described, for example, in FIG.3B.

[0098] Further, the flowchart 500 may end at 518. Although the flowchart 500 is illustrated as discrete operations, such as, 502, 504, 506, 508, 510, 512, 514, 516, and 518, the disclosure is not so limited. Accordingly, in certain embodiments, such discrete operations may be further divided into additional operations, combined into fewer operations, or eliminated, depending on the implementation without detracting from the essence of the disclosed embodiments.

[0099] Various embodiments of the disclosure may provide a non-transitory computer- readable medium and / or storage medium having stored thereon, computer-executable instructions executable by a machine and / or a computer to operate a system (for example, the system 102 of FIG. 1). Such instructions may cause the system 102 to performDOCKET NO. SYP354694WO02 operations that may include acquisition of an audio-visual input (for example, the audio- visual input 114 of FIG. 1) associated with a human speaker, wherein the audio-visual input 114 includes a first set of video frames (for example, the first set of video frames 114- 1 of FIG.1) and a set of audio spectrograms (for example, the set of audio spectrograms 114-2 of FIG.1) associated with the first set of video frames 114-1. The operations may further include generation of a second set of video frames (for example, the second set of video frames 118 of FIG. 1) based on application of a generative neural network (for example, the generative neural network 104 of FIG.1) on the acquired audio-visual input 114. The second set of video frames 118 may include a talking head of the human speaker at different time-instants which corresponds to that of the first set of video frames 114-1. The operations may further include computation of a first temporal contrastive loss 338 based on the second set of video frames 118. The operations may further include extraction of a set of visual embeddings for the first set of video frames 114-1, from the generative neural network 104. The operations may further include extraction of a set of audio embeddings based on the set of audio spectrograms. The operations may further include computation of a second temporal contrastive loss based on the extracted set of visual embeddings and the extracted set of audio embeddings. The operations may further include computation of an alignment loss between the set of audio spectrograms 114-2 and a lower-half face of the human speaker in the first set of video frames 114-1. The operations may further include training of the generative neural network 104 based on the first temporal contrastive loss, the second temporal contrastive loss, and the alignment loss.

[0100] Exemplary aspects of the disclosure may provide a system (such as, the system 102 of FIG.1) that includes circuitry (such as, the circuitry 202 of FIG.2). The circuitry 202 may be configured to acquire an audio-visual input (for example, the audio-visual input 114 of FIG. 1) associated with a human speaker, wherein the audio-visual input 114DOCKET NO. SYP354694WO02 includes a first set of video frames (for example, the first set of video frames 114-1 of FIG. 1) and a set of audio spectrograms (for example, the set of audio spectrograms 114-2 of FIG. 1) associated with the first set of video frames 114-1. The circuitry 202 may be configured to generate a second set of video frames (for example, the second set of video frames 118 of FIG.1) based on application of a generative neural network (for example, the generative neural network 104 of FIG.1) on the acquired audio-visual input 114. The second set of video frames 118 may include a talking head of the human speaker at different time-instants which corresponds to that of the first set of video frames 114-1. The circuitry 202 may be configured to compute a first temporal contrastive loss based on the second set of video frames 118. The circuitry 202 may be configured to extract a set of visual embeddings for the first set of video frames 114-1, from the generative neural network 104. The circuitry 202 may be configured to extract a set of audio embeddings based on the set of audio spectrograms. The circuitry 202 may be configured to compute a second temporal contrastive loss based on the extracted set of visual embeddings and the extracted set of audio embeddings. The circuitry 202 may be configured to compute an alignment loss between the set of audio spectrograms 114-2 and a lower-half face of the human speaker in the first set of video frames 114-1. The circuitry 202 may be configured to train the generative neural network 104 based on the first temporal contrastive loss, the second temporal contrastive loss, and the alignment loss.

[0101] In an embodiment, the circuitry 202 may be further configured to receive a dataset (for example, the dataset 116 of FIG.1) including a video clip and an audio clip. The circuitry 202 may be further configured to partition the dataset 116 into sets of audio- visual pairs associated with the human speaker, based on a time window. Each set of audio-visual pairs of the sets of audio-visual pairs includes frames in a form of a plurality of video frames and a plurality of audio spectrograms associated with the plurality of video frames. The plurality of video frames includes at least one video frame of the first set ofDOCKET NO. SYP354694WO02 video frames, and the plurality of audio spectrograms includes at least one audio spectrogram of the set of audio spectrograms.

[0102] In an embodiment, the circuitry 202 may be further configured to obtain sets of shuffled audio-visual pairs based on application of a shuffling operation on the frames in each set of audio-visual pairs of the sets of audio-visual pairs. The circuitry 202 may be further configured to acquire the audio-visual input based on application of a sampling operation on each set of shuffled audio-visual pairs of the sets of shuffled audio-visual pairs.

[0103] In an embodiment, the sampling operation includes a random selection of a video frame of the first set of video frames 114-1 and a corresponding audio spectrogram of the first set of audio spectrograms 114-2 from a set of shuffled audio-visual pairs of the set of shuffled audio-visual pairs.

[0104] In an embodiment, the second set of video frames includes a first video frame, a second video frame, and a third video frame. The first video frame, the second video frame, and the third video frame may be generated in a temporal sequence of a speech signal represented in the set of audio spectrograms.

[0105] In an embodiment, the circuitry 202 may be further configured to compute a first similarity between the first video frame and the second video frame. The circuitry 202 may be further configured to compute a second similarity between the second video frame and the third video frame. The circuitry 202 may be further configured to compute a first feature embedding based on application of a first fully connected neural network layer on the first similarity. The circuitry 202 may be further configured to compute a second feature embedding based on application of a second fully connected neural network layer on the second similarity. The circuitry 202 may be further configured to compute a first cross- entropy loss between the first feature embedding and a first ground truth embedding. The circuitry 202 may be further configured to compute a second cross-entropy loss betweenDOCKET NO. SYP354694WO02 the second feature embedding and a second ground truth embedding. The first temporal contrastive loss may be computed further based on the first cross-entropy loss and the second cross-entropy loss.

[0106] In an embodiment, the circuitry 202 may be further configured to determine, based on the audio-visual input 114, audio-visual embedding pairs between the extracted set of visual embeddings and the extracted set of audio embeddings. The circuitry 202 may be further configured to select an anchor visual embedding from the extracted set of visual embeddings. The circuitry 202 may be further configured to select an anchor audio embedding from the extracted set of audio embeddings. The circuitry 202 may be further configured to separate the audio-visual embedding pairs into first positive pairs and first negative pairs based on the anchor visual embedding. The circuitry 202 may be further configured to separate the audio-visual embedding pairs into second positive pairs and second negative pairs based on the anchor audio embedding. The circuitry 202 may be further configured to compute similarity values based on a vector similarity between each pair of the first positive pairs, the second positive pairs, the first negative pairs, and the second negative pairs. The circuitry 202 may be further configured to compute the second temporal contrastive loss based on the computed similarity values.

[0107] In an embodiment, the circuitry 202 may further be configured to create a correlation matrix based on the set of audio spectrograms 114-2 and visual embeddings for the lower-half face of the human speaker in the first set of video frames 114-1. The circuitry 202 may further be configured to compute a cross-entropy loss between the correlation matrix and a ground truth matrix. The ground truth matrix may be a doubly stochastic permutation matrix. The alignment loss may be the computed cross-entropy loss.

[0108] In an embodiment, the circuitry 202 may further be configured to compute a loss based on a weighted sum of the computed first temporal contrastive loss, the computedDOCKET NO. SYP354694WO02 second temporal contrastive loss, and the computed alignment loss. The generative neural network 104 may be trained for a number of iterations until the computed loss is below a threshold loss.

[0109] In an embodiment, the circuitry 202 may further be configured to receive, via a user interface of a user device (for example, the user device 112 of FIG.1), an input video clip associated with the human speaker. The input video clip includes initial video data and audio data associated with the initial video data. The initial video data includes a talking head of a second human speaker that is misaligned with spoken words of the audio data. The circuitry 202 may further be configured to prepare an input for the trained generative neural network 104 based on the input video clip. The circuitry 202 may further be configured to generate output video data based on application of the trained generative neural network 104 on the prepared input. The output video data may include a synthetic talking head of the second human speaker that is aligned with the spoken words of the audio data.

[0110] The present disclosure may also be embedded in a computer program product, which comprises all the features that enable the implementation of the methods described herein, and which when loaded in a computer system is able to carry out these methods. Computer program, in the present context, means any expression, in any language, code or notation, of a set of instructions intended to cause a system with information processing capability to perform a particular function either directly, or after either or both of the following: a) conversion to another language, code or notation; b) reproduction in a different material form.

[0111] While the present disclosure is described with reference to certain embodiments, it will be understood by those skilled in the art that various changes may be made, and equivalents may be substituted without departure from the scope of the present disclosure. In addition, many modifications may be made to adapt a particular situation or material toDOCKET NO. SYP354694WO02 the teachings of the present disclosure without departure from its scope. Therefore, it is intended that the present disclosure is not limited to the embodiment disclosed, but that the present disclosure will include all embodiments that fall within the scope of the appended claims.

Claims

DOCKET NO. SYP354694WO02 CLAIMS What is claimed is:

1. A system, comprising: circuitry configured to: acquire an audio-visual input associated with a human speaker, wherein the audio-visual input includes a first set of video frames and a set of audio spectrograms associated with the first set of video frames; generate a second set of video frames based on application of a generative neural network on the acquired audio-visual input, wherein the second set of video frames include a talking head of the human speaker at different time-instants which corresponds to that of the first set of video frames; compute a first temporal contrastive loss based on the second set of video frames; extract, from the generative neural network, a set of visual embeddings for the first set of video frames; extract a set of audio embeddings based on the set of audio spectrograms; compute a second temporal contrastive loss based on the extracted set of visual embeddings and the extracted set of audio embeddings; compute an alignment loss between the set of audio spectrograms and a lower-half face of the human speaker in the first set of video frames; and train the generative neural network based on the first temporal contrastive loss, the second temporal contrastive loss, and the alignment loss.

2. The system according to claim 1, wherein the circuitry is further configured to: receive a dataset comprising a video clip and an audio clip; andDOCKET NO. SYP354694WO02 partition, based on a time window, the dataset into sets of audio-visual pairs associated with the human speaker, wherein each set of audio-visual pairs of the sets of audio-visual pairs comprise frames in a form of a plurality of video frames and a plurality of audio spectrograms associated with the plurality of video frames, the plurality of video frames includes at least one video frame of the first set of video frames, and the plurality of audio spectrograms includes at least one audio spectrogram of the set of audio spectrograms.

3. The system according to claim 2, wherein the circuitry is further configured to: obtain sets of shuffled audio-visual pairs based on application of a shuffling operation on the frames in each set of audio-visual pairs of the sets of audio-visual pairs; and acquire the audio-visual input based on application of a sampling operation on each set of shuffled audio-visual pairs of the sets of shuffled audio-visual pairs.

4. The system according to claim 3, wherein the sampling operation includes a random selection of a video frame of the first set of video frames and a corresponding audio spectrogram of the first set of audio spectrograms from a set of shuffled audio-visual pairs of the set of shuffled audio-visual pairs.

5. The system according to claim 1, wherein the second set of video frames includes a first video frame, a second video frame, and a third video frame, andDOCKET NO. SYP354694WO02 the first video frame, the second video frame, and the third video frame are generated in a temporal sequence of a speech signal represented in the set of audio spectrograms.

6. The system according to claim 5, wherein the circuitry is further configured to: compute a first similarity between the first video frame and the second video frame; compute a second similarity between the second video frame and the third video frame; compute a first feature embedding based on application of a first fully connected neural network layer on the first similarity; compute a second feature embedding based on application of a second fully connected neural network layer on the second similarity; compute a first cross-entropy loss between the first feature embedding and a first ground truth embedding; and compute a second cross-entropy loss between the second feature embedding and a second ground truth embedding, wherein the first temporal contrastive loss is computed further based on the first cross-entropy loss and the second cross-entropy loss.

7. The system according to claim 1, wherein the circuitry is further configured to: determine, based on the audio-visual input, audio-visual embedding pairs between the extracted set of visual embeddings and the extracted set of audio embeddings; select an anchor visual embedding from the extracted set of visual embeddings; select an anchor audio embedding from the extracted set of audio embeddings;DOCKET NO. SYP354694WO02 separate the audio-visual embedding pairs into first positive pairs and first negative pairs based on the anchor visual embedding; separate the audio-visual embedding pairs into second positive pairs and second negative pairs based on the anchor audio embedding; compute similarity values based on a vector similarity between each pair of the first positive pairs, the second positive pairs, the first negative pairs, and the second negative pairs; and compute the second temporal contrastive loss based on the computed similarity values.

8. The system according to claim 1, wherein the circuitry is further configured to: create a correlation matrix based on the set of audio spectrograms and visual embeddings for the lower-half face of the human speaker in the first set of video frames; and compute a cross-entropy loss between the correlation matrix and a ground truth matrix, wherein the ground truth matrix is a doubly stochastic permutation matrix, and the alignment loss is the computed cross-entropy loss.

9. The system according to claim 1, wherein the circuitry is further configured to compute a loss based on a weighted sum of the computed first temporal contrastive loss, the computed second temporal contrastive loss, and the computed alignment loss, wherein the generative neural network is trained for a number of iterations until the computed loss is below a threshold loss.DOCKET NO. SYP354694WO02 10. The system according to claim 1, wherein the circuitry is further configured to: receive, via a user interface of a user device, an input video clip associated with the human speaker, wherein the input video clip includes initial video data and audio data associated with the initial video data, and the initial video data includes a talking head of a second human speaker that is misaligned with spoken words of the audio data; prepare an input for the trained generative neural network based on the input video clip; and generate output video data based on application of the trained generative neural network on the prepared input, wherein the output video data includes a synthetic talking head of the second human speaker that is aligned with the spoken words of the audio data.

11. A method, comprising: in a system: acquiring an audio-visual input associated with a human speaker, wherein the audio-visual input includes a first set of video frames and a set of audio spectrograms associated with the first set of video frames; generating a second set of video frames based on application of a generative neural network on the acquired audio-visual input, wherein the second set of video frames include a talking head of the human speaker at different time-instants which corresponds to that of the first set of video frames; computing a first temporal contrastive loss based on the second set of video frames;DOCKET NO. SYP354694WO02 extracting, from the generative neural network, a set of visual embeddings for the first set of video frames; extracting a set of audio embeddings based on the set of audio spectrograms; computing a second temporal contrastive loss based on the extracted set of visual embeddings and the extracted set of audio embeddings; computing an alignment loss between the set of audio spectrograms and a lower-half face of the human speaker in the first set of video frames; and training the generative neural network based on the first temporal contrastive loss, the second temporal contrastive loss, and the alignment loss.

12. The method according to claim 11, further comprising: receiving a dataset comprising a video clip and an audio clip; and partitioning, based on a time window, the dataset into sets of audio-visual pairs associated with the human speaker, wherein each set of audio-visual pairs of the sets of audio-visual pairs comprises frames in a form of a plurality of video frames and a plurality of audio spectrograms associated with the plurality of video frames, the plurality of video frames includes at least one video frame of the first set of video frames, and the plurality of audio spectrograms includes at least one audio spectrogram of the set of audio spectrograms.

13. The method according to claim 12, further comprising:DOCKET NO. SYP354694WO02 obtaining sets of shuffled audio-visual pairs based on application of a shuffling operation on the frames in each set of audio-visual pairs of the sets of audio-visual pairs; and acquiring the audio-visual input based on application of a sampling operation on each set of shuffled audio-visual pairs of the sets of shuffled audio-visual pairs.

14. The method according to claim 11, wherein the second set of video frames includes a first video frame, a second video frame, and a third video frame, and the first video frame, the second video frame, and the third video frame are generated in a temporal sequence of a speech signal represented in the set of audio spectrograms.

15. The method according to claim 14, further comprising: computing a first similarity between the first video frame and the second video frame; computing a second similarity between the second video frame and the third video frame; computing a first feature embedding based on application of a first fully connected neural network layer on the first similarity; computing a second feature embedding based on application of a second fully connected neural network layer on the second similarity; computing a first cross-entropy loss between the first feature embedding and a first ground truth embedding; and computing a second cross-entropy loss between the second feature embedding and a second ground truth embedding,DOCKET NO. SYP354694WO02 wherein the first temporal contrastive loss is computed further based on the first cross-entropy loss and the second cross-entropy loss.

16. The method according to claim 11, further comprising: determining, based on the audio-visual input, audio-visual embedding pairs between the extracted set of visual embeddings and the extracted set of audio embeddings; selecting an anchor visual embedding from the extracted set of visual embeddings; selecting an anchor audio embedding from the extracted set of audio embeddings; separating the audio-visual embedding pairs into first positive pairs and first negative pairs based on the anchor visual embedding; separating the audio-visual embedding pairs into second positive pairs and second negative pairs based on the anchor audio embedding; computing similarity values based on a vector similarity between each pair of the first positive pairs, the second positive pairs, the first negative pairs, and the second negative pairs; and computing the second temporal contrastive loss based on the computed similarity values.

17. The method according to claim 11, further comprising: creating a correlation matrix based on the set of audio spectrograms and visual embeddings for the lower-half face of the human speaker in the first set of video frames; andDOCKET NO. SYP354694WO02 computing a cross-entropy loss between the correlation matrix and a ground truth matrix, wherein the ground truth matrix is a doubly stochastic permutation matrix, and the alignment loss is the computed cross-entropy loss.

18. A non-transitory computer-readable medium having stored thereon, computer- executable instructions that when executed by a system, causes the system to execute operations, the operations comprising: acquiring an audio-visual input associated with a human speaker, wherein the audio-visual input includes a first set of video frames and a set of audio spectrograms associated with the first set of video frames; generating a second set of video frames based on application of a generative neural network on the acquired audio-visual input, wherein the second set of video frames include a talking head of the human speaker at different time-instants which corresponds to that of the first set of video frames; computing a first temporal contrastive loss based on the second set of video frames; extracting, from the generative neural network, a set of visual embeddings for the first set of video frames; extracting a set of audio embeddings based on the set of audio spectrograms; computing a second temporal contrastive loss based on the extracted set of visual embeddings and the extracted set of audio embeddings; computing an alignment loss between the set of audio spectrograms and a lower-half face of the human speaker in the first set of video frames; andDOCKET NO. SYP354694WO02 training the generative neural network based on the first temporal contrastive loss, the second temporal contrastive loss, and the alignment loss.

19. The non-transitory computer-readable medium according to claim 18, wherein the operations further comprise: receiving a dataset comprising a video clip and an audio clip; and partitioning, based on a time window, the dataset into sets of audio-visual pairs associated with the human speaker, wherein each set of audio-visual pairs of the sets of audio-visual pairs comprises frames in a form of a plurality of video frames and a plurality of audio spectrograms associated with the plurality of video frames, the plurality of video frames includes at least one video frame of the first set of video frames, and the plurality of audio spectrograms includes at least one audio spectrogram of the set of audio spectrograms.

20. The non-transitory computer-readable medium according to claim 19, wherein the operations further comprise: obtaining sets of shuffled audio-visual pairs based on application of a shuffling operation on the frames in each set of audio-visual pairs of the sets of audio-visual pairs; and acquiring the audio-visual input based on application of a sampling operation on each set of shuffled audio-visual pairs of the sets of shuffled audio-visual pairs.