Multimodal duration-controlled TTS (dc-TTS) system for automatic dubbing
Patent Information
- Application Number
- PCT/IB2025/052232
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-02
- Filing Date
- 2025-02-28
- Publication Date
- 2025-10-02
AI Technical Summary
Existing dubbing technologies struggle to maintain synchronization between dubbed audio and video, particularly with respect to lip movements, and fail to capture the speaking style and prosody of the original speaker when translating to a new language, leading to unnatural-sounding speech and visual artifacts.
A multimodal duration-controlled text-to-speech system that integrates visual information from reference videos to control the duration of generated audio, using neural networks for lip synchronization, speaker embedding, and cross-lingual adaptability, enabling precise audio-visual alignment and preservation of speaker characteristics.
The system achieves improved lip-sync accuracy, cross-lingual adaptability, and speaker characteristic preservation, resulting in a more natural and convincing dubbing experience with reduced manual intervention.
Smart Images

Figure IB2025052232_02102025_PF_FP_ABST
Abstract
Description
MULTIMODAL DURATION-CONTROLLED TTS (DC-TTS) SYSTEM FOR AUTOMATIC DUBBINGCROSS-REFERENCE TO RELATED APPLICATIONS / INCORPORATION BY REFERENCE
[0001] This Application also makes reference to Indian Provisional Application Ser. No. 202411015609 which was filed on March 2, 2024. The above stated Patent Applications are hereby incorporated herein by reference in their entirety.FIELD
[0002] Various embodiments of the disclosure relate to speech processing. More specifically, various embodiments of the disclosure relate to a multimodal duration- controlled text to speech (TTS) system for automatic dubbing.BACKGROUND
[0003] Automatic dubbing of video content from one language to another has become increasingly important in our globalized world. Traditional dubbing approaches often struggle to maintain synchronization between the dubbed audio and the original video, particularly with respect to lip movements. This can result in a jarring viewing experience that detracts from the content. Existing methods for addressing this issue typically rely on adjusting the speed of the synthesized speech or modifying the video playback rate. However, these techniques often lead to unnatural-sounding speech or visual artifacts in the video.
[0004] Text-to-speech (TTS) systems have made significant advancements in generating natural-sounding speech, but controlling the duration and timing of the generated audio to match video content remains challenging. Some approaches attempt to predict and control phoneme durations, but achieving precise synchronization with lip movements across different languages is still difficult. Additionally, many current dubbing systems fail to adequately capture the speaking style and prosody of the original speakerwhen translating to a new language. There is a need for improved dubbing technologies that can generate high-quality, synchronized speech across languages while maintaining naturalness and speaker characteristics.
[0005] Further limitations and disadvantages of conventional and traditional approaches will become apparent to one of skill in the art, through comparison of described systems with some aspects of the present disclosure, as set forth in the remainder of the present application and with reference to the drawings.SUMMARY
[0006] A system and method for multimodal duration-controlled text to speech (TTS) based automatic dubbing is provided substantially as shown in, and / or described in connection with, at least one of the figures, as set forth more completely in the claims.
[0007] These and other features and advantages of the present disclosure may be appreciated from a review of the following detailed description of the present disclosure, along with the accompanying figures in which like reference numerals refer to like parts throughout.BRIEF DESCRIPTION OF THE DRAWINGS
[0008] FIG. 1 is a block diagram that illustrates an exemplary network environment for multimodal duration-controlled text to speech (TTS) based automatic dubbing, in accordance with an embodiment of the disclosure.
[0009] FIG. 2 is a block diagram that illustrates an exemplary system of FIG. 1 , in accordance with an embodiment of the disclosure.
[0010] FIG. 3 is a diagram that illustrates a first exemplary architecture of the system of FIG. 1 , in accordance with an embodiment of the disclosure.
[0011] FIG. 4 is a diagram that illustrates a second exemplary architecture of the system of FIG. 1 , in accordance with an embodiment of the disclosure.
[0012] FIG. 5 is a diagram that illustrates an exemplary architecture of audio-visualcloning network, in accordance with an embodiment of the disclosure.
[0013] FIG. 6 is a diagram that illustrates an exemplary architecture of lip encoder, in accordance with an embodiment of the disclosure.
[0014] FIG. 7 is a diagram that illustrates an exemplary inference of the system of FIG. 1 , in accordance with an embodiment of the disclosure.
[0015] FIG. 8 is a flowchart that illustrates operations of an exemplary method for multimodal duration-controlled text to speech (TTS) based automatic dubbing, in accordance with an embodiment of the disclosure.DETAILED DESCRIPTION
[0016] The following described implementation may be found in a system and a method for multimodal duration-controlled text-to-speech (TTS) based automatic dubbing. Exemplary aspects of the disclosure may provide a system, which may include circuitry configured to receive a dataset associated with a human speaker. The dataset may include a reference video clip of the human speaker, a reference audio of the human speaker in a first language, and a text that includes at least one of a translation of a transcript of the reference audio in a second language or a modification of the transcript. The circuitry may generate text features based on application of a text encoder on the text. The circuitry may also generate visual features based on application of a visual feature extractor on the reference video clip. Additionally, the circuitry may generate a speaker embedding based on the reference audio and lip features associated with the human speaker based on the reference video clip. The circuitry may then generate audio tokens based on application of a neural language model on the text features, the visual features, the speaker embedding, and the lip features. Subsequently, the circuitry may generate an audio in the first language, or the second language based on application of a neural vocoder on the audio tokens.
[0017] The circuitry may train the neural language model based on the generated audioand a loss function. The neural language model may be trained for a number of iterations until a value of the loss function is below a threshold loss, where the threshold loss may correspond to a threshold time-difference between lip movements of the human speaker in the reference video and spoken words in the generated audio. In one embodiment, the circuitry may train the entire neural language model. In another embodiment, the circuitry may train only cross-attention layers associated with the neural language model.
[0018] Once the neural language model is trained, the system may utilize the model for generating audio clips based on a received input. The input may be received via a user interface of a user device, where the input may include a video clip of the human speaker and an input audio associated with the video clip. Further, the circuitry may receive an input text that may include at least one of a translation of a transcript of the input audio in the second language or a modification of the transcript. The circuitry may generate input text features based on the application of the text encoder on the input text, input visual features based on application of the visual feature extractor on the input video clip, a second speaker embedding based on the input audio, and input lip features associated with the human speaker based on the input video clip. The circuitry may then generate input audio tokens based on application of the trained neural language model on the input text features, the input visual features, the second speaker embedding, and the input lip features. Subsequently, the circuitry may generate an audio clip in the second language, or the first language based on application of the neural vocoder on the input audio tokens. The audio clip may be generated such that a time-difference between lip movements of the human speaker in the input video and spoken words of the audio clip is minimized.
[0019] Typically, various Al-based dubbing technologies used for voice conversion and speech processing may result in generation of audio via TTS. However, there is a higher likelihood that the generated audio in the target language may have a different duration compared to the corresponding segment in the source language, which may cause thegenerated audio to misalign with the source video content. This may lead to asynchronization between the source video content and the generated audio, potentially adversely affecting the quality of dubbing from audio-visual synchronization perspectives.
[0020] Various signal processing algorithms, such as Time-Domain Pitch-Synchronous Overlap and Add (TD-PSOLA), Waveform Similarity Overlap-Add (WSOLA), and similar techniques, are typically used for altering the duration of the generated audio. These algorithms may apply interpolation-based methods for sampling or up sampling video frames to alter the duration of the source video. Phoneme duration predictor networks may also be utilized in TTS to achieve duration controllability. However, all such techniques may alter or control the duration of audio / video only to a certain extent, beyond which the quality and intelligibility of the generated audio, as well as the final dubbing output, may significantly deteriorate.
[0021] The system may enable the generation of audio synchronized with a reference video (depicting a human speaker speaking) for any given text. The system may involve leveraging visual information from the reference video to exert control over the output of a text-to-speech (TTS) system. The system may translate the text into a target language. Further, the system may take the reference video and the translated text as input and generate a time-synchronized audio conditioned on the reference video. The system may also integrate video modality into a Generative Pre-Training Transformer (GPT) based Text-to-Speech (TTS) model for speech synthesis based on lip movements of the human speaker in the reference video. The system may also control duration of the generated audio in an autoregressive (AR) Large Language Model (LLM).
[0022] The system may leverage lightweight faster training by training only the randomly initialized cross-attention layers for the task of the duration controllability of the generated audio. The system may be adept at handling global scene context understanding for the synthesis of better speech style. Further, the system may have zero-shot voice cloningcapability. Specifically, the system may achieve duration controllability of the output (generated audio) while maintaining intelligibility and quality of speech, offering several advantages in applications like dubbing. Such advantages may include:1. Improved lip-sync accuracy: By utilizing visual information and lip features, the system may generate audio that closely matches the lip movements of the speaker, resulting in a more natural and convincing dubbing experience.2. Cross-lingual adaptability: The system may effectively handle translations between different languages, maintaining synchronization even when the source and target languages have different phonetic structures and speech patterns.3. Preservation of speaker characteristics: Through the use of speaker embeddings, the system may maintain the voice characteristics and speaking style of the original speaker in the dubbed audio, enhancing the authenticity of the output.4. Flexible text modification: The system may accommodate modifications to the original transcript, allowing for creative adaptations or localization of content while maintaining synchronization with the video.5. Reduced manual intervention: By automatically generating synchronized audio, the system may significantly reduce the need for manual adjustments in the dubbing process, potentially saving time and resources in content localization.6. Enhanced user experience: The improved synchronization and speech quality may result in a more immersive and enjoyable viewing experience for audiences consuming dubbed content.
[0023] FIG. 1 is a block diagram that illustrates an exemplary network environment for multimodal duration-controlled text to speech (TTS) based automatic dubbing, in accordance with an embodiment of the disclosure. With reference to FIG. 1 , there is shown a network environment 100. The network environment 100 may include a system 102, a duration-controlled Text-to-Speech (TTS) network 104, a server 106, a database 108, acommunication network 112, and a user device 114.
[0024] The system 102 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive a dataset 110, which may include a reference video clip 110- 1 of a human speaker, a reference audio 110-2 of the human speaker in a first language, and a text 110-3 that includes at least one of a translation of a transcript of the reference audio in a second language or a modification of the transcript. The system 102 may further be configured to use the dataset 110 to train the duration-controlled TTS network 104 for audio dubbing with lip synchronization. Once trained, the system 102 may deploy the duration-controlled TTS network 104 for inference. Examples of the system 102 may include, but are not limited to, a digital media player (DMP), a micro-console, a TV tuner, a digital media streamer, a media extender / regulator, a digital media hub, a computer workstation, a mainframe computer, a handheld computer, a smart appliance, a plug-in device, and / or any other computing device with content streaming functionality.
[0025] The system 102 may store the duration-controlled TTS network 104 or may be remotely connected to another system (such as the server 106) that hosts the duration- controlled TTS network 104. When hosted on another system, the system 102 may send instructions to control training or inference of the duration-controlled TTS network 104 via remote calls (e.g., API calls).
[0026] The duration-controlled TTS network 104 may include suitable logic, circuitry, interfaces, and / or code that may be configured to perform automatic audio dubbing with lip synchronization. The duration-controlled TTS network 104 may be a hybrid network, which may include multiple neural networks including a text encoder 104-1 , an audio-visual cloning network 104-2, a lip encoder 104-3, a neural language model 104-4, and a neural vocoder 104-5. Outputs from the text encoder 104-1 , the audio-visual cloning network 104- 2, the lip encoder 104-3 may be connected to an input layer of the neural language model 104-4, and output of the neural language model 104-4 may be connected to input of theneural vocoder 104-5.
[0027] Each of the multiple neural networks may be a computational network or a system of artificial neurons, arranged in a plurality of layers, as nodes. The plurality of layers of the neural network may include an input layer, one or more hidden layers, and an output layer. Each layer of the plurality of layers may include one or more nodes (or artificial neurons). Outputs of all nodes in the input layer may be coupled to at least one node of hidden layer(s). Similarly, inputs of each hidden layer may be coupled to outputs of at least one node in other layers of the neural network. Outputs of each hidden layer may be coupled to inputs of at least one node in other layers of the neural network. Node(s) in the final layer may receive inputs from at least one hidden layer to output a result. The number of layers and the number of nodes in each layer may be determined from hyperparameters of the neural network. Such hyper-parameters may be set before or after training the neural network on the training dataset.
[0028] Each neural network of the neural networks may include electronic data, which may be implemented as, for example, a software component of an application executable on the system 102. Each of the neural networks may rely on libraries, external scripts, or other logic / instructions for execution by a processing device, such as the system 102. Further, each of the neural networks may rely on code and routines to enable a computing device, such as the system 102 to perform one or more operations, such as automatic audio dubbing. In some embodiments, each of the neural networks may be implemented using hardware including a processor, a microprocessor (e.g., to perform or control performance of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, each of the neural networks may be implemented using a combination of hardware and software.
[0029] The text encoder 104-1 may include suitable logic, circuitry, interfaces, and / or code that may be configured to convert plain text strings into high-dimensional vectorrepresentations (also referred to as embeddings or. text features). The text features may be generated based on application of the text encoder 104-1 on the text 110-3. The text encoder 104-1 may include hardware or software components, including but not limited to deep learning models such as Bidirectional Encoder Representations from Transformers (BERT) or Generative Pre-trained Transformer (GPT).
[0030] The audio-visual cloning network 104-2 may be a voice cloning network that includes a visual feature extractor 506 (shown in FIG. 5) and a voice feature extractor (shown in FIG. 5 as Mel-Spectrogram feature extractor 502). The audio-visual cloning network 104-2 may also include an attention encoder 504 and a multimodal transformer 508 (also shown in FIG. 5), for instance.
[0031] The audio-visual cloning network 104-2 may be configured to process the dataset 110, which may include a reference video clip 110-1 and a reference audio 110-2. The audio-visual cloning network 104-2 may further be configured to generate a first speaker embedding associated with the human speaker based on the reference speaker audio 110-2. The audio-visual cloning network 104-2 may further be configured to generate visual features based on application of the visual feature extractor 506 on the reference scene video 110-1.
[0032] The visual feature extractor 506 may include suitable logic, circuitry, interfaces, and / or code that may be configured to generate visual features. The visual features may be generated based on application of the visual feature extractor 506 on the reference video clip 110-1. The visual feature extractor 506 may include hardware and / or software components, including deep learning models such as convolutional neural networks (CNNs) or vision transformers, or contrastively trained multimodal networks (such as video-text models). The visual feature extractor 506 may transform raw data associated with the reference video clip 110-1 , such as images or video frames, into numerical features. This may be achieved using various methods like mathematical transformations,statistics, or using pre-trained models. For example, in the context of image processing, one common method is using Convolutional Neural Networks (CNNs). CNNs are preferred for feature extraction from images because CNNs are specifically designed for processing color images and perform complex tasks such as image classification, object detection, or segmentation. The CNNs may extract complex and descriptive features with any variations such as lighting conditions, scale, and other factors in the image. In some embodiments, the visual feature extractor 506 may be a pre-trained model that provides a unified videotext representation. For instance, the pre-trained model may be VideoCLIP, as discussed in Hu et. al., “VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding”, arXiv:2109.14084, which is incorporated herein in entirety by reference. This representation may be captured by transformer model parameters for video and text. The text encoder of VideoCLIP may serve as self-supervision for videos during pre-training and as a hyper network to provide hidden states of segment textual labels for a video token.
[0033] The lip encoder 104-3 may be a neural network that may process images of lip regions (cropped from the reference video clip 110-1 ) to generate lip features. Such features may be referred to as embeddings (i.e., high dimensional feature vectors) which capture essential characteristics of the lip movements and shapes, which can be used for tasks such as lip-based biometric authentication, lipreading, and audio-visual speech enhancement. As an example, the lip encoder 104-3 may use convolutional neural networks (CNNs) that capture spatial features from the lip regions to generate the lip features from the lip regions in the images. As another example, the lip encoder 104-3 may include an Audio-Visual Hidden Unit BERT (AV-HuBERT) for lip feature extraction by leveraging both audio and visual cues from a cropped version of the reference video clip 110-1 where the audio and visual signals may be correlated. Specifically, AV-HuBERT may capture the movement of the lips during speech, combining these visual signals withauditory information to learn robust speech representations. Additionally, the lip encoder 104-3 may include an up-sampling network 604 (as shown in FIG. 6) such as a network with transposed convolution layers on the output of AV- HuBERT to generate the lip features.
[0034] The neural language model 104-4 may be a neural network that processes audio data (e.g., embeddings) and generates a sequence of discrete audio tokens representing audio information. The neural language model 104-4 may take an audio embedding, which is a numerical representation of the audio data, as input and generates audio tokens as output. An example of such a model is the Generative Spoken Language Model (GSLM). The GSLM is based on word-size continuous-valued audio tokens and is designed to generate diverse and expressive language output. Another example is the AudioLM model. The AudioLM model may map the input audio to a sequence of discrete tokens and casts audio generation as a language modeling task in this representation space. The AudioLM consists of two components: Audio Embedding and Audio Language Models. The Audio Embedding component such as w2v-BERT may a BERT-type language model trained on speech data that generates semantic tokens. The audio data may be then converted into discrete acoustic tokens via a neural codec encoder model (EnCodec). In an exemplary embodiment, the neural language model may be a transformer-based model with sequence-modeling capabilities applied to the audio domain.
[0035] The neural vocoder 104-5 may include suitable logic, circuitry, interfaces, and / or code that may be configured to generate an audio in a first language (e.g. English), or a second language (e.g., Japanese) based on an input that includes the audio tokens. The language of the generated audio may be selected by the user.
[0036] The neural vocoder 104-5 may include a deep neural network to synthesize audio samples from the input of linguistic or acoustic features (i.e., audio tokens) and is a critical component in many Text-to-Speech (TTS) systems. An example of the neural vocoder104-5 is Vocos, which is designed to synthesize audio waveforms from acoustic features. The neural vocoder 104-5 may be capable of reconstructing audio from EnCodec tokens, which are a type of audio token. For instance, a transformer-based text-to-audio model such as Bark may be used to generate audio tokens and Vocos may be used to generate the final audio from such tokens. Another example of the neural vocoder 104-5 is Hi-Fi GAN (High Fidelity Generative Adversarial Network) for speech synthesis.
[0037] The server 106 may include suitable logic, circuitry, and interfaces, and / or code that may be configured to receive the dataset 110 including the reference video clip 110- 1 of the human speaker, the reference audio 110-2 of the human speaker in the first language, and the text 110-3. The server 106 may also host the duration-controlled TTS network 104 (trained or untrained), which includes the text encoder 104-1 , the audio-visual cloning network 104-2, the lip encoder 104-3, the neural language model 104-4, and the neural vocoder 104-5.
[0038] The server 106 may be implemented as a cloud server and may execute operations through web applications, cloud applications, HTTP requests, repository operations, file transfer, and the like. Other example implementations of the server 106 may include, but are not limited to, a database server, a file server, a web server, a media server, an application server, a mainframe server, a machine learning server (enabled with or hosting, for example, a computing resource, a memory resource, and a networking resource), or a cloud computing server.
[0039] In at least one embodiment, the server 106 may be implemented as a plurality of distributed cloud-based resources by use of several technologies that are well known to those ordinarily skilled in the art. A person with ordinary skill in the art will understand that the scope of the disclosure may not be limited to the implementation of the server 106 and the system 102, as two separate entities. In certain embodiments, the functionalities of the server 106 can be incorporated in its entirety or at least partially in the system 102 withouta departure from the scope of the disclosure. In certain embodiments, the server 106 may host the database 108. Alternatively, the server 106 may be separate from the database108 and may be communicatively coupled to the database 108.
[0040] The database 108 may include suitable logic, interfaces, and / or code that may be configured to store the dataset 110 including the reference video clip 110-1 of the human speaker, the reference audio 110-2 of the human speaker in the first language, and the text 110-3. The database 108 may include multiple datasets similar to the dataset 110. The database 108 may also include generated text features, visual features, first speaker embedding, lip features, audio tokens, and audio. The database 108 may be derived from data off a relational or non-relational database, or a set of comma-separated values (csv) files in conventional or big-data storage. The database 108 may be stored or cached on a device, such as a server (e.g., the server 106) or the system 102. The device storing the database 108 may be configured to receive audio related commands or instructions from the system 102 or the server 106. In response, the device of the database 108 may be configured to retrieve and provide response of the query to the system 102 or the server 106, based on the received query.
[0041] In some embodiments, the database 108 may be hosted on a plurality of servers stored at the same or different locations. The operations of the database 108 may be executed using hardware including a processor, a microprocessor (e.g., to perform or control performance of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). In some other instances, the database 108 may be implemented using software.
[0042] The communication network 112 may include a communication medium through which the system 102 and the server 106 may communicate with one another. The communication network 112 may be one of a wired connection or a wireless connection. Examples of the communication network 112 may include, but are not limited to, theInternet, a cloud network, Cellular or Wireless Mobile Network (such as Long-Term Evolution and 5thGeneration (5G) New Radio (NR)), satellite communication system (using, for example, low earth orbit satellites), a Wireless Fidelity (Wi-Fi) network, a Personal Area Network (PAN), a Local Area Network (LAN), or a Metropolitan Area Network (MAN). Various devices in the network environment 100 may be configured to connect to the communication network 112 in accordance with various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of a Transmission Control Protocol and Internet Protocol (TIP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Zig Bee, EDGE, IEEE 802.11 , light fidelity (Li-Fi), 802.16, IEEE 802.11 s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device to device communication, cellular communication protocols, and Bluetooth (BT) communication protocols.
[0043] The user device 114 may include a user-interface through which a user may interact with the system 102, send queries, feed commands and instructions, provide reference audios and video clips for training the duration-controlled TTS network 104. The user may be a human speaker for voice cloning or a user of the system 102. The user device 114 may be fixed at a place or may be portable. Examples of the user device 114 may include, but not limited to, a smartphone, a touchpad, a GUI interface, a personal computer, a wearable device such as a smartwatch or an eXtended-Reality (XR) device, a microphone, or a display device.
[0044] In operation, the system 102 may be configured to receive the dataset 110. The dataset 110 may include the reference video clip 110-1 of the human speaker, the reference audio 110-2 of the human speaker, and the text 110-3. The reference audio 110- 2 of the human speaker may be in any of the first language or the second language. In at least one embodiment, the dataset 110 may also include a ground truth (GT) audio 302(shown in FIG. 3 and FIG. 4). The GT audio 302 may be used as a ground truth against the audio generated by the neural vocoder 104-5.
[0045] The text 110-3 may include at least one of a translation of a transcript of the reference audio 110-2 in the second language. The first language may be referred to as a source language and the second language may be referred to as a target language which may be same as or different from the source language. For instance, if the source language is English, then the target language may be Japanese. In an embodiment, the text 110-3 may include a modification of the transcript of the reference audio 110-2. By way of example, and not limitation, the modification of the transcript may include (i) a modification of a language of at least a part of the transcript from the first language to the second language, (ii) a substitution of at least a part of the transcript with a first new text in the first language, (iii) a substitution of at least a part of the transcript with a second new text in the second language, (iv) an addition of the first new text to the transcript in the first language, (v) an addition of the second new text to the transcript in the second language, or (vi) a removal of at least a part of the transcript. FIG. 3 and FIG. 4 may provide further details related to the reception of the dataset 110.
[0046] After reception of the dataset 110, the system 102 may execute an operation to generate text features based on application of the text encoder 104-1 on the text 110-3. The text features may include linguistic and semantic information extracted from the text 110-3, such as word or phoneme embeddings. The text encoder 104-1 may convert plain text strings into high-dimensional vector representations which may capture complex semantic relationships and contextual information within the text 110-3. In an embodiment, the system 102 may generate a one-hot vector-based language ID to specify the language of the text 110-3. FIG. 3 and FIG. 4 may provide further details related to the generation of the text features.
[0047] The system 102 may also execute an operation to generate visual features basedon application of the visual feature extractor 506 (which is a component of audio-visual cloning network 104-2 and is shown in FIG. 5) on the reference video clip 110-1. The visual features may capture an overall scene context and a mood of utterance (by human speaker(s)) in the reference video clip 110-1. Additionally, or alternatively, the visual features may capture facial landmarks, lip movements, facial expressions, head pose, eye gaze, or body gestures in the reference video clip 110-1. The visual feature extractor 506 may include hardware and / or software components, including but not limited to, deep learning models such as convolutional neural networks (CNNs) or vision transformers.
[0048] The system 102 may also execute an operation to generate a first speaker embedding associated with the human speaker based on the reference audio 110-2. A pre-trained model (i.e., a speaker encoder network consisting of Mel-Spectrogram Feature Extractor 502 and an Attention Encoder 504, as shown in FIG. 5) may generate the first speaker embedding from the reference audio 110-2. For instance, the pre-trained model may take an input audio waveform and generate a corresponding output in the form of a fixed length (e.g., 512-element) vector representing the first speaker embedding. The first speaker embedding may represent characteristics of the reference audio 110-2, such as pitch, energy, duration, voice, and emotion. FIG. 3 and FIG. 4 may provide further details related to the generation of the first speaker embedding.
[0049] The system 102 may also execute an operation to generate lip features associated with the human speaker based on the reference video clip 110-1. Specifically, reference video clip 110-1 may be cropped to focus on a lip region of interest (Rol). The lip Rol may be determined by identifying lip center coordinates for each frame in the reference video clip 110-1. This may enable accurate calculation of center coordinates for each frame and corresponding generation of lip features based on lip movements of the human speaker.
[0050] In some embodiments, the system 102 may employ, for the generation of lipfeatures, a masked prediction-based self-supervised training approach. This approach may leverage the AV-HuBERT (pre-trained) or another suitable model as the lip feature extractor (i.e. , a part of the lip encoder shown 104-3). The AV-HuBERT model may take both video (cropped lip region) and audio as inputs, utilizing a ResNet architecture for extracting video features and a linear layer for audio features.
[0051] The lip feature extraction process may involve first determining the lip Rol from the reference video clip 110-1. This lip Rol may be processed by the AV-HuBERT to generate lip features (i.e., the visual features). The ResNet architecture within the AV- HuBERT model may extract the visual features specifically related to lip movements, which may include information about lip shape, position, and dynamics over time. The extracted lip features may be combined with other visual features to create a comprehensive set of visual information. This may enable the system 102 to capture fine-grained details of lip movements along with broader facial and body gestures, potentially improving the accuracy of speech synthesis and lip synchronization in the generated audio.
[0052] The use of a self-supervised training approach may allow the lip-reading model to learn from large amounts of unlabeled data, potentially improving its ability to generalize across different speakers and languages. This may be particularly beneficial for the system's performance in multilingual and zero-shot scenarios. FIG. 3 and FIG. 4 may provide further details related to the generation of the lip features.
[0053] The system 102 may execute an operation to generate audio tokens based on application of the neural language model 104-4 on inputs that include the text features, the visual features, the first speaker embedding, and the lip features. In an embodiment, the inputs may also include the one-hot vector-based language ID. The neural language model 104-4 may process these inputs to generate the audio tokens. FIG. 5 may provide further details related to the generation of the audio tokens.
[0054] After generation of the audio tokens, the system 102 may execute an operationto generate audio in the first language or the second language based on application of the neural vocoder 104-5 on the audio tokens. The language of the generated audio may be selected by the user and may be same as the language of the text 110-3 (i.e. , the target language). In case of audio dubbing from the source language to the target language, the language of the generated audio may be in the target language. The generated audio may include spoken words in a voice of the human speaker (cloned from the reference audio 110-2) and the spoken words may correspond to words in the text 110-3. FIG. 3 and FIG. 4 may provide further details related to the generation of the audio.
[0055] The system 102 may execute an operation to train the neural language model 104-4 based on the generated audio and a loss function. The GT audio 302 (shown in FIG.3 and FIG. 4) may also be used in the training of the neural language model 104-4. The neural language model 104-4 may be trained for multiple iterations until a value of the loss function falls below a threshold loss. The threshold loss may correspond to a threshold time-difference between lip movements of the human speaker in the reference video and spoken words in the generated audio. During training, the neural language model 104-4 may learn to generate audio such that spoken words in the generated audio synchronize with lip movements of the human speaker. Upon satisfactory training, the neural language model 104-4 may be deployed for suitable dubbing applications. FIG. 3 and FIG. 4 may provide further details related to the training of the neural language model 104-4.
[0056] The system 102 may control duration of the generated audio for a given reference video to align with lip movements of the human speaker, even when spoken words differ or are in a different language. The system 102 may achieve improved lip synchronization and naturalness compared to traditional Text-to-Speech (TTS) methods for same-language non-parallel and cross-lingual scenarios. The system 102 may achieve global duration alignment through TTS followed by audio speed adjustment. Based on multilingual TTS with fine-tuning for duration controllability, the system 102 may functionin zero-shot scenarios or without training on multilingual videos.
[0057] FIG. 2 is a block diagram that illustrates an exemplary system of FIG. 1 , in accordance with an embodiment of the disclosure. FIG. 2 is explained in conjunction with elements from FIG. 1 . With reference to FIG. 2, there is shown a block diagram 200 of the system 102. The system 102 may include circuitry 202, a memory 204, a network interface 206, and an input / output (I / O) device 208. The I / O device 208 may include a display device 208-A. The memory 204 may include the dataset 110 and the duration-controlled TTS network 104. The network interface 206 may connect the system 102 with the server 106, via the communication network 112.
[0058] The circuitry 202 may include suitable logic, circuitry, and / or interfaces that may be configured to execute program instructions associated with different operations to be executed by the system 102. The operations may include, for instance, the dataset reception, text encoder application, visual feature extractor application, neural language model application, neural vocoder application, audio generation, neural language model training, and the like. The circuitry 202 may include one or more processing units, which may be implemented as a separate processor. In an embodiment, the one or more processing units may be implemented as an integrated processor or a cluster of processors that perform the functions of the one or more specialized processing units, collectively. The circuitry 202 may be implemented based on a number of processor technologies known in the art. Examples of implementations of the circuitry 202 may be an X86-based processor, a Graphics Processing Unit (GPU), a Reduced Instruction Set Computing (RISC) processor, an Application-Specific Integrated Circuit (ASIC) processor, a Complex Instruction Set Computing (CISC) processor, a microcontroller, a central processing unit (CPU), and / or a combination thereof.
[0059] The memory 204 may include suitable logic, circuitry, interfaces, and / or code that may be configured to store one or more instructions to be executed by the circuitry 202.The one or more instructions stored in the memory 204 may be configured to execute the different operations of the circuitry 202 (and / or the system 102). The memory 204 may be further configured to store the dataset 110, the text features, the visual features, the speaker embedding, the lip features, the audio tokens, and the audio (generated based on the audio tokens). The memory 204 may also be configured to store the duration- controlled TTS network 104. Examples of implementation of the memory 204 may include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Hard Disk Drive (HDD), a Solid-State Drive (SSD), a CPU cache, and / or a Secure Digital (SD) card.
[0060] The network interface 206 may include suitable logic, circuitry, interfaces, and / or code that may be configured to facilitate communication between the system 102 and the server 106, via the communication network 112. The network interface 206 may be implemented by use of various known technologies to support wired or wireless communication of the system 102 with the communication network 112. The network interface 206 may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, or a local buffer circuitry.
[0061] The network interface 206 may be configured to communicate via wireless communication with networks, such as the Internet, an Intranet, a wireless network, a cellular telephone network, a wireless local area network (LAN), or a metropolitan area network (MAN). The wireless communication may be configured to use one or more of a plurality of communication standards, protocols and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), wideband code division multiple access (W-CDMA), Long Term Evolution (LTE), 5thGeneration (5G) New Radio (NR), code division multiple access (CDMA), time division multiple access(TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (such as IEEE 802.11a, IEEE 802.11 b, IEEE 802.11 g or IEEE 802.11n), voice over Internet Protocol (VoIP), light fidelity (Li-Fi), Worldwide Interoperability for Microwave Access (Wi-MAX), a protocol for email, instant messaging, and a Short Message Service (SMS).
[0062] The I / O device 208 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive an input and provide an output based on the received input. For example, the I / O device 208 may receive the dataset 110. The I / O device 208 may be further configured to render the generated audio on the user interface, for instance, the user device 114.
[0063] Examples of the I / O device 208 may include, but are not limited to, a display (e.g., a touch screen), a keyboard, a mouse, a joystick, a microphone, or a speaker. Examples of the I / O device 208 may further include braille I / O devices, such as, braille keyboards and braille readers.
[0064] The display device 208-A may include suitable logic, circuitry, and interfaces that may be configured to display or render the reference video clip 110-1 , Mel-spectrogram associated with reference audio 110-2, and the text 110-3. In some embodiments, the display device 208-A may be a touch screen which may enable a user to provide a userinput via the display device 208-A. The display device 208-A may be realized through several known technologies such as, but not limited to, at least one of a Liquid Crystal Display (LCD) display, a Light Emitting Diode (LED) display, a plasma display, or an Organic LED (OLED) display technology, or other display devices. In accordance with an embodiment, the display device 208-A may refer to a display screen of a head mounted device (HMD), a smart-glass device, a see-through display, a projection-based display, an electro-chromic display, or a transparent display.
[0065] Various operations of the circuitry 202 are described further, for example, in FIG.3 and FIG. 4.
[0066] FIG. 3 is a diagram that illustrates a first exemplary architecture of the system of FIG. 1 , in accordance with an embodiment of the disclosure. FIG. 3 is explained in conjunction with elements from FIG. 1 and FIG. 2. With reference to FIG. 3, there is shown a first exemplary architecture 300 of the system 102 of FIG. 1 that includes an audio encoder 306, the audio-visual cloning network 104-2, the lip encoder 104-3, and a multimodal fusion Large Language Model (LLM) decoder 316.
[0067] The multimodal fusion LLM decoder 316 may be an exemplary implementation of the neural language model 104-4. The multimodal fusion LLM decoder 316 may include a plurality of LM layers such as LM layer 318-1 to LM layer 318-N. Exemplary operations for implementation of the multimodal duration-controlled TTS based automatic dubbing may be executed by any computing system, for example, by the system 102 of FIG. 1 or by the circuitry 202 of FIG. 2.
[0068] During operation, the circuitry 202 may receive the dataset 110, which may include the reference video clip 110-1 (interchangeably, referred to as reference scene video 110-1 , herein) of a human speaker, the reference audio 110-2 (interchangeably, referred to as reference speaker audio 110-2, herein) of the human speaker in a first language, and the text 110-3 that may include at least one of a translation of a transcript of the reference speaker audio 110-2 in a second language or a modification of the transcript.
[0069] The circuitry 202 may execute an operation to generate text features based on application of the text encoder 104-1 on the text 110-3. The text features may include linguistic and semantic information extracted from the text 110-3, such as word or phoneme embeddings. In an exemplary embodiment, the text encoder 104-1 may utilize Byte Pair Encoding (BPE) to tokenize the input text. BPE is a data compression technique that iteratively replaces the most frequent pair of bytes in a sequence with a single, unused byte. When applied to text, BPE may effectively handle out-of-vocabulary words bybreaking such words down into sub-word units. This approach may allow the text encoder 104-1 to process a wide range of languages and handle rare or unseen words more effectively. Additionally, the text encoder 104-1 may incorporate phoneme encoding as part of the feature extraction process. Phoneme encoding involves converting text into a sequence of phonemes, which are the smallest units of sound that distinguish one word from another in a language. By including phoneme-level information, the text encoder 104- 1 may capture important pronunciation details that may be crucial for accurate speech synthesis, especially in multilingual scenarios.
[0070] The combination of BPE tokenization and phoneme encoding may enable the text encoder 104-1 to generate a rich set of features that encompasses both lexical and phonetic information. These features may be particularly useful for the neural language model 104-4 in generation of audio tokens that accurately reflect the pronunciation and intonation patterns of the target language, while maintaining synchronization with the visual features extracted from the reference video clip 110-1.
[0071] In operation, the circuitry 202 may feed the reference scene video 110-1 and the reference speaker audio 110-2 to the audio-visual cloning network 104-2. The audio-visual cloning network 104-2 may execute an operation to generate visual features based on application of the visual feature extractor 506 on the reference scene video 110-1. Further, the audio-visual cloning network 104-2 may execute an operation to generate a first speaker embedding associated with the human speaker based on the reference speaker audio 110-2. The first speaker embedding may represent characteristics of the reference speaker audio 110-2, such as pitch, energy, duration, voice, and emotion. In an instance, a sample of the reference speaker audio 110-2 of the voice of the human speaker may be provided to the audio-visual cloning network 104-2. The provided reference speaker audio 110-2 may be any utterance spoken by the human speaker. To capture characteristics of the human speaker, the system 102 may employ a speaker encoder network, which mayextract the first speaker embedding.
[0072] The first speaker embedding may be later processed by a perceiver. The perceiver may include transformer or attention mechanism. In an embodiment, the perceiver may process the first speaker embedding to obtain a fixed-dimensional representation. Furthermore, to capture the overall scene and mood of the utterance, the visual features of the reference scene video 110-1 may be utilized. The visual features may be fed to the perceiver along with the first speaker embedding to capture overall speaker style, while also training the perceiver to obtain the fixed-dimensional representation. In an exemplary embodiment, the perceiver may firstly encode the first speaker embedding into a set of latent variables. Further, the encoded first speaker embedding may then be processed using a cross-attention mechanism. The crossattention mechanism may help in aggregating information from the first speaker embedding into a fixed-size representation, into a fixed-size representation. Further, fixed- size latent variables may be processed through several layers of self-attention and feedforward networks. These layers help in refining and transforming the latent variables into a more meaningful representation. Finally, the processed latent variables may be projected into a fixed-dimensional output representation.
[0073] In another operation, the circuitry 202 may crop lip-region video of the human speaker from the reference scene video 110-1 and may further apply the lip encoder 104- 3 on the lip-region video (interchangeably, referred to as reference video cropped lips 304, herein) to generate the lip features. The lip encoder 104-3 may include a lip feature extractor and a plurality of transposed-convolution layers coupled to an output of the lip feature extractor. In an instance, a pre-trained Audio-Visual Hidden Unit BERT (AV- HuBERT) model may be used as the lip encoder 104-3. The AV-HuBERT model may take both, i.e. , the reference video cropped lips 304 and the reference speaker audio 110-2 as inputs. Further, the AV-HuBERT model may utilize a Residual Networks (ResNet)architecture for extracting video features and a linear layer for audio features, i.e. , the first speaker embedding or second speaker embedding. Subsequently, the AV-HuBERT model may combine the video features and the audio features. The system 102 may utilize the features extracted by the AV-HuBERT model as the lip features. This approach may enable achieving alignment and controllability in the generated audio.
[0074] In another operation, the circuitry 202 may receive a ground truth (GT) audio 302, which may be fed into an audio encoder 306. The audio encoder 306 may convert original format of the received GT audio 302 into a processible format. For instance, the audio encoder 306 may be a multimodal saliency model (such as a Deep Audio-Visual Embedding (DAVE) encoder) that utilizes audio and visual information for predicting saliency in audio-visual data. The audio encoder 306 may map auditory information in the GT audio 302 into a feature space using 3D Convolutional Neural Networks (3D CNNs).
[0075] At 308, an operation is executed by the circuitry 202 to concatenate the text features, the visual features, the first speaker embedding, and the lip features into a first prompt. The concatenation process refers to combining multiple input features (here - text features, the visual features, the first speaker embedding, and the lip features) into a single vector or tensor before feeding it into the multimodal fusion LLM decoder 316. This technique is commonly used to enhance the representation power of the multimodal fusion LLM decoder 316 by incorporating different types of information.
[0076] At 310, the concatenated first prompt may be fed into the neural language model 104-4 for training the neural language model 104-4. This may result in an extended prompt length (modified prompt structure), and the entirety of the neural language model 104-4 may undergo fine-tuning based on the modified prompt structure.
[0077] The circuitry 202 may generate audio tokens based on application of the plurality of LM layers (LM layer 318-1 to LM layer 318-N) of the multimodal fusion LLM decoder 316 on the first prompt including the text features, the visual features, the first speakerembedding, and the lip features. In an instance, a one-hot vector-based language ID may be used to specify language of the text 110-3. The one-hot vector may be associated with the first prompt, which may be subsequently fed into the multimodal fusion LLM decoder 316 of the neural language model 104-4.
[0078] In another instance, the system 102 may be trained on multiple languages. During training, the neural language model 104-4 may learn to predict the next token autoregressively. In an exemplary embodiment, a pre-trained Discrete Variational AutoEncoder (DVAE) may be utilized to tokenize the reference speaker audio 110-2 for training of the neural language model 104-4. In a preferred embodiment, the GT audio 302 and the reference video cropped lips 304 may be taken into consideration only during the training of the neural language model 104-4.
[0079] During training, the neural language model 104-4 may learn to generate audio tokens, such that an audio 314 may be generated based on application of neural vocoder 104-5 on the audio tokens. In an embodiment, spoken words associated with the generated audio 314 may be in sync with the lip movements of the human speaker. The generated audio 314 may include spoken words in a voice of the human speaker and the spoken words may correspond to words in the text 110-3. In an instance, XTTS, a GPT-2 based TTS model (as the neural language model 104-4), may utilize DVAE encoder for generating audio tokens and High Fidelity - Generative Adversarial Network (HiFi-GAN) based neural vocoder 104-5 to generate the audio 314.
[0080] The neural language model 104-4 may be trained for a number of iterations until a value of the loss function is below a threshold loss. The threshold loss may correspond to a threshold time-difference between lip movements of the human speaker in the reference scene video 110-1 and spoken words in the generated audio 314. In an embodiment, the loss function may consider losses incurred during generation of the audio by the system 102. The losses may include, for example, a cross-entropy loss on audiotokens (ce_audio), a scaled text tokens loss (ce_text), and a scaled duration loss (duration_diff). As an example, the losses may include a sum of ce_audio with a times(ce_text) and p times the duration_diff, where a and p are constants ranging between 0 and 1 .
[0081] In another operation, the circuitry 202 may compute a duration loss 312, which may be one of the losses associated with the loss function, incurred during generation of the audio 314 by the system 102. For the computation of the duration loss 312, the circuitry 202 may firstly identify an index of an end token in a ground-truth audio and a predicted end token of the audio tokens associated with the generated audio 314. Further, the duration loss 312 may be computed based on a difference between the index of the end token in the ground-truth audio and the predicted end token of the audio tokens associated with the generated audio. The neural language model 104-4 may take into consideration the duration loss 312 while generating the audio 314, so that the generated audio 314 captures the naturalness and alignment to achieve isochrony.
[0082] At inference stage, the circuitry 202 may receive, via user interface of the user device 114, an input including a video clip of the human speaker, and an input audio associated with the video clip. Further, the circuitry 202 may receive the input text that may include at least one of a translation of a transcript of the input audio in the second language or a modification of the transcript. The circuitry 202 may then generate input text features based on the application of the text encoder 104-1 on the input text. The circuitry 202 may also generate input visual features based on application of the visual feature extractor 506 on the input video clip.
[0083] The circuitry 202 may further generate a second speaker embedding based on the input audio. The circuitry 202 may also generate input lip features associated with the human speaker based on the input video clip. Further, the circuitry 202 may generate input audio tokens based on application of the trained neural language model 104-4 on the inputtext features, the input visual features, the second speaker embedding, and the input lip features. Further, the circuitry 202 may generate an audio clip in the second language or the first language based on application of the neural vocoder 104-5 on the input audio tokens. The audio clip may be generated such that a time-difference between lip movements of the human speaker in the input video and spoken words of the audio clip is a minimum. FIG. 7 may provide further details related to the inference.
[0084] FIG. 4 is a diagram that illustrates a second exemplary processing architecture of the system of FIG. 1 , in accordance with an embodiment of the disclosure. FIG. 4 is explained in conjunction with elements from FIG. 1 , FIG. 2, and FIG. 3. With reference to FIG. 4, there is shown a second exemplary processing architecture 400 of the system 102 of FIG. 1 that illustrates exemplary operations for implementation of multimodal duration- controlled text to speech (TTS) based automatic dubbing, where the exemplary operations may be executed by any computing system, for example, by the system 102 of FIG. 1 or by the circuitry 202 of FIG. 2.
[0085] The multimodal fusion LLM decoder 316 may further include a plurality of crossattention layers including cross-attention layers 402-1 to 402-N. Cross-attention layer mechanism is a powerful mechanism that enhances the performance and capabilities of the multimodal fusion LLM decoder 316 in various natural language processing tasks.
[0086] In operation, the circuitry 202 may concatenate the text features, the visual features, and the first speaker embedding to obtain a second prompt. FIG. 3 provides further details related to the concatenation process. The circuitry 202 may further feed the lip features along with the second prompt as an input to the cross-attention layers 402-1 to 402-N. The cross-attention layers 402-1 to 402-N may directly attend to the lip features. Moreover, the cross-attention layers 402-1 to 402-N may facilitate information exchange between the second prompt, the lip features associated with the reference scene video 110-1 and internal representations of corresponding LLM, hence allowing the multimodalfusion LLM decoder 316 to adapt audio generation based on context of the reference scene video 110-1.
[0087] Further, the circuitry 202 may train the multimodal fusion LLM decoder 316 based on the second prompt and the fed lip features and may correspondingly update parameters of the cross-attention layers 402-1 to 402-N, for example, query, value, key, and number of the respective cross-attention layer. While the parameters of the crossattention layers 402-1 to 402-N are updated, parameters of layers of the neural language model 104-4 other than the cross-attention layers 402-1 to 402-N may be frozen. Training only the cross-attention layers 402-1 to 402-N instead of retraining the entire neural language model 104-4, significantly accelerates the training process, and hence efficiency of the process.
[0088] To ensure that duration of the generated audio aligns with the reference scene video 110-1 , the circuitry 202 may extend the visual features of the reference scene video 110-1 to match the expected duration. This simplifies training by enabling linear attention towards the visual features without requiring explicit duration specification. The circuitry 202 may carry out the training by applying a combination of transposed convolution and convolution to the visual features, thereby effectively learning the necessary expansion without manual intervention. An end token may then be concatenated to mark the sequence termination. During training, all other model parameters may remain fixed, while only the cross-attention layer and the transposed-convolution layers may be fine-tuned for language modeling task, i.e., for training the neural language model 104-4 in multiple languages.
[0089] In an exemplary embodiment, for a given reference scene video 110-1 of a human speaker speaking in a source language (Cs) that is required to be dubbed in a target language (Ct), the cropped lip-region (V_lip) from the reference scene video 110-1 may be retrieved. Lip features (interchangeably, referred to as lip-representation features,herein) (F_lip) may be extracted using lip-reading model (MJip), and may be further modified by training the cross-attention layers 402-1 to 402-N of the multimodal fusion LLM decoder 316. The text 110-3 in the source language (Cs) may be translated to the target language (Ct) using an off-the-shelf machine translation model. Reference speaker audio 110-2 of the human speaker (Sref) saying any arbitrary sentence, may be used to clone the voice of the human speaker. The translated text (Ct) and speaker embeddings (Sref) and language ID (Lt) may form a prompt for cross-attention layers 402-1 to 402-N of the multimodal fusion LLM decoder 316. The prompt may be fed along with the visual features to the cross-attention layers 402-1 to 402-N of the multimodal fusion LLM decoder 316, which may further generate the audio 314 (Sgen) that may align well with the lip movements of the human speaker in the reference video. The cross-attention layers 402- 1 to 402-N of the multimodal fusion LLM decoder 316 may generate the audio 314 (Sgen) conditioned on the lip features (FJip) and text (Ct).
[0090] FIG. 5 is a diagram that illustrates an exemplary architecture of audio-visual cloning network, in accordance with an embodiment of the disclosure. FIG. 5 is explained in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, and FIG. 4. With reference to FIG. 5, there is shown an exemplary architecture 500 of the audio-visual cloning network 104-2. The audio-visual cloning network 104-2 may take both the refence scene video 110-1 and the reference speaker audio 110-2 as inputs. The audio-visual cloning network 104-2 may consist of Mel-spectrogram feature extractor 502 followed by attention encoder 504, which may be configured to extract audio features.
[0091] The Mel-spectrogram feature extractor 502 may be used in audio signal processing to extract audio features from audio signals. Mel scale is a perceptual scale that approximates the way humans perceive different frequencies. By converting the frequencies to the Mel scale, a Mel-spectrogram may be created, which may better represent the human auditory perception.
[0092] In operation, the Mel-spectrogram feature extractor 502 may first acquire data from a wav file. Then, the Mel-spectrogram feature extractor 502 may utilize a Mel spectrogram API to calculate the Mel spectrogram. Further, Mel scale may be applied, which focuses more on low frequencies and suppresses the high ones, then puts all frequencies into Mel bins. Further, a Mel filter bank may be created, and the power spectrogram may be calculated. The Mel spectrogram may be generated from the power spectrogram by applying the Mel filter bank. The Mel spectrogram may then be used for feature extraction.
[0093] The attention encoder 504 may create shortcuts between the inputs (i. e. , the Mel- spectrogram and corresponding first speaker embedding. Further, the attention encoder 504 may compute attention scores between the first speaker embedding and the generated audio tokens using a learned alignment model. For each audio token that the multimodal fusion LLM decoder 316 may generate, the attention encoder 504 may calculate a score between the current audio token and the corresponding first speaker embedding. This allows the multimodal fusion LLM decoder 316 to focus on a different part of output of the attention encoder 504 for every step of the multimodal fusion LLM decoder’s own outputs.
[0094] The visual feature extractor 506 may use transformer-based pretrained visual feature extractor such as VideoClip. Both audio features (first speaker embedding) and visual features may be concatenated to form a corresponding prompt. Further, this prompt may be passed through the multimodal transformer 508. As an example, the multimodal transformer 508 may be a perceiver.
[0095] FIG. 6 is a diagram that illustrates an exemplary architecture of lip encoder, in accordance with an embodiment of the disclosure. FIG. 6 is explained in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, and FIG. 5. With reference to FIG. 6, there is shown an exemplary architecture 600 of a lip encoder (i.e. , the lip encoder 104-3 of thesystem 102). The lip encoder 104-3 may include audio-visual self-supervised (SSL) feature extractor 602, which may be followed by the up-sampling network 604.
[0096] In an embodiment, a transformer-based pre-trained Audio-Visual HuBERT (AV- HuBERT) model may be used as the audio-visual SSL feature extractor 602. The AV- HuBERT model may be used to exploit inconsistency between audio and visual modalities (of the audio features and visual features) for multi-modal video forgery detection. The model may also be used in combination with a multi-scale temporal convolutional neural network to capture a temporal correlation between corresponding audio and visual modalities.
[0097] The up-sampling network 604 may be configured to upsample the audio and visual modalities. Further, the up-sampling network 604 may include a multiresolution architecture that may jointly exploit high-resolution and large context information associated with the audio and visual modalities.
[0098] FIG. 7 is a diagram that illustrates an exemplary inference of the system of FIG.1 , in accordance with an embodiment of the disclosure. FIG. 6 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, FIG. 5, and FIG. 6. With reference to FIG. 7, there is shown a diagram illustrating an inference scenario 700, where the exemplary operations may be executed by any computing system, for example, by the system 102 of FIG. 1 or by the circuitry 202 of FIG. 2.
[0099] The system 102 may communicate with the user device 114 via the communication network 112. The user device 114 may include a user interface through which a user associated with the user device 114 may interact with the system 102. For instance, the user may send queries, feed commands and instructions, or provide reference audios and video clips for training and updating the system 102. The user device 114 may be fixed at a place or may be portable. Examples of the user device 114 may include, but are not limited to, a smartphone, a touchpad, a personal computer, or awearable display.
[0100] In operation, during inference, the system 102 may receive, via the user interface of the user device 114, an input including a video clip 704 of the human speaker and an input audio associated with the video clip 704.
[0101] Further, the system 102 may receive an input text 702. The input text 702 may include at least one of a translation of a transcript of the input audio in the second language or a modification of the transcript. In an instance, the input text 702 may be fed by the user at the user device 114. The user device 114 may then transmit the input text 702 to the system 102. Further, the system 102 may receive the input text 702 from the user device 114. In an embodiment, the system 102 may retrieve the input text 702 from the database 108 or from the memory 204. In another embodiment, the system 102 may extract the input text 702 from the input audio associated with the video clip 704. As shown, for example, the input text 702 may be in English language and may include the sentence - “But this is not a time coordinate.”
[0102] The system 102 may further generate input text features based on the application of the text encoder 104-1 on the input text 702. The input text features may include linguistic and semantic information extracted from the text 110-3, such as word or phoneme embeddings.
[0103] The system 102 may further generate input visual features based on application of the visual feature extractor 506 on the input video clip 704. The visual features may capture an overall scene context and a mood of utterance (by human speaker(s)) in the input video clip 704.
[0104] The system 102 may further generate a second speaker embedding based on the input audio. The second speaker embedding may represent characteristics of the input audio, such as pitch, energy, duration, voice, and emotion.
[0105] The system 102 may further generate lip features associated with the humanspeaker based on the input video clip 704. In an instance, the lip features may be generated through image masking technique, where input video clip 704 may be firstly subdivided into multiple frames. Further, for each frame of the multiple frames, mouth localization and frame normalization may be performed.
[0106] The system 102 may further generate input audio tokens based on application of the trained neural language model 104-4 on the input text features, the input visual features, the second speaker embedding, and the input lip features.
[0107] The system 102 may further generate an audio clip 706 in the second language or the first language based on application of the neural vocoder 104-5 on the input audio tokens. The audio clip 706 may be generated such that a time-difference between lip movements of the human speaker in the input video clip 704 and spoken words of the audio clip 706 is a minimum. As shown, for example, the generated audio clip 706 may be in Hindi language, and may include the sentence|”
[0108] In the given example, the input video clip 704 may be of 3.34 seconds. However, the generated audio clip 706 may be of 3.57 seconds. In such scenario, an updated video clip 708 of 3.57 seconds may be generated, so that the audio clip 706 may align well with the updated video clip 708. In an embodiment, starting video frames of the input video clip 704 may be concatenated at the end of the input video clip 704 for a specific time duration (for instance, here for 0.23 seconds) to generate the updated video clip 708 of 3.57 seconds. In another embodiment, terminating video frames of the input video clip 704 may be extended (for instance, here for 0.23 seconds) to generate the updated video clip 708 of 3.57 seconds.
[0109] FIG. 8 is a flowchart that illustrates operations of an exemplary method for multimodal duration-controlled text to speech (TTS) based automatic dubbing, in accordance with an embodiment of the disclosure. FIG. 8 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 4, FIG. 5, FIG. 6, and FIG. 7. With reference toFIG. 8, there is shown a flowchart 800. The flowchart 800 may include operations from 802 to 818 and may be implemented by the system 102 of FIG. 1 or by the circuitry 202 of FIG. 2. The flowchart 800 may start at 802 and proceed to 804.
[0110] At 804, a dataset may be received. The circuitry 202 may be configured to receive the dataset 110 including the reference video clip 110-1 of a human speaker, the reference audio 110-2 of the human speaker in a first language, and the text 110-3 that includes at least one of a translation of a transcript of the reference audio 110-2 in a second language or a modification of the transcript. Details related to the receipt of the dataset are further described, for example, in FIG. 3 and FIG. 4.
[0111] At 806, text features may be generated based on application of a text encoder on the text. The circuitry 202 may be configured to generate text features based on application of a text encoder 104-1 on the text 110-3. Details related to the application of the generation of the text features are further described, for example, in FIG. 3 and FIG. 4.
[0112] At 808, visual features may be generated based on application of the visual feature extractor on the reference video clip. The circuitry 202 may be configured to generate visual features based on application of the visual feature extractor 506 on the reference video clip 110-1. Details related to the generation of the visual features are further described, for example, in FIG. 3 and FIG. 4.
[0113] At 810, a first speaker embedding may be generated based on the reference audio. The circuitry 202 may be configured to generate a first speaker embedding associated with the human speaker based on the reference audio 110-2. The first speaker embedding may represent characteristics of the reference audio 110-2, such as pitch, energy, duration, voice, and emotion. Details related to the generation of the first speaker embedding are further described, for example, in FIG. 3 and FIG. 4.
[0114] At 812, lip features associated with the human speaker may be generated basedon the reference video clip. The circuitry 202 may be configured to generate lip features associated with the human speaker based on the reference video clip 110-1. The circuitry202 may crop the reference video clip 110-1 to a lip-region video of the human speaker. The circuitry 202 may further apply a lip encoder (for example, the lip encoder 104-3 of FIG. 1 ) on the lip-region video to generate the lip features. The lip encoder 104-3 may include a lip feature extractor and a plurality of transposed-convolution layers coupled to an output of the lip feature extractor. Details related to the generation of the lip features are further described, for example, in FIG. 3 and FIG. 4.
[0115] At 814, audio tokens may be generated based on application of the neural language model on the text features, the visual features, the first speaker embedding, and the lip features. The circuitry 202 may be configured to generate audio tokens based on application of the neural language model 104-4 on inputs that include the text features, the visual features, the first speaker embedding, and the lip features. Details related to the generation of the audio tokens are further described, for example, in FIG. 3 and FIG. 4.
[0116] At 816, an audio may be generated based on application of the neural vocoder the audio tokens. The circuitry 202 may be configured to generate an audio (say, the audio 314 of FIG. 3 and FIG. 4) in the first language or the second language based on application of the neural vocoder 104-5 on the audio tokens. The generated audio 314 may include spoken words in a voice of the human speaker and the spoken words may correspond to words in the text. The neural vocoder 104-5 is a signal processing device that may use feature representation to synthesize a voice waveform. Details related to the generation of the audio are further described, for example, in FIG. 3 and FIG. 4.
[0117] At 818, the neural language model 104-4 may be trained based on the generated audio and a loss function. The circuitry 202 may be configured to train the neural language model based on the generated audio 314 and a loss function. The loss function may include losses incurred during generation of the audio 314 by the system 102. The lossesmay include, but not be limited to, cross-entropy loss on audio tokens, scaled text tokens loss, and scaled duration loss. Further, the neural language model 104-4 may be trained for a number of iterations until a value of the loss function is below a threshold loss, where the threshold loss may correspond to a threshold time-difference between lip movements of the human speaker in the reference video and spoken words in the generated audio 314. Details related to the training of the neural language model are further described, for example, in FIG. 3 and FIG. 4.
[0118] Although the flowchart 800 is illustrated as discrete operations, such as, 802, 804, 806, 808, 810, 812, 814, 816, and 818, the disclosure is not so limited. Accordingly, in certain embodiments, such discrete operations may be further divided into additional operations, combined into fewer operations, or eliminated, depending on the implementation without detracting from the essence of the disclosed embodiments.
[0119] Various embodiments of the disclosure may provide a non-transitory computer- readable medium and / or storage medium having stored thereon, computer-executable instructions executable by a machine and / or a computer to operate a system (for example, the system 102 of FIG. 1 ). Such instructions may cause the system 102 to perform operations that may include receipt of a dataset (for example, the dataset 110 of FIG. 1 ) including: a reference video clip (for example, the reference video clip 110-1 of FIG. 1) of a human speaker, a reference audio (for example, the reference audio 110-2 of FIG. 1 ) of the human speaker in a first language, and a text (for example, the text 110-3 of FIG. 1 ) that includes at least one of a translation of a transcript of the reference audio 110-2 in a second language or a modification of the transcript. The operations may further include generation of text features based on application of a text encoder (for example, the text encoder 104-1 of FIG. 1 ) on the text 110-3. The operations may further include generation of visual features based on application of a visual feature extractor (for example, the visual feature extractor 506 of FIG. 5). The operations may further include generation of a firstspeaker embedding associated with the human speaker based on the reference audio110-2. The operations may further include generation of lip features associated with the human speaker based on the reference video clip 110-1. The operations may further include generation of audio tokens based on application of a neural language model (for example, the neural language model 104-4 of FIG. 1 ) on the text features, the visual features, the first speaker embedding, and the lip features. The operations may further include generation of an audio (say, the audio 314 of FIG. 3 and FIG. 4) in the first language or the second language based on application of a neural vocoder (for example, the neural vocoder 104-5 of FIG. 1 ) on the audio tokens. The operations may further include training of the neural language model 104-4 based on the generated audio 314 and a loss function.
[0120] Exemplary aspects of the disclosure may provide a system (such as, the system 102 of FIG. 1 ) that includes circuitry (such as, the circuitry 202 of FIG. 2). The circuitry 202 may be configured to receive a dataset (for example, the dataset 110 of FIG. 1 ) including: a reference video clip (for example, the reference video clip 110-1 of FIG. 1) of a human speaker, a reference audio (for example, the reference audio 110-2 of FIG. 1 ) of the human speaker in a first language, and a text (for example, the text 110-3 of FIG. 1) that includes at least one of a translation of a transcript of the reference audio 110-2 in a second language or a modification of the transcript. The circuitry 202 may be configured to generate text features based on application of a text encoder (for example, the text encoder 104-1 of FIG. 1 ) on the text 110-3. The circuitry 202 may be configured to generate visual features based on application of a visual feature extractor (for example, the visual feature extractor 506 of FIG. 5). The circuitry 202 may be configured to generate a first speaker embedding associated with the human speaker based on the reference audio 110-2. The circuitry 202 may be configured to generate lip features associated with the human speaker based on the reference video clip 110-1. The circuitry 202 may beconfigured to generate audio tokens based on application of a neural language model (for example, the neural language model 104-4 of FIG. 1 ) on the text features, the visual features, the first speaker embedding, and the lip features. The circuitry 202 may be configured to generate an audio (say, the audio 314 of FIG. 3 and FIG. 4) in the first language or the second language based on application of a neural vocoder (for example, the neural vocoder 104-5 of FIG. 1 ) on the audio tokens. The circuitry 202 may be configured to train the neural language model 104-4 based on the generated audio 314 and a loss function.
[0121] In an embodiment, the neural language model 104-4 may further be trained for a number of iterations until a value of the loss function is below a threshold loss. The threshold loss may correspond to a threshold time-difference between lip movements of the human speaker in the reference video and spoken words in the generated audio 314.
[0122] In an embodiment, the generated audio 314 may include spoken words in a voice of the human speaker and the spoken words correspond to words in the text.
[0123] In an embodiment, the circuitry 202 may further be configured to receive, via a user interface of a user device (say, the user device 114 of FIG. 1 ), an input comprising a video clip of the human speaker, and an input audio associated with the video clip. The circuitry 202 may further be configured to receive an input text that includes at least one of a translation of a transcript of the input audio in the second language or a modification of the transcript. The circuitry 202 may further be configured to generate input text features based on the application of the text encoder 104-1 on the input text. The circuitry 202 may further be configured to generate input visual features based on application of the visual feature extractor 506 on the input video clip. The circuitry 202 may further be configured to generate a second speaker embedding based on the input audio. The circuitry 202 may further be configured to generate input lip features associated with the human speaker based on the input video clip. The circuitry 202 may further be configured to generate inputaudio tokens based on application of the trained neural language model 104-4 on the input text features, the input visual features, the second speaker embedding, and the input lip features. The circuitry 202 may further be configured to generate an audio clip in the second language or the first language based on application of the neural vocoder 104-5 on the input audio tokens. The audio clip may be generated such that a time-difference between lip movements of the human speaker in the input video and spoken words of the audio clip is a minimum.
[0124] In an embodiment, the circuitry 202 may further be configured to concatenate the text features, the visual features, the first speaker embedding, and the lip features into a first prompt. The circuitry 202 may further be configured to train the neural language model 104-4 based on the first prompt.
[0125] In an embodiment, the neural language model 104-4 may include a plurality of cross-attention layers.
[0126] In an embodiment, the circuitry 202 may further be configured to concatenate the text features, the visual features, and the first speaker embedding to obtain a second prompt. The circuitry 202 may further be configured to feed the lip features as an input to the plurality of cross-attention layers of the neural language model. The circuitry 202 may further be configured to train, based on the second prompt and the fed lip features, the neural language model to update parameters of the plurality of cross-attention layers. Parameters of layers of the neural language model other than the plurality of crossattention layers may be frozen for a duration in which the neural language model 104-4 is trained.
[0127] In an embodiment, the circuitry 202 may further be configured to apply an attention encoder on the first speaker embedding to generate a voice feature. The circuitry 202 may further be configured to concatenate the voice feature and the visual features into a feature representation. The circuitry 202 may further be configured to generateoutput information based on application of a multimodal transformer-based network on the feature representation. The second prompt may be obtained based on the output information and the text features.
[0128] In an embodiment, the circuitry 202 may further be configured to crop the reference video clip to a lip-region video of the human speaker. The circuitry 202 may further be configured to apply a lip encoder on the lip-region video to generate the lip features. The lip encoder may include a lip feature extractor, and a plurality of transposed- convolution layers coupled to an output of the lip feature extractor.
[0129] In an embodiment, the circuitry 202 may further be configured to train the plurality of transposed-convolution layers based on the generated audio 314.
[0130] In an embodiment, the circuitry 202 may further be configured to compute a duration loss associated with the loss function based on a difference between an index of an end token in a ground-truth audio and a predicted end token of the audio tokens associated with the generated audio. The neural language model may be trained further based on the computed duration loss.
[0131] In an embodiment, the first language may be same as or different from the second language.
[0132] In an embodiment, the modification of the transcript may include: a modification of a language of at least a part of the transcript from the first language to the second language, a substitution of at least a part of the transcript with a first new text in the first language, a substitution of at least a part of the transcript with a second new text in the second language, an addition of the first new text to the transcript in the first language, an addition of the second new text to the transcript in the second language, and a removal of at least a part of the transcript.
[0133] The present disclosure may be realized in hardware, or a combination of hardware and software. The present disclosure may be realized in a centralized fashion,in at least one computer system, or in a distributed fashion, where different elements may be spread across several interconnected computer systems. A computer system or other apparatus adapted to carry out the methods described herein may be suited. A combination of hardware and software may be a general-purpose computer system with a computer program that, when loaded and executed, may control the computer system such that it carries out the methods described herein. The present disclosure may be realized in hardware that comprises a portion of an integrated circuit that also performs other functions.
[0134] The present disclosure may also be embedded in a computer program product, which comprises all the features that enable the implementation of the methods described herein, and which when loaded in a computer system is able to carry out these methods. Computer program, in the present context, means any expression, in any language, code or notation, of a set of instructions intended to cause a system with information processing capability to perform a particular function either directly, or after either or both of the following: a) conversion to another language, code or notation; b) reproduction in a different material form.
[0135] While the present disclosure is described with reference to certain embodiments, it will be understood by those skilled in the art that various changes may be made, and equivalents may be substituted without departure from the scope of the present disclosure. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the present disclosure without departure from its scope. Therefore, it is intended that the present disclosure is not limited to the embodiment disclosed, but that the present disclosure will include all embodiments that fall within the scope of the appended claims.
Claims
CLAIMSWhat is claimed is:1 . A system, comprising: circuitry configured to: receive a dataset comprising: a reference video clip of a human speaker; a reference audio of the human speaker in a first language; and a text that includes at least one of a translation of a transcript of the reference audio in a second language or a modification of the transcript; generate text features based on application of a text encoder on the text; generate visual features based on application of a visual feature extractor on the reference video clip; generate a first speaker embedding associated with the human speaker based on the reference audio; generate lip features associated with the human speaker based on the reference video clip; generate audio tokens based on application of a neural language model on the text features, the visual features, the first speaker embedding, and the lip features; generate an audio in the first language or the second language based on application of a neural vocoder on the audio tokens; and train the neural language model based on the generated audio and a loss function.
2. The system according to claim 1 , wherein the neural language model is trained for a number of iterations until a value of the loss function is below a threshold loss, andthe threshold loss corresponds to a threshold time-difference between lip movements of the human speaker in the reference video and spoken words in the generated audio.
3. The system according to claim 1 , wherein the generated audio includes spoken words in a voice of the human speaker and the spoken words correspond to words in the text.
4. The system according to claim 1 , wherein the circuitry is further configured to: receive, via a user interface of a user device, an input comprising a video clip of the human speaker, and an input audio associated with the video clip; receive an input text that includes at least one of a translation of a transcript of the input audio in the second language or a modification of the transcript; generate input text features based on the application of the text encoder on the input text; generate input visual features based on application of the visual feature extractor on the input video clip; generate a second speaker embedding based on the input audio; generate input lip features associated with the human speaker based on the input video clip; generate input audio tokens based on application of the trained neural language model on the input text features, the input visual features, the second speaker embedding, and the input lip features; and generate an audio clip in the second language or the first language based on application of the neural vocoder on the input audio tokens,wherein the audio clip is generated such that a time-difference between lip movements of the human speaker in the input video and spoken words of the audio clip is a minimum.
5. The system according to claim 1 , wherein the circuitry is further configured to: concatenate the text features, the visual features, the first speaker embedding, and the lip features into a first prompt; and train the neural language model based on the first prompt.
6. The system according to claim 1 , wherein the neural language model includes a plurality of cross-attention layers.
7. The system according to claim 6, wherein the circuitry is further configured to: concatenate the text features, the visual features, and the first speaker embedding to obtain a second prompt; feed the lip features as an input to the plurality of cross-attention layers of the neural language model; and train, based on the second prompt and the fed lip features, the neural language model to update parameters of the plurality of cross-attention layers, wherein parameters of layers of the neural language model other than the plurality of cross-attention layers are frozen for a duration in which the neural language model is trained.
8. The system according to claim 7, wherein the circuitry is further configured to: apply an attention encoder on the first speaker embedding to generate a voice feature;concatenate the voice feature and the visual features into a feature representation; and generate output information based on application of a multimodal transformerbased network on the feature representation, wherein the second prompt is obtained based on the output information and the text features.
9. The system according to claim 1 , wherein the circuitry is further configured to: crop the reference video clip to a lip-region video of the human speaker; and apply a lip encoder on the lip-region video to generate the lip features, wherein the lip encoder includes a lip feature extractor and a plurality of transposed-convolution layers coupled to an output of the lip feature extractor.
10. The system according to claim 9, wherein the circuitry is further configured to train the plurality of transposed-convolution layers based on the generated audio.11 . The system according to claim 1 , wherein the circuitry is further configured to compute a duration loss associated with the loss function based on a difference between an index of an end token in a ground-truth audio and a predicted end token of the audio tokens associated with the generated audio, wherein the neural language model is trained further based on the computed duration loss.
12. The system according to claim 1 , wherein the first language is same as or different from the second language.
13. The system according to claim 1 , wherein the modification of the transcript includes: a modification of a language of at least a part of the transcript from the first language to the second language, a substitution of at least a part of the transcript with a first new text in the first language, a substitution of at least a part of the transcript with a second new text in the second language, an addition of the first new text to the transcript in the first language, an addition of the second new text to the transcript in the second language, and a removal of at least a part of the transcript.
14. A method, comprising: in a system: receiving a dataset comprising: a reference video clip of a human speaker; a reference audio of the human speaker in a first language; and a text that includes at least one of a translation of a transcript of the reference audio in a second language or a modification of the transcript; generating text features based on application of a text encoder on the text; generating visual features based on application of a visual feature extractor on the reference video clip; generating a first speaker embedding associated with the human speaker based on the reference audio; generating lip features associated with the human speaker based on the reference video clip;generating audio tokens based on application of a neural language model on the text features, the visual features, the first speaker embedding, and the lip features; generating an audio in the first language or the second language based on application of a neural vocoder on the audio tokens; and training the neural language model based on the generated audio and a loss function.
15. The method according to claim 14, further comprising training the neural language model for a number of iterations until a value of the loss function is below a threshold loss, wherein the threshold loss corresponds to a threshold time-difference between lip movements of the human speaker in the reference video and spoken words in the generated audio.
16. The method according to claim 14, further comprising: receiving, via a user interface of a user device, an input comprising a video clip of the human speaker and an input audio associated with the video clip; receiving an input text that includes at least one of a translation of a transcript of the input audio in the second language or a modification of the transcript; generating input text features based on the application of the text encoder on the input text; generating input visual features based on application of the visual feature extractor on the input video clip; generating a second speaker embedding based on the input audio;generating input lip features associated with the human speaker based on the input video clip; generating input audio tokens based on application of the trained neural language model on the input text features, the input visual features, the second speaker embedding, and the input lip features; and generating an audio clip in the second language or the first language based on application of the neural vocoder on the input audio tokens, wherein the audio clip is generated such that a time-difference between lip movements of the human speaker in the input video and spoken words of the audio clip is a minimum.
17. The method according to claim 14, further comprising: concatenating the text features, the visual features, the first speaker embedding, and the lip features into a first prompt; and training the neural language model based on the first prompt.
18. The method according to claim 14, further comprising: concatenating the text features, the visual features, and the first speaker embedding to obtain a second prompt; feeding the lip features as an input to a plurality of cross-attention layers of the neural language model; and training, based on the second prompt and the fed lip features, the neural language model to update parameters of the plurality of cross-attention layers, wherein parameters of layers of the neural language model other than the plurality of cross-attention layers are frozen for a duration in which the neural language model is trained.
19. The method according to claim 18, further comprising: applying an attention encoder on the first speaker embedding to generate a voice feature; concatenating the voice feature and the visual features into a feature representation; and generating output information based on application of a multimodal transformerbased network on the feature representation, wherein the second prompt is obtained based on the output information and the text features.
20. A non-transitory computer-readable medium having stored thereon, computerexecutable instructions that when executed by a system, causes the system to execute operations, the operations comprising: receiving a dataset comprising: a reference video clip of a human speaker; a reference audio of the human speaker in a first language; a text that includes at least one of a translation of a transcript of the reference audio in a second language or a modification of the transcript; generating text features based on application of a text encoder on the text; generating visual features based on application of a visual feature extractor on the reference video clip; generating a first speaker embedding associated with the human speaker based on the reference audio; generating lip features associated with the human speaker based on the reference video clip;generating audio tokens based on application of a neural language model on the text features, the visual features, the first speaker embedding, and the lip features; generating an audio in the first language or the second language based on application of a neural vocoder on the audio tokens; and training the neural language model based on the generated audio and a loss function.