System and method for processing a video clip

The system processes a video clip by replacing the original audio with a translated audio track, synchronizing it with mouth movements, and synthesizing the audio to resemble the original speaker's voice, addressing the challenge of natural synchronization in existing technologies.

WO2025133682A1PCT designated stage expired Publication Date: 2025-06-26VORAVUTHIKUNCHAI WINN
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2023/063079
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing technologies struggle to synchronize mouth and lip movements naturally with a new audio track in a desired language, often resulting in a mechanical sound and visible mismatch.

Method used

A system and method that processes a video clip by replacing the original audio track with a translated audio track, synchronizing it with mouth movements, and synthesizing the audio to resemble the original speaker's voice using an image processor, audio processor, and rendering engine.

Benefits of technology

The solution effectively replaces the original audio with a naturally sounding translated audio, synchronized with the speaker's mouth movements, providing a seamless and authentic viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2023063079_26062025_PF_FP_ABST
    Figure IB2023063079_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The invention provides a system (100) for processing a video clip. The system (100) may comprise an image processor (110), audio processor (120) and rendering engine (130) in data communication with each of the image processor (110) and audio processor (120), either via wired or wireless communication. Preferably, the image 5 processor (110) and audio processor (120) may respectively be configured to receive and analyse the image and audio portion of the video clip. Based on the outputs from the image processor (110) and audio processor (120), the rendering engine (130) may be configured to further process the video clip. A method may also be provided herein to perform the same.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SYSTEM AND METHOD FOR PROCESSING A VIDEO CLIP

[0002] FIELD OF THE INVENTION

[0003] The invention relates to a system and method for processing a video clip.

[0004] BACKGROUND OF THE INVENTION

[0005] Sometimes, the original audio track with the original text is being replaced by another audio track with the desired language, after the video has already been captured and recorded. In this case, there is a mismatch between the audio track and the movements of the mouth, which can be disturbing and detected easily by the video viewers.

[0006] Much work has recently been focused on overcoming the above disadvantages. For example, U.S. Patent No. US 8,655,152 B2 discloses a process for presenting live action foreign language feature films in a native language by replacing the physical mouth positions of the original actors to match a newly recorded audio track in a different language with the original and / or replacement actors keeping the essence of the original dialect, while achieving the illusion that the content was originally filmed in the new voice over language.

[0007] U.S. Patent No. US 10,339,973 B2 discloses a method and system for converting a first language of a soundtrack of a person speaking in a video to a second language. In particular, the method defines an outline of a shape of a mouth opening of the person speaking a syllable of a word of the first language in the video at a given start time by selecting a predetermined number of points along a border of the mouth opening defined by the lips. A length of the spoken syllable is also measured and one or more adjacent syllables are combined to create a word. The word is translated into synonym words in the second language, the best fit synonym word is selected that most closely matches the mouth shape of the first language word, and a mouth shape adjustment script is applied to fine tune the mouth shape of the best fit synonym word.

[0008] Although the existing technologies can ensure that the mouth and lip movements are synchronized with the new audio track in the desired language, unfortunately, the new audio track does not sound natural to the viewers, as it is synthetically generated that can sound mechanical. It is also obvious to the video viewers that the new audio track is delivered by another person who has a different voice and prosodic characteristics than the original speaker.

[0009] Accordingly, there remains a need for another system and method that can resolve the shortcomings mentioned in the foregoing.

[0010] SUMMARY OF THE INVENTION

[0011] One of the objects of the invention is to provide a system and method for processing a video clip so that the original audio track therein (which is in the first language) can be replaced by another audio track in the desired language (which is different from the first language).

[0012] Another object of the invention is to provide a system and method for synchronizing the audio track in the desired language with the movements of the mouth and lips in the video clip.

[0013] Yet another object of the invention is to provide a system and method for synthesizing the audio track in the desired language with a voice resembling the original speaker’s voice.

[0014] At least one of the preceding objects is met, in whole or in part, by the invention, in which one of the embodiments of the invention describes a system for processing a video clip comprising an image processor comprising a face detection module for analysing an image in each frame of the video clip to locate a face region of a speaker; an audio processor comprising a speech diarization module for analysing an audio portion in each frame of the video clip to extract original speeches belonging to the speaker and then mel- spectrogram contained within the extracted original speeches, the original speeches being in a first language; a speech recognition module for interpreting and converting the extracted original speeches to texts; a language translation module for instantiating the translation of the texts from the first language into a second language, the first language and second language being different languages; and a text-to- speech engine for compositing the translated texts with the extracted mel- spectrogram received from the speech diarization module and converting the composited texts to translated speeches; and a rendering engine in data communication with the image processor and audio processor, wherein for each frame of the video clip, the rendering engine is configured to receive the translated speeches from the text-to-speech engine; replace the original speeches of the speaker in the video frame with the translated speeches; create visemes that are synchronized to the translated speeches; and modify the face region of the speaker based on the created visemes so that the face region movements of the speaker match the movements of the second language.

[0015] The face region may preferably be a mouth region.

[0016] The system may further comprise a recording device for capturing and recording the video clip.

[0017] The system may further comprise a face identification module for determining the speaker’s identity based on the face region located by the face detection module.

[0018] The system may further comprise a speaker identification module for determining the speaker’s identity based on the mel-spectrogram extracted by the speech diarization module.

[0019] If the face identification module and speaker identification module co-exist, the face identification module or speaker identification module may be further configured to verify whether the speaker’s identity determined by the face identification module and the speaker’s identity determined by the speaker identification module are identical.

[0020] The speech recognition module may also be further configured to determine the first language of the original speeches before converting the original speeches to texts. A further embodiment of the invention is a method for processing a video clip, the method comprising the steps of analysing an image in each frame of the video clip by using a face detection module to locate a face region of a speaker; processing an audio portion in each frame of the video clip, wherein the step of processing the audio portion in each frame of the video clip comprises analysing the audio portion in each frame of the video clip by using a speech diarization module to extract original speeches belonging to the speaker and then mel-spectrogram contained within the extracted original speeches, the original speeches being in a first language; interpreting and converting the extracted original speeches to texts by using a speech recognition module; instantiating the translation of the texts from the first language into a second language by using a language translation module, the first language and second language being different languages; and compositing the translated texts with the extracted mel-spectrogram from the speech diarization module by using a text-to- speech engine; and converting the composited texts to translated speeches by using the text-to-speech engine; and transmitting the translated speeches from the text-to- speech engine to a rendering engine which can replace the original speeches of the speaker in the video frame with the translated speeches, create visemes that are synchronized to the translated speeches, and modify the face region of the speaker based on the created visemes so that the face region movements of the speaker match the movements of the second language.

[0021] The face region may preferably be a mouth region.

[0022] The method may further comprise the step of capturing and recording the video clip by using a recording device.

[0023] The method may further comprise the step of determining the speaker’s identity by using a face recognition module, based on the face region located by the face detection module.

[0024] The method may further comprise the step of determining the speaker’s identity by using a speaker recognition module, based on the mel-spectrogram extracted by the speech diarization module.

[0025] The method may further comprise the step of verifying whether the speaker’s identity determined by the face identification module and the speaker’s identity determined by the speaker identification module are identical. The verifying step may be performed by using either the face recognition module or speaker recognition module.

[0026] The method may further comprise the step of determining the first language of the original speeches by using the speech recognition module, before the converting step.

[0027] One skilled in the art will readily appreciate that the invention is well adapted to carry out the aspects and obtain the ends and advantages mentioned, as well as those inherent therein. The embodiments described herein are not intended as limitations on the scope of the invention.

[0028] BRIEF DESCRIPTION OF DRAWINGS

[0029] For the purpose of facilitating an understanding of the invention, there is illustrated in the accompanying drawing the preferred embodiments from an inspection of which when considered in connection with the following description, the invention, its construction and operation and many of its advantages would be readily understood and appreciated.

[0030] Figure 1 is a system for processing a video clip.

[0031] Figure 2 shows an overall architecture of a method for processing a video clip, where the method comprises a training phase followed by an interference phase.

[0032] Figure 3 is a general flow chart showing the steps performed during the interference phase.

[0033] Figure 4 is a flow chart showing the initial stage of the video preprocessing process. Figure 5 is a flow chart showing the second stage of the video preprocessing process.

[0034] Figure 6 is a flow chart showing the process S54 for training a lip-syncing model by using the rendering engine.

[0035] Figure 7 is a flow chart showing the process S56 comprising the steps for processing the image portion of the video clip.

[0036] Figure 8 is a flow chart showing the process S58 comprising the steps for further processing the video clip.

[0037] Figure 9 is a flow chart showing the steps for performing the face detection procedure.

[0038] Figure 10 is a flow chart showing Step S5810 in more detail.

[0039] Figure 11 shows an overall architecture of the Whisper model with a freezing layer adopted by the speech recognition module.

[0040] Figure 12 shows an overall architecture of a non-autoregressive text-to-speech model adopted by the text-to-speech engine.

[0041] DETAILED DESCRIPTION OF THE INVENTION

[0042] Hereinafter, the invention shall be described according to the preferred embodiments of the invention and by referring to the accompanying description and drawings. It is, however, to be understood that limiting the description to the preferred embodiments of the invention and to the drawings is merely to facilitate discussion of the invention and that it is envisioned that those skilled in the art may devise various modifications without departing from the scope of the appended claims.

[0043] The terms “video clip”, “video”, “motion picture film” and “movie” are used interchangeably in the specification to refer to an audio-visual recording having a series of visual images (also called frames) and their associated sound records.

[0044] The terms “speaker” and “actor” are used interchangeably in the specification to refer to a person speaking in a video clip.

[0045] The terms “first language”, “source language” and “original language” are used interchangeably in the specification to refer to a language spoken by a person in a video clip before the video clip is being processed.

[0046] The terms “second language”, “target language” and “desired language” are used interchangeably in the specification to refer to a language different from the original language in the video clip.

[0047] The term “mouth area” and “mouth region” are used interchangeably in the specification to refer to an area covering the mouth of a speaker. The mouth area shall encompass mouth, teeth, lips, tongue, chin and jaw.

[0048] The invention provides a system and method for processing a video clip. In particular, the invention enables the video clip to be processed by replacing the original audio track therein (which is in the first language) with another audio track in the desired language, in which the audio track in the desired language is synthesized with a voice resembling the voice of the original speaker. Additionally, the invention also enables that the audio track in the desired language be synchronized with the frames in the video clip or in more particular the movements of the mouth and lips in the video clip.

[0049] Figure 1 shows a system (100) for receiving and processing a video clip according to one embodiment of the invention. The video clip may be captured and recorded by using a recording device or video capturing device. In the video clip, there may be one or more persons speaking in the first language. For ease of description, the following embodiments will be described by assuming that there is a single speaker in the video clip, but it should be appreciated that the number of speakers present in the video clip shall not be limited thereto or thereby.

[0050] As shown in Figure 1, the system (100) may comprise an image processor (110), an audio processor (120) and a rendering engine (130) in data communication with one another, either via wired or wireless communication, to process the video clip (more particularly the image and audio portions of the video clip).

[0051] The image processor (110) may comprise a face detection module (111) for receiving the video clip (more particularly the image portion of the video clip) and analysing the image in each frame of the video clip to locate the speaker and the speaker’s face or face region. The “face region” located by the face detection module (111) may be a cheek region, a temple region, side of nose, forehead, etc., but the face region located by the face detection module (111) is preferably a mouth region encompassing the mouth, teeth, lips, tongue, chin and jaw.

[0052] To locate the speaker’s face or face region in the image in each video frame, the face detection module (111) may be configured to perform operations shown in Figure 9, for instance and without limitation:

[0053] (a) To perform “person segmentation” on each video frame by deploying a machine learning model such as BodyPix model, in which the image in each video frame may be segmented into pixels that are part of the speaker and those that are not, in order to crop the speaker from the image;

[0054] (b) To track and crop the speaker’s face or face region from each video frame, while or after performing “person segmentation”. The speaker’s face or face region may be tracked and cropped from the image in each frame by using a machine learning model such as MediaPipe Holistic model; and

[0055] (c) To perform skin modelling on the cropped face or face region and subsequently locate facial features in the cropped face or face region. The facial features may encompass left eye, right eye, nose tip, mouth, left eye tragion, right eye tragion, etc.

[0056] The outputs from the face detection module (111) such as the speaker’s face or face region, are useful inputs, for instance and without limitation, for the speaker identity verification and validation purposes, as well as for creating deepfakes using a machine learning model (such as the DeepFaceLab model) to replace or overlay the face of a speaker speaking in a first language with the face of the same speaker speaking in a second language in a video frame - all of which will be described in detail in the following.

[0057] The image processor (110) may also comprise a face identification module (113). The face identification module (113) may be configured to be in data communication with the face detection module (111) for receiving the located face or face region and then determining the speaker’s identity based on the face or face region as received. For instance, the face identification module (113) may determine and verify the speaker’s identity by using a machine training model such as the MTCNN (multitask cascaded convolutional network) model or timesler / facenet-pytorch model. By deploying the machine learning model, the face identification module (113) may proceed to compare the located face or face region to a plurality of faces or face regions of previously identified speakers (which may be stored in a database (150)). Alternatively, the face identification module (113) may be configured to send a request inviting a user of the system (100) to provide an input on the speaker’s identity based on the located face or face region. The user input may be provided through a user application installed on a computing device which may be in data communication with the system (100).

[0058] In some optional embodiments, the face identification module (113) may be further configured to tag the face or face region to the speaker’s identity once determined and verified successfully.

[0059] In the event where more than one person is present in the video clip, it may then be essential to configure the face detection module (111) and face identification module (113) to perform the operations in the following sequence.

[0060] First, the face detection module (111) may segment the image in each video frame, in order to locate each person present therein. The face detection module (111) may also track and locate the face or face region of each person present. As mentioned in the foregoing, the “face region” herein is preferably a mouth region encompassing the mouth, teeth, lips, tongue, chin and jaw.

[0061] Alternatively, the face detection module (111) may analyse the movement of each located face or face region in the video frame, to determine whether the mouth region movements correspond to the movement characteristic of oral communication. If it is determined that the mouth region movements correspond to the movement characteristic of oral communication, then it is likely that that person is the speaker.

[0062] After the speaker has been detected among the persons present in the video clip, the located face or face region may be transmitted from the face detection module (111) to the face identification module (113) for determining the speaker’s identity.

[0063] In some embodiments, the face or face region may be tagged by the face identification module (113) to the speaker’s identity once determined successfully.

[0064] Audio Processor (120)

[0065] The audio processor (120) may preferably comprise a speech diarization module (121) for receiving the video clip (or more particularly the audio portion of the video clip). Upon receipt, the speech diarization module (121) may analyse the audio portion in each video frame by deploying a deep learning model (such as the pyannote / speaker-diarization model) to achieve at least the following objects:

[0066] • To extract original speeches in the video frame, in which the “original speeches” herein are associated with dialogues or lines spoken by the speaker or more than one speaker in the video frame in the first language; • To extract the mel-spectrogram contained in the extracted original speeches;

[0067] • To identify the speaker in the video frame (if a single speaker is present therein); and

[0068] • To identify the number of speakers in the video frame (if more than one speaker is present therein).

[0069] In some embodiments, the extracted mel-spectrogram may comprise prosodic features (which may comprise one or more of pitch, pitch range, intonation attitude, loudness, speaking rate and phone duration).

[0070] In some embodiments, the extracted mel-spectrogram may comprise one or both of prosodic features and quality features. The prosodic features may comprise one or more of pitch, pitch range, intonation attitude, loudness, speaking rate and phone duration, whereas the quality features may comprise one or more of phonation type, articulatory manner, voice timbre, spectral tilt, difference of amplitude between harmonics and formants, formants bandwidth, jitter and harmonic-to-noise ratio.

[0071] The audio processor (120) may also comprise a speaker identification module (123) in data communication with the speech diarization module (121). In some embodiments, the speaker identification module (123) may be configured to receive the extracted mel-spectrogram from the speech diarization module (121). Subsequently, the speaker identification module (123) may be configured to determine and verify the speaker’s identity by using a machine training model, for instance a pretrained ECAPA-TDNN (Emphasized Channel Attention, Propagation and Aggregation-Time Delay Neural Network) model using SpeechBrain or the speechbrain / spkrec-ecapa-voxcelal model. By deploying the machine learning model, the speaker identification module (123) may be configured to perform the following operations:

[0072] (a) To receive several sets of mel-spectrogram extracted from several frames within the same video clip from the speech diarization module (121), and subsequently compare the several sets of the extracted mel-spectrogram, in order to determine and verify whether they belong to the same speaker; and (b) To determine and verify the speaker’s identity by comparing the extracted mel- spectrogram to a plurality of mel-spectrograms of previously identified speakers that are stored in a database (150).

[0073] In some alternative embodiments, the speaker identification module (123) may be configured to send a request inviting a user of the system (100) to provide an input on the speaker’s identity based on the extracted mel-spectrogram. The user input may be provided through a user application installed on a computing device which may be in data communication with the system (100).

[0074] The determination results provided by both of the face identification module (113) and speaker identification module (123) can advantageously be used for the speaker identity verification and validation purposes. For instance, the determination results may be compared by the face identification module (113) or alternatively the speaker identification module (123) to verify whether the speaker’s identity determined by the face identification module (113) and the speaker’s identity determined by the speaker identification module (123) are identical. Based on the comparison, it makes possible to verify whether or not the speaker identified by the face identification module (113) and the speaker identified by the speaker identification module (123) are of the same person. Accordingly, it can substantially reduce or eliminate misidentification of the speaker in the video clip.

[0075] Optionally, the speaker identification module (123) may be further configured to tag the original speeches and the corresponding mel-spectrogram to the speaker’s identity once determined successfully. Alternatively, the tagging may be performed by the speaker identification module (123) after the speaker’s identity has been verified and validated. For instance, after it has been verified that the speaker identified by the face identification module (113) and the speaker identified by the speaker identification module (123) are of the same person, the tagging may be performed by the speaker identification module (123). The audio processor (120) may also comprise a speech recognition module (125). The speech recognition module (125) may in some embodiments be configured to receive the video clip or particularly the audio portion of the video clip. In such circumstance, the speech recognition module (125) may deploy a machine learning model (such as Whisper model with freezing layer) for extracting original speeches from the audio portion before proceeding to interpret and convert the original speeches to texts in the first language.

[0076] Alternatively, the speech recognition module (125) may be configured to receive the extracted original speeches from the speech diarization module (121) (which is in data communication with the speech recognition module (125)). Upon receipt, the speech recognition module may deploy a machine learning model (such as Whisper model with freezing layer) to interpret and convert the original speeches to texts in the first language.

[0077] Steps for interpreting and converting the original speeches to texts using the Whisper model with freezing layer will be described in more detail in Example 1 below.

[0078] Additionally, in the event where the language of the original speeches (i.e., the first language) is unknown, the speech recognition module (125) may also be configured to identify and determine the first language before proceeding to interpret and convert the original speeches to texts. For instance, the first language may be English. Once the first language is determined, the speech recognition module (125) may proceed to interpret and convert the original speeches to texts in the first language.

[0079] The audio processor (120) may also comprise a language translation module (127) in data communication with the speech recognition module (125) for receiving the texts (i.e., in the first language) from the speech recognition module (125) and instantiating the translation of the texts from the first language into a second language. It should be appreciated that the first language and second language are different languages, where the “second language” may be a default language set by the system (100) or the user of the system (100). Alternatively, the second language may be the desired language selected by the user of the system (100) among a plurality of languages. It should also be appreciated that the language translation module (127) may translate the texts from the first language into a second language by using any suitable multilingual machine translation model that is available in the market.

[0080] The audio processor (120) may also comprise a text-to-speech engine (129) that is specifically designed and built by the inventors of this invention. Preferably, the text- to-speech engine (129) may be in data communication with the language translation module (127) for receiving the translated texts in the second language. The text-to- speech engine (129) may also be in data communication with the speech diarization module (121) or alternatively the speaker identification module (123) for receiving the mel-spectrogram extracted from the original speeches.

[0081] Upon receipt, the text-to-speech engine (129) may train the translated texts and mel- spectrogram using any text-to-speech (TTS) model (such as FastSpeech or FastSpeech 2 model) or a TTS model with a neural vocoder (such as FastSpeech 2 model with UnivNet). More preferably, the text-to-speech engine (129) may train the translated texts and mel-spectrogram using a TTS model with a neural vocoder (such as FastSpeech 2 model with UnivNet, as described in Example 2). In particular, the text- to-speech engine (129) may composite the translated texts with the mel-spectrogram and then convert the composited texts to translated speeches. It should be appreciated that the translated texts and translated speeches are both in the second language. Most importantly, when the translated speeches are deployed and played, it sounds like the speeches are delivered by the speaker himself in the second language.

[0082] As mentioned in the foregoing, it is possible that more than one person is speaking in the video clip. In such circumstance, it may be necessary to configure the speech diarization module (121) to analyse the audio portion in each frame of the video clip to extract original speeches and then the mel-spectrogram contained within each of the extracted original speeches. Each of the original speeches may be associated with one mel-spectrogram. Each of the mel-spectrograms may also be transmitted from the speech diarization module (121) to the speaker identification module (123) so that the speaker’s identity can be determined and verified for each mel-spectrogram received. The original speeches and the corresponding mel-spectrograms may also be tagged by the speaker identification module (123) to the speaker’s identity once determined and verified successfully.

[0083] The rendering engine (130) may be configured to be in data communication with the image processor (110) and audio processor (120) respectively. The rendering engine (130) may also be in data communication with a database (150) where historical and new data are stored. The data communication between the rendering engine (130) and the database (150) may be advantageous in certain embodiments, as it enables the rendering engine (130) to be configured in such a manner that it retrieves the data required for performing a requested step or operation from a single location instead of multiple locations.

[0084] To establish the data communication with the respective components in the system, the rendering engine (130) may comprise one or more algorithms. The algorithms herein may be deep learning models or machine learning models. Preferably, each learning model may be trained by the rendering engine (130) using one or more video datasets (such as curated video datasets, historical video datasets or any combination thereof) to simulate the operations of the system (100) or a target component in the system (100). Each learning model may also comprise one or more connections with the components present in the system (100), thereby allowing a plurality of operations to be performed by respective target components upon deploying the trained model. These operations may include but are not limited to the following:

[0085] (a) The rendering engine (130) may retrieve a plurality of data, e.g., data associated with the speaker’s identity, original speeches (in the first language), translated speeches (in the second language), etc., from the respective components in the system (100), before providing instructions to a target component (e.g., the face detection module (111)) to perform specific tasks or steps. (b) The rendering engine (130) may replace the original speeches of the speaker in each frame of the video clip with the translated speeches in the second language.

[0086] (c) The rendering engine (130) may create visemes that are synchronized to the translated speeches and then modify the face or face region of the speaker based on the created visemes, so that the speaker’s mouth region movements match the movements of the translated speeches in the second language. To achieve these objects, the following operations may be performed by the rendering engine (130):

[0087] • To create the visemes, the rendering engine (130) may adopt a machine learning model such as a joint trained lip-syncing model with face restoration model, e.g., Wav2Lip with GAN (Generative Adversarial Network) model, Wav2Lip with GFP-GAN (Generative Face Prior- Generative Adversarial Network) model, etc.

[0088] • To accurately overlay the created visemes to the speaker’s face region, the rendering engine (130) may require additional machine learning models such as facial detecting models (e.g., the DeepFaceLab model) and facial editing models (e.g., the GAN-Based facial editing model such as the STIT (Stitch It In Time) model).

[0089] Database (150)

[0090] A database (150) may be provided and served to manage and store a plurality of data (such as the speaker’s identity, the speaker’s face, the speaker’s face region or face regions, the speaker’s mel- spectrogram, the original speeches, the translated speeches, etc.) received from the respective components in the system (100). The data may also be sorted and stored in the database (150), in some embodiments.

[0091] The database (150) may alternatively be configured in such a manner that the data may be received from the components present in the system (100), either in real time or near real time.

[0092] In some embodiments, more than one database (150) may be provided.

[0093] Feedback Module (not shown)

[0094] A feedback module (not shown) may be provided to the rendering engine (130). The feedback module may be an integral part to the rendering engine (130). Alternatively, the feedback module may be a separate part to the rendering engine (130).

[0095] Preferably, the feedback module may facilitate ongoing monitoring, processing and improvement. For instance, the feedback module may collect feedbacks periodically. Based on the collected feedbacks, the feedback module may subsequently require the rendering engine (130) to update, refine or improve the configuration and operation of the system (100) or a target component in the system (100). The rendering engine (130) may also be required to retrain each deployed model or certain deployed models based on the feedbacks, in order to update, refine or improve the performance of the deployed models.

[0096] Additionally, the feedback module may also be configured to receive feedback from a user of the system (100) after the processed video clip (i.e., the video clip that has been processed to replace the original speeches by translated speeches) is displayed to the user. The feedback may be provided in response to an invitation sent in the form of a request and accepted by the user via a user application installed on a computing device which may be in data communication with the system (100). If feedback is provided, the system (100) or in particular the rendering engine (130) may be required to repeat the entire or part of the operations to improve the quality of the processed video clip. If no feedback is to be provided, then the user may be required to provide a confirmation input through the user application, confirming that there is no feedback. The ability to review and provide feedbacks with respect to the processed video clip through the feedback module may subsequently enable a feedback loop or feedback mechanism (and so continuous learning process) to be established. In a further embodiment, a method for processing a video clip may be provided. The method may also be performed by using the system (100) described in the foregoing.

[0097] Preferably, the method may comprise two main phases, i.e., “training phase” followed by “inference phase”, as illustrated in Figure 2. In the training phase, a learning model (such as a deep learning model or a machine learning model) may be trained by the rendering engine (130) using one or more video datasets (such as curated video datasets, historical video datasets or any combination thereof), so that the resulting “trained model” may be used in the subsequent inference phase to perform the steps shown in Figure 3. These steps may include but are not limited to:

[0098] (a) Steps for processing the image portion of the video clip;

[0099] (b) Steps for processing the audio portion of the video clip; and

[0100] (c) Steps for further processing the video clip including but not limited to replacing original speeches with translated speeches and synchronizing the speaker’s face region movements with the movements of the translated speeches.

[0101] The steps (b) and (c) above may be performed in a single stage in some embodiments, thus eliminating the requirement for many different stages of a processing pipeline.

[0102] Training Phase

[0103] The rendering engine (130) may comprise a series of instructions that are executable by at least a processing module for instantiating the training phase shown in Figure 2. Upon execution of these instructions, it may cause the rendering engine (130) or other components present in the system (100) to perform steps outlined in the instructions.

[0104] Preferably, at least the processes S50, S52 and S54 are performed during the training phase, but it is to be understood that other processes than those in the foregoing may be performed during the training phase. It is also to be understood that these processes may be performed, either simultaneously, sequentially, in a pre-programmed sequence or in a user-programmed sequence, during the training phase. Process S50

[0105] Figure 4 is a flow chart showing the initial or first stage of the video preprocessing process (S50).

[0106] One or more video datasets (e.g., curated video datasets, historical video datasets or any combination thereof) may be subjected to the video preprocessing process (S50). Preferably, the curated video datasets and historical video datasets herein may each be a video clip comprising one or more speakers speaking in the first language therein.

[0107] In order to initiate the process S50, the rendering engine (130) may require a machine learning model such as the BodyPix model to be loaded to the face detection module (111), as at Step S5002, for performing the required steps and then require the outputs to be transmitted from the face detection module (111) to the rendering engine (130) or other components in the system (100).

[0108] Once the BodyPix model is successfully loaded, a plurality of video clips may be fed to the face detection module (111) directly from the database (150). Alternatively, the video clips may be fed from the database (150) to the face detection module (111) via the rendering engine (130). For ease of description, the following embodiments will be described by assuming that only one video dataset (i.e., only one video clip) is fed and there is a single speaker in the video clip as fed.

[0109] At Step S5004, each frame of the video clip may be segmented into “silent” and “non- silent” parts. The non-silent parts may preferably be identified and optionally tagged in Step S5004.

[0110] At Step S5006, the non-silent parts identified in the preceding step may each be read and analysed to check whether it is in the parameters suitable for Step S5008. Once it is checked that the non-silent part is in the suitable parameters, then the face detection module (111) may read each frame of the video clip using the readFrame method.

[0111] At Step S5008, each video frame may be subjected to a face detection procedure. The face detection procedure herein may refer to a series of processing tasks comprising prediction of the speaker and the speaker’s face or face region in the video frame. This procedure will be described in detail in the following with reference to Figure 9, but it should be appreciated that the outputs from Step S5008 may include but are not limited to the “speaker” (in the form of image), the “speaker’s face or face region” (in the form of image) and a “bounding box” enclosing the speaker’s face or face region. Preferably, the “bounding box” herein may refer to a spatial location or rectangular coordinates that defines the speaker’s face or face region in the video frame.

[0112] Optionally, the speaker and the speaker’s face or face region may also be tagged in the video frame.

[0113] At Step S5010, the video clip may be divided or trimmed into a plurality of trimmed video clips, and each of the trimmed video clips will be stored in the database (150) or a specific location or folder within the database (150). Preferably, each of the trimmed video clips may comprise at least the following characteristics:

[0114] (a) The trimmed video clip has a length of 2 seconds;

[0115] (b) The trimmed video clip has a frame rate of 25 fps; and

[0116] (c) The trimmed video clip comprises either “only non-silent parts” or “a plurality of tagged portions differentiating silent parts, non-silent parts, the speaker and the speaker’s face or face region”. In some embodiments, the non-silent parts of the trimmed video clip may or may not comprise one or both of the speaker and the speaker’s face or face region.

[0117] Process S52

[0118] Figure 5 is a flow chart showing the second stage of the video preprocessing process (S52). As shown, all video datasets may be subjected to the process S52, whereby the “video datasets” herein may be one or any combination of the “trimmed video clips” and “video clips that have not been subjected to the process S50”.

[0119] In order to initiate the process S52, the rendering engine (130) may require a machine learning model such as the BodyPix model to be loaded to the face detection module (111), as at Step S5202, for performing the required steps and then require the outputs to be transmitted from the face detection module (111) to the rendering engine (130) or other components in the system (100).

[0120] Once the BodyPix model is successfully loaded, a plurality of video clips may be fed to the face detection module (111) directly from the database (150). Alternatively, the video clips may be from the database (150) to the face detection module (111) via the rendering engine (130). For ease of description, the following embodiments will be described by assuming that each video clip as fed will be processed individually and that there is a single speaker in each video clip.

[0121] At Step S5204, each frame of the as-fed video clip may be read and analysed to check whether it is in the parameters suitable for Step S5206. Once it is checked that the video frame is in the suitable parameters, then the face detection module (111) may read each frame of the video clip using the readFrame method.

[0122] At Step S5206, each video frame may be subjected to a face detection procedure. The face detection procedure herein may refer to a series of processing tasks comprising prediction of the speaker and the speaker’s face or face region in the video frame. This procedure will be described in detail in the following with reference to Figure 9, but it should be appreciated that the outputs from Step S5206 may include but are not limited to the “speaker” (in the form of image), the “speaker’s face or face region” (in the form of image) and a “bounding box” enclosing the speaker’s face or face region. Preferably, the “bounding box” herein may refer to a spatial location or rectangular coordinates that defines the speaker’s face or face region in the video frame.

[0123] Optionally, the speaker and the speaker’s face or face region may also be tagged in the video frame.

[0124] Once successfully detecting the speaker and the speaker’s face or face region in each frame of the video clip, each frame may be extracted, written and saved in the form of an image, at Step S5208. Each image may be stored in the database (150) or a specific location or folder within the database (150). In some embodiments, each image saved to the database (150) may comprise either “the speaker”, “the speaker’s face or face region” or “a plurality of tagged portions, each of which may comprise one or both of the speaker and the speaker’s face or face region”.

[0125] After saving each frame in the form of an image in the database (150), it may proceed to analyse the audio portion in each frame of the video clip. To achieve this object, the rendering engine (130) may require a deep learning model (e.g. the pyannote / speaker- diarization model) to be loaded to the speech diarization module (121) for performing the required steps and then require the outputs generated to be transmitted from the speech diarization module (121) to the rendering engine (130) or other components in the system (100).

[0126] Once the speaker diarization module (121) deploys the pyannote / speaker-diarization model, the audio portion or more particularly the original speeches in each frame of the video clip can be extracted, as at Step S5210, and then saved in the WAV format in the database (150) or a specific location or folder within the database (150).

[0127] The mel-spectrogram may also be extracted from the audio portion in each frame of the video clip, as at Step S5212, and subsequently stored in the database (150) or a specific location or folder within the database (150).

[0128] Process S54

[0129] Figure 6 is a flow chart showing a process (S54) for training a lip-syncing model by using the rendering engine. In certain embodiments, the lip-syncing model herein may preferably be Wav2Lip model.

[0130] At Step S5402, a train test split model may first be trained. In particular, the data may be prepared and arranged in an acceptable format for train test split. For instance, the data may be grouped into “features” and “target”. Subsequently, the grouped data may be split into a training set and a testing set, in which: (a) The training set may be used to train a machine learning model or deep learning model; and

[0131] (b) The testing set may be used to test the trained model obtained from (a) above to evaluate the model performance.

[0132] Once a training dataset and testing dataset are obtained, they may be used to train an expert lip-sync discriminator and subsequently the Wav2Lip model, as at Step S5404 and S5406.

[0133] To train the expert discriminator, a discriminator model (such as the colour syncnet model) may be loaded and trained by using the training dataset. Thereafter, the trained colour syncnet model and its performance may be evaluated using the testing dataset. These training and testing steps are particularly essential in producing an accurate and strong lip-sync discriminator for the Wav2Lip model to consistently generate accurate and realistic lip motion or lip shapes.

[0134] After obtaining the lip-sync expert discriminator, it may then be trained in conjunction with a suitable speech2video (s2v) generator to produce the desired Wav2Lip model, as at Step S5406. In certain embodiments, the weight of the expert discriminator may preferably remain frozen throughout the training and the generator may be trained to minimize the discriminator sync loss. Additionally, a visual quality discriminator may be trained in conjunction with the generator and the lip-sync discriminator. Unlike the lip-sync discriminator that performs checks on the lip sync accuracy, the visual quality discriminator facilitates in producing better visual quality and penalize unrealistic face generation.

[0135] All models trained in the process S54 may be stored in the database (150) or a specific location or folder within the database (150), if it is determined that the performance of these trained models is beyond a predetermined acceptable level. Otherwise, it may be required to continue the training process until it is determined that the performance of the models is beyond a predetermined acceptable level. Interference Phase

[0136] The rendering engine (130) may also comprise a series of instructions executable by at least a processing module for instantiating the interference phase shown in Figure 2, once the training phase is completed. Upon execution of these instructions, it may cause the rendering engine (130) or other components present in the system (100) to perform steps outlined in the instructions. These steps may include but are not limited to:

[0137] (A) Steps for processing the image portion of the video clip;

[0138] (B) Steps for processing the audio portion of the video clip; and

[0139] (C) Steps for further processing the video clip including but not limited to replacing original speeches with translated speeches and synchronizing the speaker’s face region movements with the movements of the translated speeches.

[0140] Additionally, it is to be understood that in some embodiments:

[0141] Other steps than those mentioned in the foregoing may also be performed during the interference phase;

[0142] Steps (A), (B) and (C) above may be performed sequentially (which may be in a pre-programmed or a user-programmed sequence) during the interference phase; and

[0143] Steps (B) and (C) above may be performed in a single stage, thereby eliminating the need for many different stages of a processing pipeline.

[0144] (A) Steps for Processing the Image Portion of the Video Clip

[0145] Figure 7 is a flow chart showing the process S56 comprising the steps for processing the image portion of the video clip.

[0146] A video clip may first be captured and recorded by using a recording device or a video capturing device. Preferably, in the video clip, there is a single speaker speaking in the first language. However, it should be appreciated that the number of speakers present in the video clip shall not be limited thereto or thereby, as there can be more than one speaker speaking in the first language in the video clip.

[0147] After successfully capturing and recording a video clip, the image portion of the video clip may be transmitted to the face detection module (111) that will analyse the image in each frame of the video clip. To achieve this object, the face detection module (111) may be required to load the BodyPix model from the database (150), as at Step S5602. Subsequently, the face detection module (111) may read and analyse the video clip (or in particular the video frames) to check whether it is in the predetermined parameters, for instance and without limitation:

[0148] (a) The video clip may be checked if it has a 2K resolution. If it is checked that the resolution of the video clip is beyond 2K resolution, the face detection module (111) may then adjust the resolution of the video clip to 2K resolution.

[0149] (b) The video clip may be checked if it has a 360p resolution (i.e., a resolution of 640 X 360 pixels). If it is checked that the resolution of the video clip falls out of 360p resolution, the face detection module (111) may then adjust the resolution of the video clip to 360p resolution.

[0150] (c) The video clip may also be checked if it has a frame rate of 25 fps, as at Step S5604. If it is checked that the frame rate of the video clip is beyond 25 fps, the face detection module (111) may adjust or convert the frame rate of the video clip to 25 fps, as at Step S5610. However, if it is checked that the frame rate of the video clip is below 25 fps, the face detection module (111) may then trigger an alerting signal to the user of the system (100).

[0151] Once it is checked that the video clip is in the predetermined parameters, then the face detection module (111) may read each frame of the video clip by using the readFrame method, as at Step S5606. In some embodiments, the face detection module (111) may read all the frames from the video clip, either one frame at a time or all of the frames at once. In some embodiments, the face detection module (111) may read part of the video clip by specifying the time or frame interval.

[0152] At Step S5608, each video frame may be subjected to a face detection procedure. The face detection procedure herein may refer to a series of processing tasks comprising prediction of the speaker and the speaker’s face or face region in the video frame. This procedure will be described in detail in the following with reference to Figure 9, but it should be appreciated that the outputs from Step S5608 may include but are not limited to the “speaker” (in the form of image), the “speaker’s face or face region” (in the form of image) and a “bounding box” enclosing the speaker’s face or face region. Preferably, the “bounding box” herein may refer to spatial location or rectangular coordinates that defines the speaker’s face or face region in the video frame.

[0153] Optionally, the speaker and the speaker’s face or face region may also be tagged in the video frame, but it should be appreciated that the outputs from Step S5608 may all be stored in the database (150) or a specific location or folder within the database (150). In particular, once successfully detecting the speaker and the speaker’s face or face region in a video frame, each frame may be extracted and saved in the form of image. Each image may then be saved in the database (150) or a specific location or folder within the database (150). In certain embodiments, each image saved to the database (150) may comprise either “the speaker”, “the speaker’s face or face region” or “a plurality of tagged portions, each of which may comprise one or both of the speaker and the speaker’s face or face region”.

[0154] Additionally, the outputs from the face detection module (111) such as the speaker and the speaker’s face or face region may be subjected to further processing. For instance, these outputs from the face detection module (111) may be used to determine the speaker’s identity by using the face identification module (113). For instance, the face identification module (113) may determine and verify the speaker’s identity by using a machine learning model, e.g., MTCNN (multitask cascaded convolutional network) model or timesler / facenet-pytorch model. By deploying the machine learning model, the face identification module (113) may compare the speaker’s face or face region to a plurality of faces or face regions of previously identified speakers that are stored in the database (150). Alternatively, the face identification module (113) may also send a request inviting the user of the system (100) for providing an input on the speaker’s identity based on the face or face region displayed to the user. The user input may be provided through a user application installed on a computing device which may be in data communication with the system (100).

[0155] Once the speaker’s identity is determined, the speaker’s identity may optionally be tagged to the face or face region by using the face identification module (113).

[0156] In the event where more than one person is present in the video clip, it may then be essential to locate the face or face region of each person present in the video clip by using the face detection module (111). Subsequently, the face detection module (111) may analyse the movement of each located face or face region in the video frame, in order to determine whether the mouth region movements correspond to the movement characteristics of oral communication. If the mouth region movements correspond to the movement characteristic of oral communication, then it is likely that that person is the speaker.

[0157] After the speaker has been detected among the persons present in the video clip, the located face or face region may be transmitted from the face detection module (111) to the face identification module (113) for determining the speaker’s identity.

[0158] In some embodiments, the face or face region may be tagged by the face identification module (113) to the speaker’s identity once determined successfully.

[0159] Although not mentioned in the foregoing, it should also be appreciated that all outputs generated in the steps above may be stored in the database (150) or a specific location or folder within the database (150). the Audio Portion of the Video

[0160] The audio portion of the video clip may preferably be transmitted and analysed by the speech diarization module (121). In order to analyse the audio portion in each video frame, the speech diarization module (121) may deploy a deep learning model (e.g., the pyannote / speaker-diarization pipeline) for extracting the audio portion in the video frame. In particular, the original speeches and mel-spectrogram may both be extracted from the audio portion in each frame of the video clip.

[0161] Later, the extracted mel-spectrogram may be transmitted to the speaker identification module (123) which may determine and verify the speaker’s identity. In particular, the speaker identification module (123) may load and deploy a machine training model, for instance a pretrained ECAPA-TDNN (Emphasized Channel Attention, Propagation and Aggregation-Time Delay Neural Network) model using SpeechBrain or the speechbrain / spkrec-ecapa-voxcelal model. By deploying the machine training model, the speaker identification module (123) may compare the extracted mel-spectrogram to a plurality of mel- spectrograms of previously identified speakers that are stored in the database (150), to determine and verify the identity of the speaker. Alternatively, the speaker identification module (123) may also send a request inviting the user of the system (100) to provide an input on the speaker’s identity based on the extracted mel-spectrogram. The user input may be provided through a user application installed on a computing device which may be in data communication with the system (100).

[0162] The determination results provided by both of the face identification module (113) and speaker identification module (123) can advantageously be used for the speaker identity verification and validation purposes. For instance, the determination results may be compared by the face identification module (113) or alternatively the speaker identification module (123) to verify whether the speaker’s identity determined by the face identification module (113) and the speaker’s identity determined by the speaker identification module (123) are identical. Based on the comparison, it makes possible to verify whether or not the speaker identified by the face identification module (113) and the speaker identified by the speaker identification module (123) are of the same person. Accordingly, it can substantially reduce or eliminate misidentification of the speaker in the video clip. Optionally, the original speeches and corresponding mel-spectrogram may be tagged by the speaker identification module (123) to the speaker’s identity once determined successfully. Alternatively, the tagging may be performed after the speaker’s identity has been verified and validated. For instance, after it has been verified that the speaker identified by the face identification module (113) and the speaker identified by the speaker identification module (123) are of the same person, then the tagging may be performed by the speaker identification module (123).

[0163] In addition to extraction of the original speeches and corresponding mel-spectrogram, the audio portion in each video frame may also be analysed by the speech recognition module (125) to identify, interpret and convert the original speeches contained within the audio portion to texts in the first language. In some alternative embodiments, the original speeches extracted by the speech diarization module (121) may be used as the input to the speech recognition module (125). Upon receipt, the speech recognition module (125) may then proceed to interpret and convert the original speeches to texts in the first language.

[0164] In the event where the language of the original speeches (i.e., the first language) is not known, it may be required that the first language be identified and determined by the speech recognition module (125) first, before proceeding to interpret and convert the original speeches to texts in the first language.

[0165] The texts in the first language may be transmitted from the speech recognition module (125) to the language translation module (127). Upon receipt, the language translation module (127) may instantiate the translation of the texts from the first language into a second language. It should be appreciated that the first language and second language are different languages, in which the “second language” may be a default language set by the system (100) or the user of the system (100). In some alternative embodiments, the second language may be the desired language selected by the user of the system (100) among a plurality of languages.

[0166] The translated texts outputted by the language translation module (127) along with the mel-spectrogram extracted by the speech diarization module (121) may be transmitted to the text-to-speech engine (129) that is specifically designed by the inventors of this invention. Upon receipt, the text-to-speech engine (129) may composite the translated texts with the extracted mel-spectrogram and subsequently convert the composited texts to translated speeches. It should be appreciated that both the translated texts and translated speeches are in the second language. Further, by generating the translated speeches in this manner, the translated speeches when deployed and played will sound like they are delivered by the speaker himself in the second language.

[0167] In the event where more than one person is speaking in the video clip, it may then be required to analyse the audio portion in each video frame using the speech diarization module (121) to extract a plurality of original speeches and then the mel-spectrogram contained within each of the extracted original speeches. Each of the original speeches may be associated with one mel-spectrogram. Each of the mel-spectrograms may also be transmitted from the speech diarization module (121) to the speaker identification module (123) in order to determine the speaker’s identity.

[0168] Although not mentioned in the foregoing, it should also be appreciated that all outputs generated in the steps above may be stored in the database (150) or a specific location or folder within the database (150).

[0169] (C) Steps for Further Processing the Video Clip

[0170] Figure 8 is a flow chart illustrating the process S58 comprising the steps for further processing the video clip.

[0171] In some embodiments, the process S58 may be initiated by transmitting the translated speeches and the corresponding mel-spectrograms obtained in steps (B) above to the rendering engine (130). In particular, the translated speeches may be transmitted to the rendering engine (130) from the database (150) or another component in the system (100) such as the text-to-speech engine (129), as at Step S5806. The mel- spectrogram of the translated speeches may also be transmitted to the rendering engine (130), as at Step 5808. For instance, the mel-spectrogram of the translated speeches may be fed from the database (150) or another component in the system (100) such as the text-to-speech engine (129). Alternatively, the rendering engine (130) may extract the mel-spectrogram from the translated speeches upon receipt. In particular, the translated speeches may be subjected to a machine learning model such as wav2melspectrogram model which converts the translated speeches into a vector array of mel-spectrogram. The mel-spectrogram may further be divided into a plurality of portions, so that each divided portion of the mel-spectrogram may be associated with a specific frame in the video clip and subjected to further processing in Step S5810.

[0172] In the meantime, the rendering engine (130) may also be required to load one or more machine learning models from the database (150), such as the Wav2Lip with GFP- GAN model, as at Step S5802. Thereafter, at Step S5804, all data relating to the video clip (such as the image portion in each video frame, the speaker, the speaker’s face or face region, etc.) may be loaded and transmitted to the rendering engine (130) from the database (150) or another component in the system (100) (e.g., the face detection module (111)). Upon receipt, the rendering engine (130) may optionally read and analyse the data (more particularly the video frames), to check whether they are in the predetermined parameters. For instance, the video frame may be checked if it has a 360p resolution. If it is checked that the resolution of the video clip falls outside of 360p resolution, then the rendering engine (130) may adjust the resolution of the video clip to 360p resolution.

[0173] Once it is checked that the frame images are in the predetermined parameters and that the mel-spectrogram is received, then the rendering engine (130) may proceed to Step S5810, where visemes or a sequence of lip-reading frames that are synchronised to the translated speeches may be generated. Step S5810 will be described in detail in the following with reference to Figure 10, but it should be appreciated that the outputs from Step S5810 shall include a plurality of visemes in JPEG format.

[0174] At Step S5812, the rendering engine (130) may feed the visemes obtained from Step S5810 to any suitable speech2video model (such as superS2V model or particularly DeepFaceLab pre-built face extraction model) which may modify the speaker’s face region based on the received visemes, so that the speaker’s mouth region movements match the movements of the translated speech in the second language. Thereafter, the desired video clip may be obtained and displayed to the user of the system (100), as at Step S5814.

[0175] Although not mentioned in the foregoing, it should be appreciated that:

[0176] • The process S58 shall comprise Steps (B) and (C). In particular, Step S5806 and Step S58O8 of the process S58 may relate to Steps (B), whereas Steps S5802, S5804, S5810 and S5812 may relate to Steps (C). Accordingly, it is possible that the process S58 may be performed by more than one component. For instance, the process S58 may be jointly performed by the audio processor (120) and the rendering engine (130) which are in data communication with one another.

[0177] • A feedback module (not shown) may also be provided for receiving feedback from a user of the system (100) after the video clip obtained from Step S5814 is displayed to the user. The feedback may be provided in response to an invitation sent to the user in the form of a request through a user application that is installed on a computing device which may be in data communication with the system (100). If feedback is provided, then the system (100) or in particular the rendering engine (130) may repeat the entire or part of the operations in order to improve the quality of the processed video clip. If no feedback is to be provided, the user may also be required to provide a confirmation input through the user application, confirming that there is no feedback. The ability to review and provide feedback to the processed video clip through the feedback module may subsequently enable a feedback loop or feedback mechanism (and so continuous learning process) to be established.

[0178] Step S5008, Step S5206, Step S5608

[0179] As mentioned in the foregoing, the face detection procedure is performed respectively in Step S5008, Step 5206 and Step 5608. Figure 9 is a flow chart showing the steps for performing the face detection procedure.

[0180] When the video frames are in the predetermined parameters, each of them may be fed to the face detection module (111).

[0181] Upon receipt, the face detection module (111) may deploy the BodyPix model. By deploying the BodyPix model, it may segment the video frame into multiple parts, in order to predict and detect the speaker present therein. Additionally, it may require the pixels where the speaker is present to be transparent and the pixels where the speaker is not present to be opaque. Accordingly, the pixels with opacity and their coordinates can be used as a mask to identify and crop the speaker from the video frame.

[0182] The face detection module (111) may also deploy the MediaPipe Holistic model. By deploying the MediaPipe Holistic model, it may predict and estimate several regions of interest which may include the speaker’s face or face region, left hand and right hand. In some embodiments, each region of interest may be bounded by a bounding box or rectangular coordinates (xmin, Vmin. xmax, ymax). In some embodiments, the bounding box may comprise some additional paddings. For instance, about 10 to 20% padding may be applied to each side of the bounding box. In some embodiments, the intersection over union (IOU) may also be determined to measure the accuracy of the prediction on the region of interest. In particular, if it is determined that the IOU value is below a predetermined threshold value, it may be required to repeat the foregoing steps (including predicting a region of interest and determining the IOU value). If it is determined that the IOU value is equivalent to or beyond a predetermined threshold value, then the region of interest (e.g., the face or face region, left hand, or right hand) may be cropped from the image of the video frame.

[0183] Once the cropped speaker and the cropped face or face region are obtained, the face detection module (111) may proceed to perform skin modelling on the cropped face or face region, in order to detect the skin tone of the face or face region. For instance, the thresholding method may be deployed, whereby the cropped face or face region may be converted first to a grayscale image to differentiate the face or face region from the background, and then subjected to further processing to detect the skin tone pixels.

[0184] After detecting the skin tone from the cropped face or face region, it may proceed to locate facial features within the face or face region by the MediaPipe Face Detection model. The facial features herein may encompass left eye, right eye, nose tip, mouth, left eye tragion, right eye tragion, etc.

[0185] After detecting the facial features within the cropped face or face region, then the face detection module (111) may transmit all the outputs to the next component or subject all the outputs to the next processing steps. The outputs herein may include but are not limited to the “speaker” (in the form of image), the “speaker’s face or face region” (in the form of image) and a “bounding box” enclosing the speaker’s face or face region.

[0186] Step S5810

[0187] Figure 10 is a flow chart showing Step S5810 in more detail.

[0188] Once the video frame and mel-spectrogram are both received by the rendering engine (130), they may be subjected to a data generator model where batches of datasets may be generated. For instance, each of the datasets may comprise one image frame and its corresponding mel-spectrogram, while each batch may comprise one or more datasets.

[0189] For each batch, they may preferably be subjected to the following steps:

[0190] (a) First, all frames grouped within the same batch may be trained by the Wav2Lip model, where a plurality of visemes or a sequence of lip-reading frames that are synchronised to the translated speeches are created based on the corresponding mel-spectrogram. In some embodiments, each viseme generated may be in 360p resolution.

[0191] (b) Subsequently, for each frame in the same batch, the speaker’s face region may be modified and replaced by the created visemes. Additionally, each modified frame (more particularly each modified face region) may be subjected to further processing, for instance and without limitation:

[0192] (i) The GFP-GAN model may be deployed to remove noise and degradation in each modified frame, thus enhancing the quality and appearance of the modified face region in each frame. Optionally, a preview frame may also be generated. In some embodiments, the preview frame generated may be in 360p resolution.

[0193] (ii) A facial editing model (e.g., STIT model) may be deployed after the GFP- GAN model. By deploying the STIT model, each modified frame may be edited to increase smoothness while reducing jitters especially around the modified face region. In addition, each modified frame may preferably be adjusted to a higher resolution such as 720p resolution.

[0194] (c) The modified frames generated from Steps (b)(i) and (b)(ii) above may each be written and saved to the video clip, before the frames are being sent to the data generator model again.

[0195] (d) Steps (a), (b) and (c) above may be repeated until it is determined that the frame processed is the last frame. Once it is determined that the frame processed is the last frame, then a plurality of modified frames (i.e., the modified face regions) may be generated. In some embodiments, the modified frames generated may be in JPEG format.

[0196] In some alternative embodiments, the methods described in the preceding description may be converted to a series of computer-executable program instructions stored on a non-transitory computer-readable storage medium. When the program instructions are executed by either a computing device or a processing module of the computing device, it may cause the processing module to perform the steps or operation outlined above. In some alternative embodiments, a computer program product comprising a computer useable or readable storage medium having a computer readable program comprising a series of computer-executable program instructions may be provided. The computer readable program, when executed on a computing device, may cause the computing device to perform the steps or operations outlined above.

[0197] EXAMPLE

[0198] An example is provided below to illustrate different aspects and embodiments of the invention. The example is not intended in any way to limit the disclosed invention, which is limited only by the claims.

[0199] Example 1

[0200] The speech recognition module (125) described in the preceding description may adopt the Whisper model with freezing layer to interpret and convert the original speeches contained within the audio portion to texts in the first language.

[0201] Figure 11 shows the overall architecture of the Whisper model with a freezing layer.

[0202] In some exemplary embodiments, the audio portion as received may preferably be resampled to a predetermined frequency, and a log-magnitude mel-spectrogram may be computed on a predetermined window width. For feature normalization, the logmagnitude mel-spectrogram may be globally scaled to be between the values of -1 and 1 with approximately zero mean across the dataset. Later, the encoder may process this input mel-spectrogram with a stem comprising two convolution layers with a predetermined filter width and the GELU activation function whereby the second convolution layer may have a predetermined number of strides. Preferably, the stem may further comprise a freezing layer for retaining the curated datasets, historical datasets or any combination thereof.

[0203] As shown in Figure 11 , sinusoidal position embeddings may be added to the output of the stem, and transformer encoder blocks may be added thereafter. Preferably, the transformer may use pre-activation residual blocks, and a final layer normalization may be applied to the encoder output. Further, the decoder may use learned position embeddings and tied input-output token representations.

[0204] Although not mentioned in the foregoing, it is to be understood that the encoder and decoder may have the same width and number of transformer blocks. Additionally, all tasks and conditioning information may preferably be specified as a sequence of input tokens to be predicted by the decoder. For instance, the input tokens may comprise at least the following:

[0205] • A first token predicted by the model, indicating the language being spoken in the audio portion (i.e., the aforementioned “first language”) or that there is no speech in the audio portion;

[0206] • A second token specifying a task to be performed by the model, e.g., translation;

[0207] • One or a series of third tokens which are “text tokens”; and

[0208] • One or multiple pairs of fourth tokens which are “timestamp tokens” indicating the start time and end time of one sentence in the original speeches.

[0209] Example 2

[0210] The text-to-speech engine (129) described in the preceding description may adopt a non-autoregressive text-to-speech model, in order to synthesize the translated speech with better voice quality.

[0211] Figure 12 shows the overall architecture of the text-to-speech model. As shown, the text-to-speech model may be trained by a plurality of components.

[0212] In particular, a speaker encoder (702) may be provided for receiving the audio portion of the video clip or the original speeches extracted by the speech diarization module (121). Upon receipt, the speaker encoder (702) may be instantiated and trained to extract a fixed dimensional speaker embedding from the audio portion of the video clip or more particularly the original speeches included within the audio portion. A text processor (704) may also be provided for receiving the translated text from the language translation module (127) and converting the received text into phoneme or grapheme embedding sequences. The phoneme or grapheme embedding sequences may then be transferred to a conformer-based text encoder (706) for converting the phoneme or grapheme embedding sequences into the phoneme or grapheme hidden sequences.

[0213] In addition to the speaker encoder (702), text processor (704) and text encoder (706), a speech processor (708) may be provided for extracting variance information such as but not limited to phoneme duration (which represents how long the speech voice sounds), pitch (which is a key feature to convey emotions), energy (which indicates the frame-level magnitude of the mel-spectrogram and directly affects the volume and prosody of speech) and mel-spectrogram from the original speeches or texts derived from the original speech. It should be appreciated that the duration, pitch and energy values extracted may preferably be ground-truth values.

[0214] The outputs from the foregoing components, such as the speaker embedding from the speaker encoder (702), phoneme or grapheme embedding sequences from the text processor (704), phoneme or grapheme hidden sequences from the text encoder (706) and mel-spectrogram from the speech processor (708), may subsequently be fed to a variance adaptor (750) for further processing. In some exemplary embodiments, it may be preferred to add part or all of the phoneme or grapheme hidden sequences to part or all of the speaker embedding, before being transmitted to the variation adaptor (750).

[0215] The foregoing variance adaptor (750) aims to add different variance information (such as phoneme duration, pitch, energy, etc.) to the hidden sequence. Correspondingly, the variance adaptor (750) may comprise an alignment encoder (710), duration predictor (712), pitch predictor (714), energy predictor (716) and length regulator (718), in which the alignment encoder (710) herein is preferably a conformer-based alignment encoder. The alignment encoder (710) may extract the phoneme or grapheme duration, in order to improve the alignment accuracy and therefore reduce the information gap between the model input and output. Thereby, more accurately aligned hidden sequences may be obtained.

[0216] On the other hand, the duration predictor (712) may take the hidden sequences from the text encoder (706) as input and predict the duration of each phoneme or grapheme. Simultaneously, the pitch predictor (714) may add the pitch values or preferably the ground-truth pitch from the speech processor (708) to the aligned hidden sequences. The energy predictor (716) may also add the energy values or preferably ground-truth energy from the speech processor (708) to the aligned hidden sequences. It should be appreciated that the addition of the energy values to the aligned hidden sequences is similar to the addition of the pitch values to the aligned hidden sequences.

[0217] Thereafter, the outputs from the pitch predictor (714) and energy predictor (716) may be taken together with part or all of the hidden sequences that has been added with part or all of the speaker embedding as input to the length regulator (718). Based on the inputs, the length regulator (718) may be trained in order to ensure that the hidden sequences may match or adapt to the length of the audio portion of the video clip.

[0218] Lastly, the conformer-based mel-spectrogram encoder (720) may take the predicted duration provided by the duration predictor (712) as reference and convert the adapted hidden sequences into mel-spectrogram sequences in parallel. The mel-spectrogram sequences may subsequently be transmitted to a vocoder (722) for converting the mel- spectrogram sequence into the translated speech in the second language.

[0219] Although not mentioned in the foregoing, the requirement of feeding the outputs from the duration predictor (712) to the mel-spectrogram encoder (720) instead of length regulator (718) is discovered to be advantageous by the inventors of this invention, as such requirement substantially improves the quality of the voice generated. In another word, smooth and better voice quality can be obtained through this configuration. The disclosure includes as contained in the appended claims, as well as that of the foregoing description. Although this invention has been described in its preferred form with a degree of particularity, it is understood that the disclosure of the preferred form has been made only by way of example and that numerous changes in the details of construction and the combination and arrangements of parts may be resorted to without departing from the scope of the invention.

Claims

CLAIMS1. A system (100) for processing a video clip comprising: an image processor (110) comprising: a face detection module (111) for analysing an image in each frame of the video clip to locate a face region of a speaker; an audio processor (120) comprising: a speech diarization module (121) for analysing an audio portion in each frame of the video clip to extract original speeches belonging to the speaker and then mel- spectrogram contained within the extracted original speeches, the original speeches being in a first language; a speech recognition module (125) for interpreting and converting the extracted original speeches to texts; a language translation module (127) for instantiating the translation of the texts from the first language into a second language, the first language and second language being different languages; and a text-to-speech engine (129) for compositing the translated texts with the extracted mel-spectrogram received from the speech diarization module (121) and then converting the composited texts to translated speeches; and a rendering engine (130) in data communication with the image processor (110) and audio processor (120), wherein for each frame of the video clip, the rendering engine (130) is configured to: receive the translated speeches from the text-to-speech engine (129); replace the original speeches of the speaker in the video frame with the translated speeches; create visemes that are synchronized to the translated speeches; and modify the face region of the speaker based on the created visemes so that the face region movements of the speaker match the movements of the second language.

2. The system (100) according to claim 1 further comprising a face identification module (113) for determining the speaker’s identity based on the face region locatedby the face detection module (111).

3. The system (100) according to claim 1 or 2 further comprising a speaker identification module (123) for determining the speaker’s identity based on the mel- spectrogram extracted by the speech diarization module (121).

4. The system (100) according to claim 3, wherein if the face identification module (113) and speaker identification module (123) co-exist, the face identification module (113) or speaker identification module (123) is further configured to verify whether the speaker’s identity determined by the face identification module (113) and the speaker’s identity determined by the speaker identification module (123) are the same.

5. The system (100) according to claim 1, wherein the speech recognition module (125) is further configured to determine the first language of the extracted original speeches before converting the original speeches to texts.

6. The system (100) according to claim 1, wherein the face region is a mouth region.

7. The system (100) according to claim 1, wherein the mel-spectrogram comprises one or both of prosodic features and quality features.

8. The system (100) according to claim 7, wherein the prosodic features comprise one or more of pitch, pitch range, intonation attitude, loudness, speaking rate and phone duration.

9. The system (100) according to claim 7, wherein the quality features comprise one or more of phonation type, articulatory manner, voice timbre, spectral tilt, difference of amplitude between harmonics and formants, formants bandwidth, jitter and harmonic-to-noise ratio.

10. A method for processing a video clip comprising the steps of: analysing an image in each frame of the video clip by using a face detection module (111) to locate a face region of a speaker; processing an audio portion in each frame of the video clip, wherein the step of processing the audio portion in each frame of the video clip comprises: analysing the audio portion in each frame of the video clip by using a speech diarization module (121) to extract original speeches belonging to the speaker and then mel-spectrogram contained within the extracted original speeches, the original speeches being in a first language; interpreting and converting the extracted original speeches to texts by using a speech recognition module (125); instantiating the translation of the texts from the first language into a second language by using a language translation module (127), the first language and second language being different languages; and compositing the translated texts with the extracted mel-spectrogram from the speech diarization module (121) by using a text-to-speech engine (129); and converting the composited texts to translated speeches by using the text-to-speech engine (129); and transmitting the translated speeches from the text-to-speech engine (129) to a rendering engine (130) which can replace the original speeches of the speaker in the video frame with the translated speeches, create visemes that are synchronized to the translated speeches, and modify the face region of the speaker based on the created visemes so that the face region movements of the speaker match the movements of the second language.

11. The method according to claim 10 further comprising the step of determining the speaker’s identity by using a face recognition module (113), based on the face region located by the face detection module (111).

12. The method according to claim 10 or 11 further comprising the step of determining the speaker’s identity by using a speaker recognition module (123), based on the mel-spectrogram extracted by the speech diarization module (121).

13. The method according to claim 12 further comprising the step of verifying whether the speaker’s identity determined by the face identification module (113) and the speaker’s identity determined by the speaker identification module (123) are the same, wherein the verifying step is performed by using either the face recognition module (113) or speaker recognition module (123).

14. The method according to claim 10 further comprising the step of determining the first language of the extracted original speeches by using the speech recognition module (125), before the converting step.

15. The method according to claim 10, wherein the face region is a mouth region.

16. The method according to claim 10, wherein the mel-spectrogram comprises one or both of prosodic features and quality features.

17. The method according to claim 16, wherein the prosodic features comprise one or more of pitch, pitch range, intonation attitude, loudness, speaking rate and phone duration.

18. The method (500) according to claim 16, wherein the quality features comprise one or more of phonation type, articulatory manner, voice timbre, spectral tilt, difference of amplitude between harmonics and formants, formants bandwidth, jitter and harmonic-to-noise ratio.

Citation Information

Patent Citations

  • Video translation method, system and device and storage medium

    CN112562721A

  • Audio and video translator

    US20230088322A1

  • Video translation platform

    US20230325611A1

  • Methods and systems for video translation

    WO2022106654A2

  • Face-translator: end-to-end system for speech-translated lip-synchronized and voice preserving video generation

    WO2023219752A1