Lip-Synced Speaker Video for Multilingual Videoconferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Videoconferencing with simultaneous interpretation mode can create discrepancies between the interpreter's voice and the speaker's lip movements, leading to a sensation of a mediator between the speaker and the listener.
Innovation Solution
The system generates a translated speech audio signal with the voice characteristics of the speaker using an audio conversion model, and creates lip-synched speaker video by matching the lip movements with the translated speech, using a video conversion model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If simultaneous interpretation mode is used in videoconferencing, then participants can understand speech in their native language, but discrepancies arise between the interpreter's voice and the speaker's lip movements
Solution Approach 1:
The patent introduces an audio conversion model as an intermediary component that transforms the interpreter's audio signal into the speaker's voice characteristics. This mediator layer processes the translated speech to match the original speaker's vocal traits, thereby resolving the discrepancy between lip movements and audio without requiring changes to the core interpretation system
Solution Approach 2:
The system creates a copy of the speaker's voice characteristics by analyzing the original speech audio and replicating vocal features in the translated output. The audio conversion model generates a synthesized version of the speaker's voice delivering the translated speech, which is then integrated with the original video feed to maintain lip-sync accuracy
2Loss of information
If an interpreter's voice is used for translated speech, then translation accuracy is achieved, but a sensation of a mediator is created between speaker and listener
Solution Approach 1:
The patent changes the vocal parameters of the translated speech by applying the speaker's voice characteristics to the translated audio content. This parameter transformation includes matching pitch, timbre, and other acoustic features to create a seamless auditory experience that eliminates the perception of interpretation while preserving translation accuracy
3Manufacturing precision
If lip-synched speaker video is generated, then seamless experience is provided, but computational complexity increases
Solution Approach 1:
The system performs preliminary analysis of the speaker's voice characteristics during an initialization phase, storing these parameters for rapid application during live translation. This pre-processing approach reduces real-time computational complexity by preparing voice templates beforehand, allowing the audio conversion model to efficiently generate lip-synched output during the actual videoconference
Data Source
AI summary
Systems and methods for generating speaker video and audio in multiple languages for videoconferencing are provided. For example, a computing device can access a speaker speech audio signal that includes a speaker speech in a first language, a video of the speaker and a translated speech audio signal of the speaker speech in a second language. The computing device generates, based on the translated speech audio signal, a converted translated speech audio signal that includes a speech in the second language having voice characteristics in the speaker speech. The computing device further generates a lip-synched speaker video based on the video of the speaker and the converted translated speech audio signal. Lip movements in the lip-synched speaker video correspond to the converted translated speech audio signal. The converted translated speech audio signal and the lip-synched speaker video are transmitted to a video conference provider configured to host the video conference.


