Personalized Video Ad Insertion With Speech and Lip Sync
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional advertising systems require manual video editing or re-recording to incorporate advertisement changes, which is inefficient and disruptive to content delivery.
Innovation Solution
A method and apparatus that seamlessly integrates personalized advertisements into video content by using speech synthesis and lip movement generation, allowing for dynamic customization without altering the original content structure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If manual video editing or re-recording is used to incorporate advertisement changes, then advertisement customization is achieved, but content delivery efficiency deteriorates and playback continuity is disrupted
Solution Approach 1:
The patent creates a duplicate audio track containing the advertisement script that mirrors the structure and timing of the original audio track. This copied track can be independently processed and delivered without modifying the original content, enabling efficient advertisement customization while preserving the original video stream for seamless playback.
Solution Approach 2:
The patent divides the video content into separate audio and video tracks, and further segments the audio track into original content and advertisement portions. This segmentation allows independent processing of advertisement segments without affecting the overall content delivery efficiency or playback continuity.
2Adaptability or versatility
If multiple versions of video content are created for different advertisements, then personalized advertisement delivery is achieved, but system complexity increases
Solution Approach 1:
The patent implements a dynamic advertisement insertion system where advertisement tracks are generated and assigned based on user profiles and viewing context at runtime. Instead of pre-creating multiple static video versions, the system dynamically selects and inserts appropriate advertisement audio tracks into the standardized video stream, reducing system complexity while enabling personalized delivery.
Solution Approach 2:
The patent creates a universal video structure with separate audio and video tracks that can accommodate different advertisement content. The standardized framework with distinct audio tracks serves multiple purposes: original content delivery, advertisement insertion, and personalized recommendation, eliminating the need for separate processing pipelines for each advertisement type.
3Productivity
If advertisement content is inserted into video streams, then revenue generation is improved, but playback continuity may be disrupted
Solution Approach 1:
The patent introduces an intermediary advertisement audio track that acts as a mediator between the original content and the user experience. This separate audio track can be independently controlled, inserted, and removed without affecting the video stream or original audio, ensuring playback continuity while enabling revenue-generating advertisement delivery.
Solution Approach 2:
The patent performs preliminary processing of advertisement audio tracks to match the timing, duration, and structure of the original audio track before insertion. By pre-synchronizing the advertisement content with the video timeline and preparing replacement audio segments in advance, the system ensures seamless playback continuity when advertisements are inserted into the video stream.
Data Source
AI summary
Provided is a method and apparatus for video synthesis. The method includes obtaining an advertising target section from a video and sampling speech data of a target speaker from at least part of the audio track. The audio track corresponding to the target speaker is then changed into speech synthesis data, wherein the target speaker utters an advertising script related to an advertising object. In parallel, the lip movement of the target speaker in the video track is modified to match the utterance of the advertising script. This approach allows for seamless integration of personalized advertisements into video content by synchronizing both speech and visual lip movements with the advertising script.


