Animated Video Editing With Automatic Character-Text Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video editing processes require manual alignment of characters with line text on a timeline, leading to low efficiency.

Innovation Solution

A video editing method that automatically associates characters with line text using text-to-speech technology, allowing for the generation of animated videos where characters read the corresponding text synchronously.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual alignment of characters with line text is used, then precise positioning can be achieved, but video editing efficiency deteriorates

Engineering Contradiction:
Improvevideo editing efficiencyVSAvoidoperation complexity
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system performs automatic character-text alignment using text-to-speech technology and time synchronization, eliminating the need for manual positioning operations. The character automatically associates with the corresponding line text based on temporal relationships, making the system self-sufficient in the alignment task.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The manual mechanical alignment process is replaced by an automated digital system that uses text-to-speech conversion and time-based synchronization algorithms. This substitution transforms a labor-intensive manual operation into an automated computational process, significantly improving efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automatic character-text association is implemented, then editing efficiency improves, but technical complexity increases

Engineering Contradiction:
Improvevideo editing efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

Text-to-speech audio serves as an intermediary element that bridges the character and line text. By converting text to speech and synchronizing it with character actions, the system creates a temporal mediator that automatically establishes the association without requiring complex direct linking mechanisms.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary text-to-speech conversion and timing analysis before finalizing the character-text association. This preliminary processing prepares the data in advance, allowing the automatic alignment to occur smoothly during video generation without adding complexity to the final assembly process.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12499911B2Video editing method and apparatus, computer device, storage medium, and product
Publication Date: 2025.12.16 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US12499911B2 patent drawing
  • US12499911B2 patent drawing
  • US12499911B2 patent drawing

AI summary

This application provides a video editing method performed by a computer device. The video editing method includes: displaying a video editing interface; determining a target character and an input text in the video editing interface; generating an animated video and a line audio, the animated video comprising the target character, and the line audio corresponding to the text for the target character in the animated video; and synchronously playing the line audio corresponding to the text in a process of playing the animated video.