In-Video Text Translation Overlay for Continuous Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video translation methods interrupt the normal viewing experience by pausing the video to present translated text, leading to poor user experience.

Innovation Solution

A method and apparatus that convert and present translated text in real-time during video playback, using a text box region or dynamically determined region based on the video frame content, allowing seamless integration with the video without interruption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If translated text is presented during video pause, then translation information can be provided to users, but normal video watching is affected and user experience deteriorates

Engineering Contradiction:
Improvetranslation information deliveryVSAvoidvideo watching continuity
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent merges the translation function with the video playback function by overlaying translated text on the video frame. This allows translation information to be delivered during continuous video playback without requiring video pause, thus resolving the contradiction between providing translation information and maintaining video watching continuity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a new spatial dimension by displaying translated text in a target region overlaying the video frame. This dimensional approach allows translation information to be presented simultaneously with video content without interrupting playback, solving the contradiction between information delivery and viewing continuity

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If translated text is displayed prominently, then translation information is clearly presented, but the displayed text may block or affect the video content

Engineering Contradiction:
Improvetranslation information visibilityVSAvoidvideo content obstruction
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent applies local quality by determining a target region for translated text display based on the video frame content. The text is displayed in a specific region that is dynamically determined to minimize obstruction of important video content, thus balancing translation visibility with video content preservation

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses dynamic region determination where the target region for translated text is automatically adjusted based on the current video frame content. This dynamic adaptation ensures that translation text is always displayed in the least obstructive location, resolving the contradiction between clear text presentation and video content visibility

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20260032318A1Video processing method, electronic device and medium
Publication Date: 2026.01.29 DOUYIN VISION CO LTD
  • US20260032318A1 patent drawing
  • US20260032318A1 patent drawing
  • US20260032318A1 patent drawing

AI summary

Provided are a video processing method, an electronic device and a storage medium. The method includes steps described below. In a process of playing a target video, to-be-converted text in a to-be-processed video frame is determined in response to a triggering operation by a target user on the target video; and the to-be-converted text is converted into translated text of a target language type, and the translated text is presented in a target region in the to-be-processed video frame; where the target region includes a text box region to which the to-be-converted text belongs, or the target region is dynamically determined based on a picture content of the to-be-processed video frame.