Video Caption Translation Workflow With Crowdsourced Proofreading

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users on social networks have increasing demands for high-quality caption translation in videos, but existing systems lack efficient mechanisms for improving translation quality and efficiency, primarily relying on video posters for caption editing.

Innovation Solution

A video processing method and apparatus that utilizes crowd-sourcing translation, providing translators with interactive interfaces for editing captions, including proofreading pages, and displaying assessed translations within videos, tailored to translator interests and language proficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If crowd-sourcing translation is implemented, then translation quality and efficiency are improved, but system complexity increases

Engineering Contradiction:
Improvetranslation qualityVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The translation system is segmented into multiple independent components: translator terminal for translation input, proofreading terminal for quality verification, and assessment terminal for evaluation. This segmentation allows complex translation tasks to be distributed across multiple users while maintaining system manageability through clear role division.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The server acts as an intermediary that coordinates between translators, proofreaders, and assessors. It manages task allocation, collects translation results, performs automated assessments, and coordinates proofreading workflows, thereby reducing the complexity burden on individual terminals while enabling high-quality crowd-sourced translation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If automated assessment is implemented, then translation efficiency is improved, but assessment accuracy may deteriorate

Engineering Contradiction:
Improvetranslation efficiencyVSAvoidassessment accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system merges automated machine assessment with manual human proofreading to create a hybrid evaluation mechanism. The server performs automated linguistic and contextual analysis for efficiency, while human proofreaders provide nuanced quality verification, combining the speed of automation with the accuracy of human judgment.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The assessment results are fed back to translators for improvement. The system provides automated assessment feedback including linguistic accuracy, contextual appropriateness, and formatting compliance, allowing translators to learn and improve their translation quality over time while maintaining high throughput.

Inventive Principle:
Principle #23Feedback

3Productivity

If multiple translation terminals are coordinated, then translation capacity is improved, but coordination complexity increases

Engineering Contradiction:
Improvetranslation capacityVSAvoidcoordination complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The server provides universal coordination functions that handle multiple translation terminals, proofreading terminals, and assessment terminals through a unified platform. It manages task distribution, result collection, quality assessment, and workflow coordination in a multi-functional manner, enabling scalable expansion without proportionally increasing coordination complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12572757B2Video processing method, video processing apparatus, and computer-readable storage medium
Publication Date: 2026.03.10 DOUYIN VISION CO LTD
  • US12572757B2 patent drawing
  • US12572757B2 patent drawing
  • US12572757B2 patent drawing

AI summary

This disclosure relates to a video processing method, a video processing apparatus, and a computer-readable storage medium. The video processing method includes: providing a translator with a video to be translated, and providing, in-feed in the video, the translator with an interactive interface for translating an original caption in the video, on which a translation of the original caption is comprised; entering a proofreading page in response to an edit request of the translator for the translation; receiving a caption translation returned by the translator from the proofreading page; and displaying, in the video, a caption translation passing assessment.