Textless Video Scene Matching Through Automated Shot Sequence Pairing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for matching textless elements to texted elements in video content are manual, time-consuming, inefficient, and prone to human error, leading to delays and inaccuracies in the localization and distribution of video content for foreign markets.

Innovation Solution

A video processing system that automatically segments video content into shots, identifies sequences with similar durations, compares representative frames using image metrics, and pairs matching texted and textless sequences, enabling automated replacement of texted elements with textless elements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual methods are used to match textless elements to texted elements, then flexibility and adaptability are maintained, but time consumption and human error increase significantly

Engineering Contradiction:
Improvematching accuracyVSAvoidmatching time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical matching processes with automated computer-based image processing and comparison algorithms. The system uses digital image metrics, frame-by-frame analysis, and automated sequence matching to identify corresponding texted and textless shots, eliminating manual inspection while maintaining high accuracy through computational precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system creates and compares digital copies of video frames and sequences to identify matching content. By generating standardized representations of texted and textless shots and comparing these digital copies using image metrics, the system achieves rapid, accurate matching without manual intervention.

Inventive Principle:
Principle #26Copying

2Measurement precision

If automated image comparison is performed on all frames, then matching precision is maximized, but computational complexity and processing time increase

Engineering Contradiction:
Improveimage content matching precisionVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments video content into discrete shots and further divides shots into representative frames for comparison. By processing only key frames rather than every single frame, and by segmenting the comparison task into manageable units (shot-level matching followed by frame-level verification), the system achieves high precision while controlling computational complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs partial comparison by focusing on representative frames and key visual elements rather than analyzing every pixel of every frame. This selective approach maintains matching precision by concentrating computational resources on the most discriminative features while avoiding unnecessary processing of redundant information.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12394203B2Textless material scene matching in videos
Publication Date: 2025.08.19 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12394203B2 patent drawing
  • US12394203B2 patent drawing
  • US12394203B2 patent drawing

AI summary

Systems, methods, and a computer-readable medium are provided for matching textless elements to texted elements in video content. A video processing system including a textless matching system may divide a video into shots, identify shots having similar durations, identify sequences of shots having similar durations, and compare image content in representative frames of the sequences to determine whether the sequences match. When the sequences are determined to match, the sequences may be paired, wherein the first sequence may include shots with overlaid text and the second sequence may include textless version of corresponding texted shots included in the first sequence. In some examples, the video processing system may further replace the determined corresponding texted shots.