Speech-to-Text Captioning With Confidence-Based Review

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional automatic speech-to-text frameworks for video content lack accuracy and user control over captions, particularly in scenarios with burned-in captions and background music, and fail to differentiate between high and low accuracy captions.

Innovation Solution

A framework that automatically generates captions, determines their accuracy using token-level confidence values and silence intervals, and prompts users for manual review of low-accuracy captions, while detecting burned-in captions and allowing user settings for caption control.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If automated speech-to-text frameworks are used for video content, then caption generation is automated, but caption accuracy deteriorates particularly in scenarios with burned-in captions and background music

Engineering Contradiction:
Improvecaption generation automationVSAvoidcaption accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The system segments the caption generation process into multiple stages: automatic caption generation, quality assessment using multiple criteria (clarity, completeness, accuracy), and conditional routing to manual review or direct publication. This segmentation allows automated processing while maintaining quality control through staged evaluation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system implements feedback loops where generated captions are assessed against quality metrics, and low-quality captions trigger feedback to users for manual review and correction. The system continuously learns from user corrections to improve future automatic caption generation accuracy.

Inventive Principle:
Principle #23Feedback

2Device complexity

If traditional speech-to-text frameworks are used, then caption generation is simple, but user control over captions deteriorates

Engineering Contradiction:
Improvesystem simplicityVSAvoiduser control
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The system dynamically adjusts the level of user involvement based on caption quality assessment. High-quality captions are published automatically without user intervention, while low-quality captions trigger user review workflows. Users can also manually adjust quality thresholds and control their preferred level of involvement, making the system adaptable to different user needs.

Inventive Principle:
Principle #15Dynamics

3Productivity

If all captions are automatically published, then productivity is improved, but caption quality deteriorates due to lack of review

Engineering Contradiction:
Improvecaption publishing speedVSAvoidcaption quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system changes the parameter of quality assessment from binary (publish/not publish) to a continuous scale with multiple quality dimensions. Captions are evaluated on clarity, completeness, and accuracy, with dynamic quality thresholds that can be adjusted. This allows the system to publish captions efficiently when quality standards are met while maintaining reliability through configurable quality gates.

Inventive Principle:
Principle #35Parameter changes

4Reliability

If manual review is required for all captions, then caption quality is improved, but productivity deteriorates due to increased user workload

Engineering Contradiction:
Improvecaption qualityVSAvoidcaption publishing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system applies partial manual review action only when necessary. Instead of requiring review for all captions, the system generates captions automatically and only triggers manual review for those that fail quality assessment thresholds. This partial application of manual review maintains quality for problematic captions while preserving productivity for high-quality automatic captions.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12593108B2Systems and methods for automated speech-to-text captioning
Publication Date: 2026.03.31 META PLATFORMS TECHNOLOGIES LLC
  • US12593108B2 patent drawing
  • US12593108B2 patent drawing
  • US12593108B2 patent drawing

AI summary

The disclosed systems and methods may include (1) generating a caption for a spoken phrase in a video, (2) determining an accuracy rating for the caption, and (3) in response to determining that the accuracy rating is below an accuracy threshold, prompting a user to manually review the caption prior to publishing the caption. Various other methods, systems, and computer-readable media are also disclosed.