Audio Captioning System Using Segmented Human Review

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automatic captioning, subtitles, and dubbing technologies for audiovisual media face challenges in accuracy, particularly with proper nouns and idiomatic translations, and lack the emotional and vocal qualities of the original speaker.

Innovation Solution

A machine learning-based approach for generating captions, subtitles, and dubbing that includes speech-to-text conversion, temporal synchronization, and parameter determination for tone, volume, and emotion, with human editing to enhance accuracy and realism, and machine translation for multi-language support.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If speech-to-text algorithm is used for automatic captioning, then productivity is improved, but manufacturing precision (caption accuracy) deteriorates

Engineering Contradiction:
Improvecaption generation speedVSAvoidcaption accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The caption generation process is segmented into multiple stages: initial automatic speech-to-text generation, followed by selective human review and editing of specific segments (words, phrases, or sentences). This allows the system to maintain high productivity through automation while improving precision by having humans focus only on problematic segments rather than reviewing everything manually.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system incorporates feedback mechanisms where human editors review and correct automatic captions, and these corrections are fed back into the system to improve future automatic generation. The feedback loop includes machine learning models that learn from human corrections to improve their speech-to-text accuracy over time.

Inventive Principle:
Principle #23Feedback

2Productivity

If machine translation software is used for subtitle generation, then productivity is improved, but manufacturing precision (translation accuracy) deteriorates

Engineering Contradiction:
Improvesubtitle generation speedVSAvoidtranslation accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The translation process is divided into segments where machine translation handles the bulk of translation work for productivity, while human translators selectively review and edit specific segments (sentences, phrases) that require cultural nuance or idiomatic translation. This segmentation allows the system to achieve both high productivity through machine translation and high precision through targeted human review.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The translation system incorporates feedback loops where human translators review machine-generated translations and provide corrections. These corrections are fed back into the machine translation system to improve its performance over time, particularly for culturally specific terms and idiomatic expressions.

Inventive Principle:
Principle #23Feedback

3Productivity

If automatic dubbing is generated, then productivity is improved, but manufacturing precision (vocal quality realism) deteriorates

Engineering Contradiction:
Improvedubbing generation speedVSAvoidvocal quality realism
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The dubbing process is segmented into automatic generation of dubbed audio tracks and selective human review of specific segments. The machine generates the bulk of the dubbing work for productivity, while human voice artists review and refine specific portions to add emotional nuance, cultural appropriateness, and vocal realism, particularly for dialogue that requires emotional expression or cultural specificity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The dubbing system incorporates feedback mechanisms where human voice artists review machine-generated dubbed audio and provide corrections. These feedback loops allow the system to improve its ability to generate emotionally resonant and culturally appropriate dubbed voices over time.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240155205A1Method for generating captions, subtitles and dubbing for audiovisual media
Publication Date: 2024.05.09 SYNCWORDS
  • US20240155205A1 patent drawing
  • US20240155205A1 patent drawing
  • US20240155205A1 patent drawing

AI summary

The method for generating captions, subtitles and dubbing for audiovisual media uses a machine learning-based approach for automatically generating captions from the audio portion of audiovisual media, and further translates the captions to produce both subtitles and dubbing. A speech component of an audio portion of audiovisual media is converted into at least one text string which includes at least one word. Temporal start and end points for the at least one word are determined, and the at least one word is visually inserted into the video portion of the audiovisual media. The temporal start and end points for the at least one word are synchronized with corresponding temporal start and end points of the speech component of the audio portion of the audiovisual media. A latency period may be selectively inserted into broadcast of the audiovisual media such that the synchronization may be selectively adjusted during the latency period.