Automatic Speech Caption Timing Windows via Binary Score Smoothing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for adding captions to audio streams are inefficient, as they often require manual timing by volunteers, which is cumbersome and discourages participation, and automated systems provide low accuracy, leading to a shortage of captioned content.
Innovation Solution
A computer-implemented method that automatically determines the timing windows for speech sounds in an audio stream by using a speech classifier to generate raw scores, smoothing them into binary scores, and then generating timing windows for caption boxes, reducing the need for manual input and improving accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automated speech recognition systems are used to generate captions, then productivity is improved, but measurement precision deteriorates
Solution Approach 1:
The patent introduces an intermediary component (speech classifier with binary score generation) between the automated speech recognition system and the final caption output. This intermediary processes raw speech detection scores through aggregation and binary classification to produce more accurate timing windows, thereby resolving the contradiction between automated efficiency and timing precision
2Measurement precision
If manual timing by volunteers is used, then measurement precision is improved, but ease of operation deteriorates
Solution Approach 1:
The system performs self-service by automatically generating timing windows for captions without requiring manual volunteer intervention. The speech classifier and binary score aggregation process autonomously determine accurate timing windows, eliminating the need for users to manually time captions while maintaining high precision
3Measurement precision
If manual timing by volunteers is used, then measurement precision is improved, but productivity deteriorates
Solution Approach 1:
The patent replaces the mechanical human effort of manual timing with an automated computational system. The speech classifier and binary score aggregation process automatically generate precise timing windows at high speed, substituting the slow manual process while maintaining or improving measurement precision
Data Source
AI summary
The technology disclosed herein may determine timing windows for speech captions of an audio stream. In one example, the technology may involve accessing audio data comprising a plurality of segments; determining, by a processing device, that one or more of the plurality of segments comprise speech sounds; identifying a time duration for the speech sounds; and providing a user interface element corresponding to the time duration for the speech sounds, wherein the user interface element indicates an estimate of a beginning and ending of the speech sounds and is configured to receive caption text associated with the speech sounds of the audio data.


