Automatic Speech Caption Timing Windows via Binary Score Smoothing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for adding captions to audio streams are inefficient, as they often require manual timing by volunteers, which is cumbersome and discourages participation, and automated systems provide low accuracy, leading to a shortage of captioned content.

Innovation Solution

A computer-implemented method that automatically determines the timing windows for speech sounds in an audio stream by using a speech classifier to generate raw scores, smoothing them into binary scores, and then generating timing windows for caption boxes, reducing the need for manual input and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated speech recognition systems are used to generate captions, then productivity is improved, but measurement precision deteriorates

Engineering Contradiction:
Improvecaption generation efficiencyVSAvoidcaption timing accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary component (speech classifier with binary score generation) between the automated speech recognition system and the final caption output. This intermediary processes raw speech detection scores through aggregation and binary classification to produce more accurate timing windows, thereby resolving the contradiction between automated efficiency and timing precision

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual timing by volunteers is used, then measurement precision is improved, but ease of operation deteriorates

Engineering Contradiction:
Improvecaption timing accuracyVSAvoiduser convenience
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs self-service by automatically generating timing windows for captions without requiring manual volunteer intervention. The speech classifier and binary score aggregation process autonomously determine accurate timing windows, eliminating the need for users to manually time captions while maintaining high precision

Inventive Principle:
Principle #25Self-service

3Measurement precision

If manual timing by volunteers is used, then measurement precision is improved, but productivity deteriorates

Engineering Contradiction:
Improvecaption timing accuracyVSAvoidcaption generation speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces the mechanical human effort of manual timing with an automated computational system. The speech classifier and binary score aggregation process automatically generate precise timing windows at high speed, substituting the slow manual process while maintaining or improving measurement precision

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11011184B2Automatic determination of timing windows for speech captions in an audio stream
Publication Date: 2021.05.18 GOOGLE LLC
  • US11011184B2 patent drawing
  • US11011184B2 patent drawing
  • US11011184B2 patent drawing

AI summary

The technology disclosed herein may determine timing windows for speech captions of an audio stream. In one example, the technology may involve accessing audio data comprising a plurality of segments; determining, by a processing device, that one or more of the plurality of segments comprise speech sounds; identifying a time duration for the speech sounds; and providing a user interface element corresponding to the time duration for the speech sounds, wherein the user interface element indicates an estimate of a beginning and ending of the speech sounds and is configured to receive caption text associated with the speech sounds of the audio data.