Distributed Speech-to-Text Captioning for Conference Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional captioning systems in conferencing applications face challenges with accuracy due to multiple accents, languages, background noise, and poor communication link quality, often requiring excessive bandwidth and relying on human translators.

Innovation Solution

A conferencing system employing distributed Speech-To-Text (STT) conversion at endpoints, combined with language translation and captioning capabilities, which converts speech into text concurrently with audio and transfers it to participants, allowing for accurate and versatile captioning with reduced bandwidth requirements and improved noise resistance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If human translators are employed for captioning, then captioning versatility is improved, but system complexity and cost increase

Engineering Contradiction:
Improvecaptioning versatilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system enables automatic self-captioning through speech-to-text conversion at each endpoint, eliminating the need for human translators. Each participant's device independently performs captioning functions, providing service to itself and other participants without external human intervention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical human translator system with an automated speech recognition and text conversion system. The mechanical action of human listening and typing is substituted with electronic speech-to-text conversion processes that occur automatically at each endpoint.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If speech-to-text conversion is performed centrally, then captioning accuracy is improved, but bandwidth requirements increase

Engineering Contradiction:
Improvecaptioning accuracyVSAvoidbandwidth requirements
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent divides the centralized speech-to-text conversion function into distributed segments at each endpoint. Instead of one central conversion point, each participant's device performs its own speech-to-text conversion independently, segmenting the processing load across multiple locations and reducing network bandwidth requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from a single-dimensional centralized processing model to a multi-dimensional distributed processing architecture. Speech-to-text conversion occurs simultaneously at multiple endpoints rather than being funneled through a single central point, adding spatial distribution as a new dimension to the processing architecture.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If human translators listen to the conference, then captioning accuracy is improved, but noise interference increases

Engineering Contradiction:
Improvecaptioning accuracyVSAvoidnoise interference
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent introduces an intermediary speech-to-text conversion process that operates at each endpoint before text transmission. This intermediary conversion layer filters and processes speech locally, converting it to text form that is less susceptible to noise interference during transmission, thereby protecting the captioning accuracy from background noise effects.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS7830408B2Conference captioning
Publication Date: 2010.11.09 CISCO TECHNOLOGY INC
  • US7830408B2 patent drawing
  • US7830408B2 patent drawing
  • US7830408B2 patent drawing

AI summary

A system and method for providing captioning in a conference. In an illustrative embodiment, the method includes establishing a conference between a first participant and a second participant. The conference exhibits an exchange of a first type of media between the participants. A user option is provided to augment the conference with a second type of media corresponding to the first type of media. The second type of media is then generated based on one or more conference parameters in response to the signal. In a more specific embodiment, the second type of media is automatically generated. The first type of media is provided approximately concurrently with the second type of media. The second type of media, which may represent text captions, is then selectively provided to one or more participants in the conference based on predetermined preferences.