Videoconference Speech-to-Text Conversion via Intermediary Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current videoconferencing systems lack the ability to automatically convert audio speech into text information in real-time, limiting the functionality and usability of videoconferences, especially in terms of transcription and language translation.

Innovation Solution

A method is implemented in videoconferencing devices to receive audio information from remote endpoints, automatically convert speech into text, and store or display this text information, with the option to translate it into other languages, enhancing participant interaction and data management during conferences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech-to-text conversion functionality is added to videoconferencing systems, then the usability and functionality are improved, but the device complexity increases

Engineering Contradiction:
ImproveusabilityVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent introduces a separate speech-to-text conversion system that receives audio information from the videoconferencing system and processes it independently. This intermediary approach allows the core videoconferencing functionality to remain simple while adding transcription capabilities through a dedicated module that operates in parallel, thus improving usability without significantly increasing the complexity of the main system.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The videoconferencing device is enhanced to perform multiple functions: traditional video/audio communication plus speech-to-text conversion. By integrating a multi-functional capability into the existing device, the system provides both real-time communication and transcription services, improving overall usability while consolidating functions within a single platform.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If real-time speech conversion is implemented, then the productivity is improved, but the use of energy increases

Engineering Contradiction:
ImproveproductivityVSAvoiduse of energy
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The speech-to-text conversion operates periodically based on speech detection rather than continuously processing all audio. The system activates transcription processing when speech is detected and pauses during non-speech periods, maintaining high productivity for actual communication while reducing energy consumption during idle intervals.

Inventive Principle:
Principle #19Periodic action

3Loss of information

If speech-to-text conversion is added, then the loss of information is reduced, but the device complexity increases

Engineering Contradiction:
Improveloss of informationVSAvoiddevice complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system creates a text copy of the spoken audio information, providing a parallel representation of the same data. This text transcript serves as a backup or alternative form of the information, ensuring that content is preserved and can be reviewed later, thus reducing information loss while adding a relatively simple transcription layer to the existing system.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS8120638B2Speech to text conversion in a videoconference
Publication Date: 2012.02.21 ENGHOUSE INTERACTIVE
  • US8120638B2 patent drawing
  • US8120638B2 patent drawing
  • US8120638B2 patent drawing

AI summary

Various embodiments of a method for automatically converting audio speech in a videoconference into text information are described. According to one embodiment of the method, a videoconferencing device at a first endpoint in the videoconference may receive a stream of video information and audio information from a videoconferencing device at a second endpoint in the videoconference. The audio information includes speech of a participant at the second endpoint. The videoconferencing device at the first endpoint may automatically convert the speech into text information.