Conference Subtitle System for Real-Time Speech Correction Summaries

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current conference support systems face challenges in maintaining accurate speech recognition, especially when multiple speakers talk simultaneously or in noisy environments, leading to incorrect text conversions, which diverts the corrector's and viewer's attention away from the discussion, making it difficult for participants with hearing impairments to understand the ongoing conversation.

Innovation Solution

The system incorporates a recognizer to convert speech into text, a detector to identify correction operations, and a summarizer to generate summaries of ongoing discussions, allowing the subtitle generator to display summarized text during correction processes, ensuring participants with hearing impairments can follow the conversation smoothly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual correction of speech recognition results is performed, then accuracy of text conversion is improved, but attention of corrector and viewer is diverted from the ongoing discussion

Engineering Contradiction:
Improveaccuracy of text conversionVSAvoidtime to understand discussion content
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts and displays only the relevant portion of the discussion (subtitles corresponding to the correction period) separately from the full transcription, allowing correctors and viewers to focus on understanding the discussion rather than reading through all correction operations

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system prepares and displays subtitles in advance that correspond to the period when correction operations are being performed, so that when corrections are made, the relevant discussion content is already available for immediate understanding

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If full speech recognition results are displayed continuously, then complete information is provided, but difficulty in understanding ongoing discussion increases during correction operations

Engineering Contradiction:
Improvecompleteness of discussion informationVSAvoidease of understanding discussion
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent segments the speech recognition results into different parts: full transcriptions for complete information and extracted subtitles for easy understanding during corrections. This segmentation allows users to access both complete information and simplified views as needed

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The subtitle generation function acts as an intermediary that transforms the full speech recognition results into a simplified, easily understandable format specifically for the correction period, bridging the gap between complete information and ease of understanding

Inventive Principle:
Principle #24Intermediary (Mediator)

3Speed

If speech recognition is performed in noisy environments or with simultaneous speakers, then real-time transcription is achieved, but accuracy of recognition results deteriorates

Engineering Contradiction:
Improvereal-time transcription speedVSAvoidaccuracy of speech recognition
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The system employs speech recognition technology that automatically processes and transcribes speech in real-time without requiring manual intervention for each transcription task, maintaining speed while accepting that accuracy may vary in challenging acoustic environments

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10102859B2Conference support apparatus, conference support method, and computer program product
Publication Date: 2018.10.16 KK TOSHIBA
  • US10102859B2 patent drawing
  • US10102859B2 patent drawing
  • US10102859B2 patent drawing

AI summary

According to an embodiment, a conference support apparatus includes a recognizer, a detector, a summarizer, and a subtitle generator. The recognizer is configured to recognize speech in speech data and generate text data. The detector is configured to detect a correction operation on the text data, the correction operation being an operation of correcting character data that has been incorrectly converted. The summarizer is configured to generate a summary relating to the text data subsequent to a part to which the correction operation is being performed, among the text data, when the correction operation is being detected. The subtitle generator is configured to generate subtitle information corresponding to the summary when the correction operation is being detected, and configured to generate subtitle information corresponding to the text data except when the correction operation is being detected.