Conversation Assistance Device Speech Rate Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Smooth communication between native and non-native speakers is hindered due to differences in speech rates, with existing solutions either causing delays in response or requiring excessive attention on displayed keywords, leading to lag in conversation flow.

Innovation Solution

A system that includes a rate acquisition unit, voice recognition unit, and presentation unit to process and present voice recognition results in a way that subtly encourages the speaker to reduce their speech rate by making the system appear to have incorrectly recognized fast speech, thereby facilitating a natural reduction in speech rate without disrupting the conversation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the NS's speech rate is reduced to generate slow and easy-to-hear voice, then the NNS can listen to the speech slowly, but the timing at which the NNS ends listening is later than the timing at which the NS ends speech, causing the NNS to be unable to reply timely

Engineering Contradiction:
Improveease of listeningVSAvoidresponse time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system creates a copy of the speech content in the form of displayed text/keywords that can be processed independently of the audio playback timing. This allows the NNS to read the copied text representation at their own pace without being constrained by the slowed audio timing, thus maintaining response timing while improving comprehension ease.

Inventive Principle:
Principle #26Copying

2Loss of information

If keywords are displayed to the NNS to be communicated, then the NNS can see the key information, but the NNS must pay attention to both listening and displayed character information, and there is a lag time from the end of the NS's speech to the display output

Engineering Contradiction:
Improveinformation retentionVSAvoidattention demand
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system extracts only the most essential keywords from the full speech content rather than displaying the entire transcript. This selective extraction reduces the information load on the NNS, allowing them to focus on key concepts without being overwhelmed by excessive text, thus reducing attention demand while maintaining information retention.

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If the NS speaks at a normal high speech rate, then the conversation flows naturally for the NS, but the NNS finds it difficult to listen and participate in communication

Engineering Contradiction:
Improveconversation flowVSAvoidease of listening
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system introduces a visual text/keyword display as an intermediary between the NS's speech and the NNS's comprehension. This intermediary translates the auditory information into a visual format that the NNS can process more easily, maintaining the NS's natural speech rate and conversation flow while improving the NNS's ease of listening and participation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11978443B2Conversation assistance device, conversation assistance method, and program
Publication Date: 2024.05.07 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11978443B2 patent drawing
  • US11978443B2 patent drawing

AI summary

Implemented is a communication with reasonably smooth conversation even between people having different proficiency levels in a common language. Included are a voice recognition unit (14) that acquires a speech rate of a speaker and recognizes a voice on a speech content; and a call voice processing unit (12) that processes a part of a voice recognition result based on a result of comparing the acquired speech rate with a reference speech rate, and transmits a video on which a text character image of the voice recognition result having been processed is superimposed to a communication terminal TM of the speaker.