Voice Assistant Response Replay Using Text-Based Speech Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing electronic devices with speech recognition capabilities often provide unexpected responses when users cannot see the screen, making it difficult for users to correctly interpret the intended information, especially when relying on speech feedback.

Innovation Solution

An electronic device equipped with a microphone, speaker, memory, and processor that performs speech recognition, outputs a first response message, detects re-requests, recognizes the text of the first response, and generates a second response message based on a first parameter corresponding to the text, allowing for improved interaction through speech feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech recognition is implemented in electronic devices to enable hands-free control, then ease of operation is improved, but reliability deteriorates when users cannot see the screen and must rely on speech feedback alone

Engineering Contradiction:
Improvehands-free control capabilityVSAvoiduser understanding accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces text as an intermediary medium between the speech recognition system and the user. When speech feedback alone is insufficient for user understanding, the system automatically generates and displays text representations of the recognized speech and corresponding responses, providing an additional communication channel that bridges the gap between auditory feedback and user comprehension

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements a dual-feedback mechanism where both speech and text feedback are provided to the user. The text feedback serves as a visual confirmation of what was heard, allowing users to verify the accuracy of speech recognition and system responses, thereby improving reliability without compromising the hands-free operation capability

Inventive Principle:
Principle #23Feedback

2Ease of operation

If the system provides speech feedback for speech recognition results, then ease of operation is improved, but loss of information increases when the feedback does not match user intent

Engineering Contradiction:
Improvespeech-based interactionVSAvoiduser intent communication accuracy
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

Text serves as an intermediary that preserves and clarifies information that may be lost or misunderstood in speech feedback alone. The text representation provides a stable, visible record of both the user's spoken intent and the system's interpreted response, reducing information loss and ensuring accurate communication of user intent

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary processing and verification of speech recognition results before generating feedback. By converting speech to text and analyzing the text representation, the system can verify understanding accuracy and generate more precise speech feedback that better matches user intent, preventing information loss before it occurs

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12586582B2Apparatus performing based on voice recognition and artificial intelligence and method for controlling thereof
Publication Date: 2026.03.24 SAMSUNG ELECTRONICS CO LTD
  • US12586582B2 patent drawing
  • US12586582B2 patent drawing
  • US12586582B2 patent drawing

AI summary

An electronic device includes: a microphone; a speaker; a memory configured to store parameter information; and a processor configured to: perform speech recognition of a user's speech received by the microphone, control the speaker to output a first response message based on the speech recognition of the user's speech, detect a re-request for the first response message, recognize a text of the first response message, determine a second response message and a first speech signal based on a first parameter corresponding to the text, and generate the second response message comprising the determined first speech signal.