Electronic Device Speech Processing for Adaptive Dictation and Conversation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Electronic devices struggle to consistently process user utterances effectively across different modes, such as dictation and conversation, due to difficulties in selecting the appropriate processing method based on the device's state or user input.

Innovation Solution

The electronic device includes a processor and memory configured to execute operations that involve receiving user inputs through a button and microphone, processing them with automatic speech recognition (ASR) and intelligence systems, and selectively sending data to an external server for natural language understanding (NLU) or text generation based on the device's state and user interface display.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the device uses a single speech processing method, then the processing logic is simple, but the device cannot adapt to different user utterance modes (dictation vs. conversation)

Engineering Contradiction:
Improveadaptability to different user utterance modesVSAvoidspeech processing logic complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic speech processing by detecting whether a text input interface is displayed on the touchscreen. When the interface is displayed, the device activates dictation mode; when not displayed, it activates conversation mode. This dynamic adaptation allows the device to switch processing methods based on real-time operational context, resolving the contradiction between adaptability and complexity.

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If the device always performs natural language understanding, then conversation capability is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveconversation capabilityVSAvoidspeech processing time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent applies partial NLU action by performing natural language understanding only when necessary (in conversation mode when the text input interface is not displayed). In dictation mode (when the text input interface is displayed), the device skips NLU and directly processes the speech-to-text conversion. This selective application of NLU reduces unnecessary processing time and computational overhead while maintaining conversation capability when needed.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12354605B2Electronic device for processing user speech and operating method therefor
Publication Date: 2025.07.08 SAMSUNG ELECTRONICS CO LTD
  • US12354605B2 patent drawing
  • US12354605B2 patent drawing
  • US12354605B2 patent drawing

AI summary

An example electronic device includes a housing; a touchscreen display; a microphone; at least one speaker; a button disposed on a portion of the housing or set to be displayed on the touchscreen display; a wireless communication circuit; a processor; and a memory. When a user interface is not displayed on the touchscreen display, the electronic device enables a user to receive a user input through the button, receives user speech through the microphone, and then provides data on the user speech to an external server. An instruction for performing a task is received from the server. When the user interface is displayed on the touchscreen display, the electronic device enables the user to receive the user input through the button, receives user speech through the microphone, and then provides data on the user speech to the external server.