Voice Assistant Utterance Sequencing for Natural Interruptions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice assistants struggle to maintain a continuous and natural conversation with users due to user interruptions and disturbances, leading to robotic and unnatural responses.

Innovation Solution

A system and method that converts unsegmented voice inputs into textual format, classifies user inputs into predefined classes, calculates contextual pauses, and provides real-time intuitive mingling responses based on ongoing activity and user intent using an intelligent activity engine.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If the voice assistant responds immediately to user inputs, then the response time is reduced, but the conversation becomes robotic and unnatural when users are interrupted

Engineering Contradiction:
Improveresponse timeVSAvoidnatural conversation quality
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The system performs preliminary classification of user inputs into communication classes (instruction, request, command) and identifies contextual pauses before the user completes their input. This allows the voice assistant to prepare appropriate responses in advance while waiting for natural pause points, rather than responding immediately or waiting passively.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts its response timing based on detected contextual pauses and user activity state. The contextual sequencer unit sequences multiple potential responses based on the classified communication type and inserts appropriate pauses, making the conversation flow more naturally while maintaining responsiveness.

Inventive Principle:
Principle #15Dynamics

2Reliability

If the voice assistant waits for user completion, then the conversation becomes more natural, but the response time increases and user attention is lost

Engineering Contradiction:
Improvenatural conversation qualityVSAvoidconversation delay
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system classifies the user's input into communication classes (instruction, request, command) and prepares multiple contextual responses in advance using a virtual server. This preliminary preparation allows the voice assistant to respond quickly once a contextual pause is detected, rather than waiting passively for user completion.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The natural utterance analyzer unit continuously monitors the user's speech for contextual pauses and provides feedback to the contextual sequencer unit. This real-time feedback mechanism allows the system to interrupt its waiting state and deliver prepared responses at optimal moments, balancing naturalness with timeliness.

Inventive Principle:
Principle #23Feedback

3Reliability

If the voice assistant provides detailed contextual responses, then the conversation quality improves, but the complexity of processing increases

Engineering Contradiction:
Improvecontextual response qualityVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the complex task of generating contextual responses into distinct modules: communication classification unit for categorizing input type, natural utterance analyzer unit for detecting pauses, contextual sequencer unit for ordering responses, and intelligent activity engine for coordination. This segmentation allows each module to handle a specific aspect of the problem independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The virtual server acts as an intermediary that handles the computationally intensive task of generating contextual responses based on the classified communication type. The server receives the classification from the natural language understanding module and returns appropriate responses, which are then sequenced and delivered by the local system.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Device complexity

If the voice assistant loses context during interruptions, then the system remains simple, but the user must re-initiate conversations causing inconvenience

Engineering Contradiction:
Improvesystem simplicityVSAvoiduser convenience
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The natural language understanding module performs multiple functions: it converts speech to text, classifies the communication type (instruction, request, command), and identifies the user's intent. This multi-functionality allows the system to maintain context across interruptions without adding separate dedicated modules for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The contextual sequencer unit maintains a sequence of prepared responses that are delivered continuously as contextual pauses are detected. This ensures the conversation flow continues naturally even if the user is temporarily interrupted, as the system has already prepared the next relevant responses based on the classified input type.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12555569B2System to provide natural utterance by a voice assistant and method thereof
Publication Date: 2026.02.17 SAMSUNG ELECTRONICS CO LTD
  • US12555569B2 patent drawing
  • US12555569B2 patent drawing
  • US12555569B2 patent drawing

AI summary

Provided is a system to provide natural utterance by a voice assistant and method thereof, wherein the system comprises an automatic speech recognition module for converting one or more unsegmented voice inputs into textual format in real-time. Further, a natural language understanding module extracts the information and intent of the user from the converted textual inputs, wherein the natural language understanding module comprises a communication classification unit for classifying the user inputs into one or more pre-defined classes. Further, the system comprises a processing module for analyzing and processing the inputs from the natural language understanding module and activity identification module, wherein the processing module provides real-time intuitive mingling responses based on the responses, contextual pauses and ongoing activity of the user.