Voice Processing System for Multi-App Control via State Sequences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition services are limited in processing user voice inputs, unable to control multiple application programs organically and struggle to determine when to execute new applications, especially third-party apps, leading to difficulties in managing user requests effectively.

Innovation Solution

A system that includes a network interface, processor, and memory, capable of receiving user utterances, determining user intent, and executing corresponding actions across multiple applications, including third-party apps, by using natural language understanding units and automatic speech recognition to generate and provide appropriate sequences of states for electronic devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech recognition service processes only simple user voice inputs, then processing simplicity is maintained, but the ability to control multiple applications organically is lost

Engineering Contradiction:
Improveability to control multiple applicationsVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The speech recognition service divides the complex task of controlling multiple applications into separate processing modules: intent recognition module, application selection module, and state sequence generation module. Each module handles a specific aspect of the complex processing independently, making the overall system manageable while achieving high versatility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer between the simple voice input and the complex application control. This intermediary layer includes intent recognition and state sequence generation components that translate simple user commands into coordinated multi-application actions, bridging the gap between simplicity and complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If speech recognition service determines whether to execute new apps, then application execution control is improved, but difficulty in determining execution timing increases

Engineering Contradiction:
Improveapplication execution controlVSAvoidexecution timing determination
Core Design Contradiction:
Ease of operationVSDifficulty of detecting and measuring

Solution Approach 1:

The system continuously monitors the execution state of applications and uses this feedback to determine when to launch new apps. The state recognition module detects current application states and provides feedback to the decision-making module, which automatically determines the optimal timing for executing new applications based on real-time system conditions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent implements preliminary preparation of application execution sequences before user commands are fully processed. The system pre-determines potential application execution paths and prepares necessary resources in advance, so when execution timing is needed, the system can quickly activate the appropriate applications without delay or uncertainty.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If speech recognition service controls third party apps, then service coverage is expanded, but control capability over third party apps becomes insufficient

Engineering Contradiction:
Improveservice coverageVSAvoidcontrol capability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent creates a universal control framework that can manage both native and third-party applications through a common interface. The state recognition and intent execution modules are designed to work with any application type, providing consistent control capabilities across diverse application ecosystems without requiring application-specific specialized handling.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10811008B2Electronic apparatus for processing user utterance and server
Publication Date: 2020.10.20 SAMSUNG ELECTRONICS CO LTD
  • US10811008B2 patent drawing
  • US10811008B2 patent drawing
  • US10811008B2 patent drawing

AI summary

A system for processing a user utterance is provided. The system includes at least one network interface; at least one processor operatively connected to the at least one network interface; and at least one memory operatively connected to the at least one processor, wherein the at least one memory stores a plurality of specified sequences of states of at least one external electronic device, wherein each of the specified sequences is associated with a respective one of domains, wherein the at least one memory further stores instructions that, when executed, cause the at least one processor to receive first data associated with the user utterance provided via a first of the at least one external electronic device, wherein the user utterance includes a request for performing a task using the first of the at least one external device.