Intelligent Speech Processing System for Multi-Intent Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition services fail to accurately process user inputs, especially when insufficient information is provided, leading to difficulties in grasping the user's intent and failing to execute multiple application requests, as they only display results corresponding to voice inputs without processing the intent behind the utterance.

Innovation Solution

An electronic apparatus and method that uses a rule-based system to process user utterances by arranging states corresponding to the user input, involving a processor, memory, and communication circuitry to transmit data to an external server, receive responses, and display a user interface for further input, allowing for the execution of tasks and filling text input boxes with parameters, thereby enhancing the recognition of user intents and providing a more accurate service.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional speech recognition service displays only the result corresponding to voice input, then the service is simple to operate, but it fails to process user inputs requiring multiple application executions or sufficient information provision

Engineering Contradiction:
Improvecapability to process multiple application requestsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the speech recognition process into distinct stages: initial voice input reception, result display with input method selection options, and follow-up input reception. This segmentation allows the system to handle complex multi-application requests by breaking them down into manageable interaction steps, where users can select different input methods (voice, text, keyboard) at different stages of the task execution流程

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adapts its behavior based on the type of user input received. When a user requests multiple application executions or provides insufficient information, the system transitions from a simple result-display mode to an interactive mode that presents multiple input method options, allowing the service complexity to adjust according to task requirements

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If conventional speech recognition service processes only voice input, then the input method is simple, but it fails when user voice does not provide sufficient information

Engineering Contradiction:
Improveaccuracy of user intent recognitionVSAvoidinput method complexity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent implements a universal input handling framework that supports multiple input methods (voice recognition, text input, keyboard input) within a single speech recognition service. This multi-functionality allows the system to accurately recognize user intent regardless of whether the initial voice input was sufficient, by allowing users to supplement or correct input through alternative methods

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system provides feedback to users when voice input is insufficient for accurate intent recognition. After displaying the initial recognition result, the system offers feedback options that allow users to select alternative input methods or provide additional information, creating a feedback loop that improves recognition accuracy through iterative input refinement

Inventive Principle:
Principle #23Feedback

3Productivity

If conventional speech recognition service executes only single tasks, then the task execution is straightforward, but it cannot handle requests to execute multiple applications

Engineering Contradiction:
Improvetask execution efficiencyVSAvoidtask management complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements preliminary action by allowing users to declare multiple task execution requests in advance through a single voice input or text input. The system parses and stores these multiple task requests, then executes them in the predetermined sequence without requiring separate interaction steps for each task, thereby improving productivity while managing complexity through upfront task specification

Inventive Principle:
Principle #10Preliminary action

4Reliability

If conventional speech recognition service provides only basic result display, then the interface is simple, but it cannot provide service that matches complex user intent

Engineering Contradiction:
Improveservice accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary processing layer between voice input and result execution. This intermediary component analyzes the user input, determines the appropriate input method, processes the input through the selected method, and coordinates multiple tasks if needed. This intermediary structure improves service reliability by ensuring accurate intent recognition while managing processing complexity through a dedicated coordination layer

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10908763B2Electronic apparatus for processing user utterance and controlling method thereof
Publication Date: 2021.02.02 SAMSUNG ELECTRONICS CO LTD
  • US10908763B2 patent drawing
  • US10908763B2 patent drawing
  • US10908763B2 patent drawing

AI summary

A system, apparatus, and method for generating one or more path rules or combination or path rules for performing the operations of an application. According to various embodiments, a user terminal through an intelligence server may generate a path rule and may perform the operation of an application based on the path rule to provide a service. The user terminal may provide a user with the processing status of a user utterance and may receive an additional user input to process the user utterance. For example, when a keyword (or parameter) required to process the user utterance is insufficient, the user terminal may provide the user with a state requiring additional information and may output feedback corresponding to the insufficient state in order to receive the necessary information from the user.