Voice Processing App Identification via ASR and NLU Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In AI systems, there is a need to determine the appropriate application to process a user's request without explicit designation, especially when multiple apps can fulfill the request, requiring a method to identify the capable app based on voice input.
Innovation Solution
An electronic device with a network interface and processor executes automatic speech recognition (ASR) to extract text from voice input, identifies the app if possible, and uses natural language understanding (NLU) to reattempt identification if not possible, determining the app's capability and selecting it based on user preferences or history.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple applications are capable of processing a user request without explicit designation, then the system needs to implement complex identification logic using ASR and NLU to determine the appropriate app, but this increases the device complexity and processing time
Solution Approach 1:
The patent segments the app identification process into two distinct stages: first, Automatic Speech Recognition (ASR) converts voice input to text and performs initial app identification; second, if the first stage fails to identify the app, Natural Language Understanding (NLU) is activated to reattempt identification. This segmentation allows the system to handle complex identification tasks through modular, sequential processing rather than a single monolithic system.
Solution Approach 2:
The system performs preliminary app identification using ASR before attempting more complex NLU processing. By first extracting text from voice input and attempting simple keyword matching or entity recognition, the system can resolve many identification cases without needing to activate the full NLU pipeline, thus reducing overall processing complexity while maintaining high adaptability.
2Measurement precision
If the system uses both ASR and NLU to identify applications when not explicitly designated, then the app identification accuracy is improved, but the processing time and computational resources increase
Solution Approach 1:
The system applies partial action by using ASR for initial app identification, which handles many common cases quickly. Only when the ASR stage fails to identify the app does the system activate the more resource-intensive NLU process. This approach achieves high identification accuracy by applying the full processing power of NLU only when necessary, rather than always, thus reducing average processing time while maintaining precision.
Solution Approach 2:
The system implements a feedback mechanism where the result of ASR processing determines whether NLU needs to be activated. If ASR successfully identifies the app, the process terminates; if not, the feedback triggers NLU to reattempt identification. This feedback loop ensures high accuracy by systematically escalating processing depth only when needed, optimizing the balance between accuracy and processing time.
3Ease of operation
If the system attempts to identify the app through ASR first and then NLU if needed, then the ease of operation is improved by allowing implicit app selection, but the device complexity increases due to multiple processing layers
Solution Approach 1:
The patent implements a universal voice processing architecture that can handle both explicit app designation and implicit app selection through the same ASR and NLU infrastructure. The system is designed to accommodate multiple interaction modes (explicit naming, implicit reference, contextual selection) within a single multi-functional framework, allowing ease of operation across different user preferences while managing complexity through standardized processing components.
Data Source
AI summary
An electronic device and method are disclosed herein. The electronic device includes a network interface and processor. The processor implements the method, including receiving a voice input through a network interface as transmitted from a first external device, including a request to execute a function using at least one application which is not indicated in the voice input, extracting a first text from the voice input by executing automatic speech recognition (ASR), when the at least one application is identified based on the first text, transmitting, through the network interface to the first external device, second data associated with the identified at least one application for display by the first external device, and when the at least one application is not identified based at least in part on the first text, reattempting identification of the at least one application by executing natural language understanding (NLU) on the first text.


