Voice Recognition Intent Segmentation for Memory Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems face challenges in clearly determining user intent from ambiguous utterances, leading to increased memory requirements to accommodate a wide range of scenarios.
Innovation Solution
A voice recognition apparatus and method that utilize a processor to extract intents from user utterances, separate them into partial intent units using predetermined separators, and generate a final intent by combining and deleting duplicate partial intents, thereby reducing memory capacity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the voice recognition apparatus defines scenarios to match a wide range of utterances, then the recognition coverage is improved, but the memory capacity requirement increases
Solution Approach 1:
The patent segments user utterances into distinct intent components (action, target, entity) using predefined separators. This segmentation allows the system to process and match only relevant parts of utterances against scenario databases, rather than storing and comparing complete scenario definitions for every possible utterance variation, thereby reducing memory capacity requirements while maintaining wide recognition coverage.
Solution Approach 2:
The patent extracts key intent elements (action, target, entity) from complete utterances by removing unnecessary words and structures using predefined separators. This extraction process creates condensed intent representations that require less memory storage while still enabling comprehensive scenario matching, thus resolving the contradiction between recognition coverage and memory capacity.
2Measurement precision
If the voice recognition apparatus uses AI to improve voice recognition, then the recognition accuracy is improved, but the difficulty of detecting and measuring user intent from ambiguous utterances increases
Solution Approach 1:
The patent applies preliminary action by pre-defining separators and intent extraction rules before processing user utterances. This preliminary setup enables the system to automatically and consistently extract intent components from ambiguous utterances without requiring complex real-time AI analysis, thereby maintaining high recognition accuracy while reducing the difficulty of intent detection.
Solution Approach 2:
The patent uses copying by creating standardized intent templates and scenario patterns that can be repeatedly applied to different utterances. These pre-defined templates capture common intent structures, allowing the system to accurately determine user intent from ambiguous utterances by matching them against established patterns rather than performing complex analysis each time.
Data Source
AI summary
In embodiments, a voice recognition apparatus, and a method thereof, includes a microphone that extracts an utterance of a user, a memory that stores a scenario matching intent extracted from the utterance, and a processor that searches for the scenario based on the utterance and performs a voice recognition function. The processor can extract a first intent from a first utterance and extract a second intent from a second utterance. The processor can separate the first intent and the second intent into partial intent units by using separators, and generate a final intent by combining partial intents of the first intent and the second intent such that duplicate partial intents are deleted depending on definitions of the separators.


