Speech Skill Awakening via Dual Semantic Model Confidence Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional smart devices struggle to accurately distinguish between waking up a service skill and a knowledge skill based on user speech, often resulting in incorrect skill invocation.
Innovation Solution
A method and apparatus that recognize awakening text information, invoke both service skill and knowledge skill semantic models in parallel to determine confidences, and select the appropriate skill to awaken based on these confidences, reducing the likelihood of incorrect skill activation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional smart devices use single-model skill recognition, then the device complexity is low, but the skill recognition accuracy deteriorates
Solution Approach 1:
The patent combines service skill semantic models and knowledge skill semantic models into a unified recognition system. Both models process user speech inputs simultaneously and their outputs are integrated through confidence score comparison, enabling the system to distinguish between service skills (e.g., music playback) and knowledge skills (e.g., factual Q&A) with high accuracy while maintaining manageable system architecture.
Solution Approach 2:
The patent segments the skill recognition function into separate service skill semantic models and knowledge skill semantic models. Each model is specialized for its respective skill type, allowing independent optimization and processing. The segmentation enables parallel execution and confident score comparison, improving overall recognition accuracy without requiring a single monolithic complex model.
2Measurement precision
If the system invokes both service skill and knowledge skill models in parallel, then the skill awakening accuracy is improved, but the processing time increases
Solution Approach 1:
The patent performs preliminary actions by pre-loading and maintaining both service skill and knowledge skill semantic models in ready state. The models are prepared in advance with their respective skill domains and confidence scoring mechanisms, enabling immediate parallel processing when user speech is detected. This preliminary preparation eliminates model loading delays during actual skill recognition.
Solution Approach 2:
The patent employs partial action by processing only the necessary portions of user speech inputs through both models simultaneously. Instead of analyzing all possible skill categories exhaustively, the system focuses on the two primary skill types (service and knowledge) with confidence score thresholds, achieving accurate skill awakening without unnecessary processing overhead.
3Measurement precision
If the system uses confidence scores from multiple models, then the skill selection accuracy is improved, but the information processing complexity increases
Solution Approach 1:
The patent changes the parameter approach by using confidence scores as a simple numerical metric to compare and select between service and knowledge skills. Instead of complex multi-dimensional analysis, the system transforms the comparison into a straightforward parameter evaluation (confidence score magnitude), reducing information processing complexity while maintaining high skill selection accuracy through statistical confidence metrics.
Data Source
AI summary
The present invention discloses a method and apparatus for awakening skills by speech, which are applied to an electronic device. The method for awakening skills by speech includes: recognizing awakening text information corresponding to a speech request message to be processed; invoking a service skill semantic model to determine a target service field corresponding to the awakening text information and a corresponding first confidence, and invoking a knowledge skill semantic model to determine a knowledge reply answer corresponding to the awakening text information and a corresponding second confidence; and selecting to awaken one of a knowledge skill and a target service skill corresponding to the target service field based on the first confidence and the second confidence. Accordingly, the probability of erroneously awakening a skill based on the speech message can be reduced.


