Speech Skill Awakening via Dual Semantic Model Confidence Comparison

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional smart devices struggle to accurately distinguish between waking up a service skill and a knowledge skill based on user speech, often resulting in incorrect skill invocation.

Innovation Solution

A method and apparatus that recognize awakening text information, invoke both service skill and knowledge skill semantic models in parallel to determine confidences, and select the appropriate skill to awaken based on these confidences, reducing the likelihood of incorrect skill activation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional smart devices use single-model skill recognition, then the device complexity is low, but the skill recognition accuracy deteriorates

Engineering Contradiction:
Improveskill recognition accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines service skill semantic models and knowledge skill semantic models into a unified recognition system. Both models process user speech inputs simultaneously and their outputs are integrated through confidence score comparison, enabling the system to distinguish between service skills (e.g., music playback) and knowledge skills (e.g., factual Q&A) with high accuracy while maintaining manageable system architecture.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent segments the skill recognition function into separate service skill semantic models and knowledge skill semantic models. Each model is specialized for its respective skill type, allowing independent optimization and processing. The segmentation enables parallel execution and confident score comparison, improving overall recognition accuracy without requiring a single monolithic complex model.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If the system invokes both service skill and knowledge skill models in parallel, then the skill awakening accuracy is improved, but the processing time increases

Engineering Contradiction:
Improveskill awakening accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-loading and maintaining both service skill and knowledge skill semantic models in ready state. The models are prepared in advance with their respective skill domains and confidence scoring mechanisms, enabling immediate parallel processing when user speech is detected. This preliminary preparation eliminates model loading delays during actual skill recognition.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs partial action by processing only the necessary portions of user speech inputs through both models simultaneously. Instead of analyzing all possible skill categories exhaustively, the system focuses on the two primary skill types (service and knowledge) with confidence score thresholds, achieving accurate skill awakening without unnecessary processing overhead.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If the system uses confidence scores from multiple models, then the skill selection accuracy is improved, but the information processing complexity increases

Engineering Contradiction:
Improveskill selection accuracyVSAvoidinformation processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent changes the parameter approach by using confidence scores as a simple numerical metric to compare and select between service and knowledge skills. Instead of complex multi-dimensional analysis, the system transforms the comparison into a straightforward parameter evaluation (confidence score magnitude), reducing information processing complexity while maintaining high skill selection accuracy through statistical confidence metrics.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11721328B2Method and apparatus for awakening skills by speech
Publication Date: 2023.08.08 AISPEECH CO LTD
  • US11721328B2 patent drawing
  • US11721328B2 patent drawing
  • US11721328B2 patent drawing

AI summary

The present invention discloses a method and apparatus for awakening skills by speech, which are applied to an electronic device. The method for awakening skills by speech includes: recognizing awakening text information corresponding to a speech request message to be processed; invoking a service skill semantic model to determine a target service field corresponding to the awakening text information and a corresponding first confidence, and invoking a knowledge skill semantic model to determine a knowledge reply answer corresponding to the awakening text information and a corresponding second confidence; and selecting to awaken one of a knowledge skill and a target service skill corresponding to the target service field based on the first confidence and the second confidence. Accordingly, the probability of erroneously awakening a skill based on the speech message can be reduced.