Keyword Detection System Using Two-Stage Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional word detection systems face issues with erroneous keyword detection and delayed operations due to time lags in detecting keywords, particularly when multiple keywords with similar pronunciations are involved, leading to incorrect detection of similar words.
Innovation Solution
A spoken keyword detection system that calculates a second detection score based on the start and end times of detected keywords, combined with frame scores, to accurately distinguish between keywords, reducing erroneous detections and improving response times.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional word detection systems use simple keyword matching, then detection speed is fast, but detection accuracy deteriorates leading to erroneous detection of similar keywords
Solution Approach 1:
The patent segments the keyword detection process into two distinct phases: a first score calculation that performs rapid initial detection, and a second score calculation that performs accurate verification. This segmentation allows the system to maintain fast response times while improving detection accuracy by only applying the more computationally intensive verification process when necessary.
Solution Approach 2:
The patent implements preliminary action by performing the first score calculation as a quick preliminary detection step before committing to the more time-consuming second score calculation. This preliminary filtering approach reduces the overall detection time lag by quickly eliminating non-matching keywords before applying the more accurate but slower verification process.
2Measurement precision
If the system performs detailed verification to improve detection accuracy, then keyword identification precision improves, but operation delay increases
Solution Approach 1:
The verification process is segmented into two stages: a fast first score calculation for initial screening, and a detailed second score calculation for precise verification. This segmentation ensures that detailed verification is only performed on candidates that pass the initial screen, maintaining high precision while minimizing overall processing time.
Solution Approach 2:
The patent applies partial verification by performing the computationally intensive second score calculation only on keywords that meet a threshold from the first score calculation, rather than verifying all detected keywords. This partial action approach maintains high identification precision for relevant keywords while preserving operational productivity.
3Reliability
If the system detects all potential keywords, then detection completeness improves, but erroneous detection of similar words increases
Solution Approach 1:
The patent segments the detection process into broad initial detection using the first score calculation followed by precise discrimination using the second score calculation. This segmentation allows the system to maintain detection completeness by initially capturing all potential keywords, then improving distinction accuracy through the verification stage that differentiates between similar keywords.
Solution Approach 2:
The system uses feedback from the first score calculation to guide the second score calculation process. Keywords that meet a certain threshold in the first calculation are fed into the second calculation for verification, creating a feedback mechanism that improves keyword distinction accuracy while maintaining detection completeness through the two-stage approach.
Data Source
AI summary
According to one embodiment, a word detection system acquires speech data including a plurality of frames, generates the speech characteristic amount, calculates a frame score by matching a reference model based on the speech characteristic amount associated with a target word with the frames in the speech data, calculates a first score of the word from the frame score, detects the word from the speech data based on the first score, calculates a second score of the word based on time information on the start and the end of the detected word and the frame score, compares the value of the second score with the second scores of a plurality of words, and determines a word to be output based on the comparison result.


