Voice Keyword Detection Using Sub-Keyword Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice keyword detection methods face challenges in accurately distinguishing between keywords with similar pronunciation strings, leading to false detection of multiple keywords when using binary determination methods, especially when keywords like 'communication' and 'communicator' are involved.

Innovation Solution

The proposed voice keyword detection apparatus employs a sub-keyword approach by dividing keywords into sub-keywords and determining the acceptance of these sub-keywords based on their start and end times, using a composite keyword model that includes phoneme sequences, phonological representations, and pronunciation notations to correctly identify the intended keyword.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If binary determination method is used for keyword detection, then detection speed is improved, but detection accuracy deteriorates due to false detection of multiple keywords with similar pronunciation

Engineering Contradiction:
Improvedetection speedVSAvoiddetection accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent divides keywords into sub-keywords and processes them in sequential order. By segmenting the keyword detection into multiple stages (first sub-keyword detection, then second sub-keyword detection), the system maintains fast detection speed while improving accuracy through time-based differentiation of similar pronunciation strings.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If sub-keyword division and time-based determination is implemented, then detection accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvedetection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides keywords into sub-keywords and processes them in sequential order. By segmenting the keyword detection into multiple stages (first sub-keyword detection, then second sub-keyword detection), the system maintains fast detection speed while improving accuracy through time-based differentiation of similar pronunciation strings.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent dynamically adjusts the detection process based on time information. The system determines whether to accept a detected sub-keyword by comparing its detection time with the expected time range, allowing flexible adaptation to different keyword structures while maintaining a relatively simple overall system architecture.

Inventive Principle:
Principle #15Dynamics

3Reliability

If time-based acceptance determination is used for sub-keywords, then false positives are reduced, but processing time increases

Engineering Contradiction:
Improvedetection reliabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies time-based acceptance determination selectively - not all detected sub-keywords undergo full time verification. The system determines acceptance based on whether the detected sub-keyword time falls within an expected range, applying partial verification that reduces processing overhead while maintaining high reliability in distinguishing similar keywords.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10553206B2Voice keyword detection apparatus and voice keyword detection method
Publication Date: 2020.02.04 KK TOSHIBA
  • US10553206B2 patent drawing
  • US10553206B2 patent drawing
  • US10553206B2 patent drawing

AI summary

According to one embodiment, a voice keyword detection apparatus includes a memory and a circuit coupled with the memory. The circuit calculates a first score for a first sub-keyword and a second score for a second sub-keyword. The circuit detects the first and second sub-keywords based on the first and second scores. The circuit determines, when the first sub-keyword is detected from one or more first frames, to accept the first sub-keyword. The circuit determines, when the second sub-keyword is detected from one or more second frames, whether to accept the second sub-keyword based on a start time and/or an end time of the one or more first frames and a start time and/or an end time of the one or more second frames.