Keyword Detection Model Adaptation Using Segmented Acoustic Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current keyword model training for smart speakers requires a large number of utterances from various speakers, making the recording and training process costly and inefficient.

Innovation Solution

An information processing apparatus that includes a data acquisition unit, a training unit, an extraction unit, and an adaptation processing unit, which acquires and processes voice feature data to train an acoustic model and adapt it into a keyword model using a limited volume of data, extracting specific keyword, sub-word, syllable, or phoneme data for efficient keyword detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a large number of utterances from various speakers are used for keyword model training, then the accuracy of keyword detection is improved, but the cost and complexity of data collection and training increase

Engineering Contradiction:
Improvekeyword detection accuracyVSAvoidtraining process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the training process into two distinct phases: (1) training a general acoustic model using a large corpus of speech data from multiple speakers, and (2) adapting this pre-trained acoustic model to a specific keyword model using only a small amount of keyword utterance data. This segmentation allows the system to benefit from large-scale data without requiring extensive keyword-specific data collection for each application

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by pre-training the acoustic model on a large corpus of speech data before adapting it to the specific keyword detection task. This pre-training establishes a robust foundation that captures general speech patterns, which then facilitates efficient adaptation to keyword-specific patterns using minimal data

Inventive Principle:
Principle #10Preliminary action

2Reliability

If extensive keyword utterance data from multiple speakers is collected, then the keyword model performance is improved, but the time and resources required for data collection increase

Engineering Contradiction:
Improvekeyword model performanceVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary training of a general acoustic model on a large speech corpus before the actual keyword model training. This preliminary action creates a robust foundation that significantly reduces the amount of keyword-specific data needed, thereby reducing data collection time while maintaining model performance

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent effectively copies the general speech patterns learned from the large corpus into the keyword model through the adaptation process. The pre-trained acoustic model serves as a template that is then specialized for keyword detection, avoiding the need to collect and process extensive keyword-specific data from scratch

Inventive Principle:
Principle #26Copying

3Reliability

If a large volume of training data is used, then the acoustic model becomes more robust, but the computational resources and training time required increase

Engineering Contradiction:
Improveacoustic model robustnessVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The training process is segmented into two stages with different data requirements: (1) training a general acoustic model on a large corpus to achieve robustness, and (2) adapting to a specific keyword model on a small dataset for efficiency. This segmentation allows the system to achieve both robustness and training efficiency by matching data volume to task requirements

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The preliminary training of the acoustic model on a large corpus establishes robust speech pattern recognition capabilities. This preliminary action creates a foundation that requires minimal additional training data for keyword-specific adaptation, thereby improving overall training efficiency while maintaining model robustness

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11961510B2Information processing apparatus, keyword detecting apparatus, and information processing method
Publication Date: 2024.04.16 KK TOSHIBA
  • US11961510B2 patent drawing
  • US11961510B2 patent drawing
  • US11961510B2 patent drawing

AI summary

According to one embodiment, an information processing apparatus includes following units. The acquisition unit acquires first training data including a combination of a voice feature quantity and a correct phoneme label of the voice feature quantity. The training unit trains an acoustic model using the first training data in a manner to output the correct phoneme label in response to input of the voice feature quantity. The extraction unit extracts from the first training data, second training data including voice feature quantities of at least one of a keyword, a sub-word, a syllable, or a phoneme included in the keyword. The adaptation processing unit adapts the trained acoustic model using the second training data to a keyword detection model.