Keyword Detection Model Adaptation Using Segmented Acoustic Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current keyword model training for smart speakers requires a large number of utterances from various speakers, making the recording and training process costly and inefficient.
Innovation Solution
An information processing apparatus that includes a data acquisition unit, a training unit, an extraction unit, and an adaptation processing unit, which acquires and processes voice feature data to train an acoustic model and adapt it into a keyword model using a limited volume of data, extracting specific keyword, sub-word, syllable, or phoneme data for efficient keyword detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a large number of utterances from various speakers are used for keyword model training, then the accuracy of keyword detection is improved, but the cost and complexity of data collection and training increase
Solution Approach 1:
The patent segments the training process into two distinct phases: (1) training a general acoustic model using a large corpus of speech data from multiple speakers, and (2) adapting this pre-trained acoustic model to a specific keyword model using only a small amount of keyword utterance data. This segmentation allows the system to benefit from large-scale data without requiring extensive keyword-specific data collection for each application
Solution Approach 2:
The patent applies preliminary action by pre-training the acoustic model on a large corpus of speech data before adapting it to the specific keyword detection task. This pre-training establishes a robust foundation that captures general speech patterns, which then facilitates efficient adaptation to keyword-specific patterns using minimal data
2Reliability
If extensive keyword utterance data from multiple speakers is collected, then the keyword model performance is improved, but the time and resources required for data collection increase
Solution Approach 1:
The system performs preliminary training of a general acoustic model on a large speech corpus before the actual keyword model training. This preliminary action creates a robust foundation that significantly reduces the amount of keyword-specific data needed, thereby reducing data collection time while maintaining model performance
Solution Approach 2:
The patent effectively copies the general speech patterns learned from the large corpus into the keyword model through the adaptation process. The pre-trained acoustic model serves as a template that is then specialized for keyword detection, avoiding the need to collect and process extensive keyword-specific data from scratch
3Reliability
If a large volume of training data is used, then the acoustic model becomes more robust, but the computational resources and training time required increase
Solution Approach 1:
The training process is segmented into two stages with different data requirements: (1) training a general acoustic model on a large corpus to achieve robustness, and (2) adapting to a specific keyword model on a small dataset for efficiency. This segmentation allows the system to achieve both robustness and training efficiency by matching data volume to task requirements
Solution Approach 2:
The preliminary training of the acoustic model on a large corpus establishes robust speech pattern recognition capabilities. This preliminary action creates a foundation that requires minimal additional training data for keyword-specific adaptation, thereby improving overall training efficiency while maintaining model robustness
Data Source
AI summary
According to one embodiment, an information processing apparatus includes following units. The acquisition unit acquires first training data including a combination of a voice feature quantity and a correct phoneme label of the voice feature quantity. The training unit trains an acoustic model using the first training data in a manner to output the correct phoneme label in response to input of the voice feature quantity. The extraction unit extracts from the first training data, second training data including voice feature quantities of at least one of a keyword, a sub-word, a syllable, or a phoneme included in the keyword. The adaptation processing unit adapts the trained acoustic model using the second training data to a keyword detection model.


