Learning Apparatus Using Soft Utterance Voice Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge is improving the accuracy of a model for discriminating between whisper utterance voice and normal utterance voice, where the scarcity of whisper utterance voice data leads to imbalanced learning data, making it difficult to enhance model performance.
Innovation Solution
Incorporating soft utterance voice data with intermediate labels between whisper and normal utterance voice into the learning process to balance the data set, allowing the model to better distinguish between the two.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If actual whisper utterance voice data is collected for learning, then discrimination accuracy can be improved, but the amount of available learning data is insufficient
Solution Approach 1:
The patent introduces soft utterance voice data as an intermediary class between whisper utterance voice and normal utterance voice. This intermediate class serves as a bridge that helps the model understand the transition between the two extreme classes, thereby improving discrimination accuracy while using more abundant soft utterance data to compensate for the scarcity of whisper utterance data.
Solution Approach 2:
The patent changes the parameter distribution of learning data by incorporating soft utterance voice data with intermediate characteristics. This parameter change expands the data distribution from only extreme cases (whisper and normal) to include intermediate cases, enabling the model to learn more nuanced discrimination patterns.
2Device complexity
If only whisper utterance voice and normal utterance voice data are used for learning, then the learning process is simple, but discrimination accuracy cannot be sufficiently improved due to data imbalance
Solution Approach 1:
Soft utterance voice acts as an intermediary class that simplifies the learning process by providing a gradual transition between whisper and normal utterance. Instead of directly learning to distinguish between two extreme classes with imbalanced data, the model learns through an intermediate stage, making the learning process more effective while maintaining reasonable complexity.
Solution Approach 2:
The patent segments the classification task into two stages: first learning to distinguish whisper from soft utterance, then learning to distinguish soft utterance from normal utterance. This segmentation of the learning task breaks down the complex imbalanced classification problem into more manageable sub-tasks.
Data Source
AI summary
A learning device includes a learning unit learning, with a first feature value having a first feature and given a first value label, a second feature value having a second feature and given a second value label and a third feature value having a feature between the first feature and the second feature and given a value label having a value between the first value label and the second value label as teacher data, a model for estimating which of the first feature and the second feature an input feature value sequence has.


