Learning Apparatus Using Soft Utterance Voice Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The challenge is improving the accuracy of a model for discriminating between whisper utterance voice and normal utterance voice, where the scarcity of whisper utterance voice data leads to imbalanced learning data, making it difficult to enhance model performance.

Innovation Solution

Incorporating soft utterance voice data with intermediate labels between whisper and normal utterance voice into the learning process to balance the data set, allowing the model to better distinguish between the two.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If actual whisper utterance voice data is collected for learning, then discrimination accuracy can be improved, but the amount of available learning data is insufficient

Engineering Contradiction:
Improvediscrimination accuracyVSAvoidamount of learning data
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent introduces soft utterance voice data as an intermediary class between whisper utterance voice and normal utterance voice. This intermediate class serves as a bridge that helps the model understand the transition between the two extreme classes, thereby improving discrimination accuracy while using more abundant soft utterance data to compensate for the scarcity of whisper utterance data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter distribution of learning data by incorporating soft utterance voice data with intermediate characteristics. This parameter change expands the data distribution from only extreme cases (whisper and normal) to include intermediate cases, enabling the model to learn more nuanced discrimination patterns.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If only whisper utterance voice and normal utterance voice data are used for learning, then the learning process is simple, but discrimination accuracy cannot be sufficiently improved due to data imbalance

Engineering Contradiction:
Improvelearning process complexityVSAvoiddiscrimination accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

Soft utterance voice acts as an intermediary class that simplifies the learning process by providing a gradual transition between whisper and normal utterance. Instead of directly learning to distinguish between two extreme classes with imbalanced data, the model learns through an intermediate stage, making the learning process more effective while maintaining reasonable complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the classification task into two stages: first learning to distinguish whisper from soft utterance, then learning to distinguish soft utterance from normal utterance. This segmentation of the learning task breaks down the complex imbalanced classification problem into more manageable sub-tasks.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12125474B2Learning apparatus, estimation apparatus, methods and programs for the same
Publication Date: 2024.10.22 NIPPON TELEGRAPH & TELEPHONE CORP
  • US12125474B2 patent drawing
  • US12125474B2 patent drawing
  • US12125474B2 patent drawing

AI summary

A learning device includes a learning unit learning, with a first feature value having a first feature and given a first value label, a second feature value having a second feature and given a second value label and a third feature value having a feature between the first feature and the second feature and given a value label having a value between the first value label and the second value label as teacher data, a model for estimating which of the first feature and the second feature an input feature value sequence has.