Acoustic Model Training With Selective Waveform Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing acoustic model training requires significant time and cost due to the need for labeling vast amounts of voice and performance sounds, and is often compromised by noise or unnecessary sounds in the training data, limiting the types of models that can be trained.

Innovation Solution

A system and method that allows users to select specific sound waveforms for training, enabling efficient training of acoustic models by selecting desired portions of data and using a client-server architecture for training jobs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If vast amounts of voice and performance sounds are used for training, then the acoustic model can be sufficiently trained, but it requires an immense amount of time and cost for labeling

Engineering Contradiction:
Improvetraining qualityVSAvoidlabeling time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the training process into multiple stages: pre-training with large unlabeled datasets, fine-tuning with smaller labeled datasets, and selective training with user-chosen portions. This segmentation allows the system to benefit from large-scale data without requiring proportional labeling effort at each stage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary action by performing pre-training on large datasets before the actual training phase. This preliminary training establishes a baseline model that can be subsequently refined with smaller, more focused datasets, reducing the overall labeling time required.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If vast amounts of voice and performance sounds are used for training, then the acoustic model can be sufficiently trained, but it requires an immense amount of cost for labeling

Engineering Contradiction:
Improvetraining qualityVSAvoidlabeling cost
Core Design Contradiction:
Manufacturing precisionVSLoss of energy

Solution Approach 1:

The training process is divided into phases where different data quantities and labeling intensities are applied. Pre-training uses abundant unlabeled data, while fine-tuning uses smaller labeled portions, optimizing the cost-quality tradeoff.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs self-supervised learning mechanisms where the model learns from unlabeled data through self-verification and consistency checks, reducing dependency on expensive manual labeling while maintaining training quality.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If all sound waveforms are used for training, then comprehensive training is achieved, but noise and unnecessary sounds reduce training quality

Engineering Contradiction:
Improvetraining data coverageVSAvoidtraining quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent extracts and removes noise and unnecessary sounds from the training data through preprocessing filters and user-selectable portion mechanisms. This extraction process retains only the relevant sound portions while discarding harmful elements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies different quality standards to different portions of the training data. User-selectable portions allow high-quality relevant sounds to be prioritized while noise-containing segments are excluded or downweighted, creating local quality variations in the training set.

Inventive Principle:
Principle #3Local quality

4Manufacturing precision

If only companies with sufficient funds can train acoustic models, then high-quality training is possible, but the types of acoustic models are limited

Engineering Contradiction:
Improvetraining qualityVSAvoidmodel diversity
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal training framework that can be applied by various users with different resource levels. The system supports multiple training modes (full training, selective training, fine-tuning) that can be adapted to different budget constraints while maintaining reasonable training quality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system allows parameter changes in the training configuration, such as dataset size, labeling intensity, and training duration, enabling users to adjust the training process according to their available resources while still producing functional acoustic models.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250232761A1Training system and method for acoustic model
Publication Date: 2025.07.17 YAMAHA CORP
  • US20250232761A1 patent drawing
  • US20250232761A1 patent drawing
  • US20250232761A1 patent drawing

AI summary

An acoustic model training system includes a first device that is connectable to a network and that is used by a first user, and a server that is connectable to the network. The first device, under control by the first user, is configured to upload a plurality of sound waveforms to the server, select, as a first waveform set, one or more sound waveforms from the plurality of sound waveforms after or before updating the plurality of sound waveforms, and transmit to the server a first execution instruction for a first training job for an acoustic model configured to generate acoustic features. The server is configured to, based on the first execution instruction from the first device, start execution of the first training job using the first waveform set, and provide, to the first device, a trained acoustic model trained by the first training job.