Learning Data Acquisition for Voice Recognition Noise Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice/non-voice models face challenges in accurately detecting voice sections in noisy environments due to the difficulty in preparing sufficient clean voice data with appropriate signal-to-noise ratios, leading to erroneous detections and reduced accuracy.

Innovation Solution

A learning data acquisition device that calculates the influence of signal-to-noise ratio changes on voice recognition accuracy and acquires noise superimposed voice data at an optimal SN ratio, ensuring accurate learning data for model construction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If noise is artificially superimposed on clean voice to generate learning data for noisy environments, then the model can learn voice data in noisy environments, but voice that cannot be expected in a real use scene may be generated causing erroneous characteristics to be learned

Engineering Contradiction:
Improvemodel adaptability to noisy environmentsVSAvoiddetection accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies parameter changes by systematically varying the signal-to-noise ratio (SNR) as a key parameter when superimposing noise on clean voice. Instead of using a fixed SNR, the patent changes the SNR parameter across different learning data samples to cover a range of realistic noisy conditions, thereby improving model adaptability while avoiding erroneous characteristics from unrealistic noise levels

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces dynamics by making the noise superimposition process adaptive rather than static. The noise characteristics and SNR are dynamically adjusted based on target specifications, allowing the learning data generation to adapt to different use scenarios and prevent generation of unrealistic voice characteristics

Inventive Principle:
Principle #15Dynamics

2Ease of manufacture

If learning data is generated with only good SN ratio conditions, then the model can be trained on clean voice data, but the accuracy cannot be improved for noisy environments

Engineering Contradiction:
Improveease of learning data preparationVSAvoiddetection accuracy in noisy environments
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-defining target specifications that include appropriate SNR ranges before generating learning data. This preliminary setup ensures that the learning data generation process automatically produces data suitable for noisy environments without requiring manual adjustment during training, thus maintaining ease of manufacture while improving detection accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the SNR parameter from a fixed good condition to a variable parameter that spans multiple levels including noisy conditions. This parameter change allows the model to learn from both clean and noisy voice data, improving detection accuracy in noisy environments while maintaining a systematic approach to data preparation

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If noise is superimposed with small SN ratio to simulate noisy environments, then the model can learn from noisy conditions, but voice of small whispering under noisy environment may be learned causing erroneous detection

Engineering Contradiction:
Improvemodel learning from noisy conditionsVSAvoiddetection precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies parameter changes by controlling the SNR parameter within appropriate ranges defined by target specifications. Instead of using very small SNR values that create unrealistic whispering conditions, the system adjusts the SNR parameter to reflect realistic noisy environments, thereby maintaining detection precision while improving adaptability to noise

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces feedback mechanisms through target specifications that guide the noise superimposition process. The target specifications provide feedback on appropriate SNR ranges, preventing the generation of learning data with unrealistic small SNR values that would cause erroneous detection of whispering voice in noisy environments

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11942074B2Learning data acquisition apparatus, model learning apparatus, methods and programs for the same
Publication Date: 2024.03.26 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11942074B2 patent drawing
  • US11942074B2 patent drawing

AI summary

A learning data acquisition device or the like, capable of acquiring learning data by superimposing noise data on clean voice data at an appropriate SN ratio, is provided. The learning data acquisition device includes a voice recognition influence degree calculation unit and a learning data acquisition unit. The voice recognition influence degree calculation unit calculates an influence degree on voice recognition accuracy caused by a change of a signal-to-noise ratio, based on a result of voice recognition on the kth noise superimposed voice data and a result of voice recognition on the k−1th noise superimposed voice data, where K is an integer of 2 or larger, k=2, 3, . . . , K, and a signal-to-noise ratio of the the kth noise superimposed voice data is smaller than a signal-to-noise ratio of the k−1th noise superimposed voice data, and obtains a largest signal-to-noise ratio SNRapply among signal-to-noise ratios of the k−1th noise superimposed voice data when the influence degree meets a given threshold condition. The learning data acquisition unit acquires noise superimposed voice data having a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNRapply, as learning data.