Adaptive Speech Enhancement Masks for Downstream Task Performance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech enhancement methods do not guarantee optimal performance for downstream tasks such as speech recognition and speaker recognition, as the hyper-parameters of the enhancement masks are difficult to set appropriately for different tasks.

Innovation Solution

A hyper-parameter optimization system that determines an adaptive mask by calculating a hyper-parameter representing the degree to which the signal is kept, considering the nature of the downstream task, using a neural network trained on speech data with noise and downstream task labels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a fixed speech enhancement method is used, then speech quality is improved, but downstream task performance deteriorates

Engineering Contradiction:
Improvespeech qualityVSAvoiddownstream task performance
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic hyperparameter adjustment by introducing a downstream task identifier that selects different hyperparameter sets based on the specific task. The system transitions from static enhancement masks to dynamic adaptive masks that change according to the downstream task requirements, resolving the contradiction between optimized speech quality and task-specific performance

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameters of the enhancement mask by introducing multiple hyperparameter sets (α, β, γ) that are selectively applied based on the downstream task. Each task type has optimized hyperparameter values that adjust the mask characteristics, enabling the system to adapt speech enhancement to different task requirements while maintaining quality

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If hyper-parameters are manually set for each downstream task, then task performance is improved, but system complexity increases

Engineering Contradiction:
Improvedownstream task performanceVSAvoidhyper-parameter setting complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements self-service by enabling the system to automatically select appropriate hyperparameters based on the downstream task identifier. Instead of requiring manual configuration for each task, the system autonomously retrieves and applies the correct hyperparameter set from its predefined collections, reducing user burden while maintaining task optimization

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent creates a universal hyperparameter management system that handles multiple downstream tasks through a single unified framework. The system maintains multiple hyperparameter sets but provides a unified interface that automatically selects the appropriate configuration based on the task type, making the complex system easy to use across different applications

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12412592B2Hyper-parameter optimization system, method, and program
Publication Date: 2025.09.09 NEC CORP
  • US12412592B2 patent drawing
  • US12412592B2 patent drawing
  • US12412592B2 patent drawing

AI summary

A speech enhancement means 81 determines an enhancement mask generated based on a mask for speech enhancement, when a test utterance is input as speech data. A first hyper-parameter optimization means 82 determines, when the test utterance is input, a first hyper-parameter which is a hyper-parameter representing the degree to which the signal representing the test utterance is kept using the mask, and the first hyper-parameter which is set to take into account a downstream task that is processed using an enhanced test utterance. A mask generation means 83 generates an adaptive mask from the determined enhancement mask and the first hyper-parameter that enhances the test utterance for the downstream task. The mask generation means 83 generates the adaptive mask in which the first hyper-parameter is a power of the mask.