Adaptive Soft Mask Generation for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems for simultaneous recognition of multiple sources struggle to adapt to environmental changes, as soft masks are typically generated experimentally and require specific environment-based setups, limiting their adaptability and performance.
Innovation Solution
A speech recognition system that includes a sound source separating section, a mask generating section using distributions of speech and noise signals to create adaptive soft masks capable of taking continuous values between 0 and 1, and a speech recognizing section that utilizes these masks to recognize separated speech signals effectively across varying environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If soft masks are generated experimentally for each environment, then speech recognition accuracy is improved, but system adaptability to environmental changes deteriorates
Solution Approach 1:
The patent implements dynamic mask generation by continuously estimating separation reliability and updating soft masks in real-time based on current environmental conditions. The system transitions from static experimental masks to dynamic adaptive masks that automatically adjust to environmental changes, resolving the contradiction between accuracy and adaptability.
Solution Approach 2:
The patent changes the parameter of mask generation from fixed experimental values to variable values based on separation reliability estimates. By introducing reliability-based parameter adjustment, the system adapts mask characteristics to environmental conditions, achieving both high accuracy and environmental adaptability.
2Adaptability or versatility
If soft masks are generated using separation reliability distributions, then environmental adaptability is improved, but system complexity increases
Solution Approach 1:
The system performs self-service by automatically estimating separation reliability and generating appropriate soft masks without requiring external experimental calibration for each environment. The mask generation section autonomously adapts to environmental changes, reducing the need for complex external configuration while maintaining adaptability.
3Measurement precision
If experimental methods are used for each environment, then mask accuracy is improved, but time consumption for setup increases
Solution Approach 1:
The patent performs preliminary estimation of separation reliability and generates appropriate soft masks automatically before speech recognition begins. This preliminary adaptive setup eliminates the need for time-consuming experimental calibration for each environment, achieving both accuracy and time efficiency.
Data Source
AI summary
A speech recognition system according to the present invention includes a sound source separating section which separates mixed speeches from multiple sound sources from one another; a mask generating section which generates a soft mask which can take continuous values between 0 and 1 for each frequency spectral component of a separated speech signal using distributions of speech signal and noise against separation reliability of the separated speech signal; and a speech recognizing section which recognizes speeches separated by the sound source separating section using soft masks generated by the mask generating section.


