Voice Processing Device Directional Model Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice recognition systems face accuracy declines due to changes in acoustic environments, particularly in rooms where reflected sounds interfere with direct sounds, requiring extensive data collection for acoustic modeling and reverberation suppression processing.
Innovation Solution
A voice processing device that separates voice signals into incoming components by direction, updates a voice recognition model using statistics specific to each direction, and recognizes voices using a dereverberation unit to isolate direct sounds, thereby improving accuracy amidst environmental changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If reverberation suppression processing is performed to improve voice recognition accuracy, then recognition accuracy improves, but information representing acoustic environment is lost and processing complexity increases
Solution Approach 1:
The patent segments the acoustic model into multiple models corresponding to different incoming directions. The separation unit divides voice signals into incoming components by direction, and different voice recognition models are selected based on the incoming direction. This segmentation allows the system to handle reverberation differently for each direction without requiring complex global reverberation suppression processing.
Solution Approach 2:
The patent applies local quality by using direction-specific voice recognition models tailored to the characteristics of sounds coming from different directions. Each directional model is optimized for its specific direction, incorporating local acoustic characteristics rather than applying a uniform processing approach to all sounds.
2Measurement precision
If acoustic models are updated frequently to adapt to changing acoustic environments, then voice recognition accuracy improves, but processing time and computational load increase
Solution Approach 1:
The patent implements dynamic adaptation by automatically selecting appropriate voice recognition models based on the incoming direction of sounds. The system dynamically adjusts which model to use without requiring full model retraining, allowing rapid adaptation to changing acoustic environments while minimizing processing time.
Solution Approach 2:
Multiple voice recognition models for different incoming directions are prepared in advance. When a sound is received, the system simply selects the pre-prepared model corresponding to the incoming direction, avoiding the need for time-consuming real-time model generation or updates.
3Measurement precision
If voice signals are separated by incoming direction to improve recognition accuracy, then accuracy improves, but device complexity and processing complexity increase
Solution Approach 1:
The separation unit segments incoming voice signals into distinct incoming components based on their direction of arrival. This segmentation is achieved through direction estimation processing that analyzes the spatial characteristics of sounds received by the microphone array, dividing the acoustic space into different directional zones.
Solution Approach 2:
The voice recognition system achieves multi-functionality by using a single directional separation mechanism that works across all incoming directions. The same separation unit and model selection logic handle sounds from any direction, making the system universally applicable without requiring direction-specific hardware for each angle.
4Measurement precision
If reflected sound components are suppressed to improve voice recognition, then recognition accuracy improves, but information about acoustic environment changes is lost
Solution Approach 1:
The patent segments reflected sound components by their incoming directions, separating them from direct sounds. By maintaining distinct voice recognition models for different directions, the system preserves information about the acoustic environment in each direction while still suppressing reflected sounds that would interfere with direct sound recognition.
Solution Approach 2:
The system changes the parameter of model selection based on incoming direction. By adjusting which voice recognition model is active according to the direction of incoming sounds, the system adapts to acoustic environment changes without requiring full model retraining, preserving adaptability while maintaining accuracy.
Data Source
AI summary
A separation unit separates voice signals of a plurality of channels into an incoming component in each incoming direction, a selection unit selects a statistic corresponding to an incoming direction of the incoming component separated by the separation unit from a storage unit which stores a predetermined statistic and a voice recognition model for each incoming direction, an updating unit updates the voice recognition model on the basis of the statistic selected by the selection unit, and a voice recognition unit recognizes a voice of the incoming component separated using the voice recognition model.


