Sound Source Direction Estimation Device Dynamic Threshold
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound source direction estimation methods in voice recognition devices face inaccuracies due to changes in speaker position or microphone placement, leading to incorrect language identification and voice recognition results.
Innovation Solution
A method involving a sound source direction estimation device that calculates sound pressure differences between data from multiple microphones and adjusts a threshold value based on voice recognition time length to improve estimation accuracy, using a non-transitory computer-readable storage medium to execute processes for voice recognition and translation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a fixed threshold value is used for sound source direction estimation, then the device complexity is reduced, but the measurement precision deteriorates due to changes in speaker position or microphone placement
Solution Approach 1:
The patent applies dynamics by transitioning from a fixed threshold value to a dynamically adjustable threshold value that adapts to changing acoustic conditions. The threshold value is modified based on the correlation between sound pressure differences from multiple microphones, allowing the system to maintain high estimation accuracy when speaker position or microphone placement changes without requiring complex manual recalibration.
Solution Approach 2:
The patent implements feedback by using the correlation result of sound pressure differences as feedback information to automatically adjust the threshold value. The sound source direction estimation device continuously monitors the correlation between microphone signals and uses this feedback to optimize the threshold value, creating a closed-loop control system that improves measurement precision while maintaining relatively simple device architecture.
2Measurement precision
If the threshold value is dynamically adjusted based on correlation analysis, then the measurement precision is improved, but the loss of time increases due to additional calculation processes
Solution Approach 1:
The patent applies preliminary action by pre-establishing the correlation relationship between sound pressure differences from multiple microphones before performing threshold value adjustment. By calculating and storing correlation characteristics in advance, the system reduces the computational burden during real-time operation, thereby minimizing the time loss associated with dynamic threshold adjustment while maintaining high estimation accuracy.
3Measurement precision
If multiple microphones are used for sound source direction estimation, then the measurement precision is improved, but the device complexity increases due to additional hardware and processing requirements
Solution Approach 1:
The patent applies universality by designing a system where multiple microphones serve multiple functions: they simultaneously capture sound pressure information for direction estimation and generate correlation data for automatic threshold value adjustment. This multi-functional approach allows the microphone array to improve measurement precision while the same hardware components provide the feedback necessary for adaptive threshold control, reducing the need for additional specialized components.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enhances the accuracy of sound source direction estimation by dynamically updating threshold values, reducing incorrect language identification and improving voice recognition and translation processes.
Implementation Method 1
calculating a sound pressure difference between a first voice data acquired from a first microphone and a second voice data acquired from a second microphone
Data Source
AI summary
A non-transitory computer-readable storage medium storing a program that causes a processor included in a computer mounted on a sound source direction estimation device to execute a process, the process includes calculating a sound pressure difference between a first voice data acquired from a first microphone and a second voice data acquired from a second microphone and estimating a sound source direction of the first voice data and the second voice data based on the sound pressure difference, outputting an instruction to execute a voice recognition on the first voice data or the second voice data in a language corresponding to the estimated sound source direction, and controlling a reference for estimating a sound source direction based on the sound pressure difference, based on a time length of the voice data used for the voice recognition based on the instruction and a voice recognition time length.


