Self-Service Terminal Depth Filtering for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional beamforming mechanisms in self-service terminals are ineffective in distinguishing and suppressing interfering speech signals from the same direction as the desired speech signal, especially in noisy environments, leading to incorrect or no speech recognition due to the variability and randomness of speech inputs and background noise.
Innovation Solution
Implementing a mechanism that limits sound reception in depth by using distance-dependent filtering in addition to direction-dependent filtering, allowing for the attenuation of interfering sounds relative to useful sounds when the interference source is behind the desired sound source, and utilizing a staggered microphone arrangement to separate desired and unwanted signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional beamforming mechanism is used to limit signals to the sides, then lateral directional effect is achieved, but signals from consecutive sound sources in the same direction remain audible and cannot be suppressed
Solution Approach 1:
The patent extends conventional beamforming by adding a third dimension - depth filtering. While conventional beamforming handles lateral directionality (horizontal and vertical), this invention introduces distance-based filtering to create a three-dimensional acoustic spotlight, enabling the system to distinguish between sound sources at different depths and locations.
Solution Approach 2:
The patent segments the acoustic space into different regions based on distance from the microphone array. By calculating travel time differences and comparing them against stored reference values for different distances, the system creates separate filtering zones that can independently process signals from different depth ranges, allowing selective attenuation of unwanted sound sources.
2Power
If time-shift based beamforming is used to focus microphone, then directional signal amplification is achieved, but location determination fails when exact sound origin is unknown
Solution Approach 1:
The patent performs preliminary measurement of travel time differences between multiple microphones during normal operation. These measured values are stored as reference data corresponding to different distances. When a sound source is detected, the system compares the actual travel time differences against the stored references to determine the likely distance, enabling focus adjustment without requiring precise real-time location calculation.
Solution Approach 2:
The patent implements dynamic adjustment of the acoustic focus based on detected sound characteristics. The system continuously monitors travel time differences and adjusts the beamforming parameters in real-time to track moving sound sources, allowing the microphone array to adapt its focus dynamically rather than being fixed at a single location.
3Ease of operation
If speech recognition algorithms are used in noisy environments, then customer interaction is enabled, but recognition accuracy deteriorates due to mixing of speech input with background noise
Solution Approach 1:
The patent converts the harmful effect of background noise into a useful filtering mechanism. By calculating travel time differences of noise signals and comparing them against stored reference values, the system identifies noise patterns and creates corresponding attenuation filters. This transforms the noise problem into an opportunity to enhance the signal-to-noise ratio for legitimate customer commands.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enhances speech recognition accuracy in noisy environments by effectively distinguishing and suppressing interfering sounds, improving the reliability of speech recognition in self-service terminals.
Implementation Method 1
a plurality of acoustic sensors (104a, 104b) are provided
Implementation Method 2
the control device (106) is configured to process the electrical signals in order to determine a propagation time difference
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A self-service terminal (100) can have: - a product-sensing device for sensing a property of a product; - a plurality of acoustic sensors (104a, 104b); and a control device, which is designed for: superposing a signal captured by means of the plurality of acoustic sensors (104a, 104b); determining a voice pattern on the basis of the result of the superposing; outputting information on the basis of the property and on the basis of the voice pattern; wherein the superposing and the position of the plurality of acoustic sensors (104a, 104b) relative to each other are designed such that first components of the signal are attenuated relative to second components of the signal if an origin of the second components is located between the self-service terminal (100) and an origin of the first components.