Self-Service Terminal Depth Filtering for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional beamforming mechanisms in self-service terminals are ineffective in distinguishing and suppressing interfering speech signals from the same direction as the desired speech signal, especially in noisy environments, leading to incorrect or no speech recognition due to the variability and randomness of speech inputs and background noise.

Innovation Solution

Implementing a mechanism that limits sound reception in depth by using distance-dependent filtering in addition to direction-dependent filtering, allowing for the attenuation of interfering sounds relative to useful sounds when the interference source is behind the desired sound source, and utilizing a staggered microphone arrangement to separate desired and unwanted signals.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional beamforming mechanism is used to limit signals to the sides, then lateral directional effect is achieved, but signals from consecutive sound sources in the same direction remain audible and cannot be suppressed

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidinterference from consecutive sound sources
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent extends conventional beamforming by adding a third dimension - depth filtering. While conventional beamforming handles lateral directionality (horizontal and vertical), this invention introduces distance-based filtering to create a three-dimensional acoustic spotlight, enabling the system to distinguish between sound sources at different depths and locations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent segments the acoustic space into different regions based on distance from the microphone array. By calculating travel time differences and comparing them against stored reference values for different distances, the system creates separate filtering zones that can independently process signals from different depth ranges, allowing selective attenuation of unwanted sound sources.

Inventive Principle:
Principle #1Segmentation

2Power

If time-shift based beamforming is used to focus microphone, then directional signal amplification is achieved, but location determination fails when exact sound origin is unknown

Engineering Contradiction:
Improvesignal amplificationVSAvoidsound source location detection
Core Design Contradiction:
PowerVSDifficulty of detecting and measuring

Solution Approach 1:

The patent performs preliminary measurement of travel time differences between multiple microphones during normal operation. These measured values are stored as reference data corresponding to different distances. When a sound source is detected, the system compares the actual travel time differences against the stored references to determine the likely distance, enabling focus adjustment without requiring precise real-time location calculation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamic adjustment of the acoustic focus based on detected sound characteristics. The system continuously monitors travel time differences and adjusts the beamforming parameters in real-time to track moving sound sources, allowing the microphone array to adapt its focus dynamically rather than being fixed at a single location.

Inventive Principle:
Principle #15Dynamics

3Ease of operation

If speech recognition algorithms are used in noisy environments, then customer interaction is enabled, but recognition accuracy deteriorates due to mixing of speech input with background noise

Engineering Contradiction:
Improvecustomer interaction capabilityVSAvoidspeech recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent converts the harmful effect of background noise into a useful filtering mechanism. By calculating travel time differences of noise signals and comparing them against stored reference values, the system identifies noise patterns and creates corresponding attenuation filters. This transforms the noise problem into an opportunity to enhance the signal-to-noise ratio for legitimate customer commands.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enhances speech recognition accuracy in noisy environments by effectively distinguishing and suppressing interfering sounds, improving the reliability of speech recognition in self-service terminals.

Implementation Method 1

a plurality of acoustic sensors (104a, 104b) are provided

Methodology Applied
Scientific EffectAcoustic energy to electrical energy conversion:

Implementation Method 2

the control device (106) is configured to process the electrical signals in order to determine a propagation time difference

Methodology Applied
Scientific EffectSound propagation time measurement: Speed of Sound

Data Source

PatentEP4189678B1Self-service terminal and method
Publication Date: 2025.09.24 DIEBOLD NIXDORF SYST GMBH
  • EP4189678B1 patent drawingFigure 1
  • EP4189678B1 patent drawingFigure 2
  • EP4189678B1 patent drawingFigure 3

AI summary

A self-service terminal (100) can have: - a product-sensing device for sensing a property of a product; - a plurality of acoustic sensors (104a, 104b); and a control device, which is designed for: superposing a signal captured by means of the plurality of acoustic sensors (104a, 104b); determining a voice pattern on the basis of the result of the superposing; outputting information on the basis of the property and on the basis of the voice pattern; wherein the superposing and the position of the plurality of acoustic sensors (104a, 104b) relative to each other are designed such that first components of the signal are attenuated relative to second components of the signal if an origin of the second components is located between the self-service terminal (100) and an origin of the first components.