Source Separation Using Lip Motion for Kiosk Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In high public traffic settings, automated self-service kiosks face challenges in source separation of audio signals due to close adjacency of transactions, requiring effective methods to isolate target speaker voices from mixed audio signals without human intervention.
Innovation Solution
The implementation of a learning model that utilizes lip motion information from captured image data to enhance single-channel and multi-channel source separation, employing techniques like short-time Fourier transform, facial recognition, and fusion learning models to separate target speaker audio from noise and interference.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multiple queues are established at self-service kiosks to increase throughput, then productivity increases, but source separation of audio signals becomes more difficult due to close adjacency of transactions
Solution Approach 1:
The patent introduces visual dimension (lip motion information from image data) to complement the audio dimension. By transforming the source separation problem from a purely audio-based task to a multi-modal task incorporating visual facial features, the system can accurately identify target speakers even when multiple transactions occur in close proximity, thus maintaining high throughput while improving source separation accuracy.
2Ease of operation
If voice recognition technology is deployed without human intervention to enable self-service transactions, then ease of operation improves, but source separation becomes more challenging in crowded environments
Solution Approach 1:
The patent introduces lip motion information as an intermediary that bridges the gap between audio signals and speaker identification. The visual information from facial images serves as a mediator to disambiguate overlapping audio signals from multiple speakers, enabling reliable automated source separation without human intervention even in crowded transaction environments.
3Productivity
If automated self-service kiosks are placed in high public traffic locations, then productivity increases, but audio signal quality deteriorates due to noise and interference from adjacent transactions
Solution Approach 1:
The patent segments the source identification process into two independent components: audio signal processing and visual facial feature analysis. By separating these functions and then fusing their results, the system can isolate target speaker audio from noise and interference generated by adjacent transactions, maintaining high productivity in busy locations while filtering out harmful audio factors.
Data Source
AI summary
Methods and systems are provided for implementing source separation techniques, and more specifically performing source separation on mixed source single-channel and multi-channel audio signals enhanced by inputting lip motion information from captured image data, including selecting a target speaker facial image from a plurality of facial images captured over a period of interest; computing a motion vector based on facial features of the target speaker facial image; and separating, based on at least the motion vector, audio corresponding to a constituent source from a mixed source audio signal captured over the period of interest. The mixed source audio signal may be captured from single-channel or multi-channel audio capture devices. Separating audio from the audio signal may be performed by a fusion learning model comprising a plurality of learning sub-models. Separating the audio from the audio signal may be performed by a blind source separation (“BSS”) learning model.


