Source Separation Using Lip Motion for Kiosk Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In high public traffic settings, automated self-service kiosks face challenges in source separation of audio signals due to close adjacency of transactions, requiring effective methods to isolate target speaker voices from mixed audio signals without human intervention.

Innovation Solution

The implementation of a learning model that utilizes lip motion information from captured image data to enhance single-channel and multi-channel source separation, employing techniques like short-time Fourier transform, facial recognition, and fusion learning models to separate target speaker audio from noise and interference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If multiple queues are established at self-service kiosks to increase throughput, then productivity increases, but source separation of audio signals becomes more difficult due to close adjacency of transactions

Engineering Contradiction:
Improvetransaction throughputVSAvoidsource separation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces visual dimension (lip motion information from image data) to complement the audio dimension. By transforming the source separation problem from a purely audio-based task to a multi-modal task incorporating visual facial features, the system can accurately identify target speakers even when multiple transactions occur in close proximity, thus maintaining high throughput while improving source separation accuracy.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Ease of operation

If voice recognition technology is deployed without human intervention to enable self-service transactions, then ease of operation improves, but source separation becomes more challenging in crowded environments

Engineering Contradiction:
Improveself-service capabilityVSAvoidsource separation reliability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent introduces lip motion information as an intermediary that bridges the gap between audio signals and speaker identification. The visual information from facial images serves as a mediator to disambiguate overlapping audio signals from multiple speakers, enabling reliable automated source separation without human intervention even in crowded transaction environments.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated self-service kiosks are placed in high public traffic locations, then productivity increases, but audio signal quality deteriorates due to noise and interference from adjacent transactions

Engineering Contradiction:
Improvetransaction volumeVSAvoidaudio noise and interference
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent segments the source identification process into two independent components: audio signal processing and visual facial feature analysis. By separating these functions and then fusing their results, the system can isolate target speaker audio from noise and interference generated by adjacent transactions, maintaining high productivity in busy locations while filtering out harmful audio factors.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11823699B2Single-channel and multi-channel source separation enhanced by lip motion
Publication Date: 2023.11.21 ALIBABA GROUP HOLDING LTD
  • US11823699B2 patent drawing
  • US11823699B2 patent drawing
  • US11823699B2 patent drawing

AI summary

Methods and systems are provided for implementing source separation techniques, and more specifically performing source separation on mixed source single-channel and multi-channel audio signals enhanced by inputting lip motion information from captured image data, including selecting a target speaker facial image from a plurality of facial images captured over a period of interest; computing a motion vector based on facial features of the target speaker facial image; and separating, based on at least the motion vector, audio corresponding to a constituent source from a mixed source audio signal captured over the period of interest. The mixed source audio signal may be captured from single-channel or multi-channel audio capture devices. Separating audio from the audio signal may be performed by a fusion learning model comprising a plurality of learning sub-models. Separating the audio from the audio signal may be performed by a blind source separation (“BSS”) learning model.