Wearable Speech Target Classification With Camera-Guided Beamforming
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional wearable display devices process user and counterpart speech as a single speech signal, leading to low speech recognition rates during conversations.
Innovation Solution
Employ multiple cameras and microphones positioned differently to distinguish between user and counterpart utterances, using micro beamforming and deep learning to process speech differently based on identified targets, and applying directivity control to microphones.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple cameras and microphones are used to distinguish user and counterpart utterances, then speech recognition rate is improved, but device complexity increases
Solution Approach 1:
The patent segments the speech processing task by identifying different utterance targets (user vs. counterpart) and applying different processing methods to each. The processor divides the audio signal processing into separate pathways: one for user speech and another for counterpart speech, allowing optimized handling of each type while improving overall recognition accuracy.
Solution Approach 2:
The patent introduces an intermediary classification mechanism that identifies whether an utterance comes from the user or counterpart before processing. This intermediary step uses audio signal analysis to determine the utterance target, then routes the speech through appropriate processing channels, resolving the complexity issue by organizing the multi-microphone system into structured processing paths.
2Measurement precision
If different processing methods are applied to user and counterpart speech, then speech recognition accuracy is improved, but processing time increases
Solution Approach 1:
The patent performs preliminary classification of the utterance target before applying detailed processing methods. By quickly identifying whether speech comes from the user or counterpart using initial audio analysis, the system prepares the appropriate processing pathway in advance, avoiding time-consuming trial-and-error processing and enabling faster overall recognition while maintaining accuracy.
Data Source
AI summary
Various embodiments of the disclosure provide a method and a device which includes multiple cameras arranged at different positions, multiple microphones arranged at different positions, a memory, and a processor operatively connected to at least one of the multiple cameras, the multiple microphones, and the memory, wherein the processor is configured to: determine, using at least one of the multiple cameras, whether at least one of a user wearing the electronic device or a counterpart having a conversation with the user makes an utterance, configure directivity of at least one of the multiple microphones based on the determination, obtain an audio from at least one of the multiple microphones based on the configured directivity, obtain an image including a mouth shape of the user or the counterpart from at least one of the multiple cameras, and process speech of an utterance target in a different manner based on the obtained audio and the image.


