Speaker Identification via Spatial Voice-Image Overlay
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In environments with multiple speakers, such as meeting rooms or classrooms, existing technologies struggle to accurately separate voice signals from different speakers and identify the source of each signal.
Innovation Solution
An electronic device equipped with an image data receiving circuit, a voice data receiving circuit, a memory, a processor, and an image data output circuit, which receives input image and voice data, generates spatial position data for speakers, converts this data into image position data, and overlays text related to the voice data onto the image at corresponding positions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If voice signals from multiple speakers are received simultaneously, then the microphone captures all voices, but the voice signals cannot be separated and identified
Solution Approach 1:
The patent segments the voice signals by spatial position. The microphone array divides the audio space into distinct regions, and each region is associated with a specific speaker. This allows simultaneous capture of multiple speakers while maintaining separability through spatial segmentation of the voice signals.
Solution Approach 2:
The patent introduces image data and spatial position information as intermediaries between the voice signals and speaker identification. The system uses camera images to determine spatial positions of speakers and overlays this information with voice signal data, creating a mediator that enables identification without requiring direct visual recognition of speakers.
2Measurement precision
If spatial position data is converted to image position data, then speaker positions can be identified on images, but transformation accuracy depends on coordinate system alignment
Solution Approach 1:
The patent uses a unified coordinate transformation approach that works for multiple speakers and positions simultaneously. The transformation parameters are derived from the overall spatial relationship between the microphone array and camera, creating a universal mapping that applies to all speakers in the scene without requiring individual calibration for each speaker position.
Data Source
AI summary
An electronic device is disclosed. An electronic device comprises: an image data receiving circuit configured to receive, from a camera, input image data associated with an image captured by the camera; a voice data receiving circuit configured to receive input voice data associated with the voices of speakers; a memory configured to store transform parameters for projecting a space coordinate system onto an image coordinate system on the image; and a processor, which determines a speaker's spatial location from the input voice data, converts same into the speaker's image location on the image, and inserts, into an input image, text associated with the speaker's voice according to the image location, so as to generate output image data.


