Audio Processing for Virtual Meeting Room Speaker Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Virtual meeting rooms struggle to differentiate voices from multiple presenters speaking simultaneously, making it difficult for participants to identify speakers and their spoken content.
Innovation Solution
The system employs a 360-degree fisheye camera and audio processing technology that captures and maps video and audio signals, using voiceprint information and mesh vertex adjustments to enhance voice identification, allowing the server to distinguish between speakers and display their contributions effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the virtual meeting room zooms in to presenters or speakers, then the presentation clarity is improved, but the voice differentiation between multiple speakers deteriorates
Solution Approach 1:
The audio signal is segmented by speaker using voiceprint recognition technology. Each speaker's voice is identified and separated into distinct audio streams, allowing the system to maintain clear presentation while differentiating between multiple speakers through individual audio channel assignment.
Solution Approach 2:
Voiceprint recognition technology serves as an intermediary between the audio signal and the display system. It analyzes audio features to identify speakers and provides speaker identification information to the audio processing module, which then maps each speaker to a specific display region.
2Productivity
If multiple presenters speak at the same time, then the meeting productivity is improved, but the speaker identification accuracy deteriorates
Solution Approach 1:
The system performs preliminary voiceprint recognition to identify speakers before they begin speaking. Speaker identification information is pre-established and stored, allowing the audio processing module to accurately attribute spoken content to the correct speaker even when multiple presenters are present simultaneously.
Solution Approach 2:
The voiceprint recognition module continuously monitors audio signals and provides feedback on speaker identification to the audio processing module. This feedback mechanism enables real-time adjustment of audio mapping to ensure accurate speaker identification throughout the meeting.
3Difficulty of detecting and measuring
If the system uses voiceprint information to identify speakers, then the speaker differentiation is improved, but the device complexity increases
Solution Approach 1:
The voiceprint recognition module serves multiple functions: it identifies speakers, determines speaker positions, and provides differentiation information for audio mapping. This multi-functionality reduces the need for separate complex systems while achieving effective speaker differentiation.
Solution Approach 2:
The system creates a virtual copy of the physical meeting environment using voiceprint recognition and audio spatial mapping. Instead of requiring complex hardware arrangements, the system uses software-based voiceprint analysis to replicate and differentiate speaker identities in the virtual space.
Data Source
AI summary
A method for processing audio generated in a virtual meeting room (VMR) includes setting a quantity of mesh vertexes according to seats in the VMR, obtaining first voiceprint information of a presenter, the first voiceprint information comprising a frequency, an amplitude, and a phase difference of an audio signal, adjusting the frequency or amplitude of the first voiceprint information according to the quantity of the mesh vertexes, and obtaining second voiceprint information; and determining a seat of the presenter in the VMR according to the second voiceprint information. An apparatus and a non-transitory computer readable medium for processing audio as above are also disclosed.


