Audio Processing for Virtual Meeting Room Speaker Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Virtual meeting rooms struggle to differentiate voices from multiple presenters speaking simultaneously, making it difficult for participants to identify speakers and their spoken content.

Innovation Solution

The system employs a 360-degree fisheye camera and audio processing technology that captures and maps video and audio signals, using voiceprint information and mesh vertex adjustments to enhance voice identification, allowing the server to distinguish between speakers and display their contributions effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the virtual meeting room zooms in to presenters or speakers, then the presentation clarity is improved, but the voice differentiation between multiple speakers deteriorates

Engineering Contradiction:
Improvepresentation clarityVSAvoidvoice differentiation
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The audio signal is segmented by speaker using voiceprint recognition technology. Each speaker's voice is identified and separated into distinct audio streams, allowing the system to maintain clear presentation while differentiating between multiple speakers through individual audio channel assignment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Voiceprint recognition technology serves as an intermediary between the audio signal and the display system. It analyzes audio features to identify speakers and provides speaker identification information to the audio processing module, which then maps each speaker to a specific display region.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If multiple presenters speak at the same time, then the meeting productivity is improved, but the speaker identification accuracy deteriorates

Engineering Contradiction:
Improvemeeting productivityVSAvoidspeaker identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary voiceprint recognition to identify speakers before they begin speaking. Speaker identification information is pre-established and stored, allowing the audio processing module to accurately attribute spoken content to the correct speaker even when multiple presenters are present simultaneously.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The voiceprint recognition module continuously monitors audio signals and provides feedback on speaker identification to the audio processing module. This feedback mechanism enables real-time adjustment of audio mapping to ensure accurate speaker identification throughout the meeting.

Inventive Principle:
Principle #23Feedback

3Difficulty of detecting and measuring

If the system uses voiceprint information to identify speakers, then the speaker differentiation is improved, but the device complexity increases

Engineering Contradiction:
Improvespeaker differentiationVSAvoidsystem complexity
Core Design Contradiction:
Difficulty of detecting and measuringVSDevice complexity

Solution Approach 1:

The voiceprint recognition module serves multiple functions: it identifies speakers, determines speaker positions, and provides differentiation information for audio mapping. This multi-functionality reduces the need for separate complex systems while achieving effective speaker differentiation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system creates a virtual copy of the physical meeting environment using voiceprint recognition and audio spatial mapping. Instead of requiring complex hardware arrangements, the system uses software-based voiceprint analysis to replicate and differentiate speaker identities in the virtual space.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11798561B2Method, apparatus, and non-transitory computer readable medium for processing audio of virtual meeting room
Publication Date: 2023.10.24 FULIAN PRESION ELECTRONICS (TIANJIN) CO LTD
  • US11798561B2 patent drawing
  • US11798561B2 patent drawing
  • US11798561B2 patent drawing

AI summary

A method for processing audio generated in a virtual meeting room (VMR) includes setting a quantity of mesh vertexes according to seats in the VMR, obtaining first voiceprint information of a presenter, the first voiceprint information comprising a frequency, an amplitude, and a phase difference of an audio signal, adjusting the frequency or amplitude of the first voiceprint information according to the quantity of the mesh vertexes, and obtaining second voiceprint information; and determining a seat of the presenter in the VMR according to the second voiceprint information. An apparatus and a non-transitory computer readable medium for processing audio as above are also disclosed.