Meeting Role Separation Using Sound Source Angle Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing role separation systems based on voiceprint identification require offline data accumulation, making real-time separation difficult and impacting user experience.

Innovation Solution

Utilizing sound source angle data to identify and separate roles in real-time through sound source localization and identification methods, including eigenvalue decomposition and voice activity detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voiceprint identification is used for role separation, then identification accuracy is improved, but real-time processing capability deteriorates due to offline data accumulation requirements

Engineering Contradiction:
Improveidentification accuracyVSAvoidreal-time processing capability
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The patent changes the identification parameter from voiceprint (requiring time-domain signal accumulation) to sound source angle (spatial parameter). By using sound source angle data from microphone arrays, the system achieves real-time role separation without requiring offline data accumulation, thus resolving the contradiction between identification accuracy and real-time processing capability

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the traditional voiceprint identification mechanism with sound source localization technology. Instead of analyzing voice characteristics over time, the system uses spatial information from multiple microphones to identify roles in real-time, substituting a different physical measurement approach to achieve both accuracy and real-time performance

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If offline speech data is used for role separation, then identification accuracy is improved, but system complexity and data processing time increase

Engineering Contradiction:
Improveidentification accuracyVSAvoiddata processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-configuring the sound source angle calculation framework and microphone array parameters. The system prepares the spatial analysis infrastructure in advance, enabling rapid real-time role separation without time-consuming offline data processing, thus reducing data processing time while maintaining accuracy

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If voiceprint identification is implemented, then role separation accuracy is improved, but device complexity and computational requirements increase

Engineering Contradiction:
Improverole separation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the essential spatial information (sound source angle) from the complex audio signal processing task. By focusing solely on directional information from microphone arrays rather than comprehensive voiceprint analysis, the system achieves role separation accuracy with reduced computational complexity and lower system requirements

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12537020B2Role separation method, meeting summary recording method, role display method and apparatus, electronic device, and computer storage medium
Publication Date: 2026.01.27 ALIBABA GROUP HOLDING LTD
  • US12537020B2 patent drawing
  • US12537020B2 patent drawing
  • US12537020B2 patent drawing

AI summary

A role separation method, a meeting summary recording method, a role display method and apparatus, an electronic device, and a computer storage medium, relating to the field of speech processing. The role separation method comprises: obtaining sound source angle data corresponding to a speech data frame, acquired by a speech acquisition device, of a role to be separated (S102); on the basis of the sound source angle data, performing identity recognition on the role to be separated to obtain a first identity recognition result of the role to be separated (S104); and separating the role on the basis of the first identity recognition result of the role to be separated (S106). The role is separated in real time, thus making user experience smooth.