Pendant Microphone Array for DOA-Based Speaker Diarization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems are inadequate in processing audio data to accurately determine when an individual starts and stops speaking, especially in the presence of multiple speakers and background noise.

Innovation Solution

A pendant-style microphone array with multiple microphones, including dipole and omnidirectional microphones, is used to capture audio data, which is processed to determine the direction of arrival (DOA) of soundwaves, timestamped, and analyzed using machine learning and AI algorithms to identify individual voices through diarization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional audio processing systems are used, then the system complexity is low, but the accuracy of determining when individuals start and stop speaking deteriorates in noisy environments with multiple speakers

Engineering Contradiction:
Improveaccuracy of determining when individuals start and stop speakingVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The audio processing is segmented into multiple distinct stages: initial audio capture by multiple microphones, direction of arrival (DOA) calculation to determine sound source positions, timestamp assignment to audio segments, and voice identification through diarization algorithms. This segmentation allows each component to specialize in one aspect of the problem, improving overall accuracy while managing complexity through modular design

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Direction of arrival (DOA) data serves as an intermediary element that bridges the gap between raw audio capture and voice identification. The DOA information provides spatial context that helps disambiguate which speaker is speaking at any given time, acting as a mediator that enhances the accuracy of speaker attribution without requiring direct complex analysis of the audio signals themselves

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If DOA data with timestamps is used to identify speakers, then the accuracy of voice identification improves, but the processing time and computational resources increase

Engineering Contradiction:
Improveaccuracy of voice identificationVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Timestamps are assigned to audio segments and DOA data is calculated in advance during the audio capture phase, before the actual voice identification process begins. This preliminary organization of data with temporal and spatial markers allows the diarization algorithms to work more efficiently, as the data is already structured and indexed, reducing the computational burden and processing time during the identification phase

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional mechanical or rule-based audio analysis with machine learning and AI-based diarization algorithms. These intelligent systems can process DOA data and timestamps more efficiently, automatically identifying speaker patterns and voice characteristics without requiring exhaustive computational analysis, thereby reducing processing time while maintaining high accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If multiple microphones are used to capture audio data, then the ability to determine direction of arrival improves, but the device complexity and cost increase

Engineering Contradiction:
Improvedirection of arrival determination accuracyVSAvoidmicrophone array complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The microphone array is designed to serve multiple functions simultaneously: capturing audio data for recording, calculating direction of arrival for spatial awareness, providing input for timestamp synchronization, and enabling voice identification through diarization. By making the same hardware component multi-functional, the system achieves high measurement precision without proportionally increasing device complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent combines several processing functions into a unified audio processing pipeline: audio capture, DOA calculation, timestamp assignment, and voice identification are merged into an integrated system that processes all data streams simultaneously. This merging allows the multiple microphones to work together as a coordinated array rather than independent components, optimizing the complexity-performance ratio

Inventive Principle:
Principle #5Merging (Combining)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Accurately identifies and labels individual speakers in audio recordings, even in noisy environments, by associating DOA data with timestamps and voice profiles, enhancing speech-to-text transcription and conversation analysis.

Implementation Method 1

The microphone array includes a plurality of microphones. The microphones may be positioned relative to each other such that coordinates of a positional x/y/z plane can be assigned to a soundwave (or a portion of a soundwave) as the soundwave is received at the microphone array.

Methodology Applied
Scientific EffectAcoustic signal processing: Acoustics

Data Source

PatentUS20260019747A1Wearable device for user audio recording
Publication Date: 2026.01.15 META PLATFORMS INC
  • US20260019747A1 patent drawing
  • US20260019747A1 patent drawing
  • US20260019747A1 patent drawing

AI summary

An example operation includes processing audio data to determine a direction of arrival and to identify a person speaking at a particular time.