Conference Call Audio Prioritization for Overlapping Speech

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conference calls often experience confusion and inefficiency due to overlapping speech from multiple participants, particularly when users speak simultaneously, leading to delays and the need for repetitive speech.

Innovation Solution

An apparatus and method that prioritize audio data transmission based on sound type classification, using machine-learned models to identify normal speech and restrict transmission of non-prioritized sound types, such as silence, background noise, and interjections, while allowing transmission of non-restricted sounds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If audio data from multiple users is transmitted simultaneously in conference calls, then all participants can communicate, but overlapping speech causes confusion and inefficiency

Engineering Contradiction:
Improvecommunication capabilityVSAvoidcommunication clarity
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The audio transmission is segmented by prioritizing specific users' audio data over others based on speech detection. The system divides the audio stream into prioritized and non-prioritized segments, allowing selective transmission that prevents overlapping speech confusion while maintaining communication capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes the transmission parameter by introducing audio data prioritization based on speech detection. When speech is detected from a particular user, the system modifies the transmission parameters to prioritize that user's audio data, thereby eliminating overlapping speech issues without sacrificing overall communication.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If audio data prioritization is implemented based on speech detection, then communication clarity improves, but system complexity increases

Engineering Contradiction:
Improvecommunication clarityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically detecting speech and prioritizing audio data without requiring manual intervention. The speech detection mechanism autonomously identifies when a user is speaking and applies prioritization accordingly, reducing the need for complex manual control systems while maintaining clarity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system manages complexity by implementing prioritization through parameter changes in the audio transmission protocol. By modifying transmission parameters based on speech detection results, the system achieves clarity without requiring fundamental architectural changes, thus limiting the increase in system complexity.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If all sound types are transmitted equally, then complete audio information is maintained, but non-speech sounds cause interference and reduce clarity

Engineering Contradiction:
Improveaudio information completenessVSAvoidaudio clarity
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The system extracts and removes non-speech audio data from the transmission stream. By detecting speech and excluding non-speech sounds (such as background noise, music, or other audio) from the prioritized transmission, the system maintains audio information completeness for relevant content while eliminating interference that reduces clarity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies local quality by treating different audio data differently based on its content. Speech audio receives prioritized transmission quality, while non-speech audio is treated with lower priority or excluded. This localized differentiation maintains clarity for important audio information without sacrificing overall audio completeness.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12603957B2Conference calls
Publication Date: 2026.04.14 NOKIA TECHNOLOGIES OY
  • US12603957B2 patent drawing
  • US12603957B2 patent drawing
  • US12603957B2 patent drawing

AI summary

Example embodiments relate to conference calls. In a method, there may be provided audio data and associated classification data from a plurality of user devices as part of a conference call, the plurality of user devices including at least a first user device and a second user device. The classification data may be indicative of one of a plurality of predetermined sound types represented by the associated audio data in a current time frame. The method may comprise determining that the audio data provided by the first user device is associated with a dominant speaker. The method may comprise controlling transmission of the audio data as part of the conference call, including preventing transmission of audio data provided by at least the second user device based on its associated classification data being indicative of a restricted sound type and the first user device being associated with a dominant speaker.