3D Avatar Speaker Indicators for Easier Active Speaker Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Participants in 3D virtual environments for online meetings face difficulties in identifying active speakers and relevant user activities due to reduced rendering size and suboptimal navigation tools, leading to inefficiencies and loss of user engagement.

Innovation Solution

A system that automatically generates visual indicators by detecting active speakers in 3D representations and adds complementary 2D images or animations to a designated region, enhancing user interface transitions and focusing on speaker activity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a 3D environment rendering is displayed using only a portion of the display screen, then other types of renderings can be shown simultaneously, but the size of the 3D environment rendering is reduced making it more difficult to identify relevant user activity

Engineering Contradiction:
Improvedisplay arrangement flexibilityVSAvoiddifficulty in identifying active speakers
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The display is segmented into multiple regions: a first region dedicated to 3D environment rendering and a second region dedicated to displaying visual indicators of active speakers. This segmentation allows the system to simultaneously maintain a smaller 3D rendering while providing dedicated space for speaker identification, resolving the contradiction between display versatility and speaker detectability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Visual indicators (such as video feeds, images, or animations of users) act as an intermediary element that bridges the gap between the reduced-size 3D rendering and the need for clear speaker identification. These indicators provide explicit information about who is speaking without requiring users to scrutinize the small 3D environment, thus resolving the detection difficulty while preserving display flexibility.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If many users are allowed to participate in the 3D environment, then the meeting capacity is increased, but it becomes harder to identify specific conversations and people engaging in activity of interest

Engineering Contradiction:
Improvenumber of participantsVSAvoiddifficulty in identifying relevant user activity
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The system extracts the speaker identification information from the complex 3D environment and presents it separately in the second region as visual indicators. This extraction allows the system to handle many participants in the 3D space while providing a simplified, dedicated view for identifying active speakers, thus resolving the contradiction between meeting capacity and activity identifiability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system transitions from relying solely on spatial positioning in the 3D environment to a two-dimensional display region that explicitly shows speaker information. This dimensional change provides an additional channel for information presentation, allowing the system to accommodate more users in 3D space while maintaining clear speaker identification through the 2D visual indicators.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Quantity of substance

If the rendering of the 3D environment is reduced in size to accommodate other renderings, then more content can be displayed, but navigation tools are less effective and users must carefully scan the interface for relevant activity

Engineering Contradiction:
Improveamount of displayed contentVSAvoidease of navigating to find active speakers
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

Visual indicators serve as an intermediary that eliminates the need for users to navigate or scan the reduced-size 3D environment to find active speakers. The indicators directly present speaker information in the second region, making navigation unnecessary and significantly improving ease of operation while maintaining the ability to display diverse content.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Device complexity

If visual indicators are not provided for active speakers, then the user interface remains simple, but users experience loss of engagement and may miss important content

Engineering Contradiction:
Improveuser interface complexityVSAvoidmissed speaker content
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The interface is segmented into functional regions, with the second region specifically dedicated to speaker identification. This segmentation adds minimal complexity to the overall interface while effectively preventing information loss about active speakers, as the visual indicators provide clear, dedicated information without cluttering the entire interface.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12477016B2Automation of visual indicators for distinguishing active speakers of users displayed as three-dimensional representations
Publication Date: 2025.11.18 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12477016B2 patent drawing
  • US12477016B2 patent drawing
  • US12477016B2 patent drawing

AI summary

The disclosed techniques provide systems that automate visual indicators to show active speakers of a communication session who are displayed as 3D representations. Some participants of a communication session can be displayed in a user interface using 3D representations, e.g., avatars, that are each positioned within a 3D environment. The user interface may also include and number of renderings of 2D images of other participants displayed in a gallery, e.g., a display region that is designated for active speakers. When a user who is displayed as a 3D representation starts to speak, the system can detect the speaker's activity via a detection of an audio signal from the user's device. In response to the detection, the system can then automatically add a complementary image of the user to the gallery. The complementary image can help viewers navigate through complex user interface arrangements that display a large number of avatars.