Active Speaker Proxy for Sign Language Interpreters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video communication systems do not effectively allow non-verbal participants, such as those who use sign language, to appear as the active speaker when a sign language interpreter is voicing on their behalf, leading to inequitable participation in video communication sessions.

Innovation Solution

A method and system that enables a non-verbal participant to designate a sign language interpreter within the communication session, where the system highlights the video feed of the non-verbal participant when the interpreter is speaking, ensuring the non-verbal participant appears as the active speaker through visual adjustments such as a prominent feed or border, allowing equitable participation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If the sign language interpreter is displayed as the active speaker when voicing for a non-verbal participant, then the interpreter's contribution is visible, but the non-verbal participant's equitable participation is compromised

Engineering Contradiction:
Improveparticipant identity informationVSAvoidequitable participation
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system introduces a proxy mechanism where the non-verbal participant's video feed is displayed as an intermediary representation during the interpreter's speech. This mediator allows the system to attribute the spoken words to the correct participant while maintaining visual clarity, resolving the conflict between showing the interpreter's contribution and preserving participant identity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system creates a visual copy or proxy of the non-verbal participant's video feed that is displayed during the interpreter's speech. This copy serves as a representation that maintains the participant's identity and equitable participation status without requiring the actual participant to be physically present or speaking, thus resolving the contradiction between visibility and identity attribution.

Inventive Principle:
Principle #26Copying

2Ease of operation

If the non-verbal participant's video feed is highlighted during interpreter speech, then equitable participation is achieved, but system complexity increases

Engineering Contradiction:
Improveequitable participationVSAvoidvideo feed management
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system dynamically adjusts the video feed display based on the current speaker state. When the interpreter is speaking, the non-verbal participant's feed is automatically highlighted; when the participant speaks directly, their feed is highlighted normally. This dynamic behavior manages complexity by using automated state-based control rather than manual configuration, achieving equitable participation without excessive system complexity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements feedback mechanisms that monitor the current speaker state and automatically adjust the video feed highlighting accordingly. This feedback loop ensures that the correct participant is highlighted at the appropriate times without requiring complex manual management, resolving the contradiction between equitable participation and system complexity.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12184709B2Active speaker proxy presentation for sign language interpreters
Publication Date: 2024.12.31 ZOOM VIDEO COMM INC
  • US12184709B2 patent drawing
  • US12184709B2 patent drawing
  • US12184709B2 patent drawing

AI summary

Methods and systems provide for an active server proxy presentation for sign language interpreters within a video communication session. In one embodiment, a method presents a user interface for each of a number of client devices connected to a communication session, with each UI including one or more video feeds associated with participants of the communication session. The method receives an indication that a first participant is designating a second participant as a sign language interpreter who will perform voicing for the first participant. The method then determines that the second participant is performing voicing for the first participant, then presents, within the UIs of at least a subset of the client devices, a video feed associated with the first participant in a highlighted fashion concurrently to the second participant performing the voicing for the first participant.