Conference Camera Framing via Personal Device Signals

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In network-based communication sessions, particularly in conference settings, it is challenging for participants to indicate their desire to speak due to dominance by other participants, and existing methods like virtual hand buttons are not effective in shared spaces where multiple users share a single camera setup, making it difficult to discern who wishes to contribute.

Innovation Solution

A method where participants can use their personal computing devices to send an interjection request signal, which alerts the conference service and triggers the camera to automatically pan, tilt, or zoom to frame the requesting user, using ultrasonic beacons, acoustic source localization, or facial recognition to determine the user's position and adjust the camera settings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single shared camera is used in a conference room, then device complexity is reduced, but it becomes difficult to identify and frame the specific user who wishes to speak

Engineering Contradiction:
Improvecamera setupVSAvoiduser identification
Core Design Contradiction:
Device complexityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent introduces an intermediary device (mobile device or computing device) that each user possesses individually in the conference room. This device communicates with the shared camera system to convey the user's identity and location information, enabling the camera to frame the correct user without requiring multiple cameras. The intermediary device acts as a bridge between the user and the shared camera resource.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical/physical approach of using multiple cameras or manual camera operation with an automated electronic system. The shared camera receives control signals based on data from user devices (such as ultrasonic beacons, acoustic source localization, or facial recognition) to automatically pan, tilt, and zoom to frame the user who wishes to speak, eliminating the need for mechanical camera multiplication or manual intervention.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If a virtual hand button is provided in the communication application, then users can indicate their desire to speak, but it is ineffective in shared spaces where multiple users share a single camera setup

Engineering Contradiction:
Improvespeak indicationVSAvoidspeak indication effectiveness
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent enhances the virtual hand button functionality by introducing an intermediary communication channel between the user device and the camera system. When a user activates the virtual hand button, the user device sends a signal (such as an ultrasonic beacon or acoustic signal) that the shared camera system detects to determine the user's physical location and identity, making the speak indication effective even in shared camera environments.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the insufficient virtual button interface with an integrated acoustic and visual feedback system. The system uses acoustic source localization or ultrasonic detection to translate the digital button press into physical camera movement, automatically framing the user who pressed the button. This substitution of mechanical camera control with automated acoustic-optical integration resolves the ineffectiveness of the virtual button in shared spaces.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If the camera manually follows each speaker, then communication clarity is improved, but it requires constant manual intervention and cannot automatically identify who wishes to speak

Engineering Contradiction:
Improvecommunication clarityVSAvoidcamera control
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent enables the camera system to serve itself by automatically detecting and framing users who wish to speak through acoustic signals or ultrasonic beacons from their devices. The system autonomously determines which user wants to speak and adjusts the camera framing accordingly, eliminating the need for manual camera operation while maintaining clear visual focus on the active speaker.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual camera operation with an automated acoustic-optical system. The camera system uses acoustic source localization or ultrasonic detection to automatically identify the user who wishes to speak and adjusts pan, tilt, and zoom parameters accordingly. This substitution transforms the camera from a manually controlled device to an autonomous system that responds to acoustic cues from user devices.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution allows for clear framing of participants wishing to speak, ensuring all participants have an opportunity to contribute by automatically positioning the camera on the requesting user, enhancing communication equity and clarity in shared space environments.

Implementation Method 1

using ultrasonic beacons, acoustic source localization, or facial recognition to determine the user's position and adjust the camera settings

Methodology Applied
Scientific EffectAcoustic source localization:

Data Source

PatentUS11825200B2Framing an image of a user requesting to speak in a network-based communication session
Publication Date: 2023.11.21 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11825200B2 patent drawing
  • US11825200B2 patent drawing
  • US11825200B2 patent drawing

AI summary

Disclosed in some examples, are methods, systems, and machine-readable mediums for allowing participants of communication sessions joining from conference rooms to indicate a desire to speak using their own personal computing device, and, to automatically frame that user with the conference room camera using a position of the user determined automatically. For example, a user may have a communication application executing on their mobile device that is logged into the network-based conference. If the user wishes to speak, they may activate a control within the application instance executing on their mobile device. The in-room meeting device may then automatically locate the user, and may direct the camera to pan, tilt, or zoom so that the camera frames the user.