Head-Tracking Media Selection in Virtual Video Calls
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video conferencing platforms lack realism in user interactions, as they do not allow users to naturally look around while keeping their hands free for other activities, without requiring expensive equipment or new infrastructures.
Innovation Solution
A head-tracking-based media selection method for video communications in virtual environments, where 3D virtual environments are accessed by client devices with user graphical representations. The system tracks key facial landmarks to adjust virtual camera positions and orientations, allowing dynamic selection of virtual environment elements based on user head movements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If traditional video conferencing platforms are used with flat, two-dimensional user interfaces, then device complexity is reduced and ease of operation is improved, but realism and user presence are worsened
Solution Approach 1:
The patent transitions from traditional two-dimensional video conferencing interfaces to three-dimensional virtual environments. Users are represented as avatars in 3D space, allowing natural head movements and gaze tracking to control media selection. This dimensional change provides realism while maintaining ease of operation through intuitive natural interactions.
Solution Approach 2:
The patent replaces traditional mechanical input devices (mouse, keyboard, controller) with head-tracking based on facial landmark detection. The system uses camera-based vision to track head position and orientation, substituting physical mechanical inputs with natural head movements and gaze directions for media selection.
2Object-generated harmful factors
If head-tracking based media selection is implemented, then realism and user presence are improved, but device complexity increases
Solution Approach 1:
The patent uses existing cameras and computing devices to perform head-tracking and virtual environment rendering. Instead of requiring specialized expensive equipment, the system leverages universal devices (smartphones, tablets, laptops) with built-in cameras to achieve 3D avatars and head-tracking functionality, reducing overall device complexity.
Solution Approach 2:
The patent creates virtual copies (avatars) of users in three-dimensional virtual environments. These digital representations replicate user appearance and enable natural interaction without requiring physical avatars or complex hardware. The virtual environment is rendered on standard displays, avoiding the need for expensive head-mounted displays or specialized equipment.
3Object-generated harmful factors
If 3D virtual environments with head-tracking are implemented, then user presence and immersion are improved, but media selection complexity increases
Solution Approach 1:
The system automatically determines media selection based on head position and gaze direction without requiring explicit user input. The head-tracking system continuously monitors facial landmarks and automatically selects relevant media (chat, video feed, shared screen) based on where the user is looking, providing self-service media selection that reduces complexity while enhancing presence.
Data Source
AI summary
A method enabling head-tracking-based media selection for video communications implemented by a system is provided, comprising implementing a 3D virtual environment configured to be accessed by a plurality of client devices, each having a corresponding user graphical representation within the 3D virtual environment, each user graphical representation having a corresponding virtual camera; receiving head tracking metadata of a first client device that comprises 6 degrees of freedom head tracking information generated in response to tracking movement of key facial landmarks; identifying graphical elements within a field of view of the virtual camera of a second client device, wherein the identified graphical elements comprise the user graphical representation of the first user; and sending, to the second client device, the head tracking metadata and the identified graphical elements.


