Speaker Face Scaling for Video Conferencing Immersion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Video conferencing systems often struggle to maintain a 'live' feel due to difficulties in focusing on the current speaker, as raw captured images can be distracting and make it hard for viewers to concentrate on the speaker's face.
Innovation Solution
A computer system determines the distance of a speaker's face from the camera and scales the captured image based on the display size to ensure the speaker's face appears life-size, creating a more immersive experience by simulating a direct, in-person interaction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If raw captured images are sent to the display, then the videoconference can be established, but viewers have difficulty focusing on the current speaker and the live feel is reduced
Solution Approach 1:
The patent applies local quality by selectively scaling the speaker's face region differently from other parts of the image. The system identifies the speaker's face and enlarges it to life-size dimensions while keeping the rest of the scene at normal scale, allowing viewers to focus on the speaker without losing the contextual background.
Solution Approach 2:
The patent introduces a new dimension of spatial scaling by adjusting the size of the speaker's face in the image plane. By scaling the face to appear life-size on the display, the system creates a perceptual depth effect that simulates the speaker being physically present in the room, enhancing the live feel of the videoconference.
2Reliability
If the speaker's face is scaled to appear life-size on the display, then the live feel is enhanced, but the image processing complexity increases
Solution Approach 1:
The system performs preliminary action by pre-determining the scaling factor based on the display size and camera distance before rendering the final image. The computer system calculates the appropriate life-size scaling parameters in advance, which simplifies the real-time processing requirements during the videoconference.
Solution Approach 2:
The patent introduces an intermediary processing step where the computer system acts as a mediator between the camera capture and display output. The system interpolates and scales the speaker's face region using image processing algorithms that maintain visual coherence while achieving the life-size effect, balancing quality with processing complexity.
3Ease of operation
If the speaker's face size is adjusted based on distance and display size, then focus and engagement are improved, but the system complexity increases
Solution Approach 1:
The system implements feedback by continuously monitoring the camera distance to the speaker and the display size, then automatically adjusting the scaling factor accordingly. This closed-loop approach ensures the speaker's face always appears life-size regardless of camera position or display dimensions, improving focus and engagement without requiring manual intervention.
Solution Approach 2:
The patent applies parameter changes by dynamically modifying the scaling parameter based on measurable physical quantities (camera distance and display size). The system calculates the scaling factor as a function of these parameters, allowing automatic adaptation to different conference setups while maintaining the life-size appearance of the speaker's face.
Data Source
AI summary
A method can include determining a distance of a person's face from a camera that captured an image of the person's face, determining a size of a display in communication with the camera, and scaling the image based on the determined distance of the person's face and the determined size of the display.


