3D Avatar Positioning via Depth Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current virtual reality systems require users to wear VR headsets or use keyboards and mice to control avatars in virtual environments, leading to difficulties in real-time interaction and communication, especially when trying to point out content in virtual presentations, as the user's video stream is often displayed separately and can occlude important elements.
Innovation Solution
A system and method that allows users to control a virtual representation of themselves within a three-dimensional virtual world by using a two-dimensional video stream, extracting depth information to position the user's representation in a three-dimensional scene, enabling real-time interaction and movement without overlapping with other content, using voxels to create a multilayer scene where the user can be displayed in front of or behind other layers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the user's video stream is displayed in a separate window, then the user can be seen in the virtual environment, but the user cannot effectively point out or interact with specific content in the virtual scene
Solution Approach 1:
The patent transitions from displaying the user's video stream in a separate 2D window to integrating the user as a 3D avatar within the virtual environment. This dimensional integration allows the user to be spatially positioned within the scene, enabling natural pointing gestures and interactions with virtual objects while maintaining full visibility of the content.
Solution Approach 2:
The system creates a virtual representation (avatar) of the user that replicates the user's appearance and movements. This copy allows the user to interact with the virtual environment naturally while the original user remains visible in the physical world, solving the contradiction between interaction capability and content visibility.
2Ease of operation
If green screen technology is used to place the user in front of the background, then the user can be seen with the virtual content, but the user occludes important content that needs to be presented
Solution Approach 1:
The patent uses depth information and multi-layer scene composition to position the user's avatar at appropriate depths within the virtual scene. This allows the avatar to be behind certain content when needed, preventing occlusion while maintaining integration with the scene. The system creates multiple depth layers where the user can be dynamically positioned relative to different virtual objects.
3Ease of operation
If VR headsets and specialized 3D sensors are used, then the user can control the avatar in the virtual environment, but the system complexity and cost increase significantly
Solution Approach 1:
The system uses a standard 2D camera to capture the user and creates a 3D avatar representation from this 2D input. This copying approach allows the user to control their virtual representation using simple physical movements captured by a regular camera, eliminating the need for complex VR headsets and specialized 3D sensors while maintaining full avatar control capability.
4Productivity
If the user moves physically to control the avatar, then real-time interaction is improved, but the user's physical movement may not align with the virtual environment's coordinate system
Solution Approach 1:
The system continuously monitors the user's physical position and orientation captured by the camera, and uses this feedback to update the avatar's position in the virtual environment in real-time. This closed-loop approach ensures that the avatar accurately reflects the user's movements while maintaining proper alignment with the virtual scene's coordinate system through continuous spatial transformation.
Data Source
AI summary
The present disclosure relates generally to a system and method for a user to control a virtual representation of themselves within a three-dimensional virtual world. The system and method enable utilizing a two-dimensional image or video data of user with extracted depth information to position themselves in a three-dimensional scene. It also provides a control system and method for a user to control the virtual representation of themselves using the output video as a visual feedback mechanism in a three-dimensional space including the virtual representation of themselves. A user interacts with other virtual objects or items in a scene or even with other users visualized in the scene.


