Animated Avatar Generation Using Reference Data for Videoconferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current videoconferencing technologies face limitations in capturing natural interactions due to unnatural camera positions and the requirement for virtual reality equipment, which can be costly and restrictive, making it difficult to see participants clearly and accurately represent subtle gestures and modifications in real-time.
Innovation Solution
A method and system that generate an animated visual representation by selecting data units from a database of reference states measured by sensor systems, allowing for synchronization with real-time activity, enabling flexible and high-quality animation corrections and modifications, such as changing clothing or grooming, without excessive data processing, even with limited network bandwidth or missing frames.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If virtual reality equipment is used to create virtual environments for videoconferencing, then flexibility and animation quality are improved, but cost and device complexity increase significantly
Solution Approach 1:
The patent creates a 3D copy or avatar of the user that can be used in virtual environments. Instead of requiring the user to physically wear VR equipment, the system captures the user's appearance and creates a digital replica that can be animated and positioned flexibly in virtual spaces, providing adaptability without the complexity of VR headsets
Solution Approach 2:
The patent replaces the mechanical VR headset system with a camera-based capture system and software-based avatar animation system. The mechanical complexity of VR equipment is substituted with optical capture and computational rendering, achieving similar flexibility goals through different technological means
2Device complexity
If fixed cameras are used in meeting rooms to capture participants, then device complexity is reduced, but the ability to see participants clearly and capture gestures is worsened
Solution Approach 1:
The patent introduces dynamic elements to the camera system, allowing it to move and adjust its position and orientation. The camera can track participants, zoom in on speakers, and capture gestures more effectively while maintaining reasonable device complexity through automated control algorithms
Solution Approach 2:
The patent introduces a camera controller or processing system as an intermediary between the fixed camera and the final video output. This intermediary can process camera feeds, adjust framing, enhance gesture visibility, and composite multiple camera views to improve measurement precision without requiring complex physical camera systems
3Productivity
If real-time animation generation is performed without reference data, then data processing requirements are reduced, but animation quality and naturalness deteriorate
Solution Approach 1:
The patent performs preliminary actions by capturing and storing reference data about the user's appearance, features, and movements in advance. This reference library is built beforehand and can be quickly accessed during real-time animation generation, reducing processing requirements while maintaining high animation quality through comparison with stored reference data
Data Source
AI summary
Generating data to provide an animated visual representation is disclosed. A method comprises receiving input data obtained by a first sensor system measuring information about at least one target person. One data unit is selected from a database comprising a plurality of the data units. Each data unit comprises information about a reference person in a reference state measured at a previous time by the first sensor system or by a second sensor system. The information in each data unit allows generation of an animated visual representation of the reference person in the reference state. The reference state is different for each of the data units. The selected data unit and the input data are used to generate output data usable to provide an animated visual representation corresponding to the target person and synchronized with activity of the target person measured by the first sensor system.


