Avatar Breathing Rendering via Audio Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing avatar rendering technologies fail to accurately depict breathing movements during speech, often relying on unrealistic looped motion patterns or requiring complex sensors, which are not accessible to average users.
Innovation Solution
A method utilizing a respiratory prediction model to process an audio signal and predict a respiratory signal, adjusting the visualization of an avatar's breathing movements based on this signal, without the need for complex sensors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If motion capture sensors are used to achieve realistic body movement, then the realism of avatar breathing is improved, but the device complexity and cost increase significantly
Solution Approach 1:
The patent creates a virtual copy of respiratory movement data by training a machine learning model on motion capture sensor data. The trained model then generates synthetic respiratory signals from audio input, eliminating the need for physical sensors in the final application while maintaining realistic breathing animation.
Solution Approach 2:
The patent replaces the mechanical sensor-based measurement system with a computational approach using machine learning. Instead of physically measuring respiratory movement with sensors, the system uses an audio-based respiratory prediction model to generate synthetic respiratory signals that drive avatar animation.
2Measurement precision
If motion capture sensors are deployed for realistic avatar rendering, then the accuracy of breathing movement is improved, but the ease of operation deteriorates due to accessibility issues
Solution Approach 1:
The system uses the user's own audio input (speech) as the source for generating respiratory signals. The audio-based respiratory prediction model processes the user's natural speech to extract breathing patterns, making the system self-sufficient and eliminating the need for external sensor equipment that users would need to wear or set up.
Solution Approach 2:
The patent makes the respiratory prediction model accessible to all users through standard audio input devices (microphones) that are universally available. The system processes common audio speech signals to generate respiratory data, making it universally applicable without requiring specialized equipment or technical expertise.
3Reliability
If motion capture sensors are used to track breathing movement, then the realism of avatar body movement is improved, but the computational resources required increase significantly
Solution Approach 1:
The patent performs computational work in advance by training the respiratory prediction model offline using motion capture data. Once trained, the model can generate realistic respiratory signals from audio input with minimal computational overhead during runtime, reducing the energy requirements for real-time avatar animation.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Proposed concepts thus aim to provide schemes, solutions, concepts, designs, methods and systems pertaining to rending an avatar of a user. In particular, an audio signal of the user speaking is processed, with a respiratory prediction model, to predict a respiratory signal describing/indicative of breathing movement of the user whilst speaking. Using the predicted respiratory signal, a time-varying visualization of an avatar of the user speaking is adjusted. As a result, a realistic visualization of an avatar that reflects breathing movement of the user may be rendered without the need for complex sensors tracking movement of the user.