Avatar Dance Generation via Audio Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for controlling avatars in VR/AR environments require large memory databases to generate dance moves, making them unsuitable for edge devices and limiting creative responsiveness to music.

Innovation Solution

A method that extracts high-level and latent audio features from audio signals to generate joint angle distribution matrices using Gaussian parameters, allowing avatars to improvise dance steps without a predefined database, suitable for edge devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a database storing a large number of preset dance moves is maintained, then the avatar can generate dance moves, but memory usage increases and implementation on edge devices becomes difficult

Engineering Contradiction:
Improveavatar dance move generation capabilityVSAvoidmemory usage
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential rhythmic features from audio signals (beat detection, tempo extraction) rather than storing complete dance move databases. This extraction approach reduces memory requirements while maintaining the core functionality of generating dance moves that match the music rhythm.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical database storage and retrieval system with an audio analysis and synthesis system. Instead of storing preset dance moves and selecting them based on music characteristics, the system analyzes audio signals in real-time and generates corresponding dance moves through computational processing, eliminating the need for large memory databases.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If predetermined hand-crafted features are used to select dance moves from the database, then dance moves can be generated, but the avatar cannot dance creatively

Engineering Contradiction:
Improvedance move generation efficiencyVSAvoidcreative dance capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic dance move generation that adapts to the characteristics of the input music. The system extracts rhythmic features from the audio signal and uses these features to dynamically generate dance moves that match the music's tempo, beat patterns, and rhythm, rather than selecting from static preset combinations. This dynamic approach enables creative dance performances tailored to each music input.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If a large database of preset dance moves is maintained, then dance moves can be selected and recombined, but the system complexity increases

Engineering Contradiction:
Improvedance move recombination capabilityVSAvoidsystem implementation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts only the essential rhythmic characteristics from audio signals (beat timing, tempo, rhythm patterns) rather than maintaining complex databases of dance moves. This extraction focuses on the minimum necessary information to generate dance moves, significantly simplifying the system while preserving the core functionality of creating music-synchronized dance performances.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11321891B2Method for generating action according to audio signal and electronic device
Publication Date: 2022.05.03 HTC CORP
  • US11321891B2 patent drawing
  • US11321891B2 patent drawing
  • US11321891B2 patent drawing

AI summary

The disclosure provides a method for generating action according to an audio signal and an electronic device. The method includes: receiving an audio signal and extracting a high-level audio feature therefrom; extracting a latent audio feature from the high-level audio feature; in response to determining that the audio signal corresponds to a beat, obtaining a joint angle distribution matrix based on the latent audio feature; in response to determining that the audio signal corresponds to a music, obtaining a plurality of designated joint angles corresponding to a plurality of joint points based on the joint angle distribution matrix; and adjusting a joint angle of each of the joint points on the avatar according to the designated joint angles.