Personalized Digital Singer Generation With AI Voice Cloning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing virtual avatars lack personalization and interest, as they only allow for appearance changes without incorporating user talents or skills.

Innovation Solution

A metaverse personalized digital singer generation system that includes a display device, voice database host, and server-end device, utilizing 3D imaging and generative AI to create a personalized digital singer by processing user voice and facial data, training a generative pre-training model, and adjusting singing parameters to match a selected song, while providing vocal coaching.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If virtual avatar appearance is changed through clothes and bodies, then personalization is improved, but the avatar still lacks talents and skills, making personalization insufficient

Engineering Contradiction:
ImprovepersonalizationVSAvoidpersonalization depth
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments personalization into multiple dimensions: appearance customization (clothes, bodies) and talent/skill integration (singing abilities, voice characteristics). By separating these aspects, the system can independently process and integrate user data for both visual and functional personalization, resolving the contradiction between superficial and deep personalization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from two-dimensional appearance customization to three-dimensional personalization by adding the dimension of talent and skill integration. Through generative AI models, the system incorporates user voice recordings and singing performances to create digital avatars with authentic vocal characteristics, moving beyond visual changes to functional personalization.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If generative AI model is trained with user voice data, then personalization is improved, but training time and computational resources increase

Engineering Contradiction:
ImprovepersonalizationVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary voice data collection and processing during user onboarding, storing standardized voice samples and singing performances in advance. This preliminary action prepares the training data beforehand, reducing the actual training time when the digital singer needs to be generated or modified.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses a partial training approach by focusing on key voice characteristics and singing patterns rather than training on all possible vocal variations. The generative AI model is trained on essential features (pitch, timbre, rhythm) to achieve sufficient personalization without requiring exhaustive training data, thus reducing training time and computational resources.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250371995A1Metaverse Personalized Digital Singer Generation System and Method Thereof
Publication Date: 2025.12.04 SQ TECH (SHANGHAI) CORP
  • US20250371995A1 patent drawing
  • US20250371995A1 patent drawing
  • US20250371995A1 patent drawing

AI summary

A metaverse personalized digital singer generation system and a method thereof. In the system, the server-end device receives a user voice, store the user voice as a personalized voice, capture an image of a user face to generate a facial image, generate a personalized digital singer displayed in a virtual scene through a 3D imaging technology, and convert the personalized voice into voice feature vectors, and use the voice feature vectors and the personalized voice as training data, input the training data to a generative AI model to train a generative pre-training model having the personal characteristics. When the user selects an original song for singing, the original song and the user singing voice and the prompt are inputted to the generative pre-training model, the remixed song matching a style of the original song is outputted, a vocal coaching is generated and displayed based on prompt.