3D Avatar Generation from Voice Features for Impression Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for creating user-specific avatars in virtual spaces, such as the metaverse, are costly in time and resources, and fail to accurately reflect unique user elements, particularly when voice and appearance impressions mismatch.

Innovation Solution

An information processing device that acquires user voice data, analyzes it for voice features, and generates a 3D avatar based on impression word scores derived from these features, allowing for personalized avatar creation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a designer is asked to create an avatar or the user selects parts to create an avatar, then the avatar can be customized, but it involves time and financial costs

Engineering Contradiction:
Improveavatar customizationVSAvoidtime cost
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system enables users to automatically generate their own avatars by analyzing their voice characteristics. The voice acquisition unit captures user voice, the voice analysis unit extracts voice features, and the 3D avatar generation unit creates the avatar automatically, allowing users to serve themselves without designer intervention or manual part selection

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of avatar creation (designer work or user part selection) with an automated acoustic field-based system. Voice data is analyzed to extract features, which are then converted into 3D avatar parameters, substituting human creative work with automated signal processing and data transformation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If a designer is asked to create an avatar or the user selects parts to create an avatar, then the avatar can be customized, but it involves financial costs

Engineering Contradiction:
Improveavatar customizationVSAvoidfinancial cost
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The system enables users to automatically generate their own avatars by analyzing their voice characteristics. The voice acquisition unit captures user voice, the voice analysis unit extracts voice features, and the 3D avatar generation unit creates the avatar automatically, allowing users to serve themselves without designer intervention or manual part selection

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of avatar creation (designer work or user part selection) with an automated acoustic field-based system. Voice data is analyzed to extract features, which are then converted into 3D avatar parameters, substituting human creative work with automated signal processing and data transformation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Manufacturing precision

If an avatar is generated based on a face image, then the avatar reproduces the user's face, but it is difficult to reflect elements unique to the user

Engineering Contradiction:
Improveface reproduction accuracyVSAvoidunique user elements
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent changes the basis for avatar generation from visual face image parameters to acoustic voice feature parameters. By analyzing voice characteristics such as pitch, tone, and speech patterns, the system generates avatars that reflect unique user elements inherent in their voice, providing a different dimensional approach to personalization

Inventive Principle:
Principle #35Parameter changes

4Adaptability or versatility

If an avatar is made to speak using the voice of the user, then the avatar can communicate with the user's voice, but a mismatch may occur between the impression of the voice and the impression of the avatar appearance

Engineering Contradiction:
Improvevoice integrationVSAvoidimpression consistency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system uses voice analysis to extract features from the user's voice and feeds this information back into the avatar generation process. The voice analysis unit processes acoustic characteristics, and this feedback is used by the 3D avatar generation unit to create an avatar whose appearance parameters align with the voice characteristics, ensuring impression consistency

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the basis for avatar generation from visual face image parameters to acoustic voice feature parameters. By analyzing voice characteristics such as pitch, tone, and speech patterns, the system generates avatars that reflect unique user elements inherent in their voice, providing a different dimensional approach to personalization

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250371775A1Information processing device, information processing method, and recording medium
Publication Date: 2025.12.04 SONY GROUP CORP
  • US20250371775A1 patent drawing
  • US20250371775A1 patent drawing
  • US20250371775A1 patent drawing

AI summary

The present technology relates to an information processing device, an information processing method, and a recording medium that are capable of generating a 3D avatar according to characteristics of a voice of a user.An information processing device according to one aspect of the present technology acquires voice data of a user; calculates voice features based on a result of analyzing the voice data of the user; and generates a 3D avatar having an appearance according to at least one of a plurality of impression word scores calculated on the basis of the features. The present technology can be applied to processing of generating a 3D avatar.