Avatar Voice Processing via Image Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in intuitively customizing avatar voice qualities to match the impression of an avatar image, particularly due to the limited number of presets available and the high cost of creating voice data.

Innovation Solution

A voice processing device and method that extract feature values from avatar images and process voices based on these features, allowing for voice quality conversion or synthesis that matches the impression of the avatar image without requiring paired voice data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If preset voice qualities are selected from existing options, then voice processing can be performed quickly, but the number of available voice types is limited and individuality is low

Engineering Contradiction:
Improvevoice processing speedVSAvoidvoice type variety
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent transforms discrete preset voice selections into continuous voice quality adjustment by extracting feature values from avatar images and applying them to voice synthesis parameters. This allows dynamic modification of voice characteristics (pitch, timbre, tone) based on visual features, enabling unlimited voice customization rather than selecting from fixed presets.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces feature value extraction as an intermediary process between avatar image and voice output. By converting visual features into numerical representations that can modulate voice parameters, it creates a bridge that enables intuitive voice customization from image to audio without requiring direct voice data pairing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If custom voice qualities are created to match avatar impressions, then individuality is improved, but a large amount of voice data is required and costs increase

Engineering Contradiction:
Improvevoice customization capabilityVSAvoidvoice data requirement
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent copies visual feature information from avatar images and applies it to voice synthesis, rather than requiring separate voice recording sessions. By extracting features like color, shape, and texture from the avatar image and mapping them to voice characteristics, it creates voice qualities that match the avatar impression without needing corresponding voice data for each avatar.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of voice recording and data collection with an automated feature extraction and synthesis system. Instead of manually recording voice data for each avatar (mechanical process), the system automatically extracts features from images and generates matching voices through computational processing, eliminating the need for large voice datasets.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Device complexity

If fixed filter processing is applied to user voice, then processing is simple and fast, but the voice quality cannot be customized to match avatar impression

Engineering Contradiction:
Improveprocessing complexityVSAvoidvoice quality matching capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent transforms static fixed filter processing into dynamic adaptive processing. Instead of applying predetermined filters, the system dynamically adjusts voice synthesis parameters based on real-time extraction of avatar image features. This allows the processing complexity to increase only as needed to achieve the desired voice-avatar matching, rather than using overly complex fixed systems.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250140280A1Voice processing device, voice processing method, information terminal, information processing device, and computer program
Publication Date: 2025.05.01 SONY GROUP CORP
  • US20250140280A1 patent drawing
  • US20250140280A1 patent drawing
  • US20250140280A1 patent drawing

AI summary

A voice processing device that performs processing related to generation of a voice of an avatar image. The voice processing device includes: an extraction unit that extracts a feature value of an avatar image; and a processing unit that converts a voice quality of an input voice on the basis of the feature value in the feature value space or synthesizes a voice on the basis of the feature value in the feature value space. The extraction unit extracts the feature value of the avatar image by using a feature value extractor designed such that a feature value extracted from a voice and a feature value extracted from an avatar image created from a face image of a speaker who has uttered the voice share the same feature value space and are close feature values on the space.