Voice Quality Conversion Using Target Speaker Determination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice processing technologies face challenges in facilitating user-desired voice quality conversions that retain the atmosphere of conversational situations while effectively erasing the personality of other speakers' voices, making it difficult to instruct and achieve the desired voice quality conversion.

Innovation Solution

A voice processing apparatus and method that includes a voice quality determining unit to determine a target speaker determining method based on a control value, allowing for the conversion of voice quality by generating a voice that matches the desired target speaker's quality, using a voice quality model learned from multiple speakers' voices, and adjusting the frequency envelope to maintain conversation context.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If voice quality conversion is performed to erase other person's voice, then privacy protection is improved, but conversation naturalness deteriorates

Engineering Contradiction:
Improveprivacy protectionVSAvoidconversation naturalness
Core Design Contradiction:
Object-affected harmful factorsVSStability of the object's composition

Solution Approach 1:

The patent applies local quality by selectively processing only the voice quality attributes of other speakers while preserving the conversational context and structure. The voice quality conversion unit transforms the voice quality of identified other speakers into predetermined qualities (e.g., robot-like, child-like) without altering the conversation flow, timing, or interaction patterns, thus maintaining naturalness while achieving privacy protection.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes voice quality parameters (such as fundamental frequency, formant frequencies, and spectral characteristics) of other speakers' voices to transform their identities. By modifying these acoustic parameters while keeping the conversational framework intact, the system achieves both privacy protection through voice transformation and naturalness through contextual preservation.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If simple target speaker instruction is used, then ease of operation is improved, but voice quality conversion accuracy deteriorates

Engineering Contradiction:
Improveease of operationVSAvoidvoice quality conversion accuracy
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system performs self-service by automatically identifying other speakers in the conversation and determining appropriate target speakers for voice quality conversion without requiring detailed user input. The voice processing apparatus autonomously analyzes the audio stream, identifies speakers using voiceprint recognition, and selects conversion targets based on the conversation context, thereby maintaining ease of operation while achieving accurate conversion results.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent implements feedback mechanisms where the system continuously monitors the conversation, identifies speakers in real-time, and adjusts the voice quality conversion targets based on the ongoing interaction. This feedback loop ensures that the voice quality conversion accurately reflects the current conversational context while requiring minimal user intervention.

Inventive Principle:
Principle #23Feedback

3Device complexity

If voice quality conversion model is generated without conversion coefficients for multiple speaker pairs, then device complexity is reduced, but voice quality conversion versatility deteriorates

Engineering Contradiction:
Improvedevice complexityVSAvoidvoice quality conversion versatility
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by creating a single voice quality conversion model that can handle multiple speaker pairs and various conversion scenarios. The model is trained to recognize and transform different voice characteristics, allowing it to convert voices of multiple other speakers into various predetermined qualities using one unified framework, thus achieving high versatility without proportionally increasing device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses an intermediary voice quality conversion model that serves as a mediator between the input audio and the desired output. This intermediate model, trained on diverse speaker data, enables the system to handle multiple conversion cases by translating different voice qualities into a standardized representation, thereby achieving versatility while maintaining manageable complexity through the intermediary layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9711161B2Voice processing apparatus, voice processing method, and program
Publication Date: 2017.07.18 SONY GROUP CORP
  • US9711161B2 patent drawing
  • US9711161B2 patent drawing
  • US9711161B2 patent drawing

AI summary

A voice processing apparatus includes a voice quality determining unit configured to determine a target speaker determining method used for a voice quality conversion in accordance with a determining method control value for instructing the target speaker determining method of determining a target speaker whose voice quality is targeted to the voice quality conversion, and determine the target speaker in accordance with the target speaker determining method.