Voice Quality Conversion Using Target Speaker Determination
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice processing technologies face challenges in facilitating user-desired voice quality conversions that retain the atmosphere of conversational situations while effectively erasing the personality of other speakers' voices, making it difficult to instruct and achieve the desired voice quality conversion.
Innovation Solution
A voice processing apparatus and method that includes a voice quality determining unit to determine a target speaker determining method based on a control value, allowing for the conversion of voice quality by generating a voice that matches the desired target speaker's quality, using a voice quality model learned from multiple speakers' voices, and adjusting the frequency envelope to maintain conversation context.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If voice quality conversion is performed to erase other person's voice, then privacy protection is improved, but conversation naturalness deteriorates
Solution Approach 1:
The patent applies local quality by selectively processing only the voice quality attributes of other speakers while preserving the conversational context and structure. The voice quality conversion unit transforms the voice quality of identified other speakers into predetermined qualities (e.g., robot-like, child-like) without altering the conversation flow, timing, or interaction patterns, thus maintaining naturalness while achieving privacy protection.
Solution Approach 2:
The patent changes voice quality parameters (such as fundamental frequency, formant frequencies, and spectral characteristics) of other speakers' voices to transform their identities. By modifying these acoustic parameters while keeping the conversational framework intact, the system achieves both privacy protection through voice transformation and naturalness through contextual preservation.
2Ease of operation
If simple target speaker instruction is used, then ease of operation is improved, but voice quality conversion accuracy deteriorates
Solution Approach 1:
The system performs self-service by automatically identifying other speakers in the conversation and determining appropriate target speakers for voice quality conversion without requiring detailed user input. The voice processing apparatus autonomously analyzes the audio stream, identifies speakers using voiceprint recognition, and selects conversion targets based on the conversation context, thereby maintaining ease of operation while achieving accurate conversion results.
Solution Approach 2:
The patent implements feedback mechanisms where the system continuously monitors the conversation, identifies speakers in real-time, and adjusts the voice quality conversion targets based on the ongoing interaction. This feedback loop ensures that the voice quality conversion accurately reflects the current conversational context while requiring minimal user intervention.
3Device complexity
If voice quality conversion model is generated without conversion coefficients for multiple speaker pairs, then device complexity is reduced, but voice quality conversion versatility deteriorates
Solution Approach 1:
The patent applies universality by creating a single voice quality conversion model that can handle multiple speaker pairs and various conversion scenarios. The model is trained to recognize and transform different voice characteristics, allowing it to convert voices of multiple other speakers into various predetermined qualities using one unified framework, thus achieving high versatility without proportionally increasing device complexity.
Solution Approach 2:
The system uses an intermediary voice quality conversion model that serves as a mediator between the input audio and the desired output. This intermediate model, trained on diverse speaker data, enables the system to handle multiple conversion cases by translating different voice qualities into a standardized representation, thereby achieving versatility while maintaining manageable complexity through the intermediary layer.
Data Source
AI summary
A voice processing apparatus includes a voice quality determining unit configured to determine a target speaker determining method used for a voice quality conversion in accordance with a determining method control value for instructing the target speaker determining method of determining a target speaker whose voice quality is targeted to the voice quality conversion, and determine the target speaker in accordance with the target speaker determining method.


