Voice Conversion Model Preserving Vocal Identity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional sound conversion systems, such as voice conversion systems, rely heavily on professional singer databases and fail to maintain the unique characteristics of the original sound being converted, resulting in a lack of authenticity in converted singing voices.
Innovation Solution
A machine-learned sound conversion system that infers a pitch contour from musical scores and uses this information to convert spoken words into singing voices while preserving the sound signatures and vocal identity of the original speaker, employing a multi-stage process involving pitch contour inference, conversion, and combination with accompaniment music.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If conventional voice conversion systems use professional singer databases, then the converted singing voices can be generated, but the unique characteristics of the original sound are lost
Solution Approach 1:
The system creates a voice conversion model that copies the vocal characteristics from a reference audio sample and applies it to the target speech input. This allows the converted singing voice to retain the original speaker's unique sound signature while achieving the desired singing conversion, eliminating the need to rely on professional singer databases that would otherwise replace the original voice characteristics.
Solution Approach 2:
The system transforms the audio data by changing its parameters - converting speech to singing style while preserving the fundamental vocal characteristics through machine learning. The voice conversion model adjusts parameters such as pitch contour, timbre, and spectral features to achieve singing conversion while maintaining the original speaker's voice identity.
2Productivity
If conventional voice conversion systems transform speech to singing, then the singing voice can be produced, but artificial sounds are added and authenticity is reduced
Solution Approach 1:
The voice conversion model is trained to learn the natural mapping from speech to singing characteristics automatically from audio data, without requiring manual annotation of singing notes or professional singer databases. The system serves itself by learning the conversion process directly from the data, producing more authentic results that avoid artificial sound additions.
3Ease of manufacture
If conventional systems rely on annotated singing notes and professional singers, then the conversion process can be performed, but the system complexity and resource requirements increase
Solution Approach 1:
The system extracts only the essential vocal characteristics from a short reference audio sample (a few seconds of singing) and uses this extracted information to perform the voice conversion. This eliminates the need for large professional singer databases and extensive annotated singing notes, significantly reducing system complexity and resource requirements while maintaining conversion capability.
Data Source
AI summary
Systems, devices, media, and methods are presented for converting sounds in an audio stream. The systems and methods receive an audio conversion request initiating conversion of one or more sound characteristics of an audio stream from a first state to a second state. The systems and methods access an audio conversion model associated with an audio signature for the second state. The audio stream is converted based on the audio conversion model and an audio construct is compiled from the converted audio stream and a base audio segment. The compiled audio construct is presented at a client device.


