Voice Synthesizer Spectral Envelope Merging for Unison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice synthesis technologies are complex and inefficient in generating a unison effect of multiple performers, requiring large circuit scales and excessive processing loads when attempting to convert source voice characteristics for multiple voices.
Innovation Solution
A voice synthesizer that includes a data acquisition portion, envelope acquisition portion, spectrum acquisition portion, envelope adjustment portion, and voice generation portion to obtain and adjust spectral envelopes of voice segments, allowing for the generation of an output voice signal that approximates multiple voices without the need for independent conversion of each voice segment, using a simple configuration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If independent voice characteristic conversion is performed for each voice in a unison, then the phonetic entity accuracy of each voice is improved, but the device complexity and processing load increase significantly
Solution Approach 1:
The patent merges the voice characteristic conversion process by performing a single spectral envelope conversion on the mixed signal of multiple voices instead of converting each voice independently. This combining approach maintains phonetic entity accuracy while significantly reducing device complexity and processing load.
Solution Approach 2:
The patent creates a universal conversion process that handles multiple voices simultaneously through a single spectral envelope operation. This multi-functional approach allows the same conversion mechanism to serve all voices in the unison, eliminating the need for separate conversion circuits for each voice.
2Measurement precision
If independent voice characteristic conversion is performed for each voice in a unison, then the phonetic entity accuracy of each voice is improved, but the processing load becomes excessive
Solution Approach 1:
The patent merges the voice characteristic conversion process by performing a single spectral envelope conversion on the mixed signal of multiple voices instead of converting each voice independently. This combining approach maintains phonetic entity accuracy while significantly reducing device complexity and processing load.
Solution Approach 2:
The patent uses spectral envelope copying where the spectral characteristics from a reference voice are applied to the mixed signal, creating the unison effect without requiring independent conversion of each voice. This copying mechanism dramatically reduces processing requirements while maintaining accuracy.
3Device complexity
If a simple configuration is used for voice synthesis, then the device complexity is reduced, but the ability to generate accurate unison effect is compromised
Solution Approach 1:
The patent merges the voice characteristic conversion process by performing a single spectral envelope conversion on the mixed signal of multiple voices instead of converting each voice independently. This combining approach maintains phonetic entity accuracy while significantly reducing device complexity and processing load.
Solution Approach 2:
The patent introduces spectral envelope as an intermediary that bridges the simple mixed signal processing and the accurate phonetic entity representation. This intermediary allows the simple configuration to achieve accurate unison effect by manipulating the spectral characteristics of the combined signal.
Data Source
AI summary
In a voice synthesizer, an envelope acquisition portion obtains a spectral envelope of a reference frequency spectrum of a given voice. A spectrum acquisition portion obtains a collective frequency spectrum of a plurality of voices which are generated in parallel to one another. An envelope adjustment portion adjusts a spectral envelope of the collective frequency spectrum obtained by the spectrum acquisition portion so as to approximately match with the spectral envelope of the reference frequency spectrum obtained by the envelope acquisition portion. A voice generation portion generates an output voice signal from the collective frequency spectrum having the spectral envelope adjusted by the envelope adjustment portion.


