Voice Processing Apparatus Frequency-Phase Relationship
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice signal processing technologies face difficulties in generating natural voices with appropriate frequency-phase relationships, particularly for unique voices like thick or hoarse voices, due to variations in ambient components over time.
Innovation Solution
A voice processing apparatus adjusts the fundamental frequency of a target voice signal to match that of a reference voice signal in the time domain, then divides and allocates harmonic band components to corresponding frequencies, adjusting their envelopes and phases to maintain a natural frequency-phase relationship between harmonic and ambient components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If fundamental frequency is converted by shifting band components in the frequency domain, then voice characteristics can be converted, but the frequency-phase relationship between harmonic components and ambient components cannot be appropriately maintained
Solution Approach 1:
The patent segments the voice signal processing into distinct time-domain and frequency-domain operations. First, fundamental frequency conversion is performed in the time domain by adjusting the sampling rate, preserving the frequency-phase relationship. Then, in the frequency domain, the spectrum is divided into harmonic band components and processed separately, allowing voice characteristic conversion while maintaining the relationships established in the time domain.
Solution Approach 2:
The patent performs preliminary fundamental frequency adjustment in the time domain before proceeding to frequency-domain processing. By establishing the correct frequency-phase relationship first through time-domain resampling, subsequent frequency-domain operations can focus on voice characteristic conversion without disrupting the already-established phase relationships.
2Manufacturing precision
If phase of ambient component is adjusted separately from harmonic component, then natural voice can be generated, but it becomes difficult for peculiar voices like thick or hoarse voices where ambient component varies greatly with time
Solution Approach 1:
The patent merges the handling of harmonic components and ambient components by processing both through the same time-domain fundamental frequency adjustment mechanism. Instead of separately adjusting phases of different components, the unified time-domain approach automatically maintains appropriate frequency-phase relationships for all components simultaneously, reducing complexity while handling peculiar voices effectively.
3Adaptability or versatility
If harmonic band components are divided and allocated to harmonic frequencies, then voice characteristics are preserved, but complex processing is required to maintain natural frequency-phase relationships
Solution Approach 1:
The patent segments the spectrum into harmonic band components for targeted processing, allowing preservation of voice characteristics through selective manipulation of different frequency regions. This segmentation enables efficient processing by focusing operations on specific harmonic bands rather than treating the entire spectrum uniformly.
Solution Approach 2:
By performing preliminary fundamental frequency adjustment in the time domain before frequency-domain segmentation and processing, the patent establishes correct phase relationships in advance. This preliminary action simplifies subsequent processing of individual harmonic bands, as the phase relationships are already properly configured and only minor adjustments are needed during frequency-domain operations.
Data Source
AI summary
In a voice processing apparatus, a processor is configured to adjust, a fundamental frequency of a first voice signal corresponding to a voice having target voice characteristics to a fundamental frequency of a second voice signal corresponding to a voice having initial voice characteristics different from the target voice characteristics. The processor is further configured to sequentially generate a processed spectrum based on a spectrum of the first voice signal and a spectrum of the second voice signal by: dividing the spectrum of the first voice signal into a plurality of harmonic band components after the fundamental frequency of the first voice signal has been adjusted; allocating each harmonic band component of the first voice signal to each harmonic frequency associated with the fundamental frequency of the second voice signal; and adjusting an envelope and phase of each harmonic band component according to the spectrum of the second voice signal.


