Voice Analysis Synthesis Phase Coherence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice analysis/synthesis apparatuses often produce synthesized voice waveforms that give an impression of phase discrepancy, leading to a sense of phasiness or reverberation, due to the inability to maintain vertical phase coherence (VPC) during pitch scaling, especially when the scaling factor is not an integer.

Innovation Solution

The apparatus analyzes voice waveforms in units of frames, calculates phase differences between frames, and uses these differences to maintain appropriate phase relationships between frequency channels, ensuring that the phase of the synthesized voice waveform is relative to a standard frequency channel, thereby avoiding phase discrepancies by correctly unwrapping and converting phase differences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If phase scaling is performed using conventional methods with frame-based DFT/FFT, then pitch scaling and time scaling can be achieved, but vertical phase coherence is lost causing phase discrepancy and phasiness in synthesized voice

Engineering Contradiction:
Improvepitch scaling capabilityVSAvoidphase coherence accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent introduces an intermediary calculation process that computes phase differences between adjacent frames and uses these differences to reconstruct phases after pitch scaling. This intermediary phase difference calculation acts as a mediator that preserves the vertical phase coherence relationship between frequency channels even when the scaling factor is not an integer, thereby preventing phase discrepancy and phasiness in the synthesized voice.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If non-integer scaling factors are used for pitch scaling, then more flexible pitch adjustment is achieved, but phase relationships between frequency channels become incorrect causing unnatural sound

Engineering Contradiction:
Improvepitch adjustment flexibilityVSAvoidphase relationship accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements a feedback mechanism where the phase differences calculated from adjacent frames are used to continuously adjust and maintain correct phase relationships during pitch scaling. By calculating phase differences, applying the scaling factor, and then using these scaled phase differences to determine new phases, the system creates a feedback loop that ensures phase coherence is preserved even with non-integer scaling factors, preventing unnatural sound in the synthesized output.

Inventive Principle:
Principle #23Feedback

3Productivity

If frame-based processing is used for voice analysis, then computational efficiency is improved, but phase continuity between frames is disrupted causing phase wrapping issues

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidphase continuity
Core Design Contradiction:
ProductivityVSStability of the object's composition

Solution Approach 1:

The patent applies preliminary action by calculating and storing phase differences between adjacent frames before performing the actual pitch scaling operation. By pre-computing these phase differences and maintaining them as intermediate values, the system prepares the necessary information to reconstruct continuous phases after scaling, thereby preventing phase wrapping issues while maintaining the computational efficiency of frame-based processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7672835B2Voice analysis/synthesis apparatus and program
Publication Date: 2010.03.02 CASIO COMPUTER CO LTD
  • US7672835B2 patent drawing
  • US7672835B2 patent drawing
  • US7672835B2 patent drawing

AI summary

An FFT unit performs an FFT process on high-frequency-eliminated, pitch-shifted voice data for one frame. A time scaling unit calculates a frequency amplitude, a phase, a phase difference between the present and immediately preceding frames, and an unwrapped version of the phase difference for each channel from which the frequency component was obtained by the FFT, detects a reference channel based on a peak one of the frequency amplitudes, and calculates the phase of each channel in a synthesized voice based on the reference channel, using results of the calculation. An IFFT unit processes each frequency component in accordance with the calculated phase, performs an IFFT process on the resulting frequency component, and produces synthesized voice data for one frame.