Speech Quality Estimation via Pitchmark Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech synthesis technologies lack an objective and automatic method for estimating speech quality after prosody modification, particularly for pitch-synchronous methods like TD-PSOLA, which restricts the range of prosody modification and requires large memory spaces, and existing methods like PSQM and PESQ are not suitable for estimating quality changes due to prosody modifications.

Innovation Solution

A method for estimating speech quality degradation that extracts source pitchmarks, maps them to target pitchmarks, calculates degradation measures using weighting functions, and computes objective speech quality scores without synthesizing the target speech, incorporating pitch-related and duration-related degradation measures through regression or probabilistic models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If TD-PSOLA is used for prosody modification, then speech quality can be maintained within a limited modification range, but the system cannot automatically predict quality when prosody difference is large

Engineering Contradiction:
Improvespeech quality estimation accuracyVSAvoidprosody modification range
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent replaces subjective human evaluation with an objective automatic estimation system. It uses spectral analysis, pitch contour comparison, and energy distribution metrics to quantify speech quality degradation, substituting mechanical/physical measurement methods for subjective perception assessment.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces intermediate parameters including spectral distance metrics, pitch contour deviation, and energy distribution differences as mediators between source and target prosody. These intermediaries enable automatic quality prediction by bridging the gap between different prosody representations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If a large speech database is used to contain all tones and prosodies, then speech synthesis quality improves, but memory space requirements increase significantly

Engineering Contradiction:
Improvespeech synthesis qualityVSAvoidcorpus size
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts key quality-determining features from speech units including spectral characteristics, pitch contours, and energy distribution. By separating these essential features from the complete speech database, the system can estimate quality without storing or processing entire speech corpora, significantly reducing memory requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary extraction and organization of quality-relevant features from speech units before synthesis. By pre-computing spectral, pitch, and energy characteristics and storing them as compact representations, the system enables rapid quality estimation without requiring access to large amounts of raw speech data during synthesis operations.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If existing methods like PSQM or PESQ are used for quality estimation, then spectral differences can be measured, but these methods are not suitable for prosody-modified speech because the spectrum always changes

Engineering Contradiction:
Improvespectral difference measurementVSAvoidquality estimation suitability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent shifts from global spectral analysis to local feature-specific analysis. It separately evaluates pitch contour accuracy, spectral distance in critical bands, and energy distribution in formant regions. This localized approach allows quality assessment of specific prosodic parameters independent of overall spectral changes, making it suitable for prosody-modified speech.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS7801725B2Method for speech quality degradation estimation and method for degradation measures calculation and apparatuses thereof
Publication Date: 2010.09.21 IND TECH RES INST
  • US7801725B2 patent drawing
  • US7801725B2 patent drawing
  • US7801725B2 patent drawing

AI summary

A method for speech quality degradation estimation, a method for degradation measures calculation, and the apparatuses thereof are provided. The first method above estimates the speech quality of a speech signal that is modified by a pitch-synchronous prosody modification method, which comprises the following steps. First, extract at least one source pitchmark from the speech signal, and then maps the source pitchmark(s) to at least one target pitchmark(s). Finally, calculate at least one degradation measure based on the mapping between the source and the target pitchmarks. The degradation measures include several weighted pitch-related functions and duration-related functions, where the weighting functions can be calculated based on the speech signal or the pitchmark(s) mapping mentioned above.