Optimal Crossing Point for Speech Sample Concatenation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech (TTS) synthesis technologies face limitations in generating natural-sounding speech due to the need for large storage spaces and inaccurate concatenation of speech samples, which results in unnatural speech production when samples are concatenated without considering pitch and duration contexts.

Innovation Solution

A method and system for identifying an optimal crossing point for concatenating speech samples by analyzing spectral similarity and correlation in both time and frequency domains to determine the best overlap area for seamless concatenation, ensuring maximum correlation and minimizing inaccuracies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech samples are stored in large quantities to enrich the speech repository, then the quality and variety of synthesized speech is improved, but the storage space requirement and processing time increase significantly

Engineering Contradiction:
Improvespeech repository richnessVSAvoidstorage space
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent changes the parameter of spectral similarity threshold dynamically. By adjusting the threshold for determining spectral similarity between speech samples, the system can optimize the balance between repository richness and storage efficiency. Higher thresholds allow more samples to be stored with less space, while lower thresholds ensure higher quality matching for synthesis.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If speech samples are concatenated without considering pitch and duration context, then the processing speed is improved, but the naturalness of synthesized speech deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidspeech concatenation accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent applies preliminary action by pre-analyzing and storing pitch and duration information associated with each speech sample in the repository. When concatenation is needed, the system can quickly retrieve samples that match the required pitch and duration context without performing complex analysis during real-time synthesis, thus maintaining both speed and accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts the selection criteria for speech samples based on the contextual requirements of pitch and duration. By changing the parameters used for sample selection (from simple spectral matching to multi-parameter matching including pitch and duration), the system ensures natural-sounding concatenation while managing processing complexity.

Inventive Principle:
Principle #35Parameter changes

3Extent of automation

If spectral clustering techniques are used to determine spectral similarity, then the automation of speech sample selection is improved, but the measurement precision of similarity determination deteriorates

Engineering Contradiction:
Improvespeech sample selection automationVSAvoidspectral similarity determination accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where the system evaluates the quality of concatenated speech segments and uses this information to refine future sample selections. By incorporating feedback from actual synthesis results, the system can adjust its spectral similarity determination criteria to improve precision while maintaining automation.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system uses a composite approach by combining multiple techniques for determining spectral similarity, including short-time Fourier transforms (STFT), Mel-frequency cepstral coefficients (MFCC), and other spectral analysis methods. This composite methodology leverages the strengths of each technique to achieve both automation and precision in sample selection.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS9251782B2System and method for concatenate speech samples within an optimal crossing point
Publication Date: 2016.02.02 OSR ENTERPRISES
  • US9251782B2 patent drawing
  • US9251782B2 patent drawing
  • US9251782B2 patent drawing

AI summary

A method for identifying an optimal crossing point for concatenation of speech samples within an overlap area is provided. The method includes retrieving a first speech sample and a second speech sample, the second speech sample is concatenated immediately after the first speech sample is concatenated; determining a first region within the ending of the first speech sample and a second region within the beginning of the second speech sample, the first region and the second region are determined respective of relatively high spectral similarity over time between the first speech sample and the second speech sample; identifying an overlap region between the first region and the second region; determining an optimal crossing point between the first speech sample and the second speech sample, the optimal crossing point has a maximum correlation over time; and concatenating the first speech sample and the second speech sample at the optimal crossing point.