Pruning Redundant Speech Synthesis Units for Embedded Footprint Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis systems face challenges in reducing their footprint for embedded devices while maintaining voice quality, as existing pruning methods either retain overly similar units or prune representative units due to high unit appearance frequency thresholds, leading to inadequate units for low-frequency phonetic contexts during synthesis.

Innovation Solution

The method employs a delta unit appearance frequency technique in conjunction with unit appearance frequency to prune redundant speech synthesis units, ensuring only high-frequency units that are irreplaceable are preserved, utilizing a similarity factor and unit cost in the unit-selection process to guide pruning decisions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a large speech database is used to improve voice quality, then the synthesis voice quality is improved, but the system footprint becomes too large for embedded systems

Engineering Contradiction:
Improvesynthesis voice qualityVSAvoidsystem footprint
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential and representative speech units from the large database by applying pruning criteria based on unit appearance frequency and similarity metrics. This extraction process removes redundant units while preserving the core linguistic content needed for high-quality synthesis, enabling the system to fit within embedded device memory constraints.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter of database size by applying selective pruning based on unit appearance frequency thresholds and similarity factors. This parameter transformation reduces the database from a large size suitable for server systems to a compact size appropriate for embedded systems while maintaining voice quality through intelligent selection of retained units.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If bottom-up pruning is used to reduce database size, then the system footprint is reduced, but redundant units are not effectively identified because similarity criteria are independent of unit-selection strategy

Engineering Contradiction:
Improvedatabase sizeVSAvoidunit selection accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent introduces feedback by using unit appearance frequency data derived from actual synthesis operations to guide the pruning process. This feedback loop ensures that units which are frequently selected during synthesis are preserved, while rarely selected redundant units are removed, aligning the pruning criteria with the unit-selection strategy and improving the accuracy of redundant unit identification.

Inventive Principle:
Principle #23Feedback

3Quantity of substance

If up-bottom pruning based on unit appearance frequency is used to reduce database size, then the system footprint is reduced, but representative units for low-frequency phonetic contexts are pruned away

Engineering Contradiction:
Improvedatabase sizeVSAvoidcoverage of speech units
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by differentiating the treatment of units based on their appearance frequency characteristics. High-frequency units are evaluated for redundancy using similarity metrics, while low-frequency units are preserved to ensure coverage of diverse phonetic contexts. This localized approach allows the system to reduce database size while maintaining adaptability across different speech scenarios.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9520123B2System and method for pruning redundant units in a speech synthesis process
Publication Date: 2016.12.13 CERENCE OPERATING CO
  • US9520123B2 patent drawing
  • US9520123B2 patent drawing
  • US9520123B2 patent drawing

AI summary

A system and method for concatenative speech synthesis is provided. Embodiments may include accessing, using one or more computing devices, a plurality of speech synthesis units from a speech database and determining a similarity between the plurality of speech synthesis units. Embodiments may further include retrieving two or more speech synthesis units having the similarity and pruning at least one of the two or more speech synthesis units based upon, at least in part, the similarity.