Acoustic Fingerprinting for Incremental Speech Model Updates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Unit selection speech synthesis systems face inefficiencies and high computational costs when updating databases, leading to unreliable speech synthesis due to frequent linguistic and acoustic changes, as rebuilding the probabilistic language model is resource-intensive and prone to selecting invalid features.

Innovation Solution

The system generates acoustic fingerprints and uses similarity measures to update the model, allowing for incremental updates without rebuilding the entire model, ensuring persistence of the original database and accommodating frequent changes by associating updated units with similar fingerprints and existing probability estimates.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the probabilistic language model is rebuilt from an updated training utterances database, then the speech synthesis reliability is improved, but the computational cost and time consumption increase significantly

Engineering Contradiction:
Improvespeech synthesis reliabilityVSAvoidmodel rebuilding time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the language model into two parts: a persistent base model built from the original training database, and an incremental update component that processes only the updated acoustic units. This segmentation allows the system to maintain reliability through the complete model while avoiding the time cost of rebuilding the entire model from scratch.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-building the probabilistic language model from the original training utterances database and storing it in a persistent format. This preliminary model serves as a foundation that can be efficiently updated later without requiring complete reconstruction, thus reducing model rebuilding time while maintaining reliability.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If acoustic units are simply added or removed from the data store, then the system adaptability to linguistic changes is improved, but the speech synthesis reliability deteriorates due to selection of invalid features

Engineering Contradiction:
Improveadaptability to linguistic changesVSAvoidspeech synthesis reliability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where the system continuously monitors the training utterances database for updates and automatically triggers incremental model updates. This feedback loop ensures that the language model remains synchronized with the current state of the acoustic units database, maintaining reliability while adapting to linguistic changes.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent makes the language model dynamic by enabling incremental updates rather than treating it as a static structure. The model can adapt its parameters and probability distributions in response to database updates, allowing the system to maintain reliability while being adaptable to linguistic and acoustic changes.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If the entire probabilistic language model is rebuilt frequently, then the speech synthesis accuracy is improved, but the computational resources required increase

Engineering Contradiction:
Improvespeech synthesis accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by updating only the necessary portions of the language model that are affected by database changes, rather than rebuilding the entire model. This approach maintains speech synthesis accuracy for the updated units while significantly reducing computational resource consumption compared to full model reconstruction.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9424835B2Statistical unit selection language models based on acoustic fingerprinting
Publication Date: 2016.08.23 GOOGLE LLC
  • US9424835B2 patent drawing
  • US9424835B2 patent drawing
  • US9424835B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for providing statistical unit selection language modeling based on acoustic fingerprinting. The methods, systems and apparatus include the actions of obtaining a unit database of acoustic units and, for each acoustic unit, linguistic data corresponding to the acoustic unit; obtaining stored data associating each acoustic unit with (i) a corresponding acoustic fingerprint and (ii) a probability of the linguistic data corresponding to the acoustic unit occurring in a text corpus; determining that the unit database of acoustic units has been updated to include one or more new acoustic units; for each new acoustic unit in the updated unit database: generating an acoustic fingerprint for the new acoustic unit; identifying an acoustic unit that (i) has an acoustic fingerprint that is indicated as similar to the fingerprint of the new acoustic unit, and (ii) has a stored associated probability.