Acoustic Fingerprinting for Incremental Speech Model Updates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Unit selection speech synthesis systems face inefficiencies and high computational costs when updating databases, leading to unreliable speech synthesis due to frequent linguistic and acoustic changes, as rebuilding the probabilistic language model is resource-intensive and prone to selecting invalid features.
Innovation Solution
The system generates acoustic fingerprints and uses similarity measures to update the model, allowing for incremental updates without rebuilding the entire model, ensuring persistence of the original database and accommodating frequent changes by associating updated units with similar fingerprints and existing probability estimates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the probabilistic language model is rebuilt from an updated training utterances database, then the speech synthesis reliability is improved, but the computational cost and time consumption increase significantly
Solution Approach 1:
The patent segments the language model into two parts: a persistent base model built from the original training database, and an incremental update component that processes only the updated acoustic units. This segmentation allows the system to maintain reliability through the complete model while avoiding the time cost of rebuilding the entire model from scratch.
Solution Approach 2:
The patent performs preliminary actions by pre-building the probabilistic language model from the original training utterances database and storing it in a persistent format. This preliminary model serves as a foundation that can be efficiently updated later without requiring complete reconstruction, thus reducing model rebuilding time while maintaining reliability.
2Adaptability or versatility
If acoustic units are simply added or removed from the data store, then the system adaptability to linguistic changes is improved, but the speech synthesis reliability deteriorates due to selection of invalid features
Solution Approach 1:
The patent implements feedback mechanisms where the system continuously monitors the training utterances database for updates and automatically triggers incremental model updates. This feedback loop ensures that the language model remains synchronized with the current state of the acoustic units database, maintaining reliability while adapting to linguistic changes.
Solution Approach 2:
The patent makes the language model dynamic by enabling incremental updates rather than treating it as a static structure. The model can adapt its parameters and probability distributions in response to database updates, allowing the system to maintain reliability while being adaptable to linguistic and acoustic changes.
3Measurement precision
If the entire probabilistic language model is rebuilt frequently, then the speech synthesis accuracy is improved, but the computational resources required increase
Solution Approach 1:
The patent applies partial action by updating only the necessary portions of the language model that are affected by database changes, rather than rebuilding the entire model. This approach maintains speech synthesis accuracy for the updated units while significantly reducing computational resource consumption compared to full model reconstruction.
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for providing statistical unit selection language modeling based on acoustic fingerprinting. The methods, systems and apparatus include the actions of obtaining a unit database of acoustic units and, for each acoustic unit, linguistic data corresponding to the acoustic unit; obtaining stored data associating each acoustic unit with (i) a corresponding acoustic fingerprint and (ii) a probability of the linguistic data corresponding to the acoustic unit occurring in a text corpus; determining that the unit database of acoustic units has been updated to include one or more new acoustic units; for each new acoustic unit in the updated unit database: generating an acoustic fingerprint for the new acoustic unit; identifying an acoustic unit that (i) has an acoustic fingerprint that is indicated as similar to the fingerprint of the new acoustic unit, and (ii) has a stored associated probability.


