Speech Vocoder Training for Multi-Sampling-Rate Waveform Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic apparatuses face limitations in outputting voice signals of various sampling rates due to the need for separate prosody and vocoder parts, which are trained universally and result in inflexible speech waveform generation.
Innovation Solution
An electronic apparatus with a processor that includes a prosody module to extract acoustic features and a vocoder module to generate speech waveforms, capable of modifying sampling rates and training multiple vocoder learning models to accommodate different specifications and output formats.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If separate prosody parts and vocoder parts are trained for each sampling rate, then speech waveform generation accuracy for each sampling rate is improved, but device complexity and training time increase
Solution Approach 1:
The patent applies universality by training a single prosody part and a single vocoder part that can handle multiple sampling rates. The prosody part extracts acoustic features that are then processed by the vocoder part to generate speech waveforms at different sampling rates (e.g., 16kHz, 22.05kHz, 44.1kHz) without requiring separate models for each rate, thus reducing device complexity while maintaining generation accuracy.
Solution Approach 2:
The patent uses parameter changes by modifying the sampling rate parameter during the speech waveform generation process. The system accepts a target sampling rate as input and adjusts the generation process accordingly, allowing the same prosody and vocoder parts to adapt to different sampling rates through parameter adjustment rather than requiring separate trained models for each rate.
2Adaptability or versatility
If multiple prosody parts and vocoder parts are included in one electronic apparatus, then support for various sampling rates is improved, but ease of manufacture and system simplicity deteriorate
Solution Approach 1:
The patent implements universality by designing a single prosody part and a single vocoder part that can universally handle multiple sampling rates. This approach improves adaptability to various sampling rates while simplifying the system structure, making it easier to manufacture and implement compared to having separate parts for each sampling rate.
Solution Approach 2:
The patent applies merging by combining the functionality of multiple sampling rate-specific prosody and vocoder parts into a single unified system. The single prosody part and single vocoder part work together to generate speech waveforms at different sampling rates, reducing the number of components and simplifying the overall system architecture.
3Device complexity
If a single prosody part and vocoder part are used for all sampling rates, then device complexity is reduced, but speech waveform generation accuracy for specific sampling rates may deteriorate
Solution Approach 1:
The patent uses parameter changes to maintain speech waveform generation accuracy while using a single prosody part and vocoder part. The system accepts a target sampling rate as a parameter and adjusts the generation process accordingly, allowing the unified model to adapt to different sampling rates and maintain accuracy without requiring separate specialized models for each rate.
Data Source
AI summary
An electronic apparatus, a terminal apparatus, and a controlling method thereof. The electronic apparatus includes an input interface; and a processor including a prosody module configured to extract an acoustic feature and a vocoder module configured to generate a speech waveform, wherein the processor is configured to: receive a text input using the input interface; identify a first acoustic feature from the text input using the prosody module, wherein the first acoustic feature corresponds to a first sampling rate; generate a modified acoustic feature corresponding to a modified sampling rate different from the first sampling rate, based on the identified first acoustic feature; and generate a plurality of vocoder learning models by training the vocoder module based on the first acoustic feature and the modified acoustic feature.


