Speech Synthesis Cadence Prediction Missing Parts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The recorded speech editing method for speech synthesis results in unnatural synthetic speech due to discontinuous pitch component frequency at data boundaries, requiring vast storage capacity and complex data retrieval processes.
Innovation Solution
A speech synthesis device and method that selects and combines voice unit data based on predicted cadence and phoneme matching, with missing part synthesis to generate natural-sounding speech, reducing storage needs and improving processing speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If speech data is simply joined together using recorded speech editing method, then the synthesis process is simple and fast, but the synthetic speech sounds unnatural due to discontinuous pitch component frequency at boundaries
Solution Approach 1:
The patent applies preliminary action by pre-dividing speech data into voice units with clearly defined boundaries and pitch component information before synthesis. This allows the system to select and join pre-prepared units while maintaining continuous pitch, resolving the contradiction between simple joining and natural speech quality.
Solution Approach 2:
The patent applies local quality by focusing on the boundary regions between speech segments, ensuring pitch component continuity specifically at these critical junctions. This localized attention to pitch continuity at boundaries enables natural-sounding speech without requiring complete redesign of the entire synthesis process.
2Manufacturing precision
If a plurality of speech data representing different cadences are prepared for each phoneme to produce natural synthetic speech, then the speech naturalness is improved, but the storage capacity required becomes vast
Solution Approach 1:
The patent applies universality by creating voice units that can serve multiple purposes - each voice unit contains pitch component information that allows it to be adapted to different cadences and contexts. This multi-functional design reduces the need to store separate speech data for every possible cadence variation, significantly reducing storage requirements while maintaining speech naturalness.
Solution Approach 2:
The patent applies parameter changes by modifying pitch component parameters of selected voice units to match the target cadence during synthesis. Instead of storing speech data for all possible cadences, the system stores base voice units and adjusts their pitch parameters dynamically, reducing storage capacity needs while preserving speech naturalness through parameter adaptation.
3Manufacturing precision
If speech data for each phoneme with different cadences are prepared, then the speech naturalness is improved, but the amount of data for retrieval becomes vast and processing complexity increases
Solution Approach 1:
The patent applies segmentation by dividing speech into voice units with clear boundary definitions and pitch component characteristics. This segmentation allows the retrieval system to work with smaller, well-defined units rather than searching through vast amounts of phoneme-level data, reducing retrieval complexity while maintaining the ability to produce natural-sounding speech through proper unit selection and joining.
Data Source
AI summary
A simply configured speech synthesis device and the like for producing a natural synthetic speech at high speed. When data representing a message template is supplied, a voice unit editor (5) searches a voice unit database (7) for voice unit data on a voice unit whose sound matches a voice unit in the message template. Further, the voice unit editor (5) predicts the cadence of the message template and selects, one at a time, a best match of each voice unit in the message template from the voice unit data that has been retrieved, according to the cadence prediction result. For a voice unit for which no match can be selected, an acoustic processor (41) is instructed to supply waveform data representing the waveform of each unit voice. The voice unit data that is selected and the waveform data that is supplied by the acoustic processor (41) are combined to generate data representing a synthetic speech.


