Voice Synthesis Linguistic Processing Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice synthesis systems face limitations in linguistic processing due to information loss and ambiguity in textual forms, requiring operator intervention to compensate for defects, which is non-intuitive and requires significant learning to achieve desired results.
Innovation Solution
An interactive voice synthesis system that breaks down linguistic processing into elementary operations, allowing operators to control and modify parameters through a user-friendly interface, enabling intuitive manipulation and improvement of sound quality by selecting and modifying specific processing steps.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If linguistic processing is performed using conventional text-based methods, then voice synthesis can be automated, but information loss and ambiguity in textual forms reduce the quality of the synthesized voice
Solution Approach 1:
The patent segments the linguistic processing into distinct elementary operations (tokenization, phonetization, prosody assignment, etc.), allowing each operation to be independently controlled and optimized. This segmentation enables operators to intervene at specific stages to correct information loss without reprocessing the entire text.
Solution Approach 2:
The system performs preliminary linguistic processing to generate intermediate representations (phoneme sequences, prosodic contours) before final synthesis. This preliminary action allows operators to review and correct processing results at each stage, preventing information loss from propagating to the final output.
2Manufacturing precision
If operators intervene to compensate for linguistic processing defects, then voice quality improves, but the manipulation becomes non-intuitive and requires significant learning
Solution Approach 1:
By dividing linguistic processing into elementary operations with dedicated control interfaces, the system makes each operation intuitive to control. Operators can manipulate specific aspects (e.g., phonetization rules, prosody parameters) independently using straightforward controls rather than complex global parameters.
Solution Approach 2:
The system provides uniform control mechanisms across all linguistic processing stages, allowing operators to intervene consistently at any stage using the same type of interface. This homogeneity in control approach reduces learning curves and improves ease of operation.
3Device complexity
If the linguistic processing chain is treated as a black box, then the system is simpler to implement, but operators cannot control specific parameters affecting sound quality
Solution Approach 1:
The patent decomposes the linguistic processing chain into discrete, independently controllable modules. Each module (tokenization, phonetization, prosody) can be configured and controlled separately, providing operators with fine-grained control over parameters affecting sound quality while maintaining manageable system complexity through modular architecture.
Solution Approach 2:
The system allows dynamic configuration of linguistic processing parameters at each stage, enabling operators to adapt the processing chain to different synthesis requirements. This dynamic control capability provides versatility in parameter adjustment without requiring complete system redesign.
Data Source
Figure 1
Figure 2~3
Figure 4
AI summary
The present invention relates to a system of voice synthesis by concatenation of acoustic units comprising: - means (4) for linguistically processing a text so as to transform it into a string of phonemes accompanied by prosodic indications, - means (6) for synthesizing prerecorded elements by concatenation so as to restore an acoustic signal, as a function of the string of phonemes, - input and editing means (8), such that the linguistic processing means (4) comprise at least one elementary processing unit (4A, 4B, 4C) that generates intermediate results of the linguistic processing of said text, said unit being associated with an editor (8A, 8B, 8C) of the input and editing means (8), allowing an operator to modify the intermediate results and the voice synthesis system comprises means (14) for parameterizing the text on the basis of the results modified by the operator, the linguistic processing means (4) adapting the linguistic processing of the text on the basis of said parameterization.