Voice Synthesis Linguistic Processing Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice synthesis systems face limitations in linguistic processing due to information loss and ambiguity in textual forms, requiring operator intervention to compensate for defects, which is non-intuitive and requires significant learning to achieve desired results.

Innovation Solution

An interactive voice synthesis system that breaks down linguistic processing into elementary operations, allowing operators to control and modify parameters through a user-friendly interface, enabling intuitive manipulation and improvement of sound quality by selecting and modifying specific processing steps.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If linguistic processing is performed using conventional text-based methods, then voice synthesis can be automated, but information loss and ambiguity in textual forms reduce the quality of the synthesized voice

Engineering Contradiction:
Improveautomation of voice synthesisVSAvoidinformation loss in linguistic processing
Core Design Contradiction:
Extent of automationVSLoss of information

Solution Approach 1:

The patent segments the linguistic processing into distinct elementary operations (tokenization, phonetization, prosody assignment, etc.), allowing each operation to be independently controlled and optimized. This segmentation enables operators to intervene at specific stages to correct information loss without reprocessing the entire text.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary linguistic processing to generate intermediate representations (phoneme sequences, prosodic contours) before final synthesis. This preliminary action allows operators to review and correct processing results at each stage, preventing information loss from propagating to the final output.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If operators intervene to compensate for linguistic processing defects, then voice quality improves, but the manipulation becomes non-intuitive and requires significant learning

Engineering Contradiction:
Improvevoice qualityVSAvoidease of system manipulation
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

By dividing linguistic processing into elementary operations with dedicated control interfaces, the system makes each operation intuitive to control. Operators can manipulate specific aspects (e.g., phonetization rules, prosody parameters) independently using straightforward controls rather than complex global parameters.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system provides uniform control mechanisms across all linguistic processing stages, allowing operators to intervene consistently at any stage using the same type of interface. This homogeneity in control approach reduces learning curves and improves ease of operation.

Inventive Principle:
Principle #33Homogeneity

3Device complexity

If the linguistic processing chain is treated as a black box, then the system is simpler to implement, but operators cannot control specific parameters affecting sound quality

Engineering Contradiction:
Improvesystem implementation complexityVSAvoidcontrol over synthesis parameters
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent decomposes the linguistic processing chain into discrete, independently controllable modules. Each module (tokenization, phonetization, prosody) can be configured and controlled separately, providing operators with fine-grained control over parameters affecting sound quality while maintaining manageable system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system allows dynamic configuration of linguistic processing parameters at each stage, enabling operators to adapt the processing chain to different synthesis requirements. This dynamic control capability provides versatility in parameter adjustment without requiring complete system redesign.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP1960996B1Voice synthesis by concatenation of acoustic units
Publication Date: 2010.02.24 ORANGE SA
  • EP1960996B1 patent drawingFigure 1
  • EP1960996B1 patent drawingFigure 2~3
  • EP1960996B1 patent drawingFigure 4

AI summary

The present invention relates to a system of voice synthesis by concatenation of acoustic units comprising: - means (4) for linguistically processing a text so as to transform it into a string of phonemes accompanied by prosodic indications, - means (6) for synthesizing prerecorded elements by concatenation so as to restore an acoustic signal, as a function of the string of phonemes, - input and editing means (8), such that the linguistic processing means (4) comprise at least one elementary processing unit (4A, 4B, 4C) that generates intermediate results of the linguistic processing of said text, said unit being associated with an editor (8A, 8B, 8C) of the input and editing means (8), allowing an operator to modify the intermediate results and the voice synthesis system comprises means (14) for parameterizing the text on the basis of the results modified by the operator, the linguistic processing means (4) adapting the linguistic processing of the text on the basis of said parameterization.