Disordered Voice Processing Using LPC Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Individuals with articulation disorders face challenges in generating clear speech due to anatomical abnormalities, leading to misarticulation issues that current treatments like speech therapies and surgeries do not fully address, resulting in unintended pronunciation during voice generation.

Innovation Solution

A disordered voice processing method and apparatus that recognizes voice signals based on phonemes, extracts multiple voice components, processes disordered components, and synthesizes a restored voice signal using Linear Predictive Coding (LPC) to convert misarticulation into normal articulation, enabling accurate speech communication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speech therapies and surgeries are used to treat articulation disorders, then treatment is provided, but misarticulation caused by articulation disorders still occurs

Engineering Contradiction:
Improvetreatment effectivenessVSAvoidarticulation accuracy
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent replaces physical speech therapy mechanisms and surgical interventions with a digital signal processing system. The system uses voice recognition to identify disordered speech patterns, extracts vocal tract components, processes them through filtering and synthesis algorithms, and generates corrected speech output, substituting mechanical/biological treatment methods with computational processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes parameters of the voice signal through digital processing. It adjusts spectral characteristics, formant frequencies, and temporal parameters of extracted vocal tract components to transform disordered speech parameters into normal speech parameters, achieving articulation correction through parameter transformation rather than physical intervention.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If voice signal processing is performed to convert misarticulation to normal articulation, then speech accuracy is improved, but processing complexity increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the voice signal into distinct components using Linear Predictive Coding (LPC) analysis. It separates the signal into vocal tract components and glottal components, allowing independent processing of each segment. This segmentation enables targeted correction of articulation errors without requiring complex full-signal processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary processing stage between voice input and output. The voice recognition unit and component extraction unit act as intermediaries that analyze the input signal, identify disordered components, and prepare them for correction. This intermediary layer simplifies the overall process by breaking down complex transformation into manageable intermediate steps.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9646602B2Method and apparatus for improving disordered voice
Publication Date: 2017.05.09 SEOUL NATIONAL UNIVERSITY R&DB FOUNDATION
  • US9646602B2 patent drawing
  • US9646602B2 patent drawing
  • US9646602B2 patent drawing

AI summary

There is provided a method and an apparatus for processing a disordered voice. A method for processing a disordered voice according to an exemplary embodiment of the present invention includes: receiving a voice signal; recognizing the voice signal by phoneme; extracting multiple voice components from the voice signal; acquiring restored voice components by processing at least some disordered voice components of the multiple voice components by phoneme; and synthesizing a restored voice signal based on at least the restored voice components.