English Pronunciation Visualization via Speech Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional English learning systems fail to effectively visualize pronunciation and stress in English sentences, particularly for non-native speakers, as they do not accurately reflect the linguistic features of English, such as stress-timed language and non-phonetic spelling, leading to limitations in listening and speaking practice.

Innovation Solution

A speech visualization system that analyzes speech signals for frequencies, energy, and time, classifies them into segments, and assigns visualization properties to represent syllables, stress, and other linguistic features, using a combination of natural language processing and specialized properties like stress, liaisons, and schwa sounds, to generate intuitive visualization data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional visualization techniques are used to represent syllables and stress in English sentences, then the basic linguistic features can be displayed, but the visualization fails to accurately reflect the stress-timed language characteristics and non-phonetic spelling features of English

Engineering Contradiction:
Improveaccuracy of linguistic feature representationVSAvoidcomplexity of speech analysis system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The speech signal is divided into multiple segments including syllable-level segments, word-level segments, and sentence-level segments. Each segment type focuses on specific linguistic features (syllable timing, word stress, sentence rhythm) to achieve comprehensive and accurate visualization of English stress-timed characteristics without overwhelming complexity

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A speech information analysis unit acts as an intermediary between the speech signal input and visualization. This unit performs intermediate processing including speech recognition, syllable segmentation, and stress detection to transform raw speech signals into structured linguistic information that can be accurately visualized

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If detailed speech analysis is performed to capture all linguistic features, then visualization accuracy improves, but processing time and system complexity increase

Engineering Contradiction:
Improveprecision of pronunciation visualizationVSAvoidprocessing time for speech analysis
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The speech analysis process is segmented into parallel processing streams: syllable segmentation, stress detection, and rhythm analysis occur simultaneously rather than sequentially. This reduces overall processing time while maintaining comprehensive analysis precision

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Speech signals are pre-processed into standardized formats with extracted features (frequency, amplitude, duration) before main analysis. This preliminary preparation reduces the computational burden of subsequent detailed analysis and speeds up the overall processing pipeline

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If conventional syllable-timed language visualization is used, then Korean linguistic patterns are represented, but English stress-timed language features are not properly visualized

Engineering Contradiction:
Improveadaptability to different language typesVSAvoidaccuracy of English pronunciation representation
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The visualization system dynamically adapts its parameters based on the detected language type. For English, it emphasizes stress timing and pitch variation; for Korean, it focuses on syllable timing. This dynamic adjustment allows accurate representation of stress-timed English while maintaining adaptability to other language types

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes key visualization parameters including time segmentation intervals, frequency ranges, and stress detection thresholds based on the target language. This parameter adaptation enables precise visualization of English stress-timed characteristics while preserving the ability to handle syllable-timed languages like Korean

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12118898B2Voice visualization system for english learning, and method therefor
Publication Date: 2024.10.15 LEE GI HUN
  • US12118898B2 patent drawing
  • US12118898B2 patent drawing
  • US12118898B2 patent drawing

AI summary

A speech visualization system according to the present invention includes: a speech signal input unit for receiving speech signals of sentences with English pronunciations; a speech information analysis unit for analyzing speech information with frequencies, energy, and time of the speech signals and the text corresponding to the speech signals to divide the speech information into at least one or more segments; a speech information classification unit for classifying the segments of the speech information into flow units and each flow unit into at least one or more sub flow units each having at least one or more words; a visualization property assignment unit for assigning visualization properties for speech visualization to the analyzed and classified speech information; and a visualization processing unit for performing visualization processing based on the assigned visualization properties to generate speech visualization data.