Computer-Generated Speech Visualization Using Geometric Parameter Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional visual representations of speech, such as spectrograms and International Phonetic Alphabets (IPA), are not intuitive or user-friendly for language learners, making it difficult for them to distinguish between reference and their own speech recordings.

Innovation Solution

A graphical representation of speech is generated using objects where duration is represented by length, intensity by width, and pitch contour by angle, with optional color coding based on phonemes and articulation, displayed on a computing device screen.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional visual representations (spectrograms and IPA notations) are used to accurately represent speech, then measurement precision is improved, but ease of operation deteriorates because they are not intuitive or user-friendly for language learners

Engineering Contradiction:
Improveaccuracy of speech representationVSAvoiduser-friendliness for language learners
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent transforms speech parameters into intuitive visual metaphors: duration becomes length of rectangular objects, intensity becomes width, pitch contour becomes angle of inclination, and fundamental frequency offset becomes vertical position. This parameter transformation maintains measurement precision while dramatically improving ease of operation for language learners.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system creates simplified visual copies of speech characteristics using geometric shapes (rectangles) that replicate key speech parameters in an intuitive format. These visual copies preserve the essential information from spectrograms and IPA notations while presenting it in a learner-friendly manner.

Inventive Principle:
Principle #26Copying

2Loss of information

If detailed spectrogram representations are used to show speech characteristics, then information completeness is improved, but device complexity increases making the interface less user-friendly

Engineering Contradiction:
Improvecompleteness of speech informationVSAvoidcomplexity of visualization interface
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts only the most essential speech parameters (duration, intensity, pitch contour, fundamental frequency) from the complete spectrogram data and represents them using simple rectangular objects. This extraction maintains the key information needed for language learning while eliminating unnecessary complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The speech signal is segmented into discrete units represented by individual rectangular objects, with each object corresponding to a specific segment. Spaces between objects represent unvoiced periods. This segmentation organizes complex speech information into manageable, intuitive visual units.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12374352B2Methods and systems for computer-generated visualization of speech
Publication Date: 2025.07.29 SOMNIQ
  • US12374352B2 patent drawing
  • US12374352B2 patent drawing
  • US12374352B2 patent drawing

AI summary

Methods, systems and apparatuses for computer-generated visualization of speech are described herein. An example method of computer-generated visualization of speech including at least one segment includes: generating a graphical representation of an object corresponding to a segment of the speech; and displaying the graphical representation of the object on a screen of a computing device. Generating the graphical representation includes: representing a duration of the respective segment by a length of the object and representing intensity of the respective segment by a width of the object; and placing, in the graphical representation, a space between adjacent objects.