Speech-to-Text Character Shape Visualization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech-to-text conversion tools fail to visually represent how something is spoken, lacking the ability to convey emotional or acoustic properties of speech through character visualization.

Innovation Solution

A computer system that processes audio speech signals to extract acoustic features, segment them into pronunciation units, align these units with graphemes, determine speed and acoustic parameters, and generate character shapes that reflect these parameters, allowing for the visualization of speech properties like loudness, pitch, and speed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If traditional speech-to-text conversion is used, then text representation is simple and clear, but acoustic and emotional properties of speech are not visualized

Engineering Contradiction:
Improveacoustic and emotional informationVSAvoidcharacter visualization system
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extends traditional 2D character display by adding visual dimensions that represent acoustic properties. Character shapes are modified to encode pitch (vertical position), loudness (character size), and speed (spacing between characters), transforming flat text into multi-dimensional visual representations that convey speech characteristics without adding complex separate visualization layers.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

Different visual properties of characters are assigned to represent different acoustic parameters. Specifically, vertical position encodes pitch information, character size encodes loudness, and spacing between characters encodes speech speed. This local differentiation allows multiple acoustic dimensions to be represented through distinct visual channels within the same character stream.

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If character shapes are made unchangeable, then text display is consistent and readable, but speech characteristics cannot be expressed

Engineering Contradiction:
Improvecharacter shape adaptabilityVSAvoidtext readability
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent introduces dynamic character rendering where character shapes adapt based on acoustic input properties. Characters are not static but dynamically adjusted in their visual properties (position, size, spacing) according to the pitch, loudness, and speed of the spoken speech, allowing the text to reflect the speaker's emotional and acoustic state while maintaining readability through consistent recognition patterns.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10043519B2Generation of text from an audio speech signal
Publication Date: 2018.08.07 SCHLIPPE TIM
  • US10043519B2 patent drawing
  • US10043519B2 patent drawing
  • US10043519B2 patent drawing

AI summary

In one general aspect, a computer-implemented method for text generation based on an audio speech signal can include receiving the audio speech signal, extracting acoustic feature values of the speech signal at a predefined sampling frequency, mapping written words of a transcription of the audio speech signal to the units of the corresponding pronunciation objects, segmenting the audio speech signal including mapping the units of corresponding pronunciation objects to the received audio speech signal to determine a beginning time and an end time of the mapped units, aligning one or more units of the corresponding pronunciation objects to one or more graphemes based on a unit-grapheme mapping, determining a speed parameter for each aligned grapheme, determining acoustic parameters for each aligned grapheme, and generating, for each character of the aligned graphemes, a character shape representative of the speed parameter and the acoustic parameters associated with the respective grapheme.