Speech-to-Text Character Shape Visualization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech-to-text conversion tools fail to visually represent how something is spoken, lacking the ability to convey emotional or acoustic properties of speech through character visualization.
Innovation Solution
A computer system that processes audio speech signals to extract acoustic features, segment them into pronunciation units, align these units with graphemes, determine speed and acoustic parameters, and generate character shapes that reflect these parameters, allowing for the visualization of speech properties like loudness, pitch, and speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If traditional speech-to-text conversion is used, then text representation is simple and clear, but acoustic and emotional properties of speech are not visualized
Solution Approach 1:
The patent extends traditional 2D character display by adding visual dimensions that represent acoustic properties. Character shapes are modified to encode pitch (vertical position), loudness (character size), and speed (spacing between characters), transforming flat text into multi-dimensional visual representations that convey speech characteristics without adding complex separate visualization layers.
Solution Approach 2:
Different visual properties of characters are assigned to represent different acoustic parameters. Specifically, vertical position encodes pitch information, character size encodes loudness, and spacing between characters encodes speech speed. This local differentiation allows multiple acoustic dimensions to be represented through distinct visual channels within the same character stream.
2Adaptability or versatility
If character shapes are made unchangeable, then text display is consistent and readable, but speech characteristics cannot be expressed
Solution Approach 1:
The patent introduces dynamic character rendering where character shapes adapt based on acoustic input properties. Characters are not static but dynamically adjusted in their visual properties (position, size, spacing) according to the pitch, loudness, and speed of the spoken speech, allowing the text to reflect the speaker's emotional and acoustic state while maintaining readability through consistent recognition patterns.
Data Source
AI summary
In one general aspect, a computer-implemented method for text generation based on an audio speech signal can include receiving the audio speech signal, extracting acoustic feature values of the speech signal at a predefined sampling frequency, mapping written words of a transcription of the audio speech signal to the units of the corresponding pronunciation objects, segmenting the audio speech signal including mapping the units of corresponding pronunciation objects to the received audio speech signal to determine a beginning time and an end time of the mapped units, aligning one or more units of the corresponding pronunciation objects to one or more graphemes based on a unit-grapheme mapping, determining a speed parameter for each aligned grapheme, determining acoustic parameters for each aligned grapheme, and generating, for each character of the aligned graphemes, a character shape representative of the speed parameter and the acoustic parameters associated with the respective grapheme.


