Speech-to-Text Voice Visualization for Intonation Preservation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing AI systems fail to regenerate text from speech with different intonations or native speaking techniques and vice versa, losing emotion and meaning in the conversion process.

Innovation Solution

A computer-implemented method that generates personalized audio data by segmenting user input into sentences, creating voice images with pronunciation tags and wave lines, and modifying data based on these elements to preserve individual expression differences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If speech-to-text conversion is performed using existing AI systems, then text can be generated from speech, but emotion and meaning are lost in the conversion process

Engineering Contradiction:
Improveemotion and meaningVSAvoidconversion accuracy
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent segments speech conversion into multiple analysis dimensions including pitch contour analysis, pause pattern detection, and stress identification. By dividing the conversion process into these discrete analytical components, the system preserves emotional and meaningful information that would otherwise be lost in traditional speech-to-text conversion.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds visual dimensionality to speech-to-text conversion by generating voice images that display pitch contours, pause patterns, and stress markers. This transforms the conversion from a purely textual output to a multi-dimensional representation that preserves emotional and meaningful nuances through visual visualization of speech characteristics.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If speech-to-text conversion is performed without voice visualization, then conversion speed is maintained, but intonation and native speaking techniques cannot be regenerated

Engineering Contradiction:
Improveconversion speedVSAvoidintonation regeneration capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates voice images that copy and visualize the acoustic characteristics of speech including pitch contours, pause patterns, and stress markers. These visual representations serve as templates that can be used to regenerate speech with appropriate intonation and native speaking techniques, preserving the original speaker's stylistic features.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent analyzes and extracts key speech parameters including fundamental frequency contours, pause duration, and stress intensity. By identifying and preserving these critical parameters in the voice image representation, the system enables accurate regeneration of intonation and speaking style while maintaining conversion efficiency.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If voice images with pronunciation tags and wave lines are generated, then personalized output with preserved intonation is achieved, but processing complexity increases

Engineering Contradiction:
Improvepronunciation accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the voice image generation process into distinct modules: pitch contour extraction, pause pattern detection, stress identification, and visualization rendering. By organizing these functions into separate processing stages, the system achieves high pronunciation accuracy while managing complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal voice image representation that can serve multiple functions: visualizing pitch contours, displaying pause patterns, marking stress points, and guiding speech synthesis. This multi-functional representation reduces overall system complexity by using a single integrated data structure rather than separate processing systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12417762B2Speech-to-text voice visualization
Publication Date: 2025.09.16 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12417762B2 patent drawing
  • US12417762B2 patent drawing
  • US12417762B2 patent drawing

AI summary

A computer-implemented method for generating personalized audio data is disclosed. The computer-implemented method includes receiving user input data, wherein the user input data is at least one of text or audio. The computer-implemented method further includes segmenting the user input data into a set of sentences. The computer-implemented method further includes generating, for each sentence in the set of sentences, a voice image, wherein the voice image includes at least one pronunciation tag and wave line associated with a sentence. The computer-implemented method further includes modifying the user input data based, at least in part on, the wave line and pronunciation tag of the voice image.