Speech-to-Text Voice Visualization for Intonation Preservation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AI systems fail to regenerate text from speech with different intonations or native speaking techniques and vice versa, losing emotion and meaning in the conversion process.
Innovation Solution
A computer-implemented method that generates personalized audio data by segmenting user input into sentences, creating voice images with pronunciation tags and wave lines, and modifying data based on these elements to preserve individual expression differences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If speech-to-text conversion is performed using existing AI systems, then text can be generated from speech, but emotion and meaning are lost in the conversion process
Solution Approach 1:
The patent segments speech conversion into multiple analysis dimensions including pitch contour analysis, pause pattern detection, and stress identification. By dividing the conversion process into these discrete analytical components, the system preserves emotional and meaningful information that would otherwise be lost in traditional speech-to-text conversion.
Solution Approach 2:
The patent adds visual dimensionality to speech-to-text conversion by generating voice images that display pitch contours, pause patterns, and stress markers. This transforms the conversion from a purely textual output to a multi-dimensional representation that preserves emotional and meaningful nuances through visual visualization of speech characteristics.
2Productivity
If speech-to-text conversion is performed without voice visualization, then conversion speed is maintained, but intonation and native speaking techniques cannot be regenerated
Solution Approach 1:
The patent creates voice images that copy and visualize the acoustic characteristics of speech including pitch contours, pause patterns, and stress markers. These visual representations serve as templates that can be used to regenerate speech with appropriate intonation and native speaking techniques, preserving the original speaker's stylistic features.
Solution Approach 2:
The patent analyzes and extracts key speech parameters including fundamental frequency contours, pause duration, and stress intensity. By identifying and preserving these critical parameters in the voice image representation, the system enables accurate regeneration of intonation and speaking style while maintaining conversion efficiency.
3Manufacturing precision
If voice images with pronunciation tags and wave lines are generated, then personalized output with preserved intonation is achieved, but processing complexity increases
Solution Approach 1:
The patent segments the voice image generation process into distinct modules: pitch contour extraction, pause pattern detection, stress identification, and visualization rendering. By organizing these functions into separate processing stages, the system achieves high pronunciation accuracy while managing complexity through modular architecture.
Solution Approach 2:
The patent creates a universal voice image representation that can serve multiple functions: visualizing pitch contours, displaying pause patterns, marking stress points, and guiding speech synthesis. This multi-functional representation reduces overall system complexity by using a single integrated data structure rather than separate processing systems for each function.
Data Source
AI summary
A computer-implemented method for generating personalized audio data is disclosed. The computer-implemented method includes receiving user input data, wherein the user input data is at least one of text or audio. The computer-implemented method further includes segmenting the user input data into a set of sentences. The computer-implemented method further includes generating, for each sentence in the set of sentences, a voice image, wherein the voice image includes at least one pronunciation tag and wave line associated with a sentence. The computer-implemented method further includes modifying the user input data based, at least in part on, the wave line and pronunciation tag of the voice image.


