Voice Audio Text Processing with Dynamic Size and Color
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice audio data processing systems fail to effectively incorporate additional information such as emotions, volumes, and locations into the text conversion process, leading to a loss of contextual details.
Innovation Solution
A system and method that utilize a processor to obtain voice audio data, generate text based on the data, and dynamically adjust text attributes like size and color to reflect voice volumes and emotions, while also displaying text based on the location of the voice source.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If voice audio data is converted into text to indicate contents, then visualization and access to contents are improved, but additional information such as emotions, volumes, and locations are lost
Solution Approach 1:
The patent applies parameter changes by mapping voice characteristics (volume, emotion type) to text display parameters (size, color). The text size dynamically changes based on voice volume, and text color changes based on emotion type, thereby preserving additional information while maintaining text conversion benefits
Solution Approach 2:
The patent adds new dimensions to the text display by incorporating size and color attributes that correspond to voice volume and emotion type respectively. This transforms the one-dimensional text output into a multi-dimensional display that encodes additional voice characteristics
2Loss of information
If text attributes are dynamically adjusted to reflect voice characteristics, then information preservation is improved, but system complexity increases
Solution Approach 1:
The patent segments the text processing into distinct functional modules: voice audio data acquisition, text generation, emotion type determination, and display attribute mapping. Each module handles a specific aspect of the transformation process, making the overall complex system more manageable and maintainable
Solution Approach 2:
The patent introduces an intermediary processing layer that determines emotion types and maps them to color codes, and another layer that maps voice volume to text size. These intermediary functions simplify the overall system by breaking down the complex transformation into manageable steps with clear input-output relationships
Data Source
AI summary
The present disclosure may provide a voice audio data processing system. The voice audio data processing system may obtain voice audio data, which includes one or more voices, each being respectively associated with one of one or more subjects. For one of the one or more voices and the subject associated with the voice, the voice audio processing system may generate a text based on the voice audio data. The text may have one or more sizes, each size corresponding to one of one or more volumes of the voice. The text may have one or more colors, each color corresponding to one of one or more emotion types of the voice.


