Voice Audio Text Processing with Dynamic Size and Color

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice audio data processing systems fail to effectively incorporate additional information such as emotions, volumes, and locations into the text conversion process, leading to a loss of contextual details.

Innovation Solution

A system and method that utilize a processor to obtain voice audio data, generate text based on the data, and dynamically adjust text attributes like size and color to reflect voice volumes and emotions, while also displaying text based on the location of the voice source.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice audio data is converted into text to indicate contents, then visualization and access to contents are improved, but additional information such as emotions, volumes, and locations are lost

Engineering Contradiction:
Improveaccess to voice contentVSAvoidemotional and contextual information
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent applies parameter changes by mapping voice characteristics (volume, emotion type) to text display parameters (size, color). The text size dynamically changes based on voice volume, and text color changes based on emotion type, thereby preserving additional information while maintaining text conversion benefits

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent adds new dimensions to the text display by incorporating size and color attributes that correspond to voice volume and emotion type respectively. This transforms the one-dimensional text output into a multi-dimensional display that encodes additional voice characteristics

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If text attributes are dynamically adjusted to reflect voice characteristics, then information preservation is improved, but system complexity increases

Engineering Contradiction:
Improvepreservation of voice characteristicsVSAvoidtext processing system
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the text processing into distinct functional modules: voice audio data acquisition, text generation, emotion type determination, and display attribute mapping. Each module handles a specific aspect of the transformation process, making the overall complex system more manageable and maintainable

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer that determines emotion types and maps them to color codes, and another layer that maps voice volume to text size. These intermediary functions simplify the overall system by breaking down the complex transformation into manageable steps with clear input-output relationships

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250037720A1Systems and methods for voice audio data processing
Publication Date: 2025.01.30 ZHEJIANG TONGHUASHUN INTELLIGENT TECH CO LTD
  • US20250037720A1 patent drawing
  • US20250037720A1 patent drawing
  • US20250037720A1 patent drawing

AI summary

The present disclosure may provide a voice audio data processing system. The voice audio data processing system may obtain voice audio data, which includes one or more voices, each being respectively associated with one of one or more subjects. For one of the one or more voices and the subject associated with the voice, the voice audio processing system may generate a text based on the voice audio data. The text may have one or more sizes, each size corresponding to one of one or more volumes of the voice. The text may have one or more colors, each color corresponding to one of one or more emotion types of the voice.