Encoded Text Generation for Nonverbal Audio Characteristics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional mobile devices struggle to efficiently capture and convey nonverbal information, such as sentiment, tone, and inflection, in text-based communications, leading to potential misinterpretation of messages.

Innovation Solution

A system that uses machine learning to generate encoded text representations of spoken utterances, incorporating both text transcriptions and visual representations of nonverbal characteristics, enabling the capture and conveyance of rich nonverbal information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If conventional speech-to-text conversion is used, then text-based communication is simple and efficient, but nonverbal information such as sentiment, tone, and inflection is lost

Engineering Contradiction:
Improvenonverbal informationVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments nonverbal information into distinct characteristics including pitch, volume, tone of voice, inflection, and speaking rate. Each characteristic is analyzed separately by dedicated processing modules, allowing comprehensive capture of nonverbal cues while maintaining organized, manageable data structures that don't overly complicate the system architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a visual dimension to text-based communication by generating encoded text representations that incorporate nonverbal characteristics. This transforms one-dimensional text into multi-dimensional information that includes emotional and tonal context, enabling rich communication without requiring complex video or audio transmission

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If nonverbal information is captured and conveyed in text-based communications, then communication accuracy improves, but processing complexity and computational requirements increase

Engineering Contradiction:
Improvecommunication accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system replaces complex manual analysis of nonverbal cues with automated machine learning models. These models process audio characteristics and generate encoded text representations without requiring human intervention, achieving high communication accuracy while keeping processing complexity manageable through algorithmic automation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms continuous audio parameters (pitch, volume, tone) into discrete encoded text representations with specific visual characteristics. This parameter transformation allows precise capture of nonverbal information while converting complex continuous data into manageable categorical representations that are easier to process and transmit

Inventive Principle:
Principle #35Parameter changes

3Reliability

If encoded text representations with visual representations of nonverbal characteristics are generated, then misinterpretation is reduced, but generation time and computational resources increase

Engineering Contradiction:
Improvemessage interpretation reliabilityVSAvoidmessage generation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of audio characteristics during the speech-to-text conversion process itself, extracting nonverbal features simultaneously with transcription. This parallel processing approach generates encoded text representations without adding sequential delays, maintaining fast message generation while improving interpretation reliability

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250190686A1Generating encoded text based on spoken utterances using machine learning systems and methods
Publication Date: 2025.06.12 T MOBILE US INC
  • US20250190686A1 patent drawing
  • US20250190686A1 patent drawing
  • US20250190686A1 patent drawing

AI summary

Systems and methods for generating encoded text representations of spoken utterances are disclosed. Audio data is received for a spoken utterance and analyzed to identify a nonverbal characteristic, such as a sentiment, a speaking rate, or a volume. An encoded text representation of the spoken utterance is generated, comprising a text transcription and a visual representation of the nonverbal characteristic. The visual representation comprises a geometric element, such as a graph or shape, or a variation in a text attribute, such as font, font size, or color. Analysis of the audio data and/or generation of the encoded text representation can be performed using machine learning.