Text Overlay on Voice Modulation Graph for Audio Editing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio editing tools and collaboration applications do not allow users to view the text of spoken words within a modulation graph of the audio, making it difficult for users, especially those with hearing impairments, to edit audio effectively.

Innovation Solution

The method involves obtaining audio with spoken words, generating a modulation graph representative of the audio, transcribing the text from the audio, and displaying the text within the modulation graph at corresponding locations, allowing users to perform actions like editing or deleting unwanted noise based on visual cues.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users edit audio using traditional modulation graphs without text overlay, then the interface remains simple and easy to operate, but users cannot view the text of spoken words making editing difficult especially for those with hearing impairments

Engineering Contradiction:
Improveease of audio editingVSAvoidtext information loss
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent combines the modulation graph visualization with text overlay by merging two separate information representations into a single integrated display. The text representing spoken words is superimposed on the modulation graph at corresponding time positions, allowing users to view both audio waveform information and text content simultaneously in one interface element.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The text overlay acts as an intermediary between the audio signal and the user's visual perception. Instead of requiring users to listen to the audio directly, the text serves as a mediating representation that conveys the spoken words visually, enabling users with hearing impairments to access audio information through visual channels.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If users must listen to audio to edit it effectively, then no additional visual processing is needed, but users with hearing impairments cannot edit audio effectively

Engineering Contradiction:
Improveaccessibility for hearing impaired usersVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent substitutes the auditory channel with a visual channel for information representation. Instead of relying on audio playback for users to identify and edit content, the system replaces audio perception with visual text display, allowing users with hearing impairments to access and edit audio information through visual processing rather than auditory processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If text is displayed within the modulation graph, then users can visually identify words and edit more effectively, but the graph becomes more complex and harder to interpret

Engineering Contradiction:
Improveaudio editing efficiencyVSAvoidgraph complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the audio timeline into discrete time positions and places text labels at corresponding segments rather than displaying all text at once. This temporal segmentation allows the text to be distributed across the modulation graph in a way that maintains readability and doesn't create excessive visual clutter, while still providing comprehensive word identification capability.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250046331A1Visual representation of text on voice modulation graph
Publication Date: 2025.02.06 CISCO TECHNOLOGY INC
  • US20250046331A1 patent drawing
  • US20250046331A1 patent drawing
  • US20250046331A1 patent drawing

AI summary

Presented herein are techniques to display text of words spoken within a modulation graph representative of audio of the words. A method includes obtaining audio that includes words spoken by a user and generating a modulation graph representative of the audio. Text of the words spoken by the user is obtained from the audio and the text of the words is displayed within the modulation graph of the audio so the words are displayed at a location within the modulation graph that corresponds to the audio of the words being spoken. An input is received from the user to perform one or more actions with respect to the modulation graph and the one or more actions based on the input is performed.