Multi-Faceted Audio Cloud Representation for Para-Lingual Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional visual representations of text, such as word clouds, are limited in scenarios where a solely visual representation is not available or accessible, particularly in processing audio data from diverse languages and dialects, and do not effectively convey the richness of para-lingual information like speaker characteristics and conversation style.

Innovation Solution

A method for creating and rendering a multi-faceted audio cloud that uses audio segmentation, sub-word recognition, and dynamic rendering techniques to present prominent audio segments, allowing for user interaction and accommodating language agnostic methods, enabling the representation of audio data in a way that conveys linguistic and para-lingual information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional visual word clouds are used to represent text, then frequently used words can be identified through font size, but the representation becomes inaccessible or difficult to process for scenarios requiring audio data processing or for users with visual processing limitations

Engineering Contradiction:
Improveaccessibility across different data types and user needsVSAvoidloss of audio data representation capability
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent replaces the visual-mechanical system of word clouds with an audio-based system. Instead of using visual font sizes to represent word frequency, the system uses audio characteristics (volume, pitch, duration) to represent audio segment prominence, making the representation accessible to users who cannot process visual information while preserving the frequency-based weighting concept.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a multi-functional representation system that can handle both visual and audio data types. The audio cloud generator can process text data (converting to speech) or audio data directly, and can output representations suitable for different user needs, thereby achieving universal applicability across different data types and user requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If audio data is processed without language-specific constraints, then the system becomes language-agnostic and more versatile, but the complexity of handling diverse languages and dialects increases

Engineering Contradiction:
Improvelanguage agnostic capabilityVSAvoidcomplexity of processing diverse languages and dialects
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts language-specific processing requirements from the core audio processing pipeline. By separating the audio segmentation and prominence detection functions from language-specific text processing, the system achieves language-agnostic capability while managing complexity through modular design where language processing is an optional, extractable component.

Inventive Principle:
Principle #2Taking out (Extraction)

3Loss of information

If traditional word clouds are used, then visual representation is simple and straightforward, but para-lingual information such as speaker characteristics and conversation style cannot be conveyed

Engineering Contradiction:
Improvepreservation of para-lingual informationVSAvoidcomplexity of representing multiple audio facets
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent adds new dimensions to the representation by mapping audio characteristics (speaker identity, emotion, volume, pitch) to spatial and temporal dimensions in the audio cloud visualization. Instead of a single visual dimension (font size), the system uses multiple audio dimensions (panning, volume, pitch modulation) to convey para-lingual information, resolving the contradiction between information preservation and representation complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10013485B2Creating, rendering and interacting with a multi-faceted audio cloud
Publication Date: 2018.07.03 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10013485B2 patent drawing
  • US10013485B2 patent drawing
  • US10013485B2 patent drawing

AI summary

Methods and arrangements for effecting a cloud representation of audio content. An audio cloud is created and rendered, and user interaction with at least a clip portion of the audio cloud is afforded.