Real-Time Audio Word Clouds for Speaker Distinction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for summarizing audio content, such as word clouds, are limited in their ability to distinguish between different subjects and speakers within a conversation, and are typically designed for text-based formats, making them cumbersome for auditory conversations or meetings.

Innovation Solution

A computer-implemented method and system that uses speech and speaker recognition software to convert audio content into text, generate word clouds for each segment of a meeting, and visually represent key words and speakers, allowing for real-time visualization of meeting content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If a single word cloud is generated for the entire audio content, then the summary provides an overall view of the conversation, but it cannot distinguish between different speakers or subjects, reducing the clarity and utility of the summary

Engineering Contradiction:
Improvespeaker distinction informationVSAvoidsummary structure complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent divides the audio content into multiple segments based on speaker identification. Each segment generates a separate word cloud that represents only that speaker's contributions. This segmentation preserves speaker distinction information while maintaining manageable complexity through automated speaker detection and separate processing of each speaker's transcript.

Inventive Principle:
Principle #1Segmentation

2Loss of information

If speech recognition software is used to convert audio to text, then the audio content becomes analyzable, but the process is time-consuming and reduces real-time capability

Engineering Contradiction:
Improveaudio content accessibilityVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary speaker identification and segmentation of audio content before generating word clouds. By pre-organizing the audio into speaker-specific segments, the system reduces the computational burden during the word cloud generation phase, enabling faster processing and improved real-time performance.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If comprehensive speech recognition is applied to capture all spoken content, then complete transcription is achieved, but the resulting transcript becomes complex and difficult to review

Engineering Contradiction:
Improvetranscript completenessVSAvoidtranscript reviewability
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent extracts and displays only the most significant words from each speaker's transcript in visual word cloud format. This extraction process filters out less important words while preserving key information, making the transcript reviewable through visual prominence of important terms rather than reading through complete text.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system uses color coding to differentiate between speakers in the word cloud visualization. Each speaker's words are displayed in distinct colors, allowing reviewers to quickly identify which speaker said what without needing to read through complete transcripts, thereby improving reviewability while maintaining completeness.

Inventive Principle:
Principle #32Color changes

Data Source

PatentUS8825478B2Real time generation of audio content summaries
Publication Date: 2014.09.02 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8825478B2 patent drawing
  • US8825478B2 patent drawing
  • US8825478B2 patent drawing

AI summary

Audio content is converted to text using speech recognition software. The text is then associated with a distinct voice or a generic placeholder label if no distinction can be made. From the text and voice information, a word cloud is generated based on key words and key speakers. A visualization of the cloud displays as it is being created. Words grow in size in relation to their dominance. When it is determined that the predominant words or speakers have changed, the word cloud is complete. That word cloud continues to be displayed statically and a new word cloud display begins based upon a new set of predominant words or a new predominant speaker or set of speakers. This process may continue until the meeting is concluded. At the end of the meeting, the completed visualization may be saved to a storage device, sent to selected individuals, removed, or any combination of the preceding.