Word Cloud Audio Navigation via Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing solutions fail to effectively visualize and link text to media streams, making it difficult to navigate audio or video files, especially in longer recordings like conference calls, where finding specific topics is challenging without manual indexing or advanced searching.

Innovation Solution

The method involves creating a word cloud that links selected words or phrases to timestamps in audio or video streams, allowing users to click on words and navigate directly to relevant sections, with context-based summaries to pinpoint areas of interest.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual indexing with offset timestamps is used to navigate audio streams, then users can locate specific sections, but the process becomes extremely challenging for longer media files and requires significant user effort to find topics of interest

Engineering Contradiction:
Improveease of navigationVSAvoidtime to find section
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent introduces an automatic speech recognition system as an intermediary between the audio stream and the user. This mediator converts spoken content into text with timestamps, automatically creates an index, and generates a searchable database without requiring manual intervention. The intermediary handles the complex task of content analysis and organization, allowing users to simply search for topics of interest.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary actions by automatically transcribing the audio stream into text and creating an indexed database before the user needs to search. The speech recognition system processes the entire audio file in advance, extracting keywords, phrases, and timestamps, and organizing them into a searchable structure. This preliminary processing eliminates the need for users to manually index or scan through long audio files during their search.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If automatic speech recognition is used to convert audio to text for searching, then data mining becomes possible, but the system complexity increases and requires additional processing steps

Engineering Contradiction:
Improvesearch capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a multi-functional system where the automatic speech recognition engine serves multiple purposes: it transcribes audio to text, extracts keywords and phrases, generates timestamps, creates an indexed database, and enables both exact and fuzzy searching. This universal approach consolidates what would otherwise require separate processing steps into a single integrated system, managing complexity through functional consolidation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs self-service by automatically processing the audio stream through speech recognition, keyword extraction, and database indexing without requiring manual configuration or intervention. The system autonomously handles the entire pipeline from raw audio to searchable text database, adapting to different audio inputs and generating appropriate search structures automatically.

Inventive Principle:
Principle #25Self-service

3Loss of information

If existing search solutions are used to find terms in audio files, then occurrences can be located, but the relative importance of terms and their context are not reflected

Engineering Contradiction:
Improvecontext informationVSAvoidsearch efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent applies local quality by providing different levels of search results and information presentation based on user needs. The system can return simple occurrence lists for basic searches, or provide enriched results with context snippets, relevance scoring, and frequency analysis for more advanced queries. Each search result is tailored to provide the appropriate amount of contextual information locally, rather than uniformly for all searches.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system implements feedback mechanisms by analyzing search patterns and providing relevance information. The speech recognition system can identify frequently occurring terms, track their importance throughout the audio stream, and use this feedback to prioritize and present search results in order of relevance. This feedback loop allows the system to learn from usage patterns and improve the quality of search results over time.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9679567B2Word cloud audio navigation
Publication Date: 2017.06.13 AVAYA INC
  • US9679567B2 patent drawing
  • US9679567B2 patent drawing
  • US9679567B2 patent drawing

AI summary

The present invention is directed generally to linking a collection of words and/or phrases with locations in a video and/or audio stream where the words and/or phrases occur and/or associations of a collection of words and/or phrases with a call history.