Voice Data Tokenization for Bidirectional Playback Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice data management systems lack efficient navigation methods, requiring users to listen to entire messages to find specific information, as they only allow forward jumping and not backward navigation, and fail to preserve context when extracting keywords or phrases.

Innovation Solution

Tokenizing voice data based on predefined content criteria, enabling bidirectional scanning by marking pertinent points such as pauses, voice inflections, and keywords, allowing users to jump forward and backward during playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If voice data is converted to text and keywords are extracted, then information retrieval efficiency is improved, but contextual information is lost

Engineering Contradiction:
Improveinformation retrieval efficiencyVSAvoidcontextual information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent segments voice data into multiple sections marked by tokens at pause points, allowing the system to extract and navigate to specific segments containing keywords while preserving the overall structure and context of the original voice message. This segmentation enables selective retrieval without losing contextual relationships between different parts of the message.

Inventive Principle:
Principle #1Segmentation

2Ease of operation

If fixed time jumping is implemented for message navigation, then playback control flexibility is improved, but navigation precision to specific information is worsened

Engineering Contradiction:
Improveplayback control flexibilityVSAvoidnavigation precision
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent introduces tokens as intermediary markers placed at pause points throughout the voice data. These tokens serve as navigation anchors that enable precise jumping to specific information locations while maintaining flexible playback control. The tokens act as intermediaries between the user's navigation requests and the actual content locations, combining both flexibility and precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If entire voice messages must be listened to for information retrieval, then complete context understanding is improved, but time consumption is worsened

Engineering Contradiction:
Improvecontext understandingVSAvoidtime consumption
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-processing voice data to identify pause points and insert tokens at these locations before the user needs to retrieve information. This pre-segmentation and tokenization allows users to quickly navigate to specific sections containing desired information without listening to the entire message, while the token structure preserves contextual relationships for comprehensive understanding when needed.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS7478044B2Facilitating navigation of voice data
Publication Date: 2009.01.13 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7478044B2 patent drawing
  • US7478044B2 patent drawing
  • US7478044B2 patent drawing

AI summary

A system, system, and program for facilitating navigation of voice data are provided. Tokens are added to voice data based on predefined content criteria. Then, bidirectional scanning of the voice data to a next token within the voice data is enabled, such that navigation to pertinent locations within the voice data during playback is facilitated. When adding tokens to voice data, the voice data may be scanned to detect pauses, changes in voice inflection, and other vocal characteristics. Based on the detected vocal characteristics, tokens identifying ends of sentences, separations between words, and other structures are marked. In addition, when adding tokens to voice data, the voice data may be first converted to text. The text is then scanned for keywords, phrases, and types of information. Tokens are added in the voice data at locations identified within the text as meeting the predefined content criteria.