Voice Data Tokenization for Bidirectional Playback Navigation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice data management systems lack efficient navigation methods, requiring users to listen to entire messages to find specific information, as they only allow forward jumping and not backward navigation, and fail to preserve context when extracting keywords or phrases.
Innovation Solution
Tokenizing voice data based on predefined content criteria, enabling bidirectional scanning by marking pertinent points such as pauses, voice inflections, and keywords, allowing users to jump forward and backward during playback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If voice data is converted to text and keywords are extracted, then information retrieval efficiency is improved, but contextual information is lost
Solution Approach 1:
The patent segments voice data into multiple sections marked by tokens at pause points, allowing the system to extract and navigate to specific segments containing keywords while preserving the overall structure and context of the original voice message. This segmentation enables selective retrieval without losing contextual relationships between different parts of the message.
2Ease of operation
If fixed time jumping is implemented for message navigation, then playback control flexibility is improved, but navigation precision to specific information is worsened
Solution Approach 1:
The patent introduces tokens as intermediary markers placed at pause points throughout the voice data. These tokens serve as navigation anchors that enable precise jumping to specific information locations while maintaining flexible playback control. The tokens act as intermediaries between the user's navigation requests and the actual content locations, combining both flexibility and precision.
3Loss of information
If entire voice messages must be listened to for information retrieval, then complete context understanding is improved, but time consumption is worsened
Solution Approach 1:
The patent performs preliminary actions by pre-processing voice data to identify pause points and insert tokens at these locations before the user needs to retrieve information. This pre-segmentation and tokenization allows users to quickly navigate to specific sections containing desired information without listening to the entire message, while the token structure preserves contextual relationships for comprehensive understanding when needed.
Data Source
AI summary
A system, system, and program for facilitating navigation of voice data are provided. Tokens are added to voice data based on predefined content criteria. Then, bidirectional scanning of the voice data to a next token within the voice data is enabled, such that navigation to pertinent locations within the voice data during playback is facilitated. When adding tokens to voice data, the voice data may be scanned to detect pauses, changes in voice inflection, and other vocal characteristics. Based on the detected vocal characteristics, tokens identifying ends of sentences, separations between words, and other structures are marked. In addition, when adding tokens to voice data, the voice data may be first converted to text. The text is then scanned for keywords, phrases, and types of information. Tokens are added in the voice data at locations identified within the text as meeting the predefined content criteria.


