Audio Delimiter Segmentation for Keyword Search in Document Media
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in efficiently searching for keywords within long movie data embedded in PDF files, as existing methods require extensive time to locate specific portions containing the keywords.
Innovation Solution
An image processing apparatus and method that utilizes speech recognition to generate text data from audio data, determines delimiter positions to divide audio data into manageable pieces, and allows for quick playback of audio and movie data near detected keywords, enabling efficient keyword search and playback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If speech recognition technology is applied to extract keywords from audio data in PDF files, then keyword search capability is improved, but search time for long movie data becomes excessively long
Solution Approach 1:
The patent divides audio data into multiple segments based on time intervals and inserts delimiter markers at segment boundaries. This segmentation allows the search system to quickly skip through non-matching segments rather than processing the entire audio file sequentially, thereby reducing search time while maintaining keyword search capability.
Solution Approach 2:
The patent performs preliminary processing by generating text data from audio data in advance and inserting delimiter markers at predetermined time intervals before the actual search operation. This preliminary action creates a structured format that enables faster search execution, as the system can immediately jump to relevant segments without processing the entire audio file during the search phase.
2Reliability
If the entire audio data is processed to ensure complete keyword detection, then search accuracy is improved, but processing time increases significantly
Solution Approach 1:
By segmenting audio data into manageable parts with delimiter markers, the system can process and search through segments efficiently. The segmentation maintains complete keyword detection capability while allowing the system to quickly identify and focus on relevant segments, improving processing speed without sacrificing accuracy.
Solution Approach 2:
The patent introduces text data as an intermediary between audio data and keyword search. The text data serves as a bridge that enables efficient text-based search while representing the audio content, allowing accurate keyword detection without processing the entire audio file directly during search operations.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables rapid identification and playback of keyword-containing sections within audio and movie data, improving search efficiency and user experience by pre-determining delimiter positions and synchronizing playback with document data display.
Implementation Method 1
generate text data from the audio data by using a speech recognition technology
Data Source
AI summary
Regarding audio data related to document data, an image processing apparatus pertaining to the present invention generates text data by using a speech recognition technology in advance, and determines delimiter positions in the text data and the audio data in correspondence. In a keyword search, if a keyword is detected in the text data, the image processing apparatus plays the audio data from a delimiter that is immediately before the keyword.


