Audio Delimiter Segmentation for Keyword Search in Document Media

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in efficiently searching for keywords within long movie data embedded in PDF files, as existing methods require extensive time to locate specific portions containing the keywords.

Innovation Solution

An image processing apparatus and method that utilizes speech recognition to generate text data from audio data, determines delimiter positions to divide audio data into manageable pieces, and allows for quick playback of audio and movie data near detected keywords, enabling efficient keyword search and playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If speech recognition technology is applied to extract keywords from audio data in PDF files, then keyword search capability is improved, but search time for long movie data becomes excessively long

Engineering Contradiction:
Improvekeyword search capabilityVSAvoidsearch time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent divides audio data into multiple segments based on time intervals and inserts delimiter markers at segment boundaries. This segmentation allows the search system to quickly skip through non-matching segments rather than processing the entire audio file sequentially, thereby reducing search time while maintaining keyword search capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary processing by generating text data from audio data in advance and inserting delimiter markers at predetermined time intervals before the actual search operation. This preliminary action creates a structured format that enables faster search execution, as the system can immediately jump to relevant segments without processing the entire audio file during the search phase.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If the entire audio data is processed to ensure complete keyword detection, then search accuracy is improved, but processing time increases significantly

Engineering Contradiction:
Improvekeyword detection accuracyVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

By segmenting audio data into manageable parts with delimiter markers, the system can process and search through segments efficiently. The segmentation maintains complete keyword detection capability while allowing the system to quickly identify and focus on relevant segments, improving processing speed without sacrificing accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces text data as an intermediary between audio data and keyword search. The text data serves as a bridge that enables efficient text-based search while representing the audio content, allowing accurate keyword detection without processing the entire audio file directly during search operations.

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables rapid identification and playback of keyword-containing sections within audio and movie data, improving search efficiency and user experience by pre-determining delimiter positions and synchronizing playback with document data display.

Implementation Method 1

generate text data from the audio data by using a speech recognition technology

Methodology Applied
Scientific EffectSpeech recognition:

Data Source

PatentUS9093074B2Image processing apparatus, image processing program and image processing method
Publication Date: 2015.07.28 KONICA MINOLTA BUSINESS TECH INC
  • US9093074B2 patent drawing
  • US9093074B2 patent drawing
  • US9093074B2 patent drawing

AI summary

Regarding audio data related to document data, an image processing apparatus pertaining to the present invention generates text data by using a speech recognition technology in advance, and determines delimiter positions in the text data and the audio data in correspondence. In a keyword search, if a keyword is detected in the text data, the image processing apparatus plays the audio data from a delimiter that is immediately before the keyword.