Speech Retrieval Using Phoneme Code Mediator

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech retrieval systems face significant information loss and accuracy issues due to the need for conversion between character and audio formats, leading to low search result accuracy and loss of contextual information.

Innovation Solution

A speech retrieval apparatus and method that directly converts text search conditions into audio format without translation, utilizing a speech-to-speech search unit to retain and utilize speech features, thereby avoiding information loss and the negative influences of low accuracy recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech is converted into text format through automatic speech recognition for search, then text-based search can be performed, but information loss occurs and recognition accuracy limitations reduce search quality

Engineering Contradiction:
Improvetext-based search capabilityVSAvoidspeech information loss
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent introduces a phoneme code as an intermediary representation that bridges speech and text formats. Speech is converted to phoneme codes which retain speech characteristics, and text is also converted to phoneme codes for comparison. This intermediary format avoids direct speech-to-text conversion losses while enabling text-based search operations.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms both speech and text into a different parameter space (phoneme codes) rather than converting speech directly to text. This parameter transformation preserves speech features while enabling text-based query processing, resolving the contradiction between ease of text search and preservation of speech information.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If speech and text are converted into the same third format (e.g., phonemic code) for search, then unified search can be performed, but translation accuracy is not high and confusion occurs

Engineering Contradiction:
Improveunified search capabilityVSAvoidsearch accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

Instead of converting speech and text to a third format and comparing them, the patent inverts the approach by converting the text query to speech (using text-to-speech synthesis) and then performing speech-to-speech comparison. This inversion avoids the accuracy problems of translating both to phoneme codes while maintaining unified search capability.

Inventive Principle:
Principle #13The other way round (Inversion)

3Ease of operation

If only related documents of speech are used for general search, then search simplicity is maintained, but the amount of information used is very small

Engineering Contradiction:
Improvesearch simplicityVSAvoidinformation quantity
Core Design Contradiction:
Ease of operationVSQuantity of substance

Solution Approach 1:

The patent performs preliminary speech-to-speech comparison using phoneme codes before final retrieval. By pre-converting the query to speech format and performing initial matching in the speech domain, the system can quickly filter relevant results while maintaining simple operation, thereby increasing the effective information used without complicating the user interface.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP2348427B1Speech retrieval apparatus and speech retrieval method
Publication Date: 2012.11.21 RICOH CO LTD
  • EP2348427B1 patent drawingFigure 1
  • EP2348427B1 patent drawingFigure 2
  • EP2348427B1 patent drawingFigure 3

AI summary

Disclosed are a speech retrieval apparatus (100)and a speech retrieval method for searching, in an audio file database (150), for one or more target audio files by using one or more input search terms. The speech retrieval apparatus (100) comprises a related document obtaining unit (110) configured to search, in a related document database (140) where documents related to audio files in the audio file database (150) are stored, for one or more related documents by using the search terms; a correspondence audio file obtaining unit (120) configured to search, in the audio file database (150), for one or more correspondence audio files corresponding to the obtained related documents; and a speech-to-speech search unit (130) configured to search, in the audio file database (150), for the target audio files by using the obtained correspondence audio files.