Speech Retrieval Using Phoneme Code Mediator
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech retrieval systems face significant information loss and accuracy issues due to the need for conversion between character and audio formats, leading to low search result accuracy and loss of contextual information.
Innovation Solution
A speech retrieval apparatus and method that directly converts text search conditions into audio format without translation, utilizing a speech-to-speech search unit to retain and utilize speech features, thereby avoiding information loss and the negative influences of low accuracy recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If speech is converted into text format through automatic speech recognition for search, then text-based search can be performed, but information loss occurs and recognition accuracy limitations reduce search quality
Solution Approach 1:
The patent introduces a phoneme code as an intermediary representation that bridges speech and text formats. Speech is converted to phoneme codes which retain speech characteristics, and text is also converted to phoneme codes for comparison. This intermediary format avoids direct speech-to-text conversion losses while enabling text-based search operations.
Solution Approach 2:
The patent transforms both speech and text into a different parameter space (phoneme codes) rather than converting speech directly to text. This parameter transformation preserves speech features while enabling text-based query processing, resolving the contradiction between ease of text search and preservation of speech information.
2Adaptability or versatility
If speech and text are converted into the same third format (e.g., phonemic code) for search, then unified search can be performed, but translation accuracy is not high and confusion occurs
Solution Approach 1:
Instead of converting speech and text to a third format and comparing them, the patent inverts the approach by converting the text query to speech (using text-to-speech synthesis) and then performing speech-to-speech comparison. This inversion avoids the accuracy problems of translating both to phoneme codes while maintaining unified search capability.
3Ease of operation
If only related documents of speech are used for general search, then search simplicity is maintained, but the amount of information used is very small
Solution Approach 1:
The patent performs preliminary speech-to-speech comparison using phoneme codes before final retrieval. By pre-converting the query to speech format and performing initial matching in the speech domain, the system can quickly filter relevant results while maintaining simple operation, thereby increasing the effective information used without complicating the user interface.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Disclosed are a speech retrieval apparatus (100)and a speech retrieval method for searching, in an audio file database (150), for one or more target audio files by using one or more input search terms. The speech retrieval apparatus (100) comprises a related document obtaining unit (110) configured to search, in a related document database (140) where documents related to audio files in the audio file database (150) are stored, for one or more related documents by using the search terms; a correspondence audio file obtaining unit (120) configured to search, in the audio file database (150), for one or more correspondence audio files corresponding to the obtained related documents; and a speech-to-speech search unit (130) configured to search, in the audio file database (150), for the target audio files by using the obtained correspondence audio files.