Voice Search Device Phoneme Zone Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice search technologies face challenges in accurately identifying locations within voice signals, particularly when multiple speakers are involved or when the input voice is acoustically unique or difficult to recognize, leading to insufficient performance compared to string search methods.
Innovation Solution
A voice search device and method that converts search strings into phoneme sequences, derives spoken time lengths, designates likelihood calculation zones, and repeats the process to identify zones where the search string is spoken, using both monophone and triphone models to enhance accuracy and handle varying speech speeds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If voice search is performed using voice input query, then voice search functionality is provided, but search accuracy deteriorates when multiple speakers are involved or when the query voice is acoustically unique
Solution Approach 1:
The patent introduces text as an intermediary between the user's search intent and the voice signal database. Instead of directly comparing query voice with database voices (which fails with multiple speakers or unique voices), the system converts the query to text, extracts search words, and uses these as intermediaries to identify candidate zones in the database, thereby resolving the accuracy problem
2Measurement precision
If exhaustive search is performed to ensure precise voice search, then search accuracy is improved, but processing time increases
Solution Approach 1:
The patent segments the voice search process into distinct stages: text conversion, search word extraction, candidate zone determination using phoneme sequences and time lengths, and detailed voice comparison only within those candidate zones. This segmentation allows the system to avoid exhaustive search while maintaining accuracy by focusing computational resources on relevant portions of the database
Solution Approach 2:
The patent performs preliminary actions by converting the query to text and extracting search words before conducting the actual voice search. It also pre-determines candidate zones using phoneme sequences and time length information, so that the computationally intensive voice comparison is performed only on pre-identified relevant segments, significantly reducing processing time
3Adaptability or versatility
If voice search handles multiple speakers and unique voices, then adaptability is improved, but computational complexity increases
Solution Approach 1:
The patent uses text and search words as intermediaries to bridge the gap between diverse voice inputs and the search database. By converting all voice inputs to text and using extracted search words as the basis for candidate zone identification, the system handles multiple speakers and unique voices uniformly without increasing computational complexity in the voice comparison stage
Data Source
AI summary
A search string acquiring unit acquires a search string. A converting unit converts the search string into a phoneme sequence. A time length deriving unit derives the spoken time length of the voice corresponding to the search string. A zone designating unit designates a likelihood acquisition zone in a target voice signal. A likelihood acquiring device acquires a likelihood indicating how likely the likelihood acquisition interval is an interval in which voice corresponding to the search string is spoken. A repeating unit changes the likelihood acquisition zone designated by the zone designating unit, and repeats the process of the zone designating unit and the likelihood acquiring device. An identifying unit identifies, from the target voice signal, estimated intervals for which the voice corresponding to the search string is estimated to be spoken, on the basis of the likelihoods acquired for each of the likelihood acquisition zones.


