Voice Search Device Phoneme Zone Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice search technologies face challenges in accurately identifying locations within voice signals, particularly when multiple speakers are involved or when the input voice is acoustically unique or difficult to recognize, leading to insufficient performance compared to string search methods.

Innovation Solution

A voice search device and method that converts search strings into phoneme sequences, derives spoken time lengths, designates likelihood calculation zones, and repeats the process to identify zones where the search string is spoken, using both monophone and triphone models to enhance accuracy and handle varying speech speeds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If voice search is performed using voice input query, then voice search functionality is provided, but search accuracy deteriorates when multiple speakers are involved or when the query voice is acoustically unique

Engineering Contradiction:
Improvevoice search functionalityVSAvoidsearch accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces text as an intermediary between the user's search intent and the voice signal database. Instead of directly comparing query voice with database voices (which fails with multiple speakers or unique voices), the system converts the query to text, extracts search words, and uses these as intermediaries to identify candidate zones in the database, thereby resolving the accuracy problem

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If exhaustive search is performed to ensure precise voice search, then search accuracy is improved, but processing time increases

Engineering Contradiction:
Improvevoice search accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the voice search process into distinct stages: text conversion, search word extraction, candidate zone determination using phoneme sequences and time lengths, and detailed voice comparison only within those candidate zones. This segmentation allows the system to avoid exhaustive search while maintaining accuracy by focusing computational resources on relevant portions of the database

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by converting the query to text and extracting search words before conducting the actual voice search. It also pre-determines candidate zones using phoneme sequences and time length information, so that the computationally intensive voice comparison is performed only on pre-identified relevant segments, significantly reducing processing time

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If voice search handles multiple speakers and unique voices, then adaptability is improved, but computational complexity increases

Engineering Contradiction:
Improvehandling multiple speakers and unique voicesVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses text and search words as intermediaries to bridge the gap between diverse voice inputs and the search database. By converting all voice inputs to text and using extracted search words as the basis for candidate zone identification, the system handles multiple speakers and unique voices uniformly without increasing computational complexity in the voice comparison stage

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9437187B2Voice search device, voice search method, and non-transitory recording medium
Publication Date: 2016.09.06 CASIO COMPUTER CO LTD
  • US9437187B2 patent drawing
  • US9437187B2 patent drawing
  • US9437187B2 patent drawing

AI summary

A search string acquiring unit acquires a search string. A converting unit converts the search string into a phoneme sequence. A time length deriving unit derives the spoken time length of the voice corresponding to the search string. A zone designating unit designates a likelihood acquisition zone in a target voice signal. A likelihood acquiring device acquires a likelihood indicating how likely the likelihood acquisition interval is an interval in which voice corresponding to the search string is spoken. A repeating unit changes the likelihood acquisition zone designated by the zone designating unit, and repeats the process of the zone designating unit and the likelihood acquiring device. An identifying unit identifies, from the target voice signal, estimated intervals for which the voice corresponding to the search string is estimated to be spoken, on the basis of the likelihoods acquired for each of the likelihood acquisition zones.