Difference Extraction Device for Speech Recognition Dictionary Registration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing dictionary registration techniques for speech recognition systems fail to accurately identify and differentiate between unknown words and their correct notations, leading to potential misregistration of words in the dictionary, especially for compound words and notation fluctuations.

Innovation Solution

A difference extraction device that converts input notation strings into pronunciation strings and then into output notation strings using morphological analysis and acoustic score vectors, comparing these to extract differences and prevent misregistration by displaying and allowing user registration of only necessary words.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If compound words are extracted from morphological analysis results and treated as unknown words, then the quantity of candidate words for dictionary registration increases, but the precision of identifying truly unknown words decreases due to notation fluctuations

Engineering Contradiction:
Improvequantity of candidate wordsVSAvoidprecision of unknown word identification
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent introduces pronunciation strings as an intermediary representation between input notation strings and output notation strings. By converting notation strings to pronunciation strings and back, the system creates a mediating layer that helps distinguish true unknown words from notation fluctuations, resolving the contradiction between capturing all candidates and maintaining identification precision

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameter of word representation from notation-based to pronunciation-based. By using acoustic score vectors and pronunciation features as intermediate parameters, the system can compare semantic equivalence independently of notation variations, allowing it to identify true unknown words while filtering out notation fluctuations

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If all extracted differences are registered in the dictionary, then the completeness of dictionary coverage improves, but the reliability of dictionary accuracy decreases due to misregistration of notation fluctuations

Engineering Contradiction:
Improvecoverage of dictionaryVSAvoidaccuracy of dictionary registration
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent implements a feedback mechanism where pronunciation strings are converted back to output notation strings and compared with input notation strings. This feedback loop allows the system to verify whether extracted differences represent true unknown words or merely notation fluctuations, ensuring reliable dictionary registration while maintaining comprehensive coverage

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent segments the word identification process into distinct stages: input notation conversion to pronunciation, pronunciation conversion to output notation, and difference extraction through comparison. This segmentation allows each stage to focus on specific aspects, improving both coverage and accuracy of dictionary registration

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12118304B2Difference extraction device, method and program
Publication Date: 2024.10.15 KK TOSHIBA
  • US12118304B2 patent drawing
  • US12118304B2 patent drawing
  • US12118304B2 patent drawing

AI summary

According to one embodiment, a difference extraction device includes processing circuitry. The processing circuitry acquires a text in which an input notation string is described. The processing circuitry converts the input notation string into a pronunciation string. The processing circuitry executes a pronunciation string conversion process in which the pronunciation string is converted into an output notation string. The processing circuitry extracts a difference by comparing the input notation string and the output notation string with each other.