Top-Down Error Detection in Voice Recognition Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
There is a significant gap between quantitative and qualitative evaluations in voice recognition systems, with existing methods failing to accurately reflect user satisfaction despite similar word error rates, leading to inconsistent performance assessments.
Innovation Solution
A method is introduced that evaluates error rates in multiple language units (word, character, and sub-word units) using a top-down approach, allowing for a more detailed analysis of error types such as substitution, deletion, and insertion, and calculating a final error rate that considers these evaluations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If error rate is evaluated only in word unit using existing voice recognition methods, then quantitative evaluation is simple and fast, but qualitative evaluation accuracy and user satisfaction assessment are insufficient
Solution Approach 1:
The patent segments the evaluation process into multiple hierarchical levels: word-level evaluation (first language unit), character-level evaluation (second language unit), and sub-word level evaluation (third language unit). This segmentation allows the system to evaluate errors at different granularities, improving measurement precision while managing complexity through a structured top-down approach that focuses detailed analysis only where needed.
2Measurement precision
If detailed error analysis is performed in character and sub-word units, then qualitative evaluation accuracy improves, but computational resources and time consumption increase
Solution Approach 1:
The patent applies preliminary action by first evaluating error rates at the word level before proceeding to character and sub-word levels. This top-down approach allows the system to identify and focus detailed analysis only on specific error-prone areas, performing detailed character-level evaluation only when necessary rather than uniformly across all data, thus improving accuracy while preserving efficiency.
Solution Approach 2:
The patent implements local quality by applying different evaluation depths to different parts of the data. Instead of uniformly evaluating all text at the finest granularity, the system performs detailed character and sub-word level analysis only in local areas where errors are detected at coarser levels, concentrating computational resources where they provide the most value.
3Reliability
If single-level word error rate is used, then evaluation process is simple and fast, but it does not reflect user satisfaction and actual voice recognition quality
Solution Approach 1:
The patent adds another dimension to the evaluation by introducing hierarchical language units (word, character, sub-word) as additional evaluation layers. This multi-dimensional approach transforms the single-word-error-rate metric into a comprehensive evaluation framework that captures errors at multiple granularities, thereby improving reliability of performance assessment while managing complexity through the structured top-down methodology.
Data Source
AI summary
Disclosed is a method for error detection, performed by one or more processors of a computing device according to an example embodiment of the present disclosure. The method includes evaluating an error rate for a sentence to be evaluated, in a first language unit. the method includes evaluating an error rate in a second language unit which is smaller than the first language unit, based on the first language unit error.


