Speech Data Quality Assessment Using Rule-Based Error Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing quality assessment methods for speech data annotation are inadequate for automatic annotations, as they fail to effectively detect and correct errors specific to machine-generated data, leading to suboptimal performance in speech recognition for ethnic minorities due to non-uniform standards and variable annotation qualities.

Innovation Solution

A quality assessment method for automatic speech data annotation using a logical reasoning mechanism based on a rule-base, incorporating quality key indicators like word error rate, sentence error rate, bias feature error rate, and user feedback error rate, which allows for comprehensive error detection and correction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing quality assessment methods for manual data annotations are used, then quality assessment can be performed through sampling analysis or probability models, but these methods are not suitable for automatic data annotations because error causes, quality problem types and rules are completely different between automatic machine annotation and manual annotation

Engineering Contradiction:
Improveadaptability of quality assessment methodVSAvoidreliability of quality assessment
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent changes the assessment parameters from manual annotation metrics (sampling analysis, probability models) to automatic annotation-specific metrics (confidence score distribution, contradiction detection rates, error pattern frequencies). This allows the quality assessment method to adapt to automatic annotation while maintaining reliability through metrics that actually reflect machine annotation error characteristics.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

Instead of trying to apply manual annotation assessment methods to automatic annotation, the patent inverts the approach by developing assessment methods specifically tailored to automatic annotation characteristics. It assesses automatic annotation quality by detecting machine-specific error patterns rather than treating automatic annotation as a subset of manual annotation assessment.

Inventive Principle:
Principle #13The other way round (Inversion)

2Productivity

If automatic machine annotation is used to replace manpower, then productivity is improved, but data annotation errors are difficult to avoid resulting from factors such as raw data quality, model limitations and annotation complexity

Engineering Contradiction:
Improveannotation efficiencyVSAvoidannotation accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent implements feedback mechanisms where quality assessment results are used to iteratively improve the automatic annotation system. Error patterns detected through quality assessment feed back into model retraining and parameter optimization, creating a closed-loop system that continuously improves annotation accuracy while maintaining high productivity.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces quality assessment as an intermediary layer between automatic annotation and final data usage. This intermediary performs contradiction detection, confidence score validation, and error pattern analysis to filter and correct machine annotation errors before the data is used for training or analysis, thereby maintaining both productivity and accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If non-uniform standards and variable annotation qualities are present, then automatic annotation can be performed with flexibility, but applications and developments of data annotations are largely hindered

Engineering Contradiction:
Improveflexibility of annotation processVSAvoiduniformity of annotation quality
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent creates a universal quality assessment framework that can handle multiple annotation types (text, audio, video) and multiple error patterns through a single unified system. The contradiction detection mechanism and confidence score validation work across different annotation domains, providing uniform quality standards while maintaining flexibility in application-specific implementation details.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11790166B2Quality assessment method for automatic annotation of speech data
Publication Date: 2023.10.17 KUNMING UNIVERSITY
  • US11790166B2 patent drawing

AI summary

A quality assessment method for automatic annotation of speech data is provided and includes: building a base rule-base of automatically annotated speech data based on quality key indicators; reading automatically annotated speech data to be detected, and performing quality detection on the automatically annotated speech data to be detected according to the quality key indicators to thereby complete quality measurement; updating an automatically annotated speech dataset according to a result of the quality measurement; and importing the automatically annotated speech dataset after the updating into the base rule-base. The shortcomings of using traditional quality assessment methods for data annotation in automatic machine annotations can be overcome, and it can play a very positive supporting role in promoting the development of ethnic minority speech intelligence.