Speech Data Selection Model for Dialog Application Annotation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The time-consuming process of annotating speech data for dialog applications hinders the development of automatic speech recognition and spoken language understanding systems, as tens of thousands of utterances require extensive manual annotation, often taking fifty minutes to annotate one minute of speech data, and traditional random selection methods are inefficient.

Innovation Solution

A speech data selection model identifies and prioritizes utterances for annotation based on system deficiencies and needs, generating an annotation list that includes the specific files, type of annotation, and order, thereby focusing annotation efforts on improving model performance and reducing the overall time required for annotation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If random selection process is used to select speech data for annotation, then all speech data can be annotated eventually, but the annotation process is extremely time-consuming and does not address specific deficiencies in the dialog application

Engineering Contradiction:
ImproveAbility to address specific deficiencies in dialog applicationVSAvoidTime required to annotate speech data
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies local quality by selecting speech data for annotation based on specific local deficiencies in the dialog application performance. Instead of uniform random selection, the system identifies utterances that specifically address weaknesses in speech recognition accuracy or spoken language understanding for particular call types, thereby concentrating annotation efforts where they are most needed to improve overall system performance

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements preliminary action by using the speech data selection model to analyze and identify which speech data should be annotated before the actual annotation process begins. The system evaluates current model performance, identifies deficiencies in specific call types or utterance patterns, and pre-selects the most valuable utterances for annotation, allowing annotators to focus immediately on high-impact data rather than randomly selecting utterances

Inventive Principle:
Principle #10Preliminary action

2Reliability

If tens of thousands of utterances are annotated to build and train speech recognition models and spoken language understanding models, then the dialog application can function properly, but the annotation process becomes arduous and requires considerable time to complete

Engineering Contradiction:
ImprovePerformance of speech recognition and spoken language understanding modelsVSAvoidSpeed of dialog application development
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies partial or excessive action by selecting and annotating only a subset of speech data that is most critical for improving dialog application performance. The speech data selection model identifies utterances that, when annotated, will provide the maximum benefit to addressing specific model deficiencies, thereby achieving reliable model performance with less annotation work than would be required through comprehensive or random annotation approaches

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system implements self-service by automatically analyzing its own performance deficiencies and autonomously selecting which speech data requires annotation. The speech data selection model evaluates current model performance on various call types and utterance patterns, identifies where improvement is needed, and automatically generates a selection of high-priority utterances for annotation without requiring external intervention or manual assessment

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS7412383B1Reducing time for annotating speech data to develop a dialog application
Publication Date: 2008.08.12 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7412383B1 patent drawing
  • US7412383B1 patent drawing
  • US7412383B1 patent drawing

AI summary

Systems and methods for annotating speech data. The present invention reduces the time required to annotate speech data by selecting utterances for annotation that will be of greatest benefit. A selection module uses speech models, including speech recognition models and spoken language understanding models, to identify utterances that should be annotated based on criteria such as confidence scores generated by the models. These utterances are placed in an annotation list along with a type of annotation to be performed for the utterances and an order in which the annotation should proceed. The utterances in the annotation list can be annotated for speech recognition purposes, spoken language understanding purposes, labeling purposes, etc. The selection module can also select utterances for annotation based on previously annotated speech data and deficiencies in the various models.