Dynamic Section Grammar Loading for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems lack the ability to dynamically switch between domains within a single document, especially in unstructured environments like traditional telephony dictation, and are limited by the need for well-defined data fields and vocabulary constraints.

Innovation Solution

A system and method for loading and unloading dynamic grammars and language models that identify and adapt to section-based structures in documents, allowing for seamless domain switching without requiring predefined data fields or vocabulary constraints, using dynamic section identification and incremental training based on user dictation patterns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If domain-specific language models are used to improve speech recognition accuracy in specific domains, then accuracy is improved, but the system cannot switch between different domains within a single document

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiddomain switching capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system dynamically loads and unloads language models and grammars based on the current section being processed. Instead of using a static domain-specific model throughout, the system adapts by switching between different language models corresponding to different sections (e.g., narrative section model, structured section model), enabling both accuracy and domain switching capability

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The document is divided into sections with different structural characteristics (narrative vs. structured). Each section type has its own specialized language model. The system segments the processing task by identifying section boundaries and applying the appropriate language model to each section, resolving the contradiction between specialized accuracy and overall adaptability

Inventive Principle:
Principle #1Segmentation

2Productivity

If structured report organization with well-defined data fields is used to improve recognition efficiency, then efficiency is improved, but the system cannot handle unstructured environments like traditional telephony dictation

Engineering Contradiction:
Improverecognition efficiencyVSAvoidhandling unstructured documents
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system is designed to handle both structured and unstructured documents through a universal framework. It uses section identification algorithms that work regardless of document structure, and can switch between narrative language models and structured section models. This multi-functionality allows the same system to efficiently process both well-defined data fields and unstructured telephony dictation

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If vocabulary constraints are applied to improve recognition accuracy, then accuracy is improved, but the system becomes limited by predefined vocabulary

Engineering Contradiction:
Improverecognition accuracyVSAvoidvocabulary flexibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The vocabulary constraints are applied dynamically rather than statically. The system identifies the current section type and applies vocabulary constraints appropriate to that section. This allows the system to maintain high accuracy within each section type while remaining flexible across different section types and domains

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS9002710B2System and method for applying dynamic contextual grammars and language models to improve automatic speech recognition accuracy
Publication Date: 2015.04.07 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9002710B2 patent drawing
  • US9002710B2 patent drawing
  • US9002710B2 patent drawing

AI summary

The invention involves the loading and unloading of dynamic section grammars and language models in a speech recognition system. The values of the sections of the structured document are either determined in advance from a collection of documents of the same domain, document type, and speaker; or collected incrementally from documents of the same domain, document type, and speaker; or added incrementally to an already existing set of values. Speech recognition in the context of the given field is constrained to the contents of these dynamic values. If speech recognition fails or produces a poor match within this grammar or section language model, speech recognition against a larger, more general vocabulary that is not constrained to the given section is performed.