Automatic Document Section Filtering for Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automatic speech recognition systems face significant degradation due to misalignments between original dictated audio and transcribed text, particularly caused by improperly filtered document sections like headers, footers, and macros, leading to inaccuracies in language model identification, adaptation, and speaker classification, with existing solutions relying on manual rewrites that are time-consuming and costly.

Innovation Solution

An automatic document section filtering system that compares tokenized and normalized forms of original dictation and final reports to identify and filter out non-dictated sections such as headers, footers, and macros using machine-learning techniques, enabling accurate classification and alignment without manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual rewrites are used to filter non-dictated sections, then alignment accuracy between dictated audio and transcribed text is improved, but processing time and costs increase

Engineering Contradiction:
Improvealignment accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical rewriting processes with an automated computational system that uses machine learning classifiers and alignment algorithms to identify and filter non-dictated sections, thereby maintaining high alignment accuracy while eliminating the time-consuming manual intervention

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service by automatically processing documents through trained classifiers that identify non-dictated sections without requiring human transcriptionists to manually rewrite or filter content, allowing the system to serve itself in the filtering task

Inventive Principle:
Principle #25Self-service

2Productivity

If non-dictated sections are left in the text, then processing speed is maintained, but language model identification and adaptation accuracy degrade by 5%

Engineering Contradiction:
Improveprocessing speedVSAvoidlanguage model identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary filtering of non-dictated sections before the text is used for language model identification and adaptation, using trained classifiers to remove headers, footers, and other non-dictated content in advance, ensuring that only clean dictated text proceeds to the ASR processes

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the document into dictated and non-dictated sections using alignment-based filtering and machine learning classifiers, allowing selective processing where only the relevant dictated portions are used for language model tasks, thereby maintaining both speed and accuracy

Inventive Principle:
Principle #1Segmentation

3Productivity

If automatic filtering systems are implemented, then processing efficiency is improved, but system complexity increases

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements a universal filtering system that handles multiple types of non-dictated sections (headers, footers, macros, formatting elements) through a single integrated machine learning classifier framework, allowing the same system to address various filtering needs without requiring separate specialized processes for each type

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces an intermediary alignment-based filtering layer that compares recognized text with the original dictation to identify misalignments, serving as a mediator between raw transcription and clean filtered output, thereby managing complexity through a structured intermediate processing stage

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8036889B2Systems and methods for filtering dictated and non-dictated sections of documents
Publication Date: 2011.10.11 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8036889B2 patent drawing
  • US8036889B2 patent drawing
  • US8036889B2 patent drawing

AI summary

A system and method for filtering documents to determine section boundaries between dictated and non-dictated text. The system and method identifies portions of a text report that correspond to an original dictation and, correspondingly, those portions that are not part of the original dictation. The system and method include comparing tokenized and normalized forms of the original dictation and the final report, determining mismatches between the two forms, and applying machine-learning techniques to identify document headers, footers, page turns, macros, and lists automatically and accurately.