Automatic Document Section Filtering for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current automatic speech recognition systems face significant degradation due to misalignments between original dictated audio and transcribed text, particularly caused by improperly filtered document sections like headers, footers, and macros, leading to inaccuracies in language model identification, adaptation, and speaker classification, with existing solutions relying on manual rewrites that are time-consuming and costly.
Innovation Solution
An automatic document section filtering system that compares tokenized and normalized forms of original dictation and final reports to identify and filter out non-dictated sections such as headers, footers, and macros using machine-learning techniques, enabling accurate classification and alignment without manual intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual rewrites are used to filter non-dictated sections, then alignment accuracy between dictated audio and transcribed text is improved, but processing time and costs increase
Solution Approach 1:
The patent replaces manual mechanical rewriting processes with an automated computational system that uses machine learning classifiers and alignment algorithms to identify and filter non-dictated sections, thereby maintaining high alignment accuracy while eliminating the time-consuming manual intervention
Solution Approach 2:
The system enables self-service by automatically processing documents through trained classifiers that identify non-dictated sections without requiring human transcriptionists to manually rewrite or filter content, allowing the system to serve itself in the filtering task
2Productivity
If non-dictated sections are left in the text, then processing speed is maintained, but language model identification and adaptation accuracy degrade by 5%
Solution Approach 1:
The system performs preliminary filtering of non-dictated sections before the text is used for language model identification and adaptation, using trained classifiers to remove headers, footers, and other non-dictated content in advance, ensuring that only clean dictated text proceeds to the ASR processes
Solution Approach 2:
The patent segments the document into dictated and non-dictated sections using alignment-based filtering and machine learning classifiers, allowing selective processing where only the relevant dictated portions are used for language model tasks, thereby maintaining both speed and accuracy
3Productivity
If automatic filtering systems are implemented, then processing efficiency is improved, but system complexity increases
Solution Approach 1:
The patent implements a universal filtering system that handles multiple types of non-dictated sections (headers, footers, macros, formatting elements) through a single integrated machine learning classifier framework, allowing the same system to address various filtering needs without requiring separate specialized processes for each type
Solution Approach 2:
The system introduces an intermediary alignment-based filtering layer that compares recognized text with the original dictation to identify misalignments, serving as a mediator between raw transcription and clean filtered output, thereby managing complexity through a structured intermediate processing stage
Data Source
AI summary
A system and method for filtering documents to determine section boundaries between dictated and non-dictated text. The system and method identifies portions of a text report that correspond to an original dictation and, correspondingly, those portions that are not part of the original dictation. The system and method include comparing tokenized and normalized forms of the original dictation and the final report, determining mismatches between the two forms, and applying machine-learning techniques to identify document headers, footers, page turns, macros, and lists automatically and accurately.


