Secure Transcription Generation via Sensitive Content Splitting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automated clinical documentation systems face risks of data breaches when transcribing audio data containing sensitive information, particularly when done outside a firewall by contractors or quality documentation specialists, as snippets of speech can reveal personally identifiable information and other sensitive details.
Innovation Solution
Implementing a secure transcription generation process using automated speech recognition and natural language understanding to automatically identify and obscure sensitive content within transcriptions, generating obscured speech signals that maintain speaker identity and content consistency while protecting confidentiality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio data is transcribed outside of a firewall by contractors or quality documentation specialists, then transcription accuracy and manual review capability are improved, but data security and confidentiality are compromised
Solution Approach 1:
The patent introduces an intermediary system that processes audio data through automated speech recognition and natural language understanding before external access. This intermediary layer generates transcriptions without exposing the original audio data containing sensitive information, allowing external contractors to review and improve transcription accuracy while maintaining data security boundaries.
Solution Approach 2:
The system extracts only the necessary transcription text from the audio data, separating the sensitive audio content from the transcription output. By taking out just the textual representation and leaving the original audio secured, the system enables external transcription review while preventing exposure of sensitive audio information.
2Productivity
If automated speech recognition is used to generate transcriptions, then productivity is improved, but the ability to maintain speaker identity and content consistency deteriorates
Solution Approach 1:
The system implements feedback loops where natural language understanding analyzes automated speech recognition outputs and identifies inconsistencies in speaker identity attribution. The system then uses this feedback to correct errors and maintain consistent speaker identification throughout the transcription, preserving reliability while benefiting from automated processing speed.
3Object-affected harmful factors
If sensitive content is obscured in transcriptions, then data security is improved, but loss of information occurs
Solution Approach 1:
The system applies local quality by selectively obscuring only the specific sensitive portions of transcriptions (such as personally identifiable information) while leaving the rest of the content clear and accessible. This targeted approach maintains confidentiality protection for sensitive elements without causing unnecessary information loss in non-sensitive areas of the transcription.
Data Source
AI summary
A method, computer program product, and computing system for receiving an input speech signal. A transcription of the input speech signal may be generated via an automated speech recognition (ASR) system. One or more splitting points between one or more sensitive content portions and one or more non-sensitive content portions from the transcription may be identified. The input speech signal maybe split into the one or more sensitive content portions and the one or more non-sensitive content portions based upon, at least in part, the one or more splitting points, thus defining one or more sensitive content signals and one or more non-sensitive content signals.


