Voice Audio PHI Segregation Using Local Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice-activated systems lack the necessary security and privacy protections to process and store personal health information, which is classified similarly to medical data, necessitating a system to segregate and securely store such information.
Innovation Solution
A voice processing system that includes a local device with a speech-to-text transcriber, natural language processor, and machine learning classifier to identify personal health information, transmitting it to a secure personal health data ecosystem while routing non-personal health information to a less secure smart home ecosystem.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice audio is manually reviewed and transcribed to identify PHI, then identification accuracy is improved, but processing time and labor costs increase
Solution Approach 1:
The patent replaces manual mechanical review and transcription of voice audio with an automated system comprising speech-to-text conversion, natural language processing, and machine learning algorithms. This substitution maintains high PHI identification accuracy while dramatically reducing processing time and eliminating manual labor requirements.
Solution Approach 2:
The patent introduces intermediate processing layers including speech-to-text conversion services and natural language processing components that act as mediators between the raw voice audio and PHI identification. These intermediaries enable automated analysis while maintaining accuracy comparable to manual review.
2Productivity
If automated speech-to-text conversion is used to process voice audio, then processing speed is improved, but accuracy in identifying PHI deteriorates
Solution Approach 1:
The patent implements feedback mechanisms where the system continuously refines its PHI identification based on analyzed patterns and outcomes. The machine learning models are trained on labeled data and improve their accuracy over time through feedback from correct and incorrect identifications, ensuring high precision despite automated processing speed.
Solution Approach 2:
The patent performs preliminary actions by pre-training machine learning models on labeled medical transcription data and configuring natural language processing rules before actual PHI identification. This preliminary preparation ensures that when automated speech-to-text conversion occurs, the system is already optimized for accurate PHI detection from the outset.
3Measurement precision
If all voice audio data is retained for potential PHI identification, then completeness of PHI detection is improved, but data security risks and storage requirements increase
Solution Approach 1:
The patent extracts only the necessary information for PHI identification by processing voice audio through speech-to-text conversion and analyzing only relevant portions containing potential PHI. Rather than retaining all voice audio data, the system extracts and processes only the transcription data necessary for identification, reducing storage requirements and security risks while maintaining complete PHI detection.
4Measurement precision
If manual review processes are used to verify PHI identification, then identification accuracy is improved, but processing complexity and costs increase
Solution Approach 1:
The patent enables the system to perform self-service verification through machine learning models that automatically validate PHI identification against trained criteria and patterns. The system self-corrects and verifies identifications without requiring external manual review, maintaining high accuracy while reducing processing complexity and eliminating the need for additional review layers.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A system for processing voice audio includes a local device and a remote personal health data ecosystem. The local device includes (1) a local speech-to-text transcriber configured to generate voice text based on voice audio spoken by a user; (2) a local NLP configured to extract spoken phrases from the voice text; and (3) an ML classifier configured to classify the voice audio as either personal health or non-personal health voice audio. The remote personal health data ecosystem includes (1) a remote speech-to-text transcriber configured to generate personal health voice text based on the personal health voice audio; (2) a remote NLP configured to extract personal health spoken phrases from the personal health voice text; (3) a text response generator configured to generate a text response based on the personal health spoken phrases; (4) a text-to-speech translator configured to generate a voice response based on the text response.