Passive Audio Subject Monitoring via Embedding Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional health monitoring systems face challenges in accurately identifying and verifying audio events from specific subjects due to issues like misdiagnosis, limited data collection, and sensitivity to subject condition changes, particularly in pulmonary conditions, leading to incorrect health assessments and potential harm.
Innovation Solution
A system and method for passive subject-specific monitoring using audio embeddings extracted by a trained machine learning model, which compares audio segments to a match profile generated during enrollment, enabling event-independent verification and adaptation to changing conditions, thus ensuring accurate identification of audio events from specific subjects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional health monitoring systems collect audio data from subjects, then health monitoring capability is provided, but accuracy in identifying audio events from specific subjects deteriorates due to misdiagnosis and sensitivity to subject condition changes
Solution Approach 1:
The patent segments the audio verification process into distinct components: audio embedding extraction from input audio, retrieval of stored embeddings for the specific subject, comparison operations, and matching against a match profile. This segmentation allows each component to be optimized independently, improving overall accuracy in identifying audio events from specific subjects while reducing misdiagnosis.
Solution Approach 2:
The patent transforms audio data into a different parameter space using audio embeddings - converting raw audio signals into compressed numerical representations that capture essential characteristics. This parameter transformation makes the verification process more robust to subject condition changes (like respiratory variations in pulmonary conditions) while maintaining reliability in identifying genuine subject-specific audio events.
2Measurement precision
If audio embeddings are extracted and compared using machine learning models, then subject-specific verification accuracy is improved, but device complexity increases
Solution Approach 1:
The patent performs preliminary actions by pre-extracting and storing audio embeddings during an enrollment phase, before actual verification is needed. Match profiles are created in advance for each subject. This preliminary preparation reduces the computational burden during real-time verification, maintaining high subject-specific accuracy while managing device complexity by shifting computation to an offline setup phase.
Solution Approach 2:
The patent creates simplified copies of audio data in the form of audio embeddings - compressed numerical representations that capture the essential characteristics of the original audio. These embedding copies are stored and reused for multiple verification operations, reducing the need to process full audio files repeatedly and thereby managing device complexity while maintaining verification accuracy.
3Ease of operation
If passive audio recording is used for health monitoring, then unobtrusive monitoring is achieved, but reliability deteriorates due to background noise and limited data collection
Solution Approach 1:
The patent extracts only the relevant subject-specific audio events from the passive recording by comparing against stored match profiles. Instead of analyzing all recorded audio, the system extracts and verifies only those segments that match the enrolled subject's audio characteristics, filtering out background noise and irrelevant sounds. This maintains unobtrusive monitoring while improving reliability through targeted verification.
Solution Approach 2:
The patent implements feedback by continuously comparing newly extracted audio embeddings against the stored match profile and using the matching results to confirm or reject audio events as belonging to the specific subject. This feedback mechanism allows the system to reliably distinguish subject-specific audio from background noise in passive recordings, improving health monitoring accuracy without requiring active subject participation.
Data Source
AI summary
A method includes obtaining, by an electronic device, an audio segment comprising one or more audio events of a target subject. The method also includes extracting, by the electronic device, audio embeddings from the one or more audio events using an embedding model, the embedding model comprising a trained machine learning model. The method further includes comparing, by the electronic device, the extracted audio embeddings with a match profile of the target subject, the match profile generated during an enrollment stage. The method also includes generating, by the electronic device, a label for the audio segment based on whether or not the extracted audio embeddings match the match profile, wherein the label enables correlation of the audio segment with the target subject for monitoring a health condition of the target subject.


