Multi-modal Lung Capacity Measurement via Audio Video Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting respiratory illness, such as X-ray and CT scans, are expensive, require specialized setups, and are not readily available, making early detection of lung capacity changes challenging.
Innovation Solution
A method and system for multi-modal lung capacity measurement using audio and video analysis, where a user is prompted to perform specific utterances, and machine learning models analyze speech characteristics to determine lung capacity and health indicators, enabling self-assessment on mobile devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If X-ray and CT scans are used for lung capacity measurement, then measurement precision is improved, but device complexity and cost increase
Solution Approach 1:
The patent replaces complex mechanical imaging systems (X-ray, CT scanners) with acoustic and optical sensing systems using standard mobile device components. The speech-based lung capacity measurement substitutes sophisticated imaging technology with simple audio recording and analysis, while video-based respiratory rate monitoring replaces complex imaging with standard camera capabilities.
Solution Approach 2:
The patent creates functional equivalents of complex medical imaging systems using simple mobile device sensors. Instead of directly imaging lung structure, the system captures acoustic copies of breath sounds and visual copies of respiratory movements, then analyzes these copies to infer lung capacity and respiratory health indicators.
2Measurement precision
If X-ray and CT scans are used for lung capacity measurement, then measurement precision is improved, but accessibility and ease of operation worsen
Solution Approach 1:
The patent enables individuals to perform their own lung capacity assessments using standard mobile devices. The system guides users through speech-based tests and video-based respiratory monitoring, automatically analyzing the recordings to provide health indicators without requiring medical professionals or specialized facilities.
Solution Approach 2:
The patent transforms standard mobile devices into multi-functional health assessment tools. The same device used for communication and entertainment also performs lung capacity measurement, respiratory rate monitoring, and early detection of respiratory illnesses, making advanced health screening universally accessible.
3Ease of operation
If speech and video analysis is used for lung capacity measurement, then ease of operation and accessibility are improved, but measurement precision may worsen
Solution Approach 1:
The patent combines multiple sensing modalities (audio speech analysis, video respiratory movement tracking) to compensate for the limitations of individual sensors. By fusing data from different sources and analyzing multiple aspects of respiratory function simultaneously, the system achieves measurement accuracy comparable to specialized equipment.
Solution Approach 2:
The patent analyzes multiple parameters from speech and video data (acoustic features, temporal patterns, spectral characteristics, motion dynamics) to infer lung capacity. By examining changes in these parameters and their relationships, the system extracts accurate health indicators from simple mobile device recordings.
Data Source
AI summary
Determining lung capacity of includes capturing an audio waveform of the user performing an utterance presented to a user. A video of the user performing the utterance can be captured. The captured audio waveform and the video are analyzed for compliance. Based on the audio waveform, an indicator of respiratory function is determined. The indicator is compared with a reference indicator to determine health of the user. A machine learning model such as neural network can be trained to predict the indicator of the respiratory function based on input features comprising audio spectral and temporal characteristics of utterances. Determining the indicator or respiratory function can include running the trained machine learning model.


