Multi-head neural network for physiological signal estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for measuring physiological parameters like heart rate and respiratory rate often require direct contact with the body or extensive signal processing, which can be invasive or costly, and lack robustness across varying facial characteristics and conditions.
Innovation Solution
A multi-head neural network model trained on a diverse set of facial video data is used to estimate heart rate and respiratory rate from RGB video frames, allowing for non-contact, mobile, and efficient prediction of multiple physiological signals using a smartphone camera, with shared weights across prediction heads for improved robustness and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional signal processing methodologies are used to derive physiological parameters from video, then the measurement can be performed non-contact, but the measurement precision and reliability are insufficient due to noise and variability in facial characteristics
Solution Approach 1:
The patent replaces conventional signal processing methodologies (mechanical/system-based approach) with a multi-head neural network model (AI-based approach). The neural network learns to extract physiological signals directly from RGB video frames, substituting traditional signal isolation and amplification methods with deep learning-based feature extraction that is more robust to noise and facial variability
Solution Approach 2:
The patent changes the approach from processing individual color channels separately (traditional method) to utilizing multi-head neural networks that process multiple features simultaneously. The system transforms the input representation from simple green channel extraction to comprehensive RGB frame analysis with multiple prediction heads, changing how physiological parameters are derived from video data
2Measurement precision
If direct contact sensors are used to measure physiological parameters, then measurement precision is high, but ease of operation and user comfort deteriorate due to invasiveness
Solution Approach 1:
The patent replaces contact-based mechanical sensors (strain gauges, accelerometers) with a non-contact optical system using standard RGB cameras. The multi-head neural network enables accurate physiological parameter estimation from video frames without requiring physical contact with the subject, combining the convenience of non-invasive measurement with precision through AI-based analysis
Solution Approach 2:
The patent makes a single RGB camera system universal for multiple physiological measurements. Instead of requiring different specialized sensors for different parameters, the multi-head neural network enables a single camera to measure multiple physiological signals (heart rate, respiratory rate, etc.) simultaneously, achieving both convenience and accuracy
3Measurement precision
If multiple separate models are used to predict different physiological signals, then each prediction can be optimized, but device complexity and computational requirements increase
Solution Approach 1:
The patent merges multiple prediction functions into a single multi-head neural network model. Instead of using separate independent models for heart rate, respiratory rate, and other physiological parameters, the system combines them into one unified architecture with multiple output heads, reducing overall model complexity while maintaining optimization for each specific prediction through shared feature extraction
Solution Approach 2:
The multi-head neural network serves multiple prediction functions simultaneously through a single model architecture. The shared layers extract common features from RGB video frames that are useful for multiple physiological parameter predictions, while each head specializes in a specific parameter, achieving both efficiency and specialized accuracy
Data Source
AI summary
A method for estimating two or more physiological signals from a subject includes steps of a) obtaining a video input in the form of a sequence of frames of image data depicting the face and optionally the chest of the subject; b) providing the video input to a multi-head neural network model trained from a set of facial video inputs from a multitude of other subjects (such video inputs optionally including the chest), wherein the model has at least two heads and is trained to predict at least two physiological signals from a video input; and c) generating with the model data representing an estimate of the two or more physiological signals of the subject. In one embodiment the physiological signals are heart rate and respiratory rate. In one embodiment the multi-head neural network model is implemented in a smartphone having a camera which is used to capture the video input.


