Face Alignment for Vision-Based Vital Sign Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Vision-based vital monitoring systems, particularly remote photoplethysmography (rPPG), face challenges in accuracy due to relative motion of subjects and noisy camera sensors, leading to inaccuracies in face detection and tracking, which corrupt the subtle rPPG signals.

Innovation Solution

The method improves face detection and tracking by accessing video sequences, using facial landmark detection models, and applying adaptive filtering and machine-learning techniques to extract corrected motion signals, while also accounting for positional and light variations through face alignment and translation models to enhance the accuracy of vital sign estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If contactless or remote sensors are used to measure vital signs, then convenience and non-intrusiveness are improved, but measurement accuracy deteriorates due to relative motion of subjects and noisy camera sensors

Engineering Contradiction:
Improveconvenience of vital sign monitoringVSAvoidaccuracy of vital sign measurement
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary processing pipeline between the camera and vital sign measurement. This pipeline includes face detection models, facial landmark detection, motion signal extraction, and adaptive filtering algorithms that act as mediators to remove noise and correct artifacts introduced by subject motion and camera noise, thereby maintaining high measurement accuracy while preserving the contactless convenience

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces direct mechanical contact measurement (such as pulse oximeters clipped to fingers) with optical field-based measurement using cameras. The vital signs are extracted by analyzing subtle color changes in facial skin tone caused by blood volume pulsations, substituting mechanical sensing with optical field analysis to achieve non-intrusive monitoring

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If face detection and tracking are performed to extract rPPG signals, then vital sign measurement capability is improved, but signal accuracy deteriorates due to inaccuracies in face detection and tracking

Engineering Contradiction:
Improvecapability to extract vital signs from videoVSAvoidaccuracy of rPPG signal extraction
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where the system continuously monitors the quality of face detection and tracking, and adjusts processing parameters accordingly. When motion artifacts or detection inaccuracies are detected, the system applies adaptive filtering and correction algorithms to compensate for these errors, ensuring high signal accuracy while maintaining the versatility to process various video inputs

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary face detection, landmark identification, and motion estimation before extracting the rPPG signal. These preliminary actions prepare the data by identifying the exact regions of interest and pre-correcting for motion artifacts, ensuring that the subsequent signal extraction operates on optimized data that maximizes accuracy

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If adaptive filtering and machine-learning techniques are applied to correct motion signals, then tracking accuracy is improved, but computational complexity increases

Engineering Contradiction:
Improveaccuracy of facial landmark trackingVSAvoidcomputational complexity of signal processing
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex signal processing into distinct modular stages: face detection, facial landmark detection, motion signal extraction, adaptive filtering, and vital sign calculation. Each stage operates independently on specific data, allowing the system to manage computational complexity through staged processing rather than monolithic computation, while maintaining high tracking accuracy through specialized algorithms at each stage

Inventive Principle:
Principle #1Segmentation

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach significantly enhances the reliability and accuracy of vision-based vital monitoring by reducing noise and improving the tracking of facial landmarks, resulting in more precise vital sign measurements.

Implementation Method 1

rPPG typically relies on reflections off of skin on a person's face to make rPPG measurements

Methodology Applied
Scientific EffectLight reflection: Reflection

Implementation Method 2

filtering a facial motion signal determined by the FLD to extract a corrected motion signal of the user

Methodology Applied
Scientific EffectSignal filtering: Filter (electronic)

Data Source

PatentUS20240420290A1Face Alignment and Normalization For Enhanced Vision-Based Vitals Monitoring
Publication Date: 2024.12.19 SAMSUNG ELECTRONICS CO LTD
  • US20240420290A1 patent drawing
  • US20240420290A1 patent drawing
  • US20240420290A1 patent drawing

AI summary

In one embodiment, a method includes accessing a video of a user's face. The method further includes accessing, for each image frame in the video, (1) one or more facial landmarks determined by a facial landmark detection (FLD) model and (2) a corresponding determined position in the image for each facial landmark. The method further includes determining, based on the one or more facial landmarks and corresponding positions, a motion of the user's face in the captured video; extracting, from the determined motion of the user's face, a corrected motion signal of the user's face; adjusting, based on the extracted corrected motion signal of the user's face, the positions of one or more facial landmarks in the image frames; and determining, based at least in part on the adjusted positions of the facial landmarks in the sequential images of the video, one or more vital signs of the user.