Monitoring System Using Facial Cues to Capture Regions of Interest
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing camera-based monitoring systems for individuals who require care, such as infants or those with ambulatory and communication deficits, do not effectively detect eye gaze direction or provide information about objects of interest to the cared-for person.
Innovation Solution
A monitoring system that includes a camera system with an image capturing device, a memory storing a visual object library, facial expression recognition application, and eye gaze detection application, and a controller that captures image streams, detects eye gaze direction, and communicates notifications about the person's mood and objects of interest.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the monitoring system provides continuous video feeds to the caregiver, then the caregiver can observe the cared-for person, but the caregiver cannot scrutinize the video to understand the nuances of the state of the cared-for person when attending to other activities
Solution Approach 1:
The system extracts key information from the continuous video feed by detecting facial cues and eye gaze direction, then presents only this extracted information (notifications about mood and objects of interest) to the caregiver. This allows the caregiver to monitor effectively without needing to continuously scrutinize the full video feed.
Solution Approach 2:
The system introduces an intermediary processing layer between the camera and the caregiver. The controller analyzes the video feed for facial cues and eye gaze, then translates this into meaningful notifications. This intermediary layer filters and interprets the raw video data, providing actionable information without requiring constant video review.
2Device complexity
If the monitoring system only provides audio and video output, then the system is simple, but the received output does not provide information about objects of interest to the cared-for person
Solution Approach 1:
The system performs preliminary analysis of the video feed by continuously detecting facial cues and eye gaze direction before presenting information to the caregiver. This preliminary action identifies objects of interest and mood states in advance, allowing the system to proactively provide relevant information rather than reacting to what the caregiver might want to know.
Solution Approach 2:
The system serves itself by automatically analyzing the video feed for facial cues and eye gaze direction, then generating its own notifications about mood and objects of interest. This self-service capability allows the system to independently extract and communicate meaningful information without requiring constant human intervention or complex additional hardware.
3Reliability
If the caregiver continually assesses the state of the cared-for person by paying attention to the output, then the caregiver can understand the person's state, but the caregiver cannot attend to other activities
Solution Approach 1:
The system provides feedback to the caregiver in the form of notifications that summarize the cared-for person's mood and objects of interest. This feedback mechanism allows the caregiver to stay informed about the person's state without needing to continuously monitor, as the system automatically analyzes and communicates relevant information at appropriate moments.
Data Source
AI summary
A monitoring system incorporates, and method and computer program product provide a monitoring system that captures image stream of regions of interest based on identified facial cues of a person. A controller of the monitoring system receives a first image stream that encompasses a face of a person and a second image stream that encompasses surrounding object(s) and surface(s) viewable by the person. The controller detects a facial expression of the face. In response to determining that the facial expression is a mood associated expression, the controller determines, from the first image stream, an eye gaze direction of the person. The controller determines a region of interest (ROI) aligned with the eye gaze direction. The controller captures the second image stream and identifies an object contained within the ROI. The controller communicates a notification including the expression and the object within the ROI to an output device.


