Automated Image Capture Using Dual Neural Networks for Emotion Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image capture systems struggle to automatically capture desired emotions beyond smiles, such as happiness, sadness, or excitement, as they rely solely on facial expressions, missing candid moments and requiring multiple photographers and equipment, which is inefficient and often ineffective.
Innovation Solution
An automated image capture system utilizing a neural network to classify images based on both facial and body expressions, allowing users to select desired emotions and levels, with a two-layered model approach for improved computation efficiency, analyzing facial expressions first and body expressions when facial scores are below a threshold.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional smile-detection systems are used, then smiling images can be captured, but other emotions (happiness, sadness, excitement) cannot be accurately detected
Solution Approach 1:
The emotion detection system is segmented into two specialized neural networks: a facial expression network and a body expression network. Each network is optimized for detecting specific emotion cues from different body regions, allowing the system to accurately classify multiple emotions beyond just smiles by analyzing facial and body expressions separately and combining their results.
2Productivity
If multiple photographers and equipment are deployed to capture candid moments, then more emotions may be captured, but system complexity and cost increase
Solution Approach 1:
The patent replaces the mechanical system of multiple photographers and cameras with an automated computational system using neural networks. The image capture device with integrated emotion detection algorithms can automatically analyze and capture multiple emotions from a single device, eliminating the need for multiple photographers and reducing system complexity while maintaining or improving productivity.
3Device complexity
If a single neural network analyzes both facial and body expressions, then computation is simplified, but detection accuracy decreases
Solution Approach 1:
The detection system is divided into two specialized neural networks: one for facial expressions and one for body expressions. This segmentation allows each network to specialize in detecting specific emotion cues from its designated body region, improving overall detection accuracy while maintaining manageable computational complexity through modular architecture.
Solution Approach 2:
The patent merges the results from the facial expression network and body expression network to produce the final emotion classification. By combining the specialized detection capabilities of both networks, the system achieves higher accuracy than a single general-purpose network could provide, while the modular combining approach keeps the overall system complexity manageable.
4Measurement precision
If comprehensive emotion analysis is performed on all images, then emotion capture accuracy improves, but processing time increases
Solution Approach 1:
The system performs preliminary action by first analyzing facial expressions, which are more reliable indicators of emotion. Based on the facial analysis results and confidence scores, the system selectively applies the more computationally intensive body expression analysis only when needed, reducing overall processing time while maintaining high accuracy for clear facial expressions.
Solution Approach 2:
The patent applies partial action by performing comprehensive emotion analysis only on images where it is most beneficial. The system uses facial expression analysis as a primary filter and applies full body expression analysis selectively, rather than performing exhaustive analysis on all images, thereby optimizing the balance between accuracy and processing time.
Data Source
AI summary
Methods and systems are provided for performing automated capture of images based on emotion detection. In embodiments, a selection of an emotion class from among a set of emotion classes presented, via a graphical user interface, is received. The emotion class can indicate an emotion exhibited by a subject desired to be captured in an image. A set of images corresponding with a video is analyzed to identify at least one image in which a subject exhibits the emotion associated with the selected emotion class. The set of images can be analyzed using at least one neural network that classifies images in association with emotion exhibited in the images. Thereafter, the image can be presented in association with the selected emotion class.


