Image Capturing Assistant Using Computer Vision and Voice Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image capturing systems require users to manually adjust settings and use timers, leading to awkward and unnatural photo-taking experiences, especially in group settings, as users need to pose, adjust lighting, and review images for quality before capturing.

Innovation Solution

Integration of computer vision and voice recognition to determine user readiness, with assistive commands guiding users through the image capture process, including real-time adjustments for lighting and composition, using interconnected devices like smart lightbulbs to enhance image quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users manually adjust settings and use timers for image capturing, then users have control over the capturing process, but the user experience becomes awkward and unnatural

Engineering Contradiction:
Improveease of image capturing operationVSAvoidcomplexity of manual adjustment process
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system automatically detects user readiness through computer vision analysis of facial expressions and body language, and autonomously controls timing and triggering without requiring manual user input for these functions

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical timer adjustment and physical posing with automated computer vision-based detection of user readiness states, using algorithms to analyze facial expressions and determine optimal capture moments

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If users pose and adjust lighting manually before capturing, then image quality can be optimized, but the process takes excessive time and feels unnatural

Engineering Contradiction:
Improveimage qualityVSAvoidtime for posing and adjustment
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of the scene and user positioning before capture, continuously monitoring readiness states in advance to determine the optimal moment without requiring prolonged manual adjustment

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system provides real-time feedback to users through detected readiness states, using computer vision analysis to determine when proper positioning and expression are achieved, enabling rapid optimization without extended manual adjustment

Inventive Principle:
Principle #23Feedback

3Reliability

If users review images for quality before capturing, then capture quality is ensured, but the workflow becomes cumbersome and interrupts natural interaction

Engineering Contradiction:
Improveimage quality assuranceVSAvoidease of capture workflow
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system automatically performs quality assessment through continuous computer vision analysis of detected subjects, evaluating readiness criteria such as facial expressions and positioning without requiring manual user review

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual image review with automated algorithmic analysis of video frames to detect readiness states, substituting human visual inspection with machine-based quality evaluation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Quantity of substance

If conventional timer methods are used for group photos, then all users can be captured, but the process requires multiple steps and feels posed and hurried

Engineering Contradiction:
Improvenumber of users capturedVSAvoidnaturalness of group photo process
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The system automatically detects when all or sufficient users are present and ready through continuous analysis of multiple faces and body language, autonomously determining capture timing without requiring manual coordination

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system dynamically adjusts capture timing based on real-time detection of user readiness states, allowing flexible timing that adapts to when users naturally become ready rather than forcing adherence to a fixed timer schedule

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11076091B1Image capturing assistant
Publication Date: 2021.07.27 AMAZON TECH INC
  • US11076091B1 patent drawing
  • US11076091B1 patent drawing
  • US11076091B1 patent drawing

AI summary

Features are disclosed for interacting with a user to take a photo when the user is actually ready. The features combine computer vision and voice recognition to determine when the user is ready. Additional features to interact with the user to compose the image are also described. For example, the system may play audio requesting a response (e.g., “move a bit to the left” or “say cheese”). The response may include a user utterance or physical movement detected by the system. Based on the response, the image may be taken.