Camera-Based Intention Recognition for Voice Output

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current communication methods for declining incoming calls, such as text messaging, can be time-consuming and annoying, especially for individuals who are temporarily or permanently unable to speak.

Innovation Solution

An electronic device and server system that uses a camera module to detect user intentions from image data, converts this data into voice data, and outputs it, allowing for voice communication without the need for spoken words.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If text messaging is used to decline incoming calls, then communication can be maintained without speaking, but the process becomes time-consuming and annoying

Engineering Contradiction:
Improvecommunication effectivenessVSAvoidtime required to decline call
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent replaces the manual text typing process with an automated facial expression recognition system. The camera captures facial expressions, the processor analyzes them to determine user intent, and the system automatically sends pre-defined text messages or voice outputs, eliminating the need for manual typing while maintaining communication effectiveness.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables users to communicate their intent passively through facial expressions without requiring active participation in message composition. The automated recognition and message generation systems handle the communication process independently, allowing users to simply display their intent through facial expressions.

Inventive Principle:
Principle #25Self-service

2Loss of information

If manual text typing is required for communication, then precise communication is achieved, but the ease of operation deteriorates

Engineering Contradiction:
Improvecommunication accuracyVSAvoiduser effort required
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent substitutes manual text input operations with automated facial expression analysis. The system captures facial expressions via camera, processes them to determine user intent, and automatically generates appropriate messages, eliminating the need for manual typing while preserving communication accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces facial expressions as an intermediary between the user's internal intent and the external communication output. Instead of directly typing messages, users express their intent through facial expressions, which the system translates into appropriate text or voice communications.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If voice output is implemented for intention recognition, then communication efficiency improves, but device complexity increases

Engineering Contradiction:
Improvecommunication efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent leverages the existing camera module, already present in most mobile devices, for facial expression recognition. By utilizing an existing component for a new function (intention recognition), the system achieves voice-like communication efficiency without significantly increasing device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system creates a visual copy of voice communication by analyzing facial expressions that naturally occur during speech. This allows the device to recognize user intent in real-time without requiring complex new hardware, as it processes visual information that already exists in the user's natural behavior.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS9992641B2Electronic device, server, and method for outputting voice
Publication Date: 2018.06.05 SAMSUNG ELECTRONICS CO LTD
  • US9992641B2 patent drawing
  • US9992641B2 patent drawing
  • US9992641B2 patent drawing

AI summary

According to an embodiment of the present disclosure, an electronic device may include a camera module obtaining image data of a user, a controller configured to detect at least one feature corresponding to an intention of the user from the image data obtained by the camera module, to obtain a pattern based on the at least one feature, to determine text data corresponding to the pattern, and to convert at least a portion of the text data into voice data, and an output module outputting the voice data. Other various embodiments of the pattern recognition are also provided.