Electronic Device Voice-Image Integration for Context-Aware Task Execution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Electronic devices fail to organically process the current state of the device and the service being provided when receiving user utterances, as they do not effectively integrate image analysis with voice input to perform tasks associated with objects on the screen.

Innovation Solution

An electronic device with a processor, memory, microphone, and communication circuit that analyzes images, recognizes objects, and processes voice inputs to transmit data to external servers for generating path rules, which are then executed to perform tasks related to the recognized objects, providing information to the user.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the electronic device processes user utterances in isolation without integrating image analysis, then the processing logic is simpler, but the device cannot organically process the current state or services being provided

Engineering Contradiction:
Improveintegration capabilityVSAvoidprocessing logic
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges voice input processing and image analysis into a unified processing framework. The processor integrates ASR (automatic speech recognition) results with image recognition results to generate comprehensive path rules, allowing the device to understand both what the user says and what is displayed on screen simultaneously.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The processing system is designed to handle multiple types of inputs (voice, image) and multiple types of outputs (text responses, visual responses, audio responses) through a single unified processor that generates path rules. This multi-functional approach enables the device to adapt to various interaction scenarios without requiring separate processing chains.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If the electronic device separately receives user input for selecting objects on images, then the input method is simpler, but the device cannot perform tasks associated with multiple objects efficiently

Engineering Contradiction:
Improvetask execution efficiencyVSAvoiduser input method
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent replaces the mechanical interaction of pointing and clicking with visual-mechanical interaction through image analysis. The processor automatically identifies objects in the displayed image and associates them with the user's voice input, eliminating the need for manual object selection while improving task execution efficiency.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Loss of information

If the electronic device only provides results corresponding to user utterances, then the response is more direct, but the device cannot provide context-aware services

Engineering Contradiction:
Improvecontext informationVSAvoidprocessing logic
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system implements feedback by analyzing the displayed image content and using it to inform the response generation process. The processor compares the user's voice input with the current screen state, and this feedback loop ensures that the generated path rules are contextually appropriate and relevant to what is currently being displayed.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3678132B1Electronic device and server for processing user utterances
Publication Date: 2024.07.31 SAMSUNG ELECTRONICS CO LTD
  • EP3678132B1 patent drawingFigure 1.
  • EP3678132B1 patent drawingFigure 2
  • EP3678132B1 patent drawingFigure 3

AI summary

Disclosed is an electronic device including a housing, a speaker positioned at a first portion of the housing, a microphone positioned at a second portion of the housing, a touch screen display positioned at a third portion of the housing, a communication circuit positioned inside the housing or attached to the housing, a processor positioned inside the housing and operatively connected to the speaker, the microphone, the display, and the communication circuit, and a memory positioned inside the housing and operatively connected to the processor.