Conversation Meaning Extraction From Text and Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Electronic devices struggle to infer meaningful information or context from conversations that include images, making it difficult to understand the context of image-based interactions.

Innovation Solution

An electronic apparatus with an input and output interface, and a processor that extracts and analyzes text and images from conversation data, using methods like OCR, translation, and facial expression recognition to identify the meaning of the conversation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the electronic apparatus processes only text data, then the processing is simple and fast, but the apparatus cannot accurately understand conversations that include images

Engineering Contradiction:
Improveaccuracy of conversation understandingVSAvoidcomplexity of data processing
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the image processing into multiple independent modules: OCR module for text extraction, caption generation module for image description, and emotion recognition module for facial expression analysis. Each module processes specific aspects of the image separately, allowing the system to handle complex image data through divided, manageable tasks that can be executed in parallel

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The electronic apparatus is designed with multi-functionality to handle both text data and image data through a unified processing system. The processor can selectively apply different processing methods (text-only processing, OCR-based processing, caption-based processing, or emotion-based processing) depending on the input data type, making the system versatile for various conversation scenarios

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If the electronic apparatus uses multiple processing methods for images (OCR, caption generation, emotion recognition), then the understanding accuracy improves, but the processing time increases

Engineering Contradiction:
Improveaccuracy of meaning identificationVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies partial processing by selectively executing only the necessary processing methods based on the conversation context and image characteristics. The processor determines which processing methods to apply (e.g., OCR only, or OCR plus emotion recognition) rather than always executing all possible methods, reducing unnecessary processing time while maintaining sufficient accuracy

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs preliminary actions by pre-processing images to extract key information (such as detecting text regions, identifying dominant colors, or recognizing facial expressions) before the main conversation processing. This preliminary extraction of meaningful features from images reduces the computational burden during actual conversation analysis

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If the electronic apparatus replaces image areas with text, then the processing becomes simpler, but the loss of image information increases

Engineering Contradiction:
Improvesimplicity of processingVSAvoidloss of image context
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent introduces intermediary representations (captions and extracted text) that serve as mediators between the original image and the processing system. Instead of directly processing raw image pixels, the system uses these intermediary text representations that capture the essential meaning of the image, reducing processing complexity while preserving meaningful information through the caption generation and OCR processes

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables effective interpretation of image-based conversations by replacing images with text, translating languages, recognizing emotions, and extracting URL information, thereby enhancing the understanding of conversation context.

Implementation Method 1

recognizing a third text included in the image by using an optical character recognition (OCR) method

Methodology Applied
Scientific EffectOptical character recognition (OCR):

Implementation Method 2

recognize a facial expression included in the image, identify an emotion corresponding to the facial expression

Methodology Applied
Scientific EffectFacial expression analysis:

Data Source

PatentUS12632662B2Electronic apparatus and controlling method thereof
Publication Date: 2026.05.19 SAMSUNG ELECTRONICS CO LTD
  • US12632662B2 patent drawing
  • US12632662B2 patent drawing
  • US12632662B2 patent drawing

AI summary

The disclosure refers to electronic apparatuses and controlling methods thereof. In an embodiment, an electronic apparatus includes an input interface, an output interface, and a processor that is communicatively coupled to the input interface and the output interface. The processor is configured to control the input interface to receive conversation data including one or more texts and one or more images. The processor is further configured to extract a first text and an image from the conversation data. The processor is further configured to identify a meaning of the conversation data based on at least one of the first text and the image. The processor is further configured to control the output interface to output the meaning of the conversation data.