Conversation Meaning Extraction From Text and Images
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Electronic devices struggle to infer meaningful information or context from conversations that include images, making it difficult to understand the context of image-based interactions.
Innovation Solution
An electronic apparatus with an input and output interface, and a processor that extracts and analyzes text and images from conversation data, using methods like OCR, translation, and facial expression recognition to identify the meaning of the conversation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the electronic apparatus processes only text data, then the processing is simple and fast, but the apparatus cannot accurately understand conversations that include images
Solution Approach 1:
The patent segments the image processing into multiple independent modules: OCR module for text extraction, caption generation module for image description, and emotion recognition module for facial expression analysis. Each module processes specific aspects of the image separately, allowing the system to handle complex image data through divided, manageable tasks that can be executed in parallel
Solution Approach 2:
The electronic apparatus is designed with multi-functionality to handle both text data and image data through a unified processing system. The processor can selectively apply different processing methods (text-only processing, OCR-based processing, caption-based processing, or emotion-based processing) depending on the input data type, making the system versatile for various conversation scenarios
2Measurement precision
If the electronic apparatus uses multiple processing methods for images (OCR, caption generation, emotion recognition), then the understanding accuracy improves, but the processing time increases
Solution Approach 1:
The system applies partial processing by selectively executing only the necessary processing methods based on the conversation context and image characteristics. The processor determines which processing methods to apply (e.g., OCR only, or OCR plus emotion recognition) rather than always executing all possible methods, reducing unnecessary processing time while maintaining sufficient accuracy
Solution Approach 2:
The system performs preliminary actions by pre-processing images to extract key information (such as detecting text regions, identifying dominant colors, or recognizing facial expressions) before the main conversation processing. This preliminary extraction of meaningful features from images reduces the computational burden during actual conversation analysis
3Device complexity
If the electronic apparatus replaces image areas with text, then the processing becomes simpler, but the loss of image information increases
Solution Approach 1:
The patent introduces intermediary representations (captions and extracted text) that serve as mediators between the original image and the processing system. Instead of directly processing raw image pixels, the system uses these intermediary text representations that capture the essential meaning of the image, reducing processing complexity while preserving meaningful information through the caption generation and OCR processes
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables effective interpretation of image-based conversations by replacing images with text, translating languages, recognizing emotions, and extracting URL information, thereby enhancing the understanding of conversation context.
Implementation Method 1
recognizing a third text included in the image by using an optical character recognition (OCR) method
Implementation Method 2
recognize a facial expression included in the image, identify an emotion corresponding to the facial expression
Data Source
AI summary
The disclosure refers to electronic apparatuses and controlling methods thereof. In an embodiment, an electronic apparatus includes an input interface, an output interface, and a processor that is communicatively coupled to the input interface and the output interface. The processor is configured to control the input interface to receive conversation data including one or more texts and one or more images. The processor is further configured to extract a first text and an image from the conversation data. The processor is further configured to identify a meaning of the conversation data based on at least one of the first text and the image. The processor is further configured to control the output interface to output the meaning of the conversation data.


