Visual Media Multimodal Chatbot for Accurate Image and Video Interaction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing chatbots are limited to text or voice-based interactions, restricting user interaction modalities and requiring users to describe visual media, leading to potential misinterpretations and less effective responses.

Innovation Solution

A visual media-based multimodal chatbot capable of receiving and processing various types of user input, including images and videos, and outputting text, audio, image, and video responses, as well as personalized avatars, enhancing user interaction and understanding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If chatbots are limited to text or voice-based interactions, then the system complexity is reduced, but the user interaction modality and response accuracy deteriorate

Engineering Contradiction:
Improveuser interaction modalityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The chatbot system is enhanced with multi-functionality to process diverse input modalities including text, images, and videos. The system incorporates multiple processing modules: a text processing module for textual inputs, an image processing module for visual inputs, and a video processing module for temporal visual inputs. This universal design enables the chatbot to handle various interaction modalities while maintaining a unified architecture that manages complexity through modular organization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The chatbot system is segmented into distinct functional modules that process different types of inputs independently. Each module (text processing, image processing, video processing) is specialized for its specific input type, allowing the system to manage complexity through division of labor while achieving versatile interaction capabilities. The segmentation enables parallel processing paths that converge to generate comprehensive responses.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If users describe visual media in text form, then the system complexity is reduced, but the interpretation accuracy and response effectiveness deteriorate

Engineering Contradiction:
Improveinterpretation accuracyVSAvoidprocessing capability
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system introduces intermediary processing modules that act as mediators between raw visual inputs and the chatbot's understanding. The image processing module uses computer vision techniques to extract semantic information from images, while the video processing module analyzes temporal patterns and visual content. These intermediaries transform visual data into structured representations that enhance interpretation accuracy without requiring the entire system to handle raw visual data directly.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system replaces manual text description with automated visual processing mechanisms. Instead of relying on users to describe visual media, the system employs image recognition algorithms, object detection models, and video analysis techniques to automatically interpret visual inputs. This substitution of mechanical processing for human description improves accuracy while the modular architecture manages the increased processing complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20250285352A1Visual media-based multimodal chatbot
Publication Date: 2025.09.11 YONUX LLC
  • US20250285352A1 patent drawing
  • US20250285352A1 patent drawing
  • US20250285352A1 patent drawing

AI summary

Example embodiments of the present disclosure relate to a visual media-based multimodal chatbot. According to example embodiments, a method for operating a multimodal chatbot may include receiving a user input via a chatbot interface. The user input may include at least one of: a text, an audio, a first image, and a first video. The method may further include obtaining a visual media associated with the user input. The visual media may include at least one of: a second image, a second video, and an avatar associated with a person. The method may further include outputting the visual media via the chatbot interface.