Multimodal AI Query Updating With Environmental Object Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional language processing techniques fail to leverage contextual and environment-related information in generating responses to user queries, leading to inaccurate and erroneous responses.

Innovation Solution

Implementing a multimodal artificial intelligence-based communication system that utilizes object detection and extraction from environmental images, combined with a topic-aware sequence-to-sequence model to generate responses that incorporate both user queries and environmental context.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional language processing techniques are used to generate responses to user queries, then the system is simple and easy to operate, but the accuracy and contextual awareness of responses deteriorates

Engineering Contradiction:
Improveresponse accuracyVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the response generation process into multiple specialized modules: object detection module for extracting entities from images, environmental information extraction module for capturing contextual data, query understanding module for analyzing user intent, and response generation module for synthesizing accurate answers. This segmentation allows each module to specialize in specific tasks, improving overall response accuracy while maintaining manageable system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system merges multiple information sources and processing techniques into a unified response generation framework. It combines visual data from images, textual data from user queries, environmental context, and knowledge base information into a comprehensive processing pipeline. This merging enables the system to leverage diverse data types simultaneously, significantly enhancing response accuracy and contextual awareness.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If multimodal artificial intelligence techniques are implemented to incorporate environmental information, then the contextual awareness and precision of responses improves, but the device complexity and processing requirements worsens

Engineering Contradiction:
Improvecontextual precisionVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary processing of environmental information and image data before they are needed for response generation. Object detection, environmental entity extraction, and context analysis are conducted in advance, creating pre-processed data structures that can be quickly integrated with user queries. This preliminary action reduces real-time processing complexity while maintaining high contextual precision in final responses.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces intermediary processing layers between raw data input and final response generation. These intermediaries include entity extraction modules, context analysis components, and information fusion layers that translate complex multimodal data into structured representations. These intermediaries simplify the overall processing architecture by breaking down complex transformations into manageable stages, reducing device complexity while preserving contextual precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12602901B2Multimodal artificial intelligence-based communication systems
Publication Date: 2026.04.14 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12602901B2 patent drawing
  • US12602901B2 patent drawing
  • US12602901B2 patent drawing

AI summary

Methods, apparatus, and computer program products for multimodal artificial intelligence-based communication systems are provided herein. A computer-implemented method includes generating, using a first set of one or more artificial intelligence techniques, identifying information for one or more objects detected in image data associated with at least one user query; generating at least one updated version of the at least one user query by processing, using a second set of one or more artificial intelligence techniques, at least a portion of the at least one user query in conjunction with at least a portion of the identifying information for the one or more objects; generating at least one response to the at least one updated version of the at least one user query; and performing one or more automated actions based at least in part on the at least one response.