Multimodal Image Query Processing With Lightweight Context Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing large generative models (LGMs) face inefficiencies in handling multimodal inputs, requiring excessive computing resources and time to provide accurate responses, especially when handling image-based information, leading to inaccurate and delayed query answers.

Innovation Solution

A lightweight image query system utilizing multiple lightweight context models and a lightweight large generative model (LGM) processes audio and image inputs concurrently to obtain semantic and grounding information, generating query responses efficiently and accurately within a fraction of the time taken by conventional systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional large generative models are used to process multimodal inputs, then comprehensive and accurate query responses can be generated, but the system requires excessive computing resources and time

Engineering Contradiction:
Improvequery response accuracyVSAvoidcomputing resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the processing of multimodal inputs by using separate lightweight context models for different input types (audio, image, text) rather than a single large model. Each lightweight model processes its specific modality independently, and their outputs are combined to form the final query response. This segmentation reduces the computing resource consumption of each individual model while maintaining comprehensive processing capabilities.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If conventional large generative models are used to process multimodal inputs, then comprehensive and accurate query responses can be generated, but the system takes excessive time to provide responses

Engineering Contradiction:
Improvequery response accuracyVSAvoidquery response time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the query processing task into parallel segments handled by different lightweight context models for audio, image, and text inputs. These models process their respective modalities simultaneously rather than sequentially, significantly reducing the total time required to generate a comprehensive query response while maintaining accuracy through the integration of multiple specialized outputs.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs multiple lightweight context models that each perform partial processing of the overall query task. Rather than using a single large model to perform all processing, the system uses several smaller models that each contribute their specialized analysis, achieving comprehensive coverage with reduced computational overhead and faster response times.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If multiple lightweight context models are used to process different input modes concurrently, then query response time is reduced, but the system complexity increases

Engineering Contradiction:
Improvequery response speedVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent manages system complexity by clearly segmenting the architecture into distinct lightweight context models for different modalities (audio, image, text), each with a specific and simple function. This modular segmentation makes the overall system complexity manageable, as each component remains simple while their coordinated operation achieves high productivity through parallel processing.

Inventive Principle:
Principle #1Segmentation

4Use of energy by moving object

If lightweight models are used to process multimodal inputs, then computing resource consumption is reduced, but the challenge of handling diverse input modes increases

Engineering Contradiction:
Improvecomputing resource consumptionVSAvoidmultimodal input handling capability
Core Design Contradiction:
Use of energy by moving objectVSAdaptability or versatility

Solution Approach 1:

The patent achieves universality by designing a system architecture where multiple lightweight context models, each specialized in a specific modality, work together to handle diverse multimodal inputs. The integration mechanism combines the outputs of these specialized models, enabling the system to process audio, image, and text inputs effectively while keeping each individual model lightweight and resource-efficient.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250291807A1Using multimodal input and multiple lightweight models to improve query responses
Publication Date: 2025.09.18 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20250291807A1 patent drawing
  • US20250291807A1 patent drawing
  • US20250291807A1 patent drawing

AI summary

This disclosure describes a lightweight image query system that utilizes multiple lightweight models to quickly and efficiently generate responses to multimodal input queries. For example, in response to a multimodal input query, the lightweight image query system utilizes multiple lightweight context models to first obtain different types of context information based on the different input modes. The lightweight image query system then utilizes a lightweight large generative model (LGM) to quickly generate a query response using the different types of context information. By using lightweight models, including multiple lightweight context models and a lightweight LGM, the lightweight image query system can efficiently provide query responses to multimodal input queries in about half the time it takes conventional systems to return a query response.