Image-to-Text LLM Integration for Coherent Visual Responses

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional AI systems are standalone and lack integration, failing to effectively convert image recognition into coherent natural language descriptions and failing to communicate between different types of AIs, leading to inefficiencies and limitations in capability.

Innovation Solution

An integrated system that combines image analysis, content filtering, and natural language processing to generate contextually relevant and safe text responses to images, using multiple AI modules for seamless communication and information exchange.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional standalone AI systems are used for image recognition and text generation separately, then each system can be optimized for its specific task, but the systems fail to integrate information effectively and produce coherent natural language descriptions

Engineering Contradiction:
Improvecoherence of natural language descriptionsVSAvoidintegration of multiple AI systems
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges image recognition AI systems with text generation AI systems into a unified integrated system. The image analysis module processes images and passes extracted information to the text generation module, which produces coherent natural language descriptions by combining visual understanding with linguistic capabilities, thereby resolving the coherence issue while accepting the necessary integration complexity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The integrated AI system performs multiple functions through a single unified architecture: it conducts image recognition, extracts visual information, generates natural language descriptions, and provides text responses to user queries about images. This multi-functionality eliminates the need for separate standalone systems while maintaining the benefits of specialized processing

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If traditional AI systems operate independently without integration, then each system remains simple and easy to implement, but they fail to communicate between different types of AIs leading to inefficiencies

Engineering Contradiction:
Improveefficiency of AI information exchangeVSAvoidcommunication between AI modules
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary communication layer that enables efficient information exchange between the image analysis module and text generation module. This intermediary mechanism allows the systems to share extracted visual information and generate coherent responses without requiring complex direct integration, thereby improving productivity while managing communication complexity through structured data exchange protocols

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If text-only AI models are used, then the systems are simple and easy to operate, but they lack the capability to process and respond to image inputs

Engineering Contradiction:
Improvecapability to process image inputsVSAvoidsystem architecture for image processing
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the AI system into distinct functional modules: an image analysis module for processing visual inputs and a text generation module for producing natural language responses. This segmentation allows the system to handle image inputs while maintaining operational simplicity through modular architecture, where each module can be independently optimized and deployed

Inventive Principle:
Principle #1Segmentation

4Adaptability or versatility

If multiple AI modules are integrated for image-to-text conversion, then the system capabilities are broadened, but the system complexity increases

Engineering Contradiction:
Improvecontextual awareness of responsesVSAvoidnumber of AI modules
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent combines multiple AI modules (image analysis, content filtering, text generation) into a unified integrated system that processes images and generates contextually aware responses. The merging is achieved through a coordinated architecture where modules share data and work together seamlessly, enabling enhanced contextual awareness while managing overall system complexity through unified control and data flow management

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12437160B2Image-to-text large language models (LLM)
Publication Date: 2025.10.07 SNAP INC
  • US12437160B2 patent drawing
  • US12437160B2 patent drawing
  • US12437160B2 patent drawing

AI summary

Described is a system for generating a textual response from a received image by determining participation in an interaction function by a first user of an interaction system, identifying an image associated with the participation, processing data associated with the image using a first machine learning model to identify one or more features within the image, and generating a prompt based on the identified one or more features. The system then identifying instructions for a second machine learning model, processing the prompt and the instructions using the second machine learning model to generate a textual response to the image, and causing display of the textual response within the interaction function to the first user.