Image-to-Text LLM Pipeline for Safe Contextual Responses

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional AI systems are standalone and lack integration between image and text-based capabilities, leading to inefficiencies and limitations in generating coherent, contextually relevant, and safe text responses to images.

Innovation Solution

An integrated AI system that combines image analysis, content filtering, and natural language processing to generate contextually relevant and safe text responses to images, incorporating modules for image-to-text conversion, object detection, and content filtering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional standalone AI systems are used for image analysis and text generation separately, then each system can be simple and specialized, but the overall system lacks integration and produces inconsistent or unsafe responses

Engineering Contradiction:
Improveresponse safety and coherenceVSAvoidsystem integration complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges image analysis capabilities and text generation capabilities into a single integrated AI system. The system combines an image analysis module that extracts visual information with a text generation module that creates coherent responses, ensuring consistency and safety across the entire workflow rather than having separate standalone systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The integrated AI system performs multiple functions including image analysis, content filtering, and text generation within a single unified platform. This multi-functional approach allows the system to handle diverse tasks (describing images, detecting objects, filtering content, generating responses) while maintaining consistent behavior and safety standards throughout.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If an integrated AI system combines image analysis, content filtering, and text generation, then response coherence and safety improve, but computing resources and processing time increase

Engineering Contradiction:
Improvecontextual relevance of responsesVSAvoidcomputing resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system segments the image processing pipeline into distinct functional modules: image analysis module for extracting visual information, content filtering module for safety checks, and text generation module for creating responses. This segmentation allows each module to process information independently and efficiently, reducing overall computational burden while maintaining integration benefits.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The content filtering module performs preliminary safety checks on image data before it reaches the text generation module. By filtering and validating content in advance, the system prevents unnecessary processing of unsafe or inappropriate images, thereby reducing overall computing resource consumption while maintaining high reliability.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If traditional AI systems process images and generate text separately, then each processing step can be optimized independently, but the overall process is inefficient and time-consuming

Engineering Contradiction:
Improveresponse generation efficiencyVSAvoidtotal processing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The integrated AI system maintains continuous processing flow where image analysis, content filtering, and text generation occur in an uninterrupted sequence. The system continuously extracts visual information, filters content, and generates responses without idle time between steps, ensuring that the entire process operates at peak efficiency and minimizes total processing time.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS20260017468A1Image-to-text large language models (LLM)
Publication Date: 2026.01.15 SNAP INC
  • US20260017468A1 patent drawing
  • US20260017468A1 patent drawing
  • US20260017468A1 patent drawing

AI summary

Described is a system for generating a textual response from a received image by determining participation in an interaction function by a first user of an interaction system, identifying an image associated with the participation, processing data associated with the image using a first machine learning model to identify one or more features within the image, and generating a prompt based on the identified one or more features. The system then identifying instructions for a second machine learning model, processing the prompt and the instructions using the second machine learning model to generate a textual response to the image, and causing display of the textual response within the interaction function to the first user.