Multimodal Meme Template Planning for Clearer Intent Specification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional AI-based meme generation techniques rely on single-modal input, such as text prompts or user-provided images, which are restrictive and challenging for clearly specifying the intent, limiting effective meme creation.

Innovation Solution

A multi-modal input approach using a large language model (LLM) and vision language model (VLM) to generate a contextual template plan, followed by recontextualization and retrieval of a final meme template using max marginal relevance (MMR) search, combined with caption text planning to create memes from both image and text inputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If single-modal input (text prompt or user-provided image) is used for meme generation, then the system complexity is reduced, but the ability to clearly specify intent is limited

Engineering Contradiction:
Improveinput specification complexityVSAvoidintent specification capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent combines multiple input modalities (text prompts and template image selections) into a unified multi-modal input system. This allows users to specify intent more clearly by providing both textual descriptions and visual template references, resolving the contradiction between system simplicity and intent specification capability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent adds a visual dimension to the traditional text-only input by incorporating template image selections. This multi-modal approach enables users to express intent across multiple dimensions (textual and visual), improving intent specification without significantly increasing system complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If multi-modal input with LLM and VLM is used for meme generation, then intent specification capability is improved, but the device complexity increases

Engineering Contradiction:
Improveintent specification capabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the meme generation system into distinct functional modules: an LLM module for processing text prompts and generating template plans, a VLM module for processing template images, and a retrieval module for selecting final templates. This segmentation manages complexity by organizing functions into independent, manageable components.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a template plan as an intermediary representation that bridges the LLM and VLM components. This intermediate structure facilitates information exchange between different modalities while managing system complexity through a standardized interface.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If conventional single-modal input is used, then the system is easier to operate, but the productivity of meme generation is limited

Engineering Contradiction:
Improveuser interface simplicityVSAvoidmeme generation effectiveness
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent implements preliminary action by providing users with pre-curated template images to select from, rather than requiring them to search for or create templates. This preliminary preparation of visual options simplifies the user interface while significantly improving meme generation effectiveness by enabling more precise intent specification.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20260064956A1Methods and systems for meme generation with multi-modal input and planning
Publication Date: 2026.03.05 TATA CONSULTANCY SERVICES LTD
  • US20260064956A1 patent drawing
  • US20260064956A1 patent drawing
  • US20260064956A1 patent drawing

AI summary

The disclosure relates generally to methods and systems for meme generation with multi-modal input and planning. Conventional AI-based techniques either rely on input text prompt or the user-provided image as an input to generate the meme. Such input specification styles result in restrictive for clearly specifying an intent with a single modality. The present disclosure explores a multi-modal input specification style where a user can provide the input through a text prompt along with a widely popular meme template image. According to the present disclosure, the meme generation task is defined as a combination of two sub-tasks. In the first sub-task, a meme image template is retrieved from a dataset of existing meme templates using a template planning strategy. In the second sub-task, the text caption is generated for the retrieved template conditioned on the multi-modal input provided by the user through the caption planning strategy.