Multimodal LLM Display Information Acquisition System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods fail to effectively leverage display information as a primary input for multimodal Large Language Models (LLMs), limiting the versatility and efficiency of AI applications.

Innovation Solution

A method and apparatus that utilize a multimodal LLM as a central information processing hub, integrating display information, audio, and text inputs, with specified output locations and formats, enabling human-like operations and efficient data transmission.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If display information is extracted and transmitted to multimodal LLM in real-time, then the versatility and application areas of AI are enriched, but the system complexity and computational resources increase

Engineering Contradiction:
Improveapplication areas of AIVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a display information acquisition module as an intermediary that captures display content from the frame buffer and pre-processes it before transmitting to the multimodal LLM. This intermediary layer simplifies the overall system architecture by centralizing the display information extraction logic and providing a standardized interface between the display system and the AI model.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system is divided into distinct functional modules: a display information acquisition module that extracts visual content, a multimodal LLM processing module that analyzes the content, and a control instruction translator module that converts AI output into actionable commands. This segmentation allows each module to be optimized independently while maintaining overall system versatility.

Inventive Principle:
Principle #1Segmentation

2Productivity

If display information is transmitted at high sampling frequency, then the information processing capability is enhanced, but the data transmission volume and processing load increase

Engineering Contradiction:
Improveinformation processing capabilityVSAvoiddata transmission volume
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent implements selective sampling where the display information acquisition module captures frames at variable frequencies based on the actual content changes. When the display content is static or changes minimally, the sampling frequency is reduced or paused, transmitting only essential information to the multimodal LLM, thus optimizing the balance between processing capability and data volume.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If multiple information modalities are integrated into a unified model, then the efficiency of processing different tasks is improved, but the model complexity and computational requirements increase

Engineering Contradiction:
Improveefficiency of processing different tasksVSAvoidmodel complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent employs a multimodal large language model that serves as a universal processor for multiple information types including display images, audio, and text. This single unified model handles diverse tasks such as screen content analysis, voice command processing, and text generation, replacing what would otherwise require multiple specialized models, thus improving efficiency while managing complexity through a consolidated architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP4654047A1Method and apparatus applying multimodal artificial intelligence
Publication Date: 2025.11.26 BEIJING GOOSE FACTORY TECH CO LTD
  • EP4654047A1 patent drawingFigure 1~2
  • EP4654047A1 patent drawingFigure 3
  • EP4654047A1 patent drawing

AI summary

Multimodal Large Language Models (LLMs) have emerged as a significant breakthrough in enhancing societal productivity. For instance, multimodal LLMs such as OpenAI's GPT-4.0 are expanding their multifaceted information interaction modalities, encompassing forms such as text, speech, images, and video. These powerful LLMs are anticipated to continuously evolve and improve. Rapidly leveraging the capabilities of these LLMs to further improve societal productivity has become a crucial research direction across various fields. A primary objective of LLMs is to process diverse tasks using a unified model or framework. Consequently, in application domains, the development of more versatile methods to harness the features and capabilities of multimodal LLMs is of substantial value. For humans, display devices represent the primary means of obtaining information from electronic devices, enabling various systems to increasingly adapt to human habits. If artificial intelligence can acquire a comparable volume of information from display devices as humans do, utilizing LLMs, it will significantly expand Al application domains and could further alleviate human workload.