A computing system and method for performing a visual language model

TW202634486APending Publication Date: 2026-08-16LITE ON TECH CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TW115100600
Authority / Receiving Office
TW · TW
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-10-07
Filing Date
2026-01-07
Publication Date
2026-08-16

Smart Images

  • Figure TWG2TA001073354_001
    Figure TWG2TA001073354_001
  • Figure TWG2TA001073354_002
    Figure TWG2TA001073354_002
  • Figure TWG2TA001073354_003
    Figure TWG2TA001073354_003
Patent Text Reader

Abstract

A computing system for performing a visual language model includes an edge device and a server is provided. The edge device includes an image sensor and a processor. The image sensor is configured to capture an original image. The processor is configured to downscale the image, divide it into patches, embed the patches into vectors to generate image tokens, and encode the tokens. The server is configured to receive and decode the tokens, extract image features using a vision transformer, provide a user interface for receiving a prompt, process the prompt using a language model to generate text features, and analyze both image and text features to produce a multi-modal output.
Need to check novelty before this filing date? Find Prior Art