Virtual batches in large language model inferences

TW202634500APending Publication Date: 2026-08-16GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TW114150109
Authority / Receiving Office
TW · TW
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-12-31
Filing Date
2025-12-19
Publication Date
2026-08-16

Smart Images

  • Figure TWG2TA001073273_001
    Figure TWG2TA001073273_001
  • Figure TWG2TA001073273_002
    Figure TWG2TA001073273_002
  • Figure TWG2TA001073273_003
    Figure TWG2TA001073273_003
Patent Text Reader

Abstract

This document describes systems and techniques directed at virtual batches in large language model (LLM) inferences. An LLM, at least partially deployed on an electronic device, generates a dependency map for a plurality of inference tokens. The inference tokens can be based on a same input, different inputs, or a mixture of both. The dependency map indicates sequential or otherwise logical dependence of each inference token. The LLM can further generate a plurality of virtual batches based on the dependency map. The plurality of virtual batches includes masked portions indicating positions in one or more of the plurality of virtual batches that do not have an active token reference.
Need to check novelty before this filing date? Find Prior Art