Virtual batches in large language model inferences
TW202634500APending Publication Date: 2026-08-16GOOGLE LLC
View PDF 0 Cites 0 Cited by
Patent Information
- Application Number
- TW114150109
- Authority / Receiving Office
- TW · TW
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-12-31
- Filing Date
- 2025-12-19
- Publication Date
- 2026-08-16
Smart Images

Figure TWG2TA001073273_001 
Figure TWG2TA001073273_002 
Figure TWG2TA001073273_003
Abstract
This document describes systems and techniques directed at virtual batches in large language model (LLM) inferences. An LLM, at least partially deployed on an electronic device, generates a dependency map for a plurality of inference tokens. The inference tokens can be based on a same input, different inputs, or a mixture of both. The dependency map indicates sequential or otherwise logical dependence of each inference token. The LLM can further generate a plurality of virtual batches based on the dependency map. The plurality of virtual batches includes masked portions indicating positions in one or more of the plurality of virtual batches that do not have an active token reference.
Need to check novelty before this filing date? Find Prior Art