Efficiently serving machine learning model computations with high throughput and low latency
By separating pre-filling and generation workloads and optimizing hardware and service configurations, the problems of low processor utilization and high latency in existing machine learning models are solved, achieving efficient and low-latency model inference.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2024-10-25
- Publication Date
- 2026-06-26
AI Technical Summary
Existing technologies cannot effectively decouple when processing the pre-filling and generation of machine learning models, resulting in low processor utilization, insufficient throughput, and excessive latency, making it difficult to meet the needs of different applications.
By separating pre-populated and generated workloads and optimizing service configurations based on the needs of each workload, different hardware and partitioning strategies are used, such as processing tasks in an interleaved or decomposed service infrastructure, to optimize batch size and hardware resource allocation.
It improves processor utilization, reduces latency and cost, enhances overall system efficiency and throughput, and adapts to the needs of different applications, especially interactive and offline inference tasks.
Smart Images

Figure CN122295650A_ABST