Efficiently serving machine learning model computations with high throughput and low latency

By separating pre-filling and generation workloads and optimizing hardware and service configurations, the problems of low processor utilization and high latency in existing machine learning models are solved, achieving efficient and low-latency model inference.

CN122295650APending Publication Date: 2026-06-26GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GOOGLE LLC
Filing Date
2024-10-25
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing technologies cannot effectively decouple when processing the pre-filling and generation of machine learning models, resulting in low processor utilization, insufficient throughput, and excessive latency, making it difficult to meet the needs of different applications.

Method used

By separating pre-populated and generated workloads and optimizing service configurations based on the needs of each workload, different hardware and partitioning strategies are used, such as processing tasks in an interleaved or decomposed service infrastructure, to optimize batch size and hardware resource allocation.

Benefits of technology

It improves processor utilization, reduces latency and cost, enhances overall system efficiency and throughput, and adapts to the needs of different applications, especially interactive and offline inference tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122295650A_ABST
    Figure CN122295650A_ABST
Patent Text Reader

Abstract

An example method includes: receiving an input request to process multiple input sequences using a machine learning sequence processing model to generate multiple output sequences corresponding to the multiple input sequences; generating multiple initial attention tensors for the multiple input sequences, wherein: one or more corresponding initial attention tensors are generated in parallel for the input elements of each corresponding input sequence; generating the one or more corresponding initial attention tensors in one or more batches having a first batch size using a pre-filling system including one or more pre-filling computing devices, and executing one or more layers of the machine learning sequence processing model; and using the multiple initial attention tensors, regressively generating multiple output elements for each of the multiple output sequences in one or more batches having a second batch size, wherein: the multiple output elements are generated using a generation system including one or more generation computing devices.
Need to check novelty before this filing date? Find Prior Art