大语言模型推理的智能并发控制方法及系统

By using a data-driven regression model to predict the concurrency control parameters of a large language model inference service, this technology solves the problems of high computational resource consumption and difficult tuning in existing technologies. It achieves adaptive performance optimization and flexible configuration to adapt to different hardware platforms and load scenarios.

CN121858254BActive Publication Date: 2026-07-17NANJING UNIV OF POSTS & TELECOMM +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF POSTS & TELECOMM
Filing Date
2026-03-19
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing large language model inference services suffer from high computational resource consumption, complex concurrent requests, and difficulties in performance tuning. Current technologies rely on inefficient model internal structure analysis and manual parameter tuning, making them difficult to adapt to different hardware platforms and diverse inference workload scenarios.

Method used

By using a data-driven approach, single-objective or multi-objective regression models can be trained to predict the intelligent configuration of concurrent control parameters, avoiding dependence on the internal structure of the model and achieving adaptive optimization and flexible trade-offs in performance.

Benefits of technology

It improves the overall throughput performance of large language model inference services, reduces latency, improves hardware resource utilization efficiency, adapts to different model sizes and hardware platforms, and supports single-objective and multi-objective optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121858254B_ABST
    Figure CN121858254B_ABST
Patent Text Reader

Abstract

本发明公开了一种大语言模型推理的智能并发控制方法及系统,该方法针对待部署模型与推理场景,枚举可行配置组合并基于单目标或多目标回归模型,实现对吞吐量、首令牌延迟和令牌生成延迟等性能指标的预测;单目标优化选择性能最优配置,多目标优化采用帕累托优化方法获得权衡解集,从而实现并发控制参数的自动化配置推荐。本发明无需依赖模型内部结构参数和试错调参,也无需对Transformer各层的计算负载和通信开销进行显式分析,具有良好的通用性,可有效提升大语言模型推理服务的性能与资源利用效率。
Need to check novelty before this filing date? Find Prior Art