大语言模型推理的智能并发控制方法及系统
By using a data-driven regression model to predict the concurrency control parameters of a large language model inference service, this technology solves the problems of high computational resource consumption and difficult tuning in existing technologies. It achieves adaptive performance optimization and flexible configuration to adapt to different hardware platforms and load scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-03-19
- Publication Date
- 2026-07-17
AI Technical Summary
Existing large language model inference services suffer from high computational resource consumption, complex concurrent requests, and difficulties in performance tuning. Current technologies rely on inefficient model internal structure analysis and manual parameter tuning, making them difficult to adapt to different hardware platforms and diverse inference workload scenarios.
By using a data-driven approach, single-objective or multi-objective regression models can be trained to predict the intelligent configuration of concurrent control parameters, avoiding dependence on the internal structure of the model and achieving adaptive optimization and flexible trade-offs in performance.
It improves the overall throughput performance of large language model inference services, reduces latency, improves hardware resource utilization efficiency, adapts to different model sizes and hardware platforms, and supports single-objective and multi-objective optimization.
Smart Images

Figure CN121858254B_ABST