Optimization Method for Joint Deployment of Hybrid Expert Model Inference in Heterogeneous GPU Environments
By optimizing expert copy placement and mixed-precision quantization in a heterogeneous GPU environment, and combining a local-first and load-aware weighted round-robin routing mechanism, the problem of load imbalance and high communication overhead in MoE model inference is solved, achieving efficient coordination of computing, storage and communication resources, and improving inference efficiency and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV OF SCI & TECH
- Filing Date
- 2026-04-23
- Publication Date
- 2026-05-26
AI Technical Summary
In heterogeneous GPU environments, the inference efficiency of hybrid expert models is low, mainly because the problems of unbalanced expert activation frequency, device heterogeneity, and strong coupling between expert placement, quantization and runtime routing have not been systematically resolved, resulting in load imbalance, system bottlenecks and high communication overhead.
By coordinating and optimizing expert replica placement and mixed-precision quantization in the offline phase, and introducing a replica routing mechanism that prioritizes local resources and uses load-aware weighted round-robin in the online phase, the system achieves joint and coordinated utilization of computing, storage, and communication resources in a heterogeneous GPU environment, reducing end-to-end inference latency and improving system throughput.
It significantly reduces end-to-end inference latency, improves overall throughput, and maintains model accuracy, thereby enhancing the inference efficiency and stability of the MoE model in a heterogeneous GPU environment.
Smart Images

Figure CN122086631A_ABST
Abstract
Citation Information
Patent Citations
Text data reasoning method and device based on hybrid expert model
CN119443279A
Hybrid parallel and dynamic scheduling method of hybrid expert model based on 3D near-memory processing
CN120687215A
Multi-edge device collaborative reasoning method and system oriented to hybrid expert large model
CN121300997A
Heterogeneous computing power cooperative scheduling system and method for mixed precision training
CN121579206A
EDGE DEPLOYMENT OF A MIXTURE OF EXPERTS (MoE) ARCHITECTURE
US20250342370A1