Optimization Method for Joint Deployment of Hybrid Expert Model Inference in Heterogeneous GPU Environments

By optimizing expert copy placement and mixed-precision quantization in a heterogeneous GPU environment, and combining a local-first and load-aware weighted round-robin routing mechanism, the problem of load imbalance and high communication overhead in MoE model inference is solved, achieving efficient coordination of computing, storage and communication resources, and improving inference efficiency and stability.

CN122086631APending Publication Date: 2026-05-26NANJING UNIV OF SCI & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV OF SCI & TECH
Filing Date
2026-04-23
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In heterogeneous GPU environments, the inference efficiency of hybrid expert models is low, mainly because the problems of unbalanced expert activation frequency, device heterogeneity, and strong coupling between expert placement, quantization and runtime routing have not been systematically resolved, resulting in load imbalance, system bottlenecks and high communication overhead.

Method used

By coordinating and optimizing expert replica placement and mixed-precision quantization in the offline phase, and introducing a replica routing mechanism that prioritizes local resources and uses load-aware weighted round-robin in the online phase, the system achieves joint and coordinated utilization of computing, storage, and communication resources in a heterogeneous GPU environment, reducing end-to-end inference latency and improving system throughput.

Benefits of technology

It significantly reduces end-to-end inference latency, improves overall throughput, and maintains model accuracy, thereby enhancing the inference efficiency and stability of the MoE model in a heterogeneous GPU environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122086631A_ABST
    Figure CN122086631A_ABST
Patent Text Reader

Abstract

This invention discloses a joint deployment optimization method for hybrid expert model inference in heterogeneous GPU environments. The method first uses offline trajectory statistics to determine expert activation frequencies and collects resource profiles of heterogeneous devices. Then, it constructs a joint optimization model with the objective of minimizing the weighted sum of inference latency surrogate terms and quantization error penalty terms. Under memory constraints, a pruned greedy search algorithm is used to solve for the placement location, number, and quantization accuracy of each expert's replicas. During the online inference phase, a routing mechanism combining local priority and load-aware weighted round-robin is used for dynamic token distribution. This invention can simultaneously address inference latency, memory usage, and accuracy preservation in heterogeneous GPU environments, improving the processing efficiency of popular experts, reducing cross-device communication overhead and slow device trailing effects, and is suitable for distributed large-scale model inference deployment scenarios.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Text data reasoning method and device based on hybrid expert model

    CN119443279A

  • Hybrid parallel and dynamic scheduling method of hybrid expert model based on 3D near-memory processing

    CN120687215A

  • Multi-edge device collaborative reasoning method and system oriented to hybrid expert large model

    CN121300997A

  • Heterogeneous computing power cooperative scheduling system and method for mixed precision training

    CN121579206A

  • EDGE DEPLOYMENT OF A MIXTURE OF EXPERTS (MoE) ARCHITECTURE

    US20250342370A1