基于存算一体的GPU透明自适应LLM推理与传输优化的方法及系统
By optimizing the transparent conversion and dynamic scheduling between GPU and PIM, the problem of low inference efficiency in heterogeneous architecture is solved, achieving efficient memory management and transmission, and adapting to the needs of long context inference.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UESTC (SHENZHEN) ADVANCED RES INST
- Filing Date
- 2026-05-20
- Publication Date
- 2026-07-17
AI Technical Summary
The existing GPU+PIM heterogeneous architecture suffers from problems such as architectural opacity, significant DLA overhead, and static scheduling strategies that cannot adapt to runtime dynamics in large-scale long-context deployments, leading to heterogeneous load imbalance, rigid memory management, and low inference and transmission efficiency.
By setting up an address translation intermediate unit in the computer system, a hardware-level mapping layer is constructed to establish a transparent conversion between the GPU and PIM. Combined with a hierarchical time correction closed-loop dynamic scheduler and a dual-table remapping mechanism, adaptive unloading of operators and dynamic optimization of memory layout are achieved.
It achieves transparent decoupling of the GPU-PIM architecture, improves inference and transmission efficiency, solves the problem of uneven load on heterogeneous resources, and enables uninterrupted memory layout reconfiguration, which has engineering application value.
Smart Images

Figure CN122220303B_ABST