基于存算一体的GPU透明自适应LLM推理与传输优化的方法及系统

By optimizing the transparent conversion and dynamic scheduling between GPU and PIM, the problem of low inference efficiency in heterogeneous architecture is solved, achieving efficient memory management and transmission, and adapting to the needs of long context inference.

CN122220303BActive Publication Date: 2026-07-17UESTC (SHENZHEN) ADVANCED RES INST

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
UESTC (SHENZHEN) ADVANCED RES INST
Filing Date
2026-05-20
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

The existing GPU+PIM heterogeneous architecture suffers from problems such as architectural opacity, significant DLA overhead, and static scheduling strategies that cannot adapt to runtime dynamics in large-scale long-context deployments, leading to heterogeneous load imbalance, rigid memory management, and low inference and transmission efficiency.

Method used

By setting up an address translation intermediate unit in the computer system, a hardware-level mapping layer is constructed to establish a transparent conversion between the GPU and PIM. Combined with a hierarchical time correction closed-loop dynamic scheduler and a dual-table remapping mechanism, adaptive unloading of operators and dynamic optimization of memory layout are achieved.

Benefits of technology

It achieves transparent decoupling of the GPU-PIM architecture, improves inference and transmission efficiency, solves the problem of uneven load on heterogeneous resources, and enables uninterrupted memory layout reconfiguration, which has engineering application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122220303B_ABST
    Figure CN122220303B_ABST
Patent Text Reader

Abstract

本发明申请实施例公开了一种基于存算一体的GPU透明自适应LLM推理与传输优化的方法及系统,方法包括:步骤1:在计算机系统中的存储模块与地址转换单元之间设置地址转换中间单元,完成图形处理单元的逻辑请求到存内计算单元的物理执行的透明转换;步骤2:构建层级时间校正的闭环动态调度器,对静态卸载模型进行实时校正;步骤3:在存储模块中构建基于活跃映射表和影子映射表的动态地址的双表重映射机制;步骤4:执行大语言模型的预填充阶段与解码阶段的两阶段推理,实现图形处理单元与存内计算单元异构架构的透明协同与自适应优化。本方法及系统能够提高大语言模型推理效率及传输效率。
Need to check novelty before this filing date? Find Prior Art