基于库操作系统的大模型推理方法

By using a dedicated large model inference library based on a library operating system, the system startup and memory management of the edge devices are optimized, solving the problem of insufficient computing power of the edge devices, realizing efficient hardware and software co-optimization, and improving the efficiency and resource utilization of large model inference.

CN122414397APending Publication Date: 2026-07-17ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610558249.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-24
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Traditional large model inference methods lack sufficient computing power on edge devices, resulting in computational response delays, high energy consumption, and a lack of customized processing, making it difficult to meet the requirements of fast startup and low energy consumption.

Method used

Based on the library operating system, a dedicated large-model inference library is designed. By analyzing the requirements of inference scenarios, a minimized application image is built, the operating system startup process and memory management are optimized, and inference calculations are performed by combining graph optimization technology and autoregressive states to achieve hardware and software co-optimization.

Benefits of technology

It significantly reduces resource consumption, improves the efficiency of large-scale model inference on the device side, optimizes startup time and memory management, and meets the multi-dimensional optimization needs of device side applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122414397A_ABST
    Figure CN122414397A_ABST
Patent Text Reader

Abstract

本发明公开了一种基于库操作系统的大模型推理方法。本发明围绕大模型推理全流程优化展开:首先分析推理场景,明确系统组件,输出配置文件并建立依赖关系,生成最小化应用镜像;接着优化系统与推理环境启动流程,在推理阶段,先处理输入文本生成Token id序列,再分批次推理并优化计算图;随后进入自回归状态,将Token id序列映射为高维张量输入Transformer层,结合KV Cache完成注意力计算;最后同步推理状态机状态,捕获终止信号后将Token id序列解码为结构化文本响应。本发明支持用户按需定制大模型推理应用镜像,剔除冗余组件优化资源占用,搭配深度优化的推理算子提升效率,适配多场景端侧需求。
Need to check novelty before this filing date? Find Prior Art