基于库操作系统的大模型推理方法
By using a dedicated large model inference library based on a library operating system, the system startup and memory management of the edge devices are optimized, solving the problem of insufficient computing power of the edge devices, realizing efficient hardware and software co-optimization, and improving the efficiency and resource utilization of large model inference.
Patent Information
- Application Number
- CN202610558249.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-24
- Publication Date
- 2026-07-17
AI Technical Summary
Traditional large model inference methods lack sufficient computing power on edge devices, resulting in computational response delays, high energy consumption, and a lack of customized processing, making it difficult to meet the requirements of fast startup and low energy consumption.
Based on the library operating system, a dedicated large-model inference library is designed. By analyzing the requirements of inference scenarios, a minimized application image is built, the operating system startup process and memory management are optimized, and inference calculations are performed by combining graph optimization technology and autoregressive states to achieve hardware and software co-optimization.
It significantly reduces resource consumption, improves the efficiency of large-scale model inference on the device side, optimizes startup time and memory management, and meets the multi-dimensional optimization needs of device side applications.
Smart Images

Figure CN122414397A_ABST