An edge chip-based inference large model system running method
By employing Ascend-Native-Docker images and fine-grained memory pool management on edge chips, the problem of insufficient memory for large models in edge computing is solved, enabling stable and efficient operation of large inference models on edge chips, supporting continuous dialogue and low-latency interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGNAN ELECTROMECHANICAL DESIGN INST
- Filing Date
- 2026-04-23
- Publication Date
- 2026-07-17
AI Technical Summary
In edge computing scenarios, the memory capacity and bandwidth of existing AI inference chips such as Ascend 310B4 cannot meet the needs of large language models, resulting in memory overflow and computational stagnation, making it impossible to achieve continuous dialogue and efficient interaction.
Containerized deployment is achieved using the Ascend-Native-Docker image. Through precise planning of static and dynamic memory pools, combined with warm-up testing and concurrency control strategies, resource allocation and management are optimized to avoid memory contention and ensure stable system operation.
It enables stable and efficient operation of large models on edge chips, supports continuous dialogue and low-latency interaction, reduces hardware costs and improves system reliability.
Smart Images

Figure CN122412142A_ABST