A low-latency scheduling method, system and medium for real-time interactive AI services
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU AOGONG INFORMATION TECH CO LTD
- Filing Date
- 2026-04-28
- Publication Date
- 2026-05-29
AI Technical Summary
Existing AI service systems cannot effectively distinguish the latency differences between different tenant levels and task types when handling mixed loads of high concurrency, multi-tenancy, and multiple task types. This leads to resource waste or insufficient allocation, an inability to balance latency and throughput, and difficulty in meeting the high real-time and high interactivity requirements of real-time interactive AI services.
By dividing the GPU memory space into a first cache area and a second cache area, the priority coefficient of task requests is dynamically analyzed, the cache area for high-priority tasks is pre-allocated, and the memory capacity is predicted and allocated based on historical data and task request rate, thereby optimizing the number of parallel processing tasks and realizing dynamic adjustment and optimization of resources.
It effectively reduced the response latency of high-priority tasks, improved the service response speed and efficiency of real-time interactive tasks, met high performance requirements, achieved balanced processing of high-priority and low-priority tasks, and improved the robustness and throughput efficiency of the system.
Smart Images

Figure CN122111625A_ABST
Abstract
Citation Information
Patent Citations
Industrial personal computer and multi-graphics card collaborative parallel operation acceleration system
CN121029352A
Dynamic batch processing and delay optimization method for deep learning model reasoning service
CN121119152A
Multi-task parallel processing method for user problems under AI platform
CN121210104A
Display card multi-task cooperative processing system based on dynamic video memory allocation
CN121277631A
Multi-tenant visual large model reasoning resource dynamic allocation and isolation method
CN121722549A