A low-latency scheduling method, system and medium for real-time interactive AI services

CN122111625AActive Publication Date: 2026-05-29JIANGSU AOGONG INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU AOGONG INFORMATION TECH CO LTD
Filing Date
2026-04-28
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing AI service systems cannot effectively distinguish the latency differences between different tenant levels and task types when handling mixed loads of high concurrency, multi-tenancy, and multiple task types. This leads to resource waste or insufficient allocation, an inability to balance latency and throughput, and difficulty in meeting the high real-time and high interactivity requirements of real-time interactive AI services.

Method used

By dividing the GPU memory space into a first cache area and a second cache area, the priority coefficient of task requests is dynamically analyzed, the cache area for high-priority tasks is pre-allocated, and the memory capacity is predicted and allocated based on historical data and task request rate, thereby optimizing the number of parallel processing tasks and realizing dynamic adjustment and optimization of resources.

Benefits of technology

It effectively reduced the response latency of high-priority tasks, improved the service response speed and efficiency of real-time interactive tasks, met high performance requirements, achieved balanced processing of high-priority and low-priority tasks, and improved the robustness and throughput efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122111625A_ABST
    Figure CN122111625A_ABST
Patent Text Reader

Abstract

The application discloses a low-delay scheduling method and system for real-time interactive AI services and a medium, the method comprising: determining an initial priority coefficient according to each tenant level and task type; dynamically analyzing the priority coefficient of each task request according to the waiting time; predicting and dynamically allocating the video memory space capacity of a first cache area; obtaining the real-time residual amount of the video memory capacity of the predicted and allocated cache area, dynamically analyzing the upper limit value of the number of task requests processed in parallel, and based on the number of currently waiting task requests and the preset regulation mode, analyzing the number of currently optimal task requests processed in parallel to balance resource utilization and delay waiting time. The application effectively solves the delay problem of existing AI services when processing high real-time and high-interactive tasks, improves the service response speed and efficiency of real-time interactive tasks, and meets the strict performance index standards of various high-demand AI services.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Industrial personal computer and multi-graphics card collaborative parallel operation acceleration system

    CN121029352A

  • Dynamic batch processing and delay optimization method for deep learning model reasoning service

    CN121119152A

  • Multi-task parallel processing method for user problems under AI platform

    CN121210104A

  • Display card multi-task cooperative processing system based on dynamic video memory allocation

    CN121277631A

  • Multi-tenant visual large model reasoning resource dynamic allocation and isolation method

    CN121722549A