This invention belongs to the field of distributed large
language model inference scheduling technology, and relates to a multi-model dynamic scheduling method for high-
concurrency dialogue requests. The method includes: collecting the network card hard interrupt
timestamp sequence at the streaming response
bus to establish a streaming output time interval vector; extracting the differential fluctuation envelope of the streaming output time interval vector and comparing it with a reference
transmission delay line to separate the channel congestion index; calculating the real-time routing suppression coefficient based on the
queue backlog length and the channel congestion index and reconstructing the cost-aware routing matrix; compressing the sampling window when concurrent
data traffic is overloaded, and sending a redirection
signal to shift the concurrent
data stream when the real-time routing suppression coefficient exceeds the limit. This invention eliminates the need for active probing, decouples the calculation of congestion and network
jitter in situ, effectively eliminates long-
tail response latency, and improves
throughput.