The invention relates to the technical field of AI reasoning, and provides a request
failover method for a distributed AI reasoning scene, which comprises the following steps: after receiving a user reasoning request by an API (Application Program Interface) service, forwarding the request to a reasoning engine instance through a gateway, encoding an original text prompt into input tokens, executing Prefilling
processing, and entering a Decoding stage to continuously generate tokens after a result is generated. In the process, the
system detects the running state of the instance in real time, if reasoning interruption is caused by network fluctuation, hardware faults or
video memory OOM or other abnormalities, generated tokens can be obtained and spliced with original input tokens into newininput tokens, and the newininput tokens are forwarded to the healthy instance through a scheduling module. According to the new instance, Prefilling is carried out on the newputtokens by utilizing a KV cache technology, the remaining tokens are continuously pushed based on a Prefilling result, and then the newly generated tokens are decoded into a text chunk and are returned to a
client in a streaming manner until an
inference session is completed. According to the method,
user perception reasoning abnormity is avoided, retry time consumption is reduced, the service SLA is effectively improved, and continuity and high efficiency of the reasoning process are guaranteed.