This application provides a method, apparatus, and electronic device for matching batch requests with computing nodes. Relating to the field of
computer technology, it proposes a method based on the fixed
concurrency capability of computing nodes in an AI
chip system. Requests belonging to the same expert are divided into multiple request blocks, and a scheduling
queue is determined for these request blocks. The larger the request volume of a block, the higher its priority in the scheduling
queue. When an unmatchable request block occurs, under the constraint of a pre-stored expert copy corresponding to the computing node, an augmenting
path search is used to match the already matched original request block from its original computing node to a redundant computing node that matches the request block. The unmatchable request block is then scheduled to be matched back to the original computing node, which has a pre-stored expert copy for
processing the unmatchable request block. This improves the effective utilization of hardware resources, reduces the dropout rate, and increases
throughput.