Node management method and device

By dynamically adjusting the node type in the inference cluster and configuring the target node to a high-load type, the problem of low throughput in the inference cluster is solved, and node load balancing and performance improvement are achieved.

WO2026153335A1PCT designated stage Publication Date: 2026-07-23HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2026-01-14
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

In existing technologies, the throughput of inference clusters is low, which seriously affects the performance of inference clusters.

Method used

By identifying target nodes from first-type and second-type nodes in the inference cluster and configuring them to be of the high-load type, the number of high-load nodes is increased, and low-load nodes are fully utilized, thereby improving the processing rate of high-load nodes and the utilization rate of low-load nodes.

Benefits of technology

Without increasing the total number of nodes, the throughput of the inference cluster was improved, performance bottlenecks and lags on heavily loaded nodes were avoided, node load was balanced, and the overall performance of the inference cluster was improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2026072407_23072026_PF_FP_ABST
    Figure CN2026072407_23072026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed is a node management method and device, relating to the technical field of computers. The method is applied to a reasoning cluster, the reasoning cluster comprises nodes of a first type and nodes of a second type, the nodes of the first type are used for executing prefill tasks, and the nodes of the second type are used for executing decode tasks. During execution of a reasoning request by the reasoning cluster, a target node can be determined from among the nodes of the first type and the nodes of the second type, the type of the target node being the one of the first type and the second type having a small load, and the type of the target node being configured as the one of the first type and the second type having a large load. In this way, it is possible not only to increase the number of nodes of the type having a large load without increasing the total number of nodes in the reasoning cluster, thereby improving the rate at which nodes of the type having a large load process reasoning tasks, but also to achieve full utilization of nodes of the type having a small load in the reasoning cluster, thereby improving the utilization rate of nodes of the type having a small load.
Need to check novelty before this filing date? Find Prior Art