Cluster system and scheduling method for cluster system

WO2026201054A1PCT designated stage Publication Date: 2026-10-01HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/086212
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-26
Publication Date
2026-10-01

Smart Images

  • Figure CN2026086212_01102026_PF_FP_ABST
    Figure CN2026086212_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A cluster system and a scheduling method for a cluster system. The method comprises: a cluster management device (100) sends a first message to a first computer device (500) (1510), wherein the first message is used for indicating replacement of a first AI chip in the first computer device; the first computer device receives the first message, deletes a mapping relationship between a first port and the first AI chip from a first AI model training list, and establishes a mapping relationship between a second port and a second AI chip, wherein the first port includes a port in the first computer device that communicates with the first AI chip, the second port includes a port in the first computer device that can communicate with the second AI chip, and the first AI model training list is stored in the first computer device (1520); and the cluster management device sends model training data to AI chips in the first AI model training list (1530). The scheduling method can reduce resource waste caused during scheduling of the cluster system.
Need to check novelty before this filing date? Find Prior Art

Description

Cluster systems and their scheduling methods

[0001] This application claims priority to Chinese Patent Application No. 202510397343.4, filed on March 28, 2025, entitled "Cluster System and Scheduling Method for Cluster System", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer technology, specifically to a cluster system and a scheduling method for the cluster system. Background Technology

[0003] Artificial intelligence (AI) is the main direction of current technological development. AI network models are the foundation of artificial intelligence. Through training with AI cluster systems, more accurate and powerful AI network models can be trained quickly.

[0004] As AI models become increasingly complex, larger, and have more parameters, AI training often requires AI cluster systems with tens of thousands of GPUs. Consequently, the probability of AI cluster system failures is also rising. Real-world analysis shows that over 50% of failures are due to cluster infrastructure malfunctions, such as one to two graphics processing units (GPUs) failing weekly.

[0005] In existing technologies, when an AI chip in a computer device in an AI cluster system fails, the entire computer device will have to be replaced, and other working AI chips will also be unusable, resulting in a waste of equipment resources. Summary of the Invention

[0006] This application provides a cluster system and a scheduling method for the cluster system, which can reduce the waste of equipment resources caused during the scheduling process of the cluster system.

[0007] In a first aspect, a scheduling method for a cluster system is provided. The cluster system includes a cluster management device and at least one computer device, each computer device including at least one AI chip. The method includes: the cluster management device sending a first message to a first computer device, the first message indicating the replacement of a first AI chip in the first computer device, the first computer device being one of the at least one computer device; the first computer device receiving the first message, deleting a mapping relationship between a first port and the first AI chip from a first AI model training list, and establishing a mapping relationship between a second port and a second AI chip, the first port including ports in the first computer device that communicate with the first AI chip, the second port including ports in the first computer device that can communicate with the second AI chip, the first AI model training list being stored in the first computer device; and the cluster management device sending model training data to the AI ​​chips in the first AI model training list.

[0008] This application provides a scheduling method for a cluster system. When an AI chip in a computer device malfunctions, only the malfunctioning AI chip can be replaced, instead of replacing the entire computer device. This improves the reliability of the cluster system, reduces the waste of equipment resources during the cluster system scheduling process, saves the deployment cost of AI chips, and enhances product competitiveness.

[0009] In conjunction with the first aspect, in some implementations of the first aspect, the second AI chip belongs to a second computer device, which is one of the at least one computer device. The method further includes: the first computer device establishing a mapping relationship between the second port and the first AI chip in the first AI model training list; the cluster management device sending a second message to the second computer device, the second message instructing the second computer device to select a backup AI chip to replace the first AI chip, the second AI chip belonging to the backup AI chip; the second computer device receiving the second message, deleting the mapping relationship between the third port and the second AI chip in the second AI model training list, and establishing a mapping relationship between the fourth port and the second AI chip, the third port including a port in the second computer device that communicates with the second AI chip, the fourth port including a port in the second computer device that can communicate with the first computer device, and the second AI model training list being stored in the second computer device.

[0010] The second computer device is a different computer device from the first computer device in the cluster system.

[0011] The first computer device receives the first message, deletes the mapping relationship between the first port and the first AI chip from the first AI model training list, and establishes a mapping relationship between the second port and the second AI chip. The second port includes the port in the first computer device that can communicate with the second AI chip, that is, the second port can communicate with the second computer device.

[0012] In some possible implementations, the second message can be used to instruct a second computer device to select an idle backup AI chip to replace the first AI chip. After receiving the second message, the second computer device selects the second AI chip to replace the first AI chip. In other possible implementations, the cluster management device can also specify the use of the second AI chip to replace the first AI chip via the second message.

[0013] Each computer device in the cluster system may include a host, a bus switching chip, and an AI chip. The second AI model training list can be stored in the bus switching chip of the second computer device.

[0014] The second computer device may include at least one third port, and the second computer device may delete any one or more third ports and the mapping relationship between the second AI chip from the second AI model training list.

[0015] The second computer device may include at least one fourth port, and the second computer device may establish a mapping relationship between any one or more fourth ports and the second AI chip in the second AI model training list.

[0016] In some possible implementation scenarios, the second computer device can be a backup computer device, that is, at least some of the AI ​​chips in the second computer device are backup AI chips.

[0017] This application provides a scheduling method for a cluster system. The backup AI chip can be located in a different computer device than the faulty AI chip. When the first AI chip in the first computer device fails, the second AI chip in the second computer device can replace the faulty first AI chip, without having to replace the entire computer device. This improves the reliability of the cluster system and reduces the waste of equipment resources caused during the cluster system scheduling process. At the same time, having the backup AI chip and the faulty AI chip in different computer devices facilitates the management of the cluster system.

[0018] In conjunction with the first aspect, in some implementations of the first aspect, before the cluster management device sends the first message to the first computer device, the method further includes: the cluster management device receiving a third message, the third message being used to indicate that the first AI chip has malfunctioned or that the first AI chip needs to be upgraded.

[0019] In some possible implementation scenarios, if the first AI chip in the first computer device malfunctions, the first computer device can send a third message to the cluster management device to indicate that the first AI chip has malfunctioned. Upon receiving the third message, the cluster management device sends a first message to the first computer device, instructing it to replace the first AI chip in the first computer device.

[0020] In some other possible implementation scenarios, the cluster management device receives a third message indicating that the first AI chip needs to be upgraded, and the cluster management device sends a first message to the first computer device to indicate that the first AI chip in the first computer device needs to be replaced.

[0021] This application provides a scheduling method for a cluster system, applicable to scenarios such as AI chip failure or AI chip upgrade. It allows for the replacement of only the faulty or upgradeable AI chip without replacing the entire computer equipment, thereby improving the reliability of the cluster system, reducing the waste of equipment resources during cluster system scheduling, saving the deployment cost of AI chips, and enhancing product competitiveness.

[0022] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the first computer device assigning a device identifier (ID) to the second AI chip, wherein the device ID of the second AI chip is the same as or different from the device ID of the first AI chip.

[0023] In some possible implementation scenarios, the device ID of the second AI chip can be the same as that of the first AI chip. That is, after replacing the first AI chip with the second AI chip, the second AI chip takes over the device ID of the first AI chip. Subsequently, when the first AI chip recovers, it can be rejoined to the cluster system. The first AI chip can use the new device ID, or it can replace the second AI chip and restore the original device ID.

[0024] In other possible implementation scenarios, the device ID of the second AI chip can be different from that of the first AI chip. That is, after replacing the first AI chip with the second AI chip, the first computer device assigns a different device ID to the second AI chip. Subsequently, after the first AI chip recovers, it can rejoin the cluster system. The first AI chip can use the new device ID or directly revert to its original device ID.

[0025] This application provides a scheduling method for a cluster system. A first computer device can assign a device ID to a second AI chip that is the same as or different from the first AI chip. When the device ID of the second AI chip is the same as the device ID of the first AI chip, ID allocation resources can be saved. When the device ID of the second AI chip is different from the device ID of the first AI chip, the process of rejoining the cluster system after the faulty AI chip recovers can be simplified.

[0026] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the cluster management device receiving a fourth message, the fourth message being used to instruct the first AI chip to return to normal; the cluster management device sending a fifth message to the first computer device, the fifth message being used to instruct the resumption of management of the first AI chip; the first computer device receiving the fifth message and establishing a mapping relationship between the first port and the first AI chip in the first AI model training list.

[0027] After receiving the fifth message, the first computer device can determine that the first AI chip has returned to normal, for example, by completing fault recovery or upgrading. The first computer device then establishes a mapping relationship between the first port and the first AI chip in the first AI model training list, thereby re-adding the first AI chip to the cluster system.

[0028] Each computer device in the cluster system may include a host, a bus switching chip, and an AI chip. The first AI model training list may be stored in the bus switching chip of the first computer device.

[0029] This application provides a scheduling method for a cluster system, which can realize the replacement and restoration of AI chips in computer devices at the chip level, improve the reliability of the cluster system, reduce the waste of equipment resources caused by the cluster system scheduling process, save the deployment cost of AI chips, and improve product competitiveness.

[0030] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: the first computer device deleting the mapping relationship between the second port and the second AI chip from the first AI model training list.

[0031] In some possible implementations, the first AI chip and the second AI chip both belong to the first computer device. After the first AI chip recovers, it rejoins the cluster system. The first computer device establishes a mapping relationship between the first port and the first AI chip in the first AI model training list, and deletes the mapping relationship between the second port and the second AI chip. In this scenario, the first port and the second port can be the same port or different ports.

[0032] In some possible implementations, the first AI chip belongs to the first computer device, and the second AI chip belongs to the second computer device. After the first AI chip recovers and rejoins the cluster system, the first computer device can either establish a mapping relationship between the first port and the first AI chip only in the first AI model training list, without deleting the mapping relationship between the second port and the second AI chip, or simultaneously establish the mapping relationship between the first port and the first AI chip in the first AI model training list and delete the mapping relationship between the second port and the second AI chip. Whether the first computer device deletes the mapping relationship between the second port and the second AI chip depends on whether the first computer device continues to manage the second AI chip after the first AI chip recovers and rejoins the cluster system. If the first computer device continues to manage the second AI chip after the first AI chip recovers and rejoins the cluster system, the mapping relationship between the second port and the second AI chip does not need to be deleted; if the first computer device no longer manages the second AI chip after the first AI chip recovers and rejoins the cluster system, the mapping relationship between the second port and the second AI chip can be deleted.

[0033] This application provides a scheduling method for a cluster system, which can realize the replacement and restoration of AI chips in computer devices at the chip level, improve the reliability of the cluster system, reduce the waste of equipment resources caused by the cluster system scheduling process, save the deployment cost of AI chips, and improve product competitiveness.

[0034] In conjunction with the first aspect, in some implementations of the first aspect, the second AI chip belongs to a second computer device, which is one of the at least one computer device. The method further includes: the cluster management device sending a sixth message to the second computer device, the sixth message indicating the resumption of the second computer device's management of the second AI chip; the second computer device receiving the sixth message, establishing a mapping relationship between a third port and the second AI chip in a second AI model training list, and deleting the mapping relationship between a fourth port and the second AI chip, wherein the third port includes a port in the second computer device that communicates with the second AI chip, the fourth port includes a port in the second computer device that can communicate with the first computer device, and the second AI model training list is stored in the second computer device.

[0035] In some possible implementation scenarios, after the first AI chip recovers and rejoins the cluster system, the first computer device no longer manages the second AI chip. In this case, the mapping relationship between the second port and the second AI chip can be deleted, and the second computer device resumes management of the second AI chip. After receiving the sixth message, the second computer device establishes a mapping relationship between the third port and the second AI chip in the second AI model training list, deletes the mapping relationship between the fourth port and the second AI chip, and resumes management of the second AI chip.

[0036] This application provides a scheduling method for a cluster system, which can realize the replacement and restoration of AI chips in computer devices at the chip level, improve the reliability of the cluster system, reduce the waste of equipment resources caused by the cluster system scheduling process, save the deployment cost of AI chips, and improve product competitiveness.

[0037] In conjunction with the first aspect, in some implementations of the first aspect, the first AI model training list includes a first AI chip address space and a first backup AI chip address space, wherein the first AI chip address space is used to indicate the address range of the first AI chip, and the first backup AI chip address space is used to indicate the address range of the second AI chip.

[0038] For example, when a first computer device receives a first message, it deletes the mapping relationship between a first port and a first AI chip in a first AI model training list and establishes a mapping relationship between a second port and a second AI chip. This may include the first computer device deleting the mapping relationship between the first port and the address space of the first AI chip in the first AI model training list and establishing a mapping relationship between the second port and the address space of the first backup AI chip.

[0039] This application provides a scheduling method for a cluster system. The AI ​​model training list in the computer device includes the AI ​​chip address space and the AI ​​chip backup address space. It can continue to manage the second AI chip after the first AI chip is restored, thereby improving the flexibility of cluster system scheduling and enhancing product competitiveness.

[0040] Secondly, a cluster system is provided, including a cluster management device and at least one computer device. Each computer device includes at least one artificial intelligence (AI) chip. The cluster management device is configured to send a first message to a first computer device, the first message indicating the replacement of a first AI chip in the first computer device, wherein the first computer device is one of the at least one computer device. The first computer device is configured to receive the first message, delete the mapping relationship between a first port and the first AI chip in a first AI model training list, and establish a mapping relationship between a second port and the second AI chip. The first port includes a port in the first computer device that communicates with the first AI chip, and the second port includes a port in the first computer device that can communicate with the second AI chip. The first AI model training list is stored in the first computer device. The cluster management device is further configured to send model training data to the AI ​​chips in the first AI model training list.

[0041] In conjunction with the second aspect, in some implementations of the second aspect, the second AI chip belongs to a second computer device, which is one of the at least one computer device. The first computer device is further configured to establish a mapping relationship between the second port and the first AI chip in the first AI model training list. The cluster management device is further configured to send a second message to the second computer device, which instructs the second computer device to select a backup AI chip to replace the first AI chip. The second AI chip belongs to the backup AI chip. The second computer device is configured to receive the second message, delete the mapping relationship between the third port and the second AI chip in the second AI model training list, and establish a mapping relationship between the fourth port and the second AI chip. The third port includes a port in the second computer device that communicates with the second AI chip, and the fourth port includes a port in the second computer device that can communicate with the first computer device. The second AI model training list is stored in the second computer device.

[0042] In conjunction with the second aspect, in some implementations of the second aspect, the cluster management device is further configured to receive a third message, which indicates that the first AI chip has malfunctioned or that the first AI chip needs to be upgraded.

[0043] In conjunction with the second aspect, in some implementations of the second aspect, the first computer device is further configured to assign a device identifier ID to the second AI chip, wherein the device ID of the second AI chip may be the same as or different from the device ID of the first AI chip.

[0044] In conjunction with the second aspect, in some implementations of the second aspect, the cluster management device is further configured to receive a fourth message, the fourth message being used to instruct the first AI chip to return to normal operation; the cluster management device is further configured to send a fifth message to the first computer device, the fifth message being used to instruct the resumption of management of the first AI chip; the first computer device is further configured to receive the fifth message and establish a mapping relationship between the first port and the first AI chip in the first AI model training list.

[0045] In conjunction with the second aspect, in some implementations of the second aspect, the first computer device is further configured to remove the mapping relationship between the second port and the second AI chip from the first AI model training list.

[0046] In conjunction with the second aspect, in some implementations of the second aspect, the second AI chip belongs to a second computer device, which is one of the at least one computer device. The cluster management device is further configured to send a sixth message to the second computer device, the sixth message being used to instruct the second computer device to resume management of the second AI chip. The second computer device is configured to receive the sixth message, establish a mapping relationship between a third port and the second AI chip in the second AI model training list, and delete the mapping relationship between a fourth port and the second AI chip. The third port includes ports in the second computer device that communicate with the second AI chip, and the fourth port includes ports in the second computer device that can communicate with the first computer device. The second AI model training list is stored in the second computer device.

[0047] In conjunction with the second aspect, in some implementations of the second aspect, the first AI model training list includes a first AI chip address space and a first backup AI chip address space, wherein the first AI chip address space is used to indicate the address range of the first AI chip, and the first backup AI chip address space is used to indicate the address range of the second AI chip.

[0048] The beneficial effects of the second aspect and any possible implementation of the second aspect correspond to the beneficial effects of the first aspect and any possible implementation of the first aspect, which will not be elaborated further.

[0049] Thirdly, embodiments of this application provide a computer-readable storage medium storing computer program instructions that, when executed by a cluster system according to the second aspect or any possible implementation thereof, cause the cluster system to perform the method as described in the first aspect or any possible implementation thereof. Attached Figure Description

[0050] Figure 1 is a schematic diagram of a cluster system provided in an embodiment of this application.

[0051] Figure 2 is a schematic diagram of another cluster system provided in an embodiment of this application.

[0052] Figure 3 is a schematic diagram of the architecture of another AI cluster system provided in an embodiment of this application.

[0053] Figure 4 is a schematic diagram of the architecture of another AI cluster system provided in an embodiment of this application.

[0054] Figure 5 is a schematic diagram of the address space of a CPU and an AI chip provided in an embodiment of this application.

[0055] Figure 6 is a schematic diagram of the PCIE ECAM address space of an AI chip provided in an embodiment of this application.

[0056] Figure 7 is an exemplary flowchart of a scheduling method for a cluster system provided in an embodiment of this application.

[0057] Figure 8 is a schematic diagram of a backup AI chip replacement for a faulty or upgradeable AI chip provided in an embodiment of this application.

[0058] Figure 9 is a schematic diagram of AI chip recovery provided in an embodiment of this application.

[0059] Figure 10 is a schematic diagram of the architecture of another AI cluster system provided in an embodiment of this application.

[0060] Figure 11 is a schematic diagram of the architecture of another AI cluster system provided in an embodiment of this application.

[0061] Figure 12 is an exemplary flowchart of another scheduling method for a cluster system provided in an embodiment of this application.

[0062] Figure 13 is a schematic diagram of another backup AI chip replacement for a faulty or upgradeable AI chip provided in an embodiment of this application.

[0063] Figure 14 is a schematic diagram of another AI chip recovery provided in an embodiment of this application.

[0064] Figure 15 is an exemplary flowchart of another scheduling method for a cluster system provided in an embodiment of this application.

[0065] Figure 16 is an exemplary flowchart of another scheduling method for a cluster system provided in an embodiment of this application.

[0066] Figure 17 is an exemplary flowchart of another scheduling method for a cluster system provided in an embodiment of this application. Detailed Implementation

[0067] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort should fall within the scope of protection of this application.

[0068] In the embodiments of this application, the words "exemplary," "for example," etc., are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design that is described as "exemplary" in this application should not be construed as being more preferred or advantageous than other embodiments or design options. Specifically, the use of the term "exemplary" is intended to present the concept in a concrete manner.

[0069] The business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0070] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0071] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0072] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0073] To facilitate understanding of the embodiments of this application, some definitions involved in this application will be briefly explained first.

[0074] 1. Resumable from breakpoint: When the AI ​​cluster system performs AI calculations, the computer equipment will periodically save the training results of each iteration at checkpoints (CKPTs) so that if training is interrupted or an error occurs, training can start from the CKPT without starting from scratch, reducing the waste of time and computing resources.

[0075] 2. Peripheral Component Interconnect Express (PCIe) Configuration Space: This is a small block of memory reserved for each device on the Peripheral Component Interconnect (PCI) bus, which contains basic information about the device and its control registers.

[0076] Figure 1 is a schematic diagram of a cluster system provided in an embodiment of this application.

[0077] Artificial intelligence (AI) is the main direction of current technological development. Network models are the foundation of AI. Through training with AI cluster systems, more accurate and powerful network models can be trained quickly.

[0078] The cluster system shown in Figure 1 includes a cluster management device 100 and N computer devices, where N ≥ 1. These N computer devices are interconnected via a bus switching chip B, enabling communication between them. The bus switching chip B can be located within a network switch. Each computer device includes a host, a bus switching chip A, and at least one AI chip. The host is a central processing unit (CPU) based system, and the host and AI chips are interconnected via bus switching chip A. Each host can interconnect with multiple AI chips. The computer devices in the cluster system can be servers, specifically AI servers.

[0079] AI cluster systems provide AI computing power and can be used in various AI application scenarios, including AI network model training and inference. For example, an AI cluster system may consist of thousands or tens of thousands of computer devices, with at least some AI chips in all these devices participating in the training of the same AI network model. The cluster management device 100 distributes AI tasks to the computer devices, and the AI ​​chips on these devices execute the AI ​​training tasks.

[0080] As AI models become increasingly complex, larger, and have more parameters, AI training often requires AI clusters with tens of thousands of GPUs (over 10,000 CPUs), leading to a higher probability of cluster failures. Real-world analysis shows that over 50% of failures are due to cluster infrastructure malfunctions, such as 1-2 GPUs failing weekly. As shown in Figure 2, to improve system stability, backup computers can be added to the AI ​​cluster system. During AI computation, these computers periodically save the training results for each iteration using CKPT. Currently, when a computer in an AI cluster system fails, the mainstream recovery method is interrupted resuming the process. The AI ​​task is interrupted, and the backup computer replaces the failed one to continue executing subsequent AI tasks. However, in this method, if one AI chip in a computer device fails, the entire computer device will have to be replaced, and other working AI chips will also be unusable, resulting in a waste of equipment resources. Secondly, existing technologies can only replace the entire computer device, requiring a large number of backup computers, which is costly. Furthermore, as the AI ​​network model becomes larger, frequent failures will cause CKPT to roll back frequently, resulting in extremely low utilization of AI cluster computing power.

[0081] This application proposes a scheduling method for an AI chip-level cluster system. When an AI chip in a computer device fails, only the faulty AI chip can be replaced instead of replacing the entire computer device. This improves the reliability of the cluster system, reduces resource waste during cluster system scheduling, saves deployment costs for AI chips, and enhances product competitiveness.

[0082] Figure 3 is a schematic diagram of the architecture of another AI cluster system provided in an embodiment of this application.

[0083] This application embodiment uses an AI cluster system including a cluster management device 100 and a computer device 200 as an example to specifically describe the architecture of the AI ​​cluster system. The cluster management device 100 includes a cluster device management module 111, a cluster task scheduling module 112, and a cluster fault management module 113. The computer device 200 includes a host 210, a bus switching chip 220, an AI chip 231, and an AI chip 232. The host 210 is interconnected with the AI ​​chips 231 and 232 via the bus switching chip 220. The host 210 includes an AI chip management module 211, a task execution module 214, an AI chip fault management module 215, an AI chip driver 216, a CPU 217, and a bus switching routing driver 218. The AI ​​chip management module 211 includes a routing management submodule 212 and a device management submodule 213. The bus switching chip 220 includes ports 21, 22, and 24, a routing engine 221, an AI model training list 222, and a device configuration window 223. AI chip 231 includes device configuration information 234, AI chip 232 includes device configuration information 235, and each AI chip also includes its own registers and memory space. CPU 217 communicates with bus switching chip 220 through port 24, AI chip 231 communicates with bus switching chip 220 through port 21, and AI chip 232 communicates with bus switching chip 220 through port 22.

[0084] The various modules within the cluster management device 100 and the computer equipment can work together to achieve hot backup functionality for the AI ​​chip, including but not limited to AI chip failure scenarios, AI chip upgrade scenarios, or computer equipment expansion scenarios requiring additional AI chips. The functions of each module are described in detail below:

[0085] The cluster device management module 111 is responsible for the management of all chips in an AI cluster system. This module manages AI chip 231, AI chip 232 and bus switching chip 220 through AI chip management module 211.

[0086] The cluster task scheduling module 112 is responsible for the task scheduling of the entire AI cluster system. It is responsible for distributing AI computing tasks to each computer device, and then the task execution module 214 of the computer device distributes them to each AI chip for execution.

[0087] The cluster fault management module 113 is responsible for the fault management of the entire AI cluster system. It is responsible for collecting information on AI chip faults or computer device faults in a computer device, marking the AI ​​chip or computer device as faulty, and removing the faulty AI chip or computer device from the cluster system; it is responsible for adding backup AI chips and backup computer devices to the AI ​​cluster system; it is responsible for adding AI chips and computer devices that have recovered to normal to the cluster system; and it is responsible for returning backup AI chips to backup computer devices.

[0088] The AI ​​chip management module 211 is responsible for managing the bus switching chip 220, AI chip 231, and AI chip 232. The AI ​​chip management module 211 includes a routing management submodule 212 and a device management submodule 213. The routing management submodule 212 manages the bus switching chip 220, controlling routing rules between AI chips by configuring the AI ​​model training list 222 of the routing engine 221. The device management submodule 213 manages AI chips 231 and 232, including resource switching between AI chips 231 and 232 in the AI ​​chip driver 216 on the host 210 side, and identity (ID) allocation.

[0089] The task execution module 214 is responsible for distributing the computing tasks of the computer device to each AI chip for execution.

[0090] The AI ​​chip fault management module 215 is responsible for reporting fault information of AI chip 231 and / or AI chip 232.

[0091] The routing engine 221 is responsible for data forwarding between AI chip 231, AI chip 232 and CPU 217, between AI chip 231 and AI chip 232, and between CPU 217. It forwards data based on routing rules.

[0092] AI model training list 222 stores the mapping relationship between the ports of bus switching chip 220 and the address space of AI chip in computer device 200. By dynamically adjusting the mapping relationship between the ports and the AI ​​chip address space, the data forwarding direction of routing engine 221 can be controlled. Routing management submodule 212 modifies AI model training list 222 through bus switching routing driver 218.

[0093] Device configuration window 223 is the access address space for device configuration information 234 and device configuration information 235, and is allocated on the bus switching chip 220 for each AI chip. Device configuration window 223 includes the address space of AI chip 231, the address space of AI chip 232, the address space of backup AI chip 231, and the address space of backup AI chip 232. The AI ​​chip address space is used to indicate the address range of the AI ​​chip, and the backup AI chip address space is used to indicate the address range of the backup AI chip that replaces the AI ​​chip. When AI chip 231 participates in model training as a normal AI chip, the address space of AI chip 231 is used to obtain device configuration information 234; when AI chip 231 fails or needs to be upgraded, the address space of backup AI chip 231 is used to obtain the device configuration information of the backup AI chip that replaces AI chip 231. When AI chip 232 participates in model training as a normal AI chip, the address space of AI chip 232 is used to obtain device configuration information 235; when AI chip 232 malfunctions or needs to be upgraded, the address space of backup AI chip 232 is used to obtain device configuration information of the backup AI chip that replaces AI chip 232.

[0094] Device configuration information 234 includes management information of AI chip 231, and device configuration information 235 includes management information of AI chip 232.

[0095] AI chip driver 216 is used to initialize AI chip 231 through device configuration information 234, initialize AI chip 232 through device configuration information 235, establish management plane and business plane channels between CPU 217 and AI chips 231 and 232, provide other software with driver interfaces for managing AI chips 231 and 232, assign device IDs to AI chips 231 and 232, and provide task execution driver interfaces for AI chips 231 and 232 for use by other software.

[0096] It should be understood that the number of CPUs, bus switching chips and AI chips included in Figure 3 is only an example, and this application does not limit the specific number of CPUs, bus switching chips and AI chips in a computer device.

[0097] Figure 4 is a schematic diagram of the architecture of another AI cluster system provided in an embodiment of this application.

[0098] Figure 4 is similar to Figure 3, except that the bus switching chip 220 includes ports 21, 22, and 24, a routing engine 221, an AI model training list 222, and a device configuration window 225. The device configuration window 225 includes the address spaces of AI chip 231 and AI chip 232.

[0099] The descriptions of the remaining modules can be found in Figure 3, and will not be repeated here.

[0100] Figure 5 is a schematic diagram of the address space of a CPU and an AI chip provided in an embodiment of this application.

[0101] Within the entire AI cluster system, all AI chips are uniformly addressed, with each CPU and AI chip having its own address space. Each CPU allocates a PCIe Enhanced Configuration Access Mechanism (ECAM) space for each AI chip located within the same computer device. The PCIe ECAM space includes the address range of the AI ​​chip required for interconnection between the CPU and the AI ​​chip.

[0102] For example, a computer device includes P CPUs and M AI chips, where P ≥ 1 and M ≥ 1. The address space of each of the P CPUs (CPU-1 to CPU-P) includes the PCIE ECAM space of the M AI chips. Optionally, it may also include the PCIE ECAM space of backup AI chips. As shown in Figure 6, taking CPU-1 as an example, in the address space of CPU-1, the m-th AI chip corresponds to a PCIE ECAM space for AI chip m and a PCIE ECAM space for a backup AI chip m. When the m-th AI chip is in a normal state, the PCIE ECAM space of AI chip m is used to obtain the device configuration information of the m-th AI chip. When the m-th AI chip is in an abnormal state, the PCIE ECAM space of the backup AI chip m is used to obtain the device configuration information of the backup AI chip that replaces the m-th AI chip, where 1 ≤ m ≤ M. An abnormal state can be, for example, a fault state or a state awaiting upgrade.

[0103] It should be understood that the address space of the AI ​​chip and the PCIe ECAM space of the AI ​​chip in this application embodiment are different concepts. The address space of the AI ​​chip includes the address range corresponding to the AI ​​chip and is a characteristic attribute of the AI ​​chip. The PCIe ECAM space of the AI ​​chip is located in the CPU and is used to help the CPU obtain the device configuration information of the AI ​​chip.

[0104] The management process of AI chips will be described in detail below with reference to Figures 3 to 6.

[0105] After the CPU on the host side starts up, it accesses the PCIe ECAM space of the AI ​​chip within the CPU. The access command is sent to the bus switching chip, and then routed to the AI ​​chip through the bus switching chip to obtain the device configuration information of the AI ​​chip. For example, in computer device 200, when CPU 217 and AI chip 231 are powered on together, CPU 217 accesses the PCIe ECAM space of AI chip 231 and sends a device configuration information request message to bus switching chip 220 through port 24. The routing engine 221 on bus switching chip 220 forwards the device configuration information request message that hits the address space of AI chip 231 to port 21 according to the routing rules of AI model training list 222, thereby obtaining the device configuration information 234 of AI chip 231.

[0106] After obtaining the device configuration information 234 of AI chip 231, AI chip driver 216 can create a device symbol and device ID for AI chip 231. For example, the device symbol can be Device0, and the device ID can be 231. Alternatively, AI chip driver 216 can create a device ID only for AI chip 231, which can be Device231. Taking Device231 as an example, AI chip management module 211, task execution module 214, and AI chip fault management module 215 on host 210 can all manage AI chip 231 based on this device ID and issue tasks to AI chip 231. The management process for AI chip 232 is similar and will not be repeated here. The device ID created by AI chip driver 216 for AI chip 232 can be Device232.

[0107] In addition to the management process of the AI ​​chip, the data from the host-side AI chip driver to the AI ​​chip for issuing AI computing tasks can also be routed and mapped from the AI ​​chip address space to the AI ​​chip's registers and memory space through the bus switching chip. For example, in computer device 200, when the task execution module 214 issues an AI computing task through the AI ​​chip driver 216, the task issuance instruction is issued to the bus switching chip 220 through port 24. The routing engine 221 on the bus switching chip 220 forwards the task issuance instruction to port 21, and then sends it from port 21 to the AI ​​chip 231.

[0108] It should be understood that in communication scenarios between AI chips, such as the communication between AI chip 231 and AI chip 232 in computer device 200, data transmission is also performed through a bus switching chip. For example, when AI chip 231 in computer device 200 needs to send a message to AI chip 232 in computer device 200, the data is sent to bus switching chip 220 through port 21. Bus switching chip 220 determines that the destination address of the data matches the address space of AI chip 232, and then sends the data to port 22 according to the routing rules of AI model training list 222. The data is then sent to AI chip 232 through port 22.

[0109] Figure 7 is an exemplary flowchart of a scheduling method for a cluster system provided in an embodiment of this application. Figure 7 corresponds to Figures 3 and 4.

[0110] 710, AI chip malfunction or upgrade information reported.

[0111] Taking Figures 3 and 4 as examples, if the AI ​​chip 232 in computer device 200 malfunctions, the fault is reported to the AI ​​chip fault management module 215 via the AI ​​chip driver 216. The AI ​​chip fault management module 215 then reports the fault information to the cluster fault management module 113, which in turn reports the fault information to the cluster device management module 111. The upgrade process for the AI ​​chip 232 is similar to the fault reporting process described above, and will not be repeated here.

[0112] 720, Modify the AI ​​model training list.

[0113] The cluster device management module 111 obtains the fault information of AI chip 232 from the cluster fault management module 113, and then selects an idle backup AI chip to replace the faulty AI chip. Optionally, the cluster device management module 111 can also determine whether the current fault of AI chip 232 affects the AI ​​cluster system's execution of AI calculations based on the fault level. If it does, an idle backup AI chip is selected to replace the faulty AI chip. For example, the backup AI chip can be AI chip 231, i.e., AI chip 231 is used to replace AI chip 232. The AI ​​model training list 222 stores the mapping relationship between ports and AI chip address spaces. By dynamically adjusting the mapping relationship between ports and address spaces, the data forwarding direction of the routing engine can be controlled, thereby realizing the replacement of the faulty AI chip with a backup AI chip. The specific process can be seen in the description of Figure 8.

[0114] The backup AI chip and the faulty AI chip can belong to the same computer device or different computer devices; this application does not impose any restrictions on this. In some possible application scenarios, the AI ​​cluster system includes a backup computer device, and at least some of the AI ​​chips in the backup computer device are backup AI chips.

[0115] 730, AI chip device information switching and resource reconstruction.

[0116] After adjusting the routing mapping relationship between the completion port and the AI ​​chip address space on the computer device 200, the backup AI chip needs to be initialized on the computer device 200 by calling the interface of the AI ​​chip driver 216.

[0117] In one possible implementation scenario, the AI ​​chip driver 216 can retain the device ID of the faulty AI chip 232, but the backup AI chip replaces the device ID of the faulty AI chip 232. The AI ​​chip management module 211 and the task execution module 214 continue to use the device ID of the AI ​​chip 232 to manage and distribute AI computing tasks to the backup AI chip. For example, if the backup AI chip is AI chip 231, after replacing the faulty AI chip 232 with AI chip 231, the AI ​​chip management module 211 and the task execution module 214 continue to use the device ID of the AI ​​chip 232 to manage and distribute AI computing tasks to the AI ​​chip 231. It should be understood that if the AI ​​chip driver 216 has already assigned a device ID to the AI ​​chip 231 before the backup AI chip 231 is initialized, the device ID of the AI ​​chip 231 can be deleted.

[0118] In another possible implementation scenario, the AI ​​chip driver 216 can assign a different device ID to the backup AI chip than the AI ​​chip 232. For example, after initializing the AI ​​chip 231, the device ID assigned to the AI ​​chip 231 is Device 231. The AI ​​chip management module 211 and the task execution module 214 use Device 231 to manage and distribute AI computing tasks to the AI ​​chip 231.

[0119] 740, Recovery of the faulty AI chip.

[0120] Once the faulty AI chip 232 is restored to normal, it can be added back to the original computer device 200 and AI cluster system. For details, please refer to Figure 9.

[0121] The scheduling method provided in this application embodiment does not cause the entire computer equipment to fail when an AI chip fails in an AI cluster system. It does not require replacing the computer equipment with the faulty AI chip with the entire computer equipment. Instead, it only requires replacing the faulty AI chip with a backup AI chip, thus saving equipment resources.

[0122] Figure 8 is a schematic diagram of a backup AI chip replacement for a faulty or upgradeable AI chip provided in an embodiment of this application. Figure 8 corresponds to step 720 in Figure 7.

[0123] 810, the cluster fault management module 113 reports the fault information of the AI ​​chip 232 to the cluster device management module 111, and the cluster device management module 111 obtains the fault information of the AI ​​chip 232 from the cluster fault management module 113.

[0124] It should be understood that the cluster fault management module 113 can also report the chip upgrade information of the AI ​​chip 232 to the cluster device management module 111. The scheduling process of the cluster system in the chip upgrade scenario is basically the same as that in the chip fault scenario, and will not be repeated here.

[0125] 820. The cluster device management module 111 determines whether the current failure of the AI ​​chip 232 affects the AI ​​cluster system's execution of AI calculations based on the fault level. If it does, it selects an idle backup AI chip to replace the faulty AI chip. For example, the cluster device management module 111 can send a first message to the computer device 200 to instruct the computer device 200 to replace the faulty AI chip 232.

[0126] 830, The AI ​​chip management module 211 of the computer device 200 selects an idle AI chip 231 as a backup AI chip to replace the faulty AI chip 232.

[0127] 840, the routing management submodule 212 of the AI ​​chip management module 211 changes the AI ​​model training list 222 through the bus switching routing driver 218.

[0128] For the architecture diagram shown in Figure 3, the mapping relationship between port 24 and the address space of AI chip 232 can be deleted by configuring AI model training list 222, and a mapping relationship between port 24 and the address space of backup AI chip 232 can be established. When a message hits the address space of backup AI chip 232 according to the routing rules of AI model training list 222, it can access AI chip 231 through port 21. It should be understood that the address space of backup AI chip 232 corresponds to AI chip 231, and this correspondence can be configured before or after the failure of AI chip 232. Optionally, the mapping relationship between port 22 and the address space of AI chip 232 in AI model training list 222 can also be deleted.

[0129] For the architecture diagram shown in Figure 4, the mapping relationship between port 24 and the address space of AI chip 232 can be deleted by configuring AI model training list 222, and the mapping relationship between port 24 and the address space of AI chip 231 can be established. Alternatively, the mapping relationship between port 22 and the address space of AI chip 232 can be deleted, and the mapping relationship between port 21 and the address space of AI chip 231 can be established.

[0130] Figure 9 is a schematic diagram of AI chip recovery provided in an embodiment of this application. Figure 9 corresponds to step 740 in Figure 7.

[0131] 910. After the faulty AI chip 232 returns to normal, it can be added back to the original computer device 200 and AI cluster system. The cluster fault management module 113 reports the information that the AI ​​chip 232 has returned to normal to the cluster device management module 111, and the cluster device management module 111 obtains the information that the AI ​​chip 232 has returned to normal from the cluster fault management module 113.

[0132] 920, the cluster device management module 111 sends a fifth message to the computer device 200, instructing the computer device 200 to resume management of the AI ​​chip 232.

[0133] 930, The AI ​​chip management module 211 of computer device 200 changes the AI ​​model training list 222 via bus switching routing driver 218.

[0134] For the architecture diagram shown in Figure 3, the mapping relationship between port 24 and the address space of the backup AI chip 232 can be established by configuring the AI ​​model training list 222, deleting the mapping relationship between port 24 and the address space of the backup AI chip 232.

[0135] For the architecture diagram shown in Figure 4, the mapping relationship between port 24 and the address space of AI chip 232 can be established by configuring AI model training list 222. Optionally, the mapping relationship between port 24 and the address space of AI chip 231 can be deleted.

[0136] 940. Initialize AI chip 232 through AI chip driver 216. If AI chip 231 is used as a backup to replace the device ID of AI chip 232, delete the device ID of AI chip 231 in computer device 200 and restore the device ID of AI chip 232. Alternatively, a new device ID can be reassigned to AI chip 232.

[0137] Figure 10 is a schematic diagram of the architecture of another AI cluster system provided in an embodiment of this application.

[0138] The AI ​​cluster system includes a cluster management device 100, a computer device 300, and a computer device 400. The cluster management device 100 includes a cluster device management module 111, a cluster task scheduling module 112, and a cluster fault management module 113. Computer devices 300 and 400 are interconnected via a bus switching chip 520. The bus switching chip 520 includes ports 51 and 52, a routing engine 521, and an AI model training list 522.

[0139] Computer device 300 includes host 310, bus switching chip 320, AI chip 331 and AI chip 332. The AI ​​chip in computer device 300 is interconnected with CPU 317 on the host 310 side through bus switching chip 320.

[0140] The host 310 includes an AI chip management module 311, a task execution module 314, an AI chip fault management module 315, an AI chip driver 316, a CPU 317, and a bus switching and routing driver 318. The AI ​​chip management module 311 includes a routing management submodule 312 and a device management submodule 313. The bus switching chip 320 includes ports 31, 32, 33, and 34, a routing engine 321, an AI model training list 322, and a device configuration window 323. The CPU 317 communicates with the bus switching chip 320 via port 34, the AI ​​chip 331 communicates with the bus switching chip 320 via port 31, and the AI ​​chip 332 communicates with the bus switching chip 320 via port 32. AI chip 331 includes device configuration information 334, AI chip 332 includes device configuration information 335, and each AI chip also includes its own registers and memory space.

[0141] Computer device 400 includes host 410, bus switching chip 420, AI chip 431 and AI chip 432. The AI ​​chip in computer device 400 is interconnected with CPU 417 on the host 410 side through bus switching chip 420.

[0142] The host 410 includes an AI chip management module 411, a task execution module 414, an AI chip fault management module 415, an AI chip driver 416, a CPU 417, and a bus switching and routing driver 418. The AI ​​chip management module 411 includes a routing management submodule 412 and a device management submodule 413. The bus switching chip 420 includes ports 41, 42, 43, and 44, a routing engine 421, and a device configuration window 423. The CPU 417 communicates with the bus switching chip 420 via port 44, the AI ​​chip 431 communicates with the bus switching chip 420 via port 41, and the AI ​​chip 432 communicates with the bus switching chip 420 via port 42. AI chip 431 includes device configuration information 434, and AI chip 432 includes device configuration information 435. Each AI chip also includes its own registers and memory space.

[0143] The various modules within the cluster management device 100 and the computer equipment can work together to achieve hot backup functionality for the AI ​​chip, including but not limited to AI chip failure scenarios, AI chip upgrade scenarios, or computer equipment expansion scenarios requiring additional AI chips. The functions of each module are described in detail below:

[0144] The cluster device management module 111 is responsible for the management of all chips in an AI cluster system. This module manages AI chip 331, AI chip 332 and bus switching chip 320 through AI chip management module 311, and manages AI chip 431, AI chip 432 and bus switching chip 420 through AI chip management module 411.

[0145] The cluster task scheduling module 112 is responsible for the task scheduling of the entire AI cluster system. It is responsible for distributing AI computing tasks to each computer device, and then the task execution module 314 of the computer device 300 distributes them to each AI chip of the computer device 300 for execution. The task execution module 414 of the computer device 400 distributes them to each AI chip of the computer device 400 for execution.

[0146] The cluster fault management module 113 is responsible for the fault management of the entire AI cluster system. It is responsible for collecting information on AI chip faults or computer device faults in a computer device, marking the AI ​​chip or computer device as faulty, and removing the faulty AI chip or computer device from the cluster system; it is responsible for adding backup AI chips and backup computer devices to the AI ​​cluster system; it is responsible for adding AI chips and computer devices that have recovered to normal to the cluster system; and it is responsible for returning backup AI chips to backup computer devices.

[0147] The AI ​​chip management module 311 is responsible for managing the bus switching chip 320, AI chip 331, and AI chip 332. The AI ​​chip management module 311 includes a routing management submodule 312 and a device management submodule 313. The routing management submodule 312 manages the bus switching chip 320, controlling the routing rules between AI chips by configuring the AI ​​model training list 322 of the routing engine 321. The device management submodule 313 manages AI chips 331 and 332, including resource switching and ID allocation between AI chips 331 and 332 in the AI ​​chip driver 316 on the host 310 side.

[0148] The task execution module 314 is responsible for distributing the computing tasks of the computer device 300 to each AI chip of the computer device 300 for execution.

[0149] The AI ​​chip fault management module 315 is responsible for reporting fault or upgrade information of AI chip 331 and / or AI chip 332.

[0150] Routing engine 321 is responsible for data forwarding between AI chip 331, AI chip 332 and CPU 317, between AI chip 331 and AI chip 332, and between CPU 317. It forwards data based on routing rules.

[0151] AI model training list 322 stores the mapping relationship between the ports of bus switching chip 320 and the address space of AI chip in computer device 300. By dynamically adjusting the mapping relationship between the ports and the AI ​​chip address space, the data forwarding direction of routing engine 321 can be controlled. Routing management submodule 312 modifies AI model training list 322 through bus switching routing driver 318.

[0152] The device configuration window 323 includes the address space of AI chip 331, the address space of AI chip 332, the address space of backup AI chip 331, and the address space of backup AI chip 332. The AI ​​chip address space indicates the address range of the AI ​​chip, and the backup AI chip address space indicates the address range of the backup AI chip that replaces the AI ​​chip. When AI chip 331 participates in model training as a normal AI chip, the address space of AI chip 331 is used to obtain device configuration information 334; when AI chip 331 malfunctions or needs to be upgraded, the address space of backup AI chip 331 is used to obtain device configuration information of the backup AI chip that replaces AI chip 331. When AI chip 332 participates in model training as a normal AI chip, the address space of AI chip 332 is used to obtain device configuration information 335; when AI chip 332 malfunctions or needs to be upgraded, the address space of backup AI chip 332 is used to obtain device configuration information of the backup AI chip that replaces AI chip 332.

[0153] The AI ​​chip management module 411 is responsible for managing the bus switching chip 420, AI chip 431, and AI chip 432. The AI ​​chip management module 411 includes a routing management submodule 412 and a device management submodule 413. The routing management submodule 412 manages the bus switching chip 420, controlling routing rules between AI chips by configuring the AI ​​model training list 422 of the routing engine 421. The device management submodule 413 manages AI chips 431 and 432, including resource switching and ID allocation between AI chips 431 and 432 in the AI ​​chip driver 416 on the host 410 side.

[0154] The task execution module 414 is responsible for distributing the computing tasks of the computer device 400 to each AI chip of the computer device 400 for execution.

[0155] The AI ​​chip fault management module 415 is responsible for reporting fault or upgrade information of AI chip 431 and / or AI chip 432.

[0156] The routing engine 421 is responsible for data forwarding between AI chips 431, 432 and CPU 417, between AI chips 431 and 432, and between CPU 417. It forwards data based on routing rules.

[0157] AI model training list 422 stores the mapping relationship between the ports of bus switching chip 420 and the address space of AI chip in computer device 400. By dynamically adjusting the mapping relationship between the ports and the AI ​​chip address space, the data forwarding direction of routing engine 421 can be controlled. Routing management submodule 412 modifies AI model training list 422 through bus switching routing driver 418.

[0158] The device configuration window 423 includes the address space of AI chip 431, the address space of AI chip 432, the address space of backup AI chip 431, and the address space of backup AI chip 432. When AI chip 431 participates in model training as a normal AI chip, the address space of AI chip 431 is used to obtain device configuration information 434; when AI chip 431 malfunctions or needs to be upgraded, the address space of backup AI chip 431 is used to obtain device configuration information of the backup AI chip that replaces AI chip 431. When AI chip 432 participates in model training as a normal AI chip, the address space of AI chip 432 is used to obtain device configuration information 435; when AI chip 432 malfunctions or needs to be upgraded, the address space of backup AI chip 432 is used to obtain device configuration information of the backup AI chip that replaces AI chip 432.

[0159] AI model training list 522 stores the mapping relationship between the ports of bus switching chip 520 and the address space of AI chip. By dynamically adjusting the mapping relationship between ports, the data forwarding direction of routing engine 521 can be controlled.

[0160] It should be understood that the number of CPUs, bus switching chips and AI chips included in Figure 10 is only an example, and this application does not limit the specific number of CPUs, bus switching chips and AI chips in a computer device.

[0161] Figure 11 is a schematic diagram of the architecture of another AI cluster system provided in an embodiment of this application.

[0162] Figure 11 is similar to Figure 10, except that the bus switching chip 320 includes ports 31, 32, 33, and 34, a routing engine 321, an AI model training list 322, and a device configuration window 325. The bus switching chip 420 includes ports 41, 42, 43, and 44, a routing engine 421, and a device configuration window 425. The device configuration window 325 includes the address spaces of AI chip 331 and AI chip 332. The device configuration window 425 includes the address spaces of AI chip 431 and AI chip 432.

[0163] The descriptions of the remaining modules can be found in Figure 10, and will not be repeated here.

[0164] The management process of the AI ​​chip is described in detail below with reference to Figures 10 and 11.

[0165] After the CPU on the host side starts up, it accesses the PCIe ECAM space of the AI ​​chip within the CPU. The access command is sent to the bus switching chip, and then routed to the AI ​​chip through the bus switching chip to obtain the device configuration information of the AI ​​chip. For example, in computer device 300, when CPU 317 and AI chip 331 are powered on together, CPU 317 accesses the PCIe ECAM space of AI chip 331 and sends a device configuration information request message to bus switching chip 320 through port 34. The routing engine 321 on bus switching chip 320 forwards the device configuration information request message that hits the address space of AI chip 331 to port 31 according to the routing rules of AI model training list 322, thereby obtaining the device configuration information 334 of AI chip 331.

[0166] After obtaining the device configuration information 334 of AI chip 331, AI chip driver 316 creates a device ID for AI chip 331: Device331. AI chip management module 311, task execution module 314, and AI chip fault management module 315 on host 310 all manage AI chip 331 based on this device ID and issue tasks to AI chip 331. The management process for AI chip 332 is similar and will not be repeated here. The device ID created by AI chip driver 316 for AI chip 332 can be Device332.

[0167] In addition to the management process of the AI ​​chip, the data from the host-side AI chip driver to the AI ​​chip for issuing AI computing tasks can also be routed and mapped from the AI ​​chip address space to the AI ​​chip's registers and memory space through the bus switching chip. For example, in computer device 300, when the task execution module 314 issues an AI computing task through the AI ​​chip driver 316, the task issuance instruction is issued to the bus switching chip 320 through port 34. The routing engine 321 on the bus switching chip 320 forwards the task issuance instruction to port 31, and then sends it from port 31 to the AI ​​chip 331.

[0168] It should be understood that in communication scenarios between AI chips, such as communication between AI chips 331 and 332 in computer device 300, and communication between AI chips 331 and 432 in computer device 400, data transmission is also performed through a bus switching chip. For example, when AI chip 331 in computer device 300 needs to send a message to AI chip 332 in computer device 300, the data is sent to the bus switching chip 320 through port 31. The bus switching chip 320 determines that the destination address of the data matches the address space of AI chip 332, and then sends the data to port 32 according to the routing rules of the AI ​​model training list 322. The data is then sent to AI chip 332 through port 32.

[0169] Figure 12 is an exemplary flowchart of another scheduling method for a cluster system provided in an embodiment of this application. Figure 12 corresponds to Figures 10 and 11.

[0170] 1210, AI chip malfunction or upgrade information reported.

[0171] Taking Figures 10 and 11 as examples, if the AI ​​chip 332 in computer device 300 malfunctions, the fault is reported to the AI ​​chip fault management module 315 via the AI ​​chip driver 316. The AI ​​chip fault management module 315 then reports the fault information to the cluster fault management module 113, which in turn reports the fault information to the cluster device management module 111. The upgrade process for the AI ​​chip 332 is similar to the fault reporting process described above, and will not be repeated here.

[0172] 1220, Modify the AI ​​model training list.

[0173] The cluster device management module 111 obtains the fault information of AI chip 332 from the cluster fault management module 113, and then selects an idle backup AI chip to replace the faulty AI chip. Optionally, the cluster device management module 111 can also determine whether the current fault of AI chip 332 affects the AI ​​cluster system's execution of AI calculations based on the fault level. If it does, it selects an idle backup AI chip to replace the faulty AI chip. For example, the backup chip can be AI chip 331, AI chip 431, or AI chip 432. The AI ​​model training list stores the mapping relationship between ports and AI chip address spaces. By dynamically adjusting the mapping relationship between ports and address spaces, the data forwarding direction of the routing engine can be controlled, thereby realizing the replacement of the faulty AI chip with a backup AI chip. The specific process can be seen in Figure 13.

[0174] The backup AI chip and the faulty AI chip can belong to the same computer device or different computer devices; this application does not impose any restrictions on this. In some possible application scenarios, the AI ​​cluster system includes a backup computer device, and at least some of the AI ​​chips in the backup computer device are backup AI chips.

[0175] 1230, AI chip device information switching and resource reconstruction.

[0176] After adjusting the routing mapping relationship between the completion port and the AI ​​chip address space on the computer device 300, the backup AI chip needs to be initialized on the computer device 300 by calling the interface of the AI ​​chip driver 316.

[0177] In one possible scenario, the AI ​​chip driver 316 can retain the device ID of the faulty AI chip 332, but a backup AI chip replaces the device ID of AI chip 332. The AI ​​chip management module 311 and the task execution module 314 continue to use the device ID of AI chip 332 to manage and distribute AI computing tasks to the backup AI chip. For example, if the backup AI chip is AI chip 432, after replacing the faulty AI chip 332 with AI chip 432, the AI ​​chip management module 311 and the task execution module 314 continue to use the device ID of AI chip 332 to manage and distribute AI computing tasks to AI chip 432.

[0178] In another possible implementation scenario, the AI ​​chip driver 316 can assign a different device ID to the backup AI chip than the AI ​​chip 332. For example, after initializing the backup AI chip 432, the device ID assigned to the AI ​​chip 432 is Device432. The AI ​​chip management module 311 and the task execution module 314 use Device432 to manage and distribute AI computing tasks to the AI ​​chip 432.

[0179] In some possible implementation scenarios, if the backup AI chip and the faulty AI chip belong to different computer devices, for example, the backup AI chip is AI chip 432 and AI chip 432 belongs to computer device 400, after adjusting the routing mapping between the completion port and the AI ​​chip address space on computer device 400, the device ID of AI chip 432 on computer device 400 can be deleted by calling the interface of AI chip driver 416 through AI chip management module 411 on computer device 400.

[0180] 1240, Recovery of the faulty AI chip.

[0181] Once the faulty AI chip 332 is restored to normal, it can be added back to the original computer device 300 and AI cluster system. For details, please refer to the description in Figure 14.

[0182] The scheduling method provided in this application embodiment does not cause the entire computer equipment to fail when an AI chip fails in an AI cluster system. It does not require replacing the computer equipment with the faulty AI chip with the entire computer equipment. Instead, it only requires replacing the faulty AI chip with a backup AI chip, thus saving equipment resources.

[0183] Figure 13 is a schematic diagram of another backup AI chip replacement for a faulty or upgradeable AI chip provided in an embodiment of this application. Figure 13 corresponds to step 1220 in Figure 12.

[0184] 1310, the cluster fault management module 113 reports the fault information of the AI ​​chip 332 to the cluster device management module 111, and the cluster device management module 111 obtains the fault information of the AI ​​chip 332 from the cluster fault management module 113.

[0185] It should be understood that the cluster fault management module 113 can also report the chip upgrade information of the AI ​​chip 332 to the cluster device management module 111. The scheduling process of the cluster system in the chip upgrade scenario is basically the same as that in the chip fault scenario, and will not be repeated here.

[0186] 1320, the cluster device management module 111 determines whether the current failure of AI chip 332 affects the AI ​​cluster system's execution of AI calculations based on the fault level. If it does, it selects an idle backup AI chip to replace the faulty AI chip. For example, the cluster device management module 111 can send a second message to the computer device 400 to instruct the computer device 400 to select a backup AI chip to replace the faulty AI chip 332.

[0187] 1330, computer device 400 receives the second message, and AI chip management module 411 selects an idle AI chip 432 in computer device 400 as a backup AI chip to replace the faulty AI chip 332.

[0188] 1340, The routing management submodule 412 of the AI ​​chip management module 411 changes the AI ​​model training list 422 through the bus switching routing driver 418.

[0189] For the architecture diagram shown in Figure 10, by configuring the AI ​​model training list 422, the mapping relationship between port 44 and the address space of AI chip 432 is deleted, and a mapping relationship between port 43 and the address space of AI chip 432 is established. Optionally, a mapping relationship between port 44 and the address space of backup AI chip 432 can also be established.

[0190] For the architecture diagram shown in Figure 11, by configuring the AI ​​model training list 422, the mapping relationship between port 44 and the address space of AI chip 432 is deleted, and the mapping relationship between port 43 and the address space of AI chip 432 is established.

[0191] 1350, the cluster device management module 111 sends a first message to the computer device 300, instructing the computer device 300 to replace the faulty AI chip 332.

[0192] 1360, The AI ​​chip management module 311 of computer device 300 changes the AI ​​model training list 322 via bus switching routing driver 318.

[0193] For the architecture diagram shown in Figure 10, by configuring the AI ​​model training list 322, the mapping relationship between port 34 and the address space of AI chip 332 is deleted, and a mapping relationship between port 34 and the address space of backup AI chip 332 is established, as are the mapping relationships between port 33 and the address space of backup AI chip 332, and the address space of AI chip 432. Optionally, the mapping relationship between port 32 and the address space of AI chip 332 can also be deleted. It should be understood that establishing the mapping relationship between port 33 and the address space of AI chip 432 can include establishing the mapping relationship between port 33 and port 51, port 51 and port 52, port 52 and port 43, and the mapping relationship between port 43 and the address space of AI chip 432. It should be understood that both the address space of backup AI chip 332 and the address space of AI chip 432 indicate the address range of AI chip 432, but the prefixes they have may be different.

[0194] For the architecture diagram shown in Figure 11, by configuring the AI ​​model training list 322, the mapping relationship between port 32 and the address space of AI chip 332 is deleted, the mapping relationship between port 33 and the address space of AI chip 332 is established, and the mapping relationship between port 33 and the address space of AI chip 432 is established.

[0195] Figure 14 is a schematic diagram of another AI chip recovery method provided in an embodiment of this application. Figure 14 corresponds to step 1240 in Figure 12.

[0196] 1410. After the faulty AI chip 332 returns to normal, it can be added back to the original computer device 300 and AI cluster system. The cluster fault management module 113 reports the information that the AI ​​chip 332 has returned to normal to the cluster device management module 111, and the cluster device management module 111 obtains the information that the AI ​​chip 332 has returned to normal from the cluster fault management module 113.

[0197] 1420, the cluster device management module 111 sends a fifth message to the computer device 300, instructing the computer device 300 to resume management of the AI ​​chip 332.

[0198] 1430, The AI ​​chip management module 311 of the computer device 300 changes the AI ​​model training list 322 via bus switching routing driver 318.

[0199] For the architecture diagram shown in Figure 10, the mapping relationship between port 34 and the address space of backup AI chip 332 can be deleted by configuring AI model training list 322, and the mapping relationship between port 34 and the address space of backup AI chip 332 can be established. Optionally, the mapping relationship between port 33 and the address space of backup AI chip 332 can also be deleted, as well as the mapping relationship between port 33 and the address space of AI chip 432.

[0200] For the architecture diagram shown in Figure 11, the mapping relationship between port 32 and the address space of AI chip 332 can be established by configuring AI model training list 322, and the mapping relationship between port 33 and the address space of AI chip 332 can be deleted. Optionally, the mapping relationship between port 33 and the address space of AI chip 432 can also be deleted.

[0201] 1440, the cluster device management module 111 sends a sixth message to the computer device 400, instructing the computer device 400 to resume management of the AI ​​chip 432.

[0202] 1450, The AI ​​chip management module 411 of computer device 400 changes the AI ​​model training list 422 via bus switching routing driver 418.

[0203] For the architecture diagram shown in Figure 10, the mapping relationship between port 43 and the address space of AI chip 432 can be deleted by configuring AI model training list 422, and the mapping relationship between port 44 and the address space of AI chip 432 can be established. Optionally, the mapping relationship between port 44 and the address space of backup AI chip 432 can also be deleted.

[0204] For the architecture diagram shown in Figure 11, the mapping relationship between port 43 and the address space of AI chip 432 can be deleted by configuring the AI ​​model training list 422, and the mapping relationship between port 44 and the address space of AI chip 432 can be established.

[0205] 1460. The AI ​​chip 332 is initialized through the AI ​​chip driver 316. If the backup AI chip 432 replaces the device ID of the AI ​​chip 332, the device ID of the backup AI chip 432 can be deleted in the computer device 300 to restore the device ID of the AI ​​chip 332. Alternatively, a new device ID can be reassigned to the AI ​​chip 332.

[0206] 1470. In computer device 400, the backup AI chip 432 is initialized through AI chip driver 416, and the device ID of the backup AI chip 432 is restored. Alternatively, a new device ID can be reassigned to the AI ​​chip 432.

[0207] Figure 15 is an exemplary flowchart of another scheduling method for a cluster system provided in an embodiment of this application.

[0208] The cluster system includes a cluster management device 100 and at least one computer device, each computer device including at least one AI chip, and a first computer device 500 is one of the at least one computer device.

[0209] 1510, the cluster management device 100 sends a first message to the first computer device 500, the first message being used to instruct the replacement of the first AI chip in the first computer device 500.

[0210] The cluster management device 100 can be the cluster management device 100 shown in Figure 3, Figure 4, Figure 10 or Figure 11, and the first computer device 500 can be the computer device 200 in Figure 3 or Figure 4, or the computer device 300 or computer device 400 in Figure 10 or Figure 11.

[0211] The first AI chip can be a faulty AI chip in the first computer device 500, or it can be an AI chip that needs to be upgraded. For example, if the first computer device 500 is computer device 200 in Figure 3 or 4, the first AI chip can be AI chip 231 or AI chip 232. If the first computer device 500 is computer device 300 in Figure 10 or 11, the first AI chip can be AI chip 331 or AI chip 332. If the first computer device 500 is computer device 400 in Figure 10 or 11, the first AI chip can be AI chip 431 or AI chip 432.

[0212] 1520, the first computer device 500 receives the first message, deletes the mapping relationship between the first port and the first AI chip in the first AI model training list, and establishes the mapping relationship between the second port and the second AI chip.

[0213] The first port includes a port in the first computer device that communicates with the first AI chip, and the second port includes a port in the first computer device that can communicate with the second AI chip. The first AI model training list is stored in the first computer device 500. The first port and the second port can be the same port or different ports. The first AI chip and the second AI chip can belong to the same computer device or different computer devices.

[0214] In one possible implementation scenario, the first computer device 500 is the computer device 200 in Figure 3 or Figure 4, the first AI chip is AI chip 232, and the second AI chip is AI chip 231. The first AI model training list is AI model training list 222, the first port includes any one or more of port 24 and port 22, and the second port includes any one or more of port 24 and port 21. After receiving the first message, the computer device 200 can delete the mapping relationship between port 24 and the address space of AI chip 232 in the AI ​​model training list 222 and establish a mapping relationship between port 24 and the address space of AI chip 231; alternatively, it can also delete the mapping relationship between port 22 and the address space of AI chip 232 and establish a mapping relationship between port 21 and the address space of AI chip 231.

[0215] In another possible implementation scenario, the first computer device 500 is the computer device 300 in Figure 10 or Figure 11, the first AI chip is AI chip 332, and the second AI chip is AI chip 432. The first AI model training list is AI model training list 322, the first port includes any one or more of port 34 and port 32, and the second port includes port 33. After receiving the first message, the computer device 300 can delete the mapping relationship between port 32 and the address space of AI chip 332 in AI model training list 322, establish the mapping relationship between port 33 and the address space of AI chip 332, and establish the mapping relationship between port 33 and the address space of AI chip 432.

[0216] At 1530, the cluster management device 100 sends model training data to the AI ​​chips in the first AI model training list.

[0217] In one possible implementation scenario, the first computer device 500 is the computer device 200 in Figure 3 or Figure 4, the first AI chip is AI chip 232, the second AI chip is AI chip 231, the first AI model training list is AI model training list 222, and the cluster management device 100 sends model training data to AI chip 231 for AI chip 231 to participate in the training of AI models.

[0218] In another possible implementation scenario, the first computer device 500 is the computer device 300 in Figure 10 or Figure 11, the first AI chip is AI chip 332, the second AI chip is AI chip 432, the first AI model training list is AI model training list 322, and the cluster management device 100 sends model training data to AI chip 432 for AI chip 432 to participate in the training of AI models.

[0219] Figure 16 is an exemplary flowchart of another scheduling method for a cluster system provided in an embodiment of this application.

[0220] Steps 1501 to 1505 may be included before step 1510. The first AI chip and the second AI chip belong to different computer devices; for example, the first AI chip belongs to the first computer device 500, and the second AI chip belongs to the second computer device 600.

[0221] 1501, Cluster management device 100 receives a third message, which indicates that the first AI chip in the first computer device 500 has malfunctioned or needs to be upgraded.

[0222] The third step can be found in step 710 of Figure 7, step 810 of Figure 8, step 1210 of Figure 12, and step 1310 of Figure 13. This application will not repeat the details.

[0223] 1503, the cluster management device 100 sends a second message to the second computer device 600, which instructs the second computer device 600 to select a backup AI chip to replace the first AI chip.

[0224] Optionally, the cluster management device 100 can specify the use of a second AI chip to replace the first AI chip, or the second computer device 600 can select an idle backup AI chip to replace the first AI chip.

[0225] 1505, the second computer device 600 receives the second message, deletes the mapping relationship between the third port and the second AI chip in the second AI model training list, and establishes a mapping relationship between the fourth port and the second AI chip. The third port includes the port in the second computer device 600 that communicates with the second AI chip, and the fourth port includes the port in the second computer device 600 that can communicate with the first computer device 500. The second AI model training list is stored in the second computer device 600.

[0226] In one possible implementation scenario, the first computer device 500 is computer device 300 in Figure 10 or Figure 11, and the second computer device 600 is computer device 400 in Figure 10 or Figure 11. The first AI chip is AI chip 332, and the second AI chip is AI chip 432. The first AI model training list is AI model training list 322, and the second AI model training list is AI model training list 422. The third port includes any one or more of ports 44 and 42, and the fourth port includes port 43. After receiving the second message, computer device 400 can delete the mapping relationship between port 44 and the address space of AI chip 432 in AI model training list 422 and establish a mapping relationship between port 43 and the address space of AI chip 432; alternatively, it can also delete the mapping relationship between port 42 and the address space of AI chip 432 and establish a mapping relationship between port 43 and the address space of AI chip 432.

[0227] 1510, the cluster management device 100 sends a first message to the first computer device 500, the first message being used to instruct the replacement of the first AI chip in the first computer device 500.

[0228] The cluster management device 100 can be the cluster management device 100 shown in Figure 3, Figure 4, Figure 10 or Figure 11, the first computer device 500 can be the computer device 300 in Figure 10 or Figure 11, and the second computer device 600 can be the computer device 400 in Figure 10 or Figure 11.

[0229] The first AI chip can be a malfunctioning AI chip in the first computer device 500, or it can be an AI chip that needs to be upgraded. For example, the first AI chip can be AI chip 331 or AI chip 332, and the second AI chip can be AI chip 431 or AI chip 432.

[0230] 1520, the first computer device 500 receives the first message, deletes the mapping relationship between the first port and the first AI chip in the first AI model training list, and establishes the mapping relationship between the second port and the second AI chip.

[0231] The first port includes the port in the first computer device that communicates with the first AI chip, and the second AI chip is a backup AI chip. The second port includes the port in the first computer device that can communicate with the second AI chip, that is, the second port can communicate with the second computer device 600, and the first AI model training list is stored in the first computer device 500.

[0232] For example, the first computer device 500 is the computer device 300 in Figure 10 or Figure 11, the first AI chip is AI chip 332, and the second AI chip is AI chip 432. The first AI model training list is AI model training list 322, the first port includes any one or more of port 34 and port 32, and the second port includes port 33. After receiving the first message, the computer device 300 can delete the mapping relationship between port 32 and the address space of AI chip 332 in the AI ​​model training list 322, establish the mapping relationship between port 33 and the address space of AI chip 332, and establish the mapping relationship between port 33 and the address space of AI chip 432.

[0233] Figure 17 is an exemplary flowchart of another scheduling method for a cluster system provided in an embodiment of this application.

[0234] At 1530, the cluster management device 100 receives a fourth message, which is used to instruct the first AI chip to return to normal.

[0235] The cluster management device 100 can be the cluster management device 100 shown in Figure 3, Figure 4, Figure 10 or Figure 11, and the first computer device 500 can be the computer device 200 in Figure 3 or Figure 4, or the computer device 300 or computer device 400 in Figure 10 or Figure 11.

[0236] The first AI chip can be a faulty AI chip in the first computer device 500, or it can be an AI chip that needs to be upgraded. For example, if the first computer device 500 is computer device 200 in Figure 3 or 4, the first AI chip can be AI chip 231 or AI chip 232. If the first computer device 500 is computer device 300 in Figure 10 or 11, the first AI chip can be AI chip 331 or AI chip 332. If the first computer device 500 is computer device 400 in Figure 10 or 11, the first AI chip can be AI chip 431 or AI chip 432.

[0237] 1540, the cluster management device 100 sends a fifth message to the first computer device 500, which is used to instruct the resumption of management of the first AI chip.

[0238] The fifth message can be found in step 920 of Figure 9 and step 1420 of Figure 14, and will not be repeated here.

[0239] 1550, the first computer device 500 receives the fifth message and establishes a mapping relationship between the first port and the first AI chip in the first AI model training list.

[0240] In one possible implementation scenario, the first computer device 500 is the computer device 200 in Figure 3 or Figure 4, the first AI chip is AI chip 232, and the second AI chip is AI chip 231. The first AI model training list is AI model training list 222, the first port includes any one or more of port 24 and port 22, and the second port includes any one or more of port 24 and port 21. After receiving the fifth message, the computer device 200 can establish a mapping relationship between port 24 and the address space of AI chip 232 in the AI ​​model training list 222, or it can also establish a mapping relationship between port 22 and the address space of AI chip 232.

[0241] In another possible implementation scenario, the first computer device 500 is the computer device 300 in Figure 10 or Figure 11, the first AI chip is AI chip 332, and the second AI chip is AI chip 432. The first AI model training list is AI model training list 322, the first port includes any one or more of port 34 and port 32, and the second port includes port 33. After receiving the fifth message, the computer device 300 can establish a mapping relationship between port 34 and the address space of AI chip 332 in the AI ​​model training list 322, or establish a mapping relationship between port 32 and the address space of AI chip 332.

[0242] 1560. In some possible implementations, the first AI chip and the second AI chip belong to different computer devices; the first AI chip belongs to the first computer device 500, and the second AI chip belongs to the second computer device 600. The cluster management device 100 sends a sixth message to the second computer device 600, which instructs the second computer device 600 to resume management of the second AI chip.

[0243] For example, the first computer device 500 is computer device 300 in Figure 10 or Figure 11, and the second computer device 600 is computer device 400 in Figure 10 or Figure 11. The first AI chip is AI chip 332, and the second AI chip is AI chip 432. The cluster management device 100 sends a sixth message to the computer device 400, which instructs the computer device 400 to resume management of the AI ​​chip 432.

[0244] 1570, the second computer device 600 receives the sixth message, establishes a mapping relationship between the third port and the second AI chip in the second AI model training list, and deletes the mapping relationship between the fourth port and the second AI chip. The third port includes the port in the second computer device 600 that communicates with the second AI chip, and the fourth port includes the port in the second computer device 600 that can communicate with the first computer device 500. The second AI model training list is stored in the second computer device 600.

[0245] In one possible implementation scenario, the first computer device 500 is computer device 300 in Figure 10 or Figure 11, and the second computer device 600 is computer device 400 in Figure 10 or Figure 11. The first AI chip is AI chip 332, and the second AI chip is AI chip 432. The first AI model training list is AI model training list 322, and the second AI model training list is AI model training list 422. The third port includes any one or more of ports 44 and 42, and the fourth port includes port 43. After receiving the sixth message, computer device 400 can establish a mapping relationship between port 44 and the address space of AI chip 432 in AI model training list 422, and delete the mapping relationship between port 43 and the address space of AI chip 432; alternatively, it can also establish a mapping relationship between port 42 and the address space of AI chip 432, and delete the mapping relationship between port 43 and the address space of AI chip 432.

[0246] This application also provides a computer-readable storage medium storing computer program instructions that, when executed by the cluster system shown in Figures 3, 4, 10, or 11, cause the cluster system to perform the method described in Figures 7-9 and 12-14.

[0247] Those skilled in the art will recognize that the modules and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0248] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0249] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.

[0250] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0251] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0252] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0253] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A scheduling method for a cluster system, characterized in that, The cluster system includes a cluster management device and at least one computer device, each computer device including at least one artificial intelligence (AI) chip, and the method includes: The cluster management device sends a first message to a first computer device, the first message being used to instruct the replacement of the first AI chip in the first computer device, the first computer device being one of the at least one computer devices; The first computer device receives the first message, deletes the mapping relationship between the first port and the first AI chip in the first AI model training list, and establishes a mapping relationship between the second port and the second AI chip. The first port includes the port in the first computer device that communicates with the first AI chip, and the second port includes the port in the first computer device that can communicate with the second AI chip. The first AI model training list is stored in the first computer device. The cluster management device sends model training data to the AI ​​chips in the first AI model training list.

2. The method according to claim 1, characterized in that, The second AI chip belongs to a second computer device, which is one of the at least one computer device, and the method further includes: The first computer device establishes a mapping relationship between the second port and the first AI chip in the first AI model training list; The cluster management device sends a second message to the second computer device, the second message being used to instruct the second computer device to select a backup AI chip to replace the first AI chip, the second AI chip being the backup AI chip; The second computer device receives the second message, deletes the mapping relationship between the third port and the second AI chip in the second AI model training list, and establishes a mapping relationship between the fourth port and the second AI chip. The third port includes the port in the second computer device that communicates with the second AI chip, and the fourth port includes the port in the second computer device that can communicate with the first computer device. The second AI model training list is stored in the second computer device.

3. The method according to claim 1 or 2, characterized in that, Before the cluster management device sends the first message to the first computer device, the method further includes: The cluster management device receives a third message, which indicates that the first AI chip has malfunctioned or that the first AI chip needs to be upgraded.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The first computer device assigns a device identifier ID to the second AI chip, and the device ID of the second AI chip may be the same as or different from the device ID of the first AI chip.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: The cluster management device receives a fourth message, which is used to instruct the first AI chip to return to normal. The cluster management device sends a fifth message to the first computer device, the fifth message being used to instruct the resumption of management of the first AI chip; The first computer device receives the fifth message and establishes a mapping relationship between the first port and the first AI chip in the first AI model training list.

6. The method according to claim 5, characterized in that, The method further includes: The first computer device removes the mapping relationship between the second port and the second AI chip from the first AI model training list.

7. The method according to claim 6, characterized in that, The second AI chip belongs to a second computer device, which is one of the at least one computer device, and the method further includes: The cluster management device sends a sixth message to the second computer device, the sixth message being used to instruct the second computer device to resume management of the second AI chip; The second computer device receives the sixth message, establishes a mapping relationship between the third port and the second AI chip in the second AI model training list, and deletes the mapping relationship between the fourth port and the second AI chip. The third port includes the port in the second computer device that communicates with the second AI chip, and the fourth port includes the port in the second computer device that can communicate with the first computer device. The second AI model training list is stored in the second computer device.

8. The method according to any one of claims 1 to 7, characterized in that, The first AI model training list includes a first AI chip address space and a first backup AI chip address space. The first AI chip address space is used to indicate the address range of the first AI chip, and the first backup AI chip address space is used to indicate the address range of the second AI chip.

9. A cluster system, characterized in that, It includes cluster management equipment and at least one computer device, each computer device including at least one artificial intelligence (AI) chip; The cluster management device is used to send a first message to a first computer device, the first message being used to instruct the replacement of a first AI chip in the first computer device, the first computer device being one of the at least one computer device; The first computer device is used to receive the first message, delete the mapping relationship between the first port and the first AI chip in the first AI model training list, and establish the mapping relationship between the second port and the second AI chip. The first port includes the port in the first computer device that communicates with the first AI chip, and the second port includes the port in the first computer device that can communicate with the second AI chip. The first AI model training list is stored in the first computer device. The cluster management device is also used to send model training data to the AI ​​chips in the first AI model training list.

10. The system according to claim 9, characterized in that, The second AI chip belongs to a second computer device, which is one of the at least one computer devices; The first computer device is also used to establish a mapping relationship between the second port and the first AI chip in the first AI model training list; The cluster management device is further configured to send a second message to the second computer device, the second message being configured to instruct the second computer device to select a backup AI chip to replace the first AI chip, the second AI chip being the backup AI chip; The second computer device is used to receive the second message, delete the mapping relationship between the third port and the second AI chip in the second AI model training list, and establish the mapping relationship between the fourth port and the second AI chip. The third port includes the port in the second computer device that communicates with the second AI chip, and the fourth port includes the port in the second computer device that can communicate with the first computer device. The second AI model training list is stored in the second computer device.

11. The system according to claim 9 or 10, characterized in that, The cluster management device is also used to receive a third message, which indicates that the first AI chip has malfunctioned or that the first AI chip needs to be upgraded.

12. The system according to any one of claims 9 to 11, characterized in that, The first computer device is also used to assign a device identifier ID to the second AI chip, wherein the device ID of the second AI chip may be the same as or different from the device ID of the first AI chip.

13. The system according to any one of claims 9 to 12, characterized in that, The cluster management device is also used to receive a fourth message, which is used to instruct the first AI chip to return to normal. The cluster management device is also used to send a fifth message to the first computer device, the fifth message being used to instruct the resumption of management of the first AI chip; The first computer device is further configured to receive the fifth message and establish a mapping relationship between the first port and the first AI chip in the first AI model training list.

14. The system according to claim 13, characterized in that, The first computer device is also used to remove the mapping relationship between the second port and the second AI chip from the first AI model training list.

15. The system according to claim 14, characterized in that, The second AI chip belongs to a second computer device, which is one of the at least one computer devices; The cluster management device is also used to send a sixth message to the second computer device, the sixth message being used to instruct the second computer device to resume management of the second AI chip; The second computer device is used to receive the sixth message, establish a mapping relationship between the third port and the second AI chip in the second AI model training list, and delete the mapping relationship between the fourth port and the second AI chip. The third port includes the port in the second computer device that communicates with the second AI chip, and the fourth port includes the port in the second computer device that can communicate with the first computer device. The second AI model training list is stored in the second computer device.

16. The system according to any one of claims 9 to 15, characterized in that, The first AI model training list includes a first AI chip address space and a first backup AI chip address space. The first AI chip address space is used to indicate the address range of the first AI chip, and the first backup AI chip address space is used to indicate the address range of the second AI chip.

17. A computer-readable storage medium, characterized in that, Includes computer program instructions, which, when executed by the cluster system according to any one of claims 9 to 16, enable the cluster system to perform the method according to any one of claims 1 to 8.