Fault Handling Method, Device, Electronic Device and Storage Medium for Heterogeneous Acceleration Card
By identifying and replacing the failed heterogeneous accelerator card in the integrated architecture cabinet, the problem of fault handling in the existing technology affecting the operation of the entire cabinet business is solved, and rapid and impactless fault handling and business recovery are achieved.
Patent Information
- Application Number
- CN202510332602.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-20
AI Technical Summary
When the heterogeneous accelerator card equipment in the integrated cabinet fails, the existing technology needs to power off the entire heterogeneous accelerator card equipment, affecting the business operation of all computing resource pools in the cabinet and bringing a greater business impact to users.
By determining whether the power-down instruction of the faulty heterogeneous acceleration card has been received, if received, the heterogeneous acceleration card to be removed is removed when the heterogeneous acceleration card to be removed is in the assigned state, and the target heterogeneous acceleration card in the unassigned state is determined from at least one heterogeneous acceleration card device of the integrated architecture cabinet, the target heterogeneous acceleration card to be removed is replaced with the target heterogeneous acceleration card, and the target heterogeneous acceleration card is allocated to the computing resource pool corresponding to the heterogeneous acceleration card to be removed.
It realizes that the replacement of the fault accelerator card is completed while minimizing the impact on the business functions of the entire system as much as possible, and the normal operation of the business is restored in a timely manner, reducing the maintenance pressure of operation and maintenance personnel.
Smart Images

Figure CN119883750B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a method, apparatus, electronic device, and storage medium for handling faults of heterogeneous acceleration cards. Background Art
[0002] Traditional computing-centered architectures have difficulty meeting the requirements of emerging application scenarios such as artificial intelligence, machine learning, and intelligent computing for high performance, high flexibility, and high scalability. Therefore, the data center architecture is undergoing a profound transformation, that is, a transformation from a computing-centered architecture to a data-centered converged architecture.
[0003] In the design of the converged architecture, resource decoupling has become a key design principle. By decomposing the entire system into multiple independent pooled modules, such as a computing resource pool, a heterogeneous acceleration card resource pool, and an IO resource pool, etc., flexible configuration and efficient utilization of resources can be achieved. These pooled modules usually exist in the form of Boxes, making the expansion and upgrade of the system more convenient.
[0004] In a design of a converged architecture whole cabinet for IO expansion capabilities, the system is cleverly divided into an IOBox and multiple device Boxes. The IO Box, as the centralized management point of IO resources, is responsible for handling all IO-related operations, such as data transmission, device identification, and resource allocation, etc. The device Box contains multiple Host Boxes and heterogeneous acceleration card Boxes, which are used to provide computing resources and heterogeneous acceleration capabilities respectively.
[0005] However, although this converged architecture performs well in resource allocation and scalability, there are still some challenges in the maintenance and management of heterogeneous acceleration card resources. Especially when a heterogeneous acceleration card fails and needs to be replaced, how to quickly repair and replace the faulty device without affecting the business functions of the entire system has become an urgent problem to be solved. Summary of the Invention
[0006] The present invention provides a method, apparatus, electronic device, and storage medium for handling faults of heterogeneous acceleration cards, so as to at least solve the technical problem that when a heterogeneous acceleration card device in a converged architecture whole cabinet fails, the existing processing method requires powering down the entire heterogeneous acceleration card device, affecting the business operations of all computing resource pools in the whole cabinet and bringing a great business impact to users.
[0007] The present invention provides a method for fault handling of a heterogeneous acceleration card. The method includes: determining whether a power-down instruction for a faulty heterogeneous acceleration card is received, where the power-down instruction for the faulty heterogeneous acceleration card includes the heterogeneous acceleration card to be removed; if the power-down instruction for the faulty heterogeneous acceleration card is received, then in the case where the heterogeneous acceleration card to be removed is in an allocated state, removing the heterogeneous acceleration card to be removed, and determining a target heterogeneous acceleration card in an unallocated state from at least one heterogeneous acceleration card device in the integrated cabinet of the fusion architecture; using the target heterogeneous acceleration card to replace the heterogeneous acceleration card to be removed, and allocating the target heterogeneous acceleration card to the computing resource pool corresponding to the heterogeneous acceleration card to be removed.
[0008] The present invention also provides a device for fault handling of a heterogeneous acceleration card, including: a determination module, configured to determine whether a power-down instruction for a faulty heterogeneous acceleration card is received, where the power-down instruction for the faulty heterogeneous acceleration card includes the heterogeneous acceleration card to be removed; a fault handling module, configured to, if the power-down instruction for the faulty heterogeneous acceleration card is received, then in the case where the heterogeneous acceleration card to be removed is in an allocated state, remove the heterogeneous acceleration card to be removed, and determine a target heterogeneous acceleration card in an unallocated state from at least one heterogeneous acceleration card device in the integrated cabinet of the fusion architecture; a replacement module, configured to use the target heterogeneous acceleration card to replace the heterogeneous acceleration card to be removed, and allocate the target heterogeneous acceleration card to the computing resource pool corresponding to the heterogeneous acceleration card to be removed.
[0009] The present invention also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above methods for fault handling of a heterogeneous acceleration card when executing the computer program.
[0010] The present invention also provides a computer-readable storage medium, in which a computer program is stored, where the computer program implements the steps of any of the above methods for fault handling of a heterogeneous acceleration card when executed by a processor.
[0011] Through the present invention, it is determined whether a power-down instruction for a faulty heterogeneous acceleration card is received, where the power-down instruction for the faulty heterogeneous acceleration card includes the heterogeneous acceleration card to be removed. If the power-down instruction for the faulty heterogeneous acceleration card is received, then when the heterogeneous acceleration card to be removed is in an allocated state, the heterogeneous acceleration card to be removed is removed, and a target heterogeneous acceleration card in an unallocated state is determined from at least one heterogeneous acceleration card device in the integrated cabinet of the fusion architecture. The target heterogeneous acceleration card is used to replace the heterogeneous acceleration card to be removed, and the target heterogeneous acceleration card is allocated to the computing resource pool corresponding to the heterogeneous acceleration card to be removed. Therefore, the technical problem that when a heterogeneous acceleration card device in the integrated cabinet of the fusion architecture fails, the existing processing method requires powering down the entire heterogeneous acceleration card device, affecting the business operations of all computing resource pools in the integrated cabinet and bringing a great business impact to users is solved, and the technical effect of being able to complete the replacement operation of the faulty acceleration card while minimizing the impact on the business functions of the entire system and promptly restoring the normal operation of the business and reducing the maintenance pressure on the operation and maintenance personnel is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0013] Figure 1 Schematic flowchart of a method for handling faults of a heterogeneous acceleration card provided by an embodiment of the present invention;
[0014] Figure 2 Schematic diagram of the hardware architecture of a fusion architecture system according to an embodiment of the present invention;
[0015] Figure 3 Schematic diagram of the power supply structure of a heterogeneous acceleration card according to an embodiment of the present invention;
[0016] Figure 4 Schematic flowchart of a method for handling faults of a heterogeneous acceleration card according to an embodiment of the present invention;
[0017] Figure 5 Schematic diagram of a device for handling faults of a heterogeneous acceleration card provided by an embodiment of the present invention;
[0018] Figure 6 Schematic diagram of the structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0020] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0021] Before specifically introducing the fault handling method of the heterogeneous acceleration card of the present invention, a brief introduction is made to the existing technology handling method when a heterogeneous acceleration card fails and needs to be replaced. The resources of the heterogeneous acceleration card are mainly managed by the Pooled System Management Controller (PSMC) in the IO Box for resource allocation and fault recording. When a heterogeneous acceleration card fails, the PSMC can determine the location of the faulty device by querying the allocation relationship table according to the fault information reported by the Host and record the fault log.
[0022] However, this solution has obvious limitations in the process of replacing faulty devices. Since the entire heterogeneous acceleration card Box needs to be powered off for replacement, this will cause the business operations of all Host resource pools in the entire cabinet to be interrupted, bringing a greater business impact to users.
[0023] Therefore, to solve this problem, it is necessary to develop a more efficient and reliable heterogeneous acceleration card resource maintenance and management solution to achieve rapid location, repair and replacement of faulty devices, while ensuring that the business continuity of the entire system is not affected.
[0024] Specifically, the embodiments of the present invention provide a fault handling method for a heterogeneous acceleration card, as Figure 1 shown.
[0025] The embodiments of the present invention are a fusion architecture system for enhancing heterogeneous computing capabilities. The structure of the fusion architecture system is as Figure 2As shown in the figure, there is 1 IO resource pool, 2 heterogeneous acceleration card devices (i.e., heterogeneous acceleration card BOX0 and heterogeneous acceleration card BOX1), and a Host resource pool (including Host Box0, Host Box1, Host Box2, Host Box3, Host Box4, Host Box5, Host Box6, Host Box7, Host Box8) in the integrated architecture whole cabinet. Each Box is designed with 1 BMC (Baseboard Management Controller) module, which is responsible for the management of the corresponding Box. All BMCs are connected under the same network switch to form an internal local area network. The BMC in the IO Box serves as the pooled management controller PSMC. The PSMC can communicate with the BMCs of each device Box through the network to implement the management and control functions of the entire system.
[0026] The IO resource pool contains 4 layers of PCIe Switch boards. Each layer of board has two PCIe Switch chips, which expand, network, and allocate the PCIe resources of multiple CPUs in the HostBox. The PCIe Switch chip is an IO chip used to expand a group of PCIe signals into multiple groups of PCIe signals for expanding the CPU IO resources. Each PCIe Switch chip leads out 5 CDFP ports (high-speed signal ports). The entire IO resource pool can finally provide 40 (8 * 5) CDFP ports externally. Each device Box is also configured with corresponding CDFP ports, and the device Box is connected to the IO resource pool through Cable cables.
[0027] The PSMC interacts with the BMCs and PCIe Switch firmware of each device Box, summarizes each mapping relationship, and then establishes an allocation relationship table between the heterogeneous acceleration card BOX and the Host Box. Through this allocation relationship table, it can clearly show which heterogeneous acceleration cards in which heterogeneous acceleration card Boxes are allocated under a Host resource pool. The allocation relationship of the embodiment of the present invention is as follows:
[0028] {
[0029] "HostBox_ID": "1",
[0030] "Switch_CDFP_ID": "CDFP0-1",
[0031] "Switch_Port": "0x000040",
[0032] "Device":
[0033] {
[0034] "AcceleratorCardBox_ID": "0",
[0035] "Slot": "AcceleratorCard0",
[0036] "Switch_CDFP_ID": "CDFP1-1",
[0037] "Switch_Port": "1P96",
[0038] "Box_CDFP_ID": "CDFP0",
[0039] "BusNumber": XX,
[0040] "DeviceNumber": XX,
[0041] "FunctionNumber": XX,
[0042] "AssetInfo": {
[0043] "BiosVersion": "XX",
[0044] "BuildDate": "XX",
[0045] "DeviceID": "XX",
[0046] "FwVersion": "XX",
[0047] "Manufacturer": "XX",
[0048] "Model": "XX",
[0049] "PartNumber": "XX",
[0050] "SerialNumber": "XX",
[0051] "SubSystemID": "XX",
[0052] "UUID": "XX",
[0053] "vendorID": "XX"
[0054] }
[0055] }
[0057] }
[0058] The PSMC provides a web page and a Redfish interface for allocating and releasing heterogeneous acceleration card resources. After that, the PSMC updates the allocation relationship table between the heterogeneous acceleration cards and the Host BOX in the heterogeneous acceleration card device to ensure the accuracy of the allocation relationship.
[0059] In step S101, it is determined whether a power-off instruction for a faulty heterogeneous acceleration card is received. The power-off instruction for the faulty heterogeneous acceleration card includes the heterogeneous acceleration card to be removed.
[0060] Optionally, in some embodiments, the power-off instruction for the faulty heterogeneous acceleration card further includes the first number of the heterogeneous acceleration card to be removed and the second number of the heterogeneous acceleration card device to which the heterogeneous acceleration card to be removed belongs.
[0061] The system first detects whether there is an input of a power-off instruction for a faulty heterogeneous acceleration card. The operation and maintenance personnel send a power-off instruction for the faulty heterogeneous acceleration card to the PSMC through IPMI (Intelligent Platform Management Interface) commands, Redfish interfaces or web pages.
[0062] Specifically, when the system detects that a certain heterogeneous acceleration card fails, the operation and maintenance personnel confirm the heterogeneous acceleration card to be removed and issue a power-off instruction for the faulty heterogeneous acceleration card. After receiving the power-off instruction for the faulty heterogeneous acceleration card, the integrated chassis of the fusion architecture will parse it and extract key information, including the heterogeneous acceleration card to be removed, the first number of the heterogeneous acceleration card to be removed, and the second number of the heterogeneous acceleration card device to which it belongs.
[0063] Among them, the first number and the second number can be obtained by viewing the fault log recorded by the PSMC or other means. Obtained through the fault log method: When the basic input / output system BIOS detects a PCIe fault of the heterogeneous acceleration card, it will report to the PSMC to record the fault log. The BMC of the heterogeneous acceleration card Box reports to the PSMC to record the fault log after detecting a fault of the heterogeneous acceleration card through an out-of-band method, or obtained through other means: The in-band monitoring under the host operating system Host OS detects a fault of the heterogeneous acceleration card, etc., so as to obtain the first number of the heterogeneous acceleration card to be removed and the second number of the heterogeneous acceleration card device to which it belongs.
[0064] Through the above technical solution, by determining the first number of the heterogeneous acceleration card to be removed, the system can accurately identify and locate the heterogeneous acceleration card to be powered off, avoiding the risk of accidentally powering off other normally working acceleration cards due to misoperation. During the fault troubleshooting and recovery process, the detailed power-off instruction information can help technicians locate the faulty heterogeneous acceleration card faster and take appropriate measures for repair.
[0065] In step S102, if a power-down instruction for a faulty heterogeneous accelerator card is received, when the heterogeneous accelerator card to be removed is in an allocated state, the heterogeneous accelerator card to be removed is removed, and a target heterogeneous accelerator card in an unallocated state is determined from at least one heterogeneous accelerator card device in the integrated cabinet of the fusion architecture.
[0066] It should be understood that the fusion architecture system extracts the information of the heterogeneous accelerator card to be removed according to the received power-down instruction for the faulty heterogeneous accelerator card, and checks the status of the heterogeneous accelerator card to be removed.
[0067] If the heterogeneous accelerator card to be removed is in an allocated state, the integrated cabinet of the fusion architecture removes the heterogeneous accelerator card to be removed from its currently allocated computing resource pool (the computing resource pool is the Figure 2 Host resource pool in it).
[0068] After the heterogeneous accelerator card to be removed is successfully powered down and removed, the integrated cabinet of the fusion architecture determines a target heterogeneous accelerator card in an unallocated state from at least one heterogeneous accelerator card device to replace the heterogeneous accelerator card to be removed.
[0069] Optionally, in some embodiments, after receiving the power-down instruction for the faulty heterogeneous accelerator card, it further includes: based on the first number and the second number, determining whether the heterogeneous accelerator card to be removed has a corresponding allocation relationship with at least one computing resource pool of the integrated cabinet of the fusion architecture; if the heterogeneous accelerator card to be removed has a corresponding allocation relationship with at least one computing resource pool of the integrated cabinet of the fusion architecture, it is determined that the heterogeneous accelerator card to be removed is in an allocated state, otherwise, it is determined that the heterogeneous accelerator card to be removed is in an unallocated state.
[0070] Specifically, the fusion architecture system determines whether the heterogeneous accelerator card to be removed has a corresponding allocation relationship with at least one computing resource pool of the integrated cabinet of the fusion architecture based on the extracted first number and second number.
[0071] If the heterogeneous accelerator card to be removed has a corresponding allocation relationship with at least one computing resource pool, it is determined that the accelerator card is in an allocated state. If the heterogeneous accelerator card to be removed has no corresponding allocation relationship with any computing resource pool, it is determined that the accelerator card is in an unallocated state.
[0072] Through the above technical solution, according to the judgment result, the system can distinguish whether the heterogeneous accelerator card to be removed is in an allocated state or an unallocated state, which helps the system to manage the heterogeneous accelerator card resources more precisely, avoid resource conflicts and waste, and automatically allocate the unallocated heterogeneous accelerator card to the computing resource pool corresponding to the original heterogeneous accelerator card to be removed during operation to meet the real-time changing performance requirements.
[0073] In step S103, the target heterogeneous acceleration card is used to replace the heterogeneous acceleration card to be removed, and the target heterogeneous acceleration card is allocated to the computing resource pool corresponding to the heterogeneous acceleration card to be removed.
[0074] Specifically, the fusion architecture system replaces the position of the heterogeneous acceleration card to be removed with the target heterogeneous acceleration card, and allocates the target heterogeneous acceleration card to the computing resource pool originally corresponding to the heterogeneous acceleration card to be removed, so that the target acceleration card will assume the responsibilities originally borne by the faulty acceleration card and provide acceleration capabilities for relevant computing instances or tasks.
[0075] Optionally, in some embodiments, after allocating the target heterogeneous acceleration card to the computing resource pool corresponding to the heterogeneous acceleration card to be removed, it includes: updating the allocation relationship table of the target heterogeneous acceleration card in the fusion architecture whole cabinet.
[0076] It can be understood that after allocating the target heterogeneous acceleration card to the computing resource pool corresponding to the heterogeneous acceleration card to be removed, the system obtains the allocation relationship table of the heterogeneous acceleration cards in the fusion architecture whole cabinet. This table is usually stored in the internal database or configuration file of the system. The system finds the record of the heterogeneous acceleration card to be removed in the allocation relationship table according to the previous replacement and allocation operations, and replaces it with the information of the target heterogeneous acceleration card, including the number of the target heterogeneous acceleration card, the number of the heterogeneous acceleration card device to which it belongs, etc.
[0077] The fusion architecture system can save the updated allocation relationship table back to the internal database or configuration file to ensure that subsequent operations and scheduling of the system are based on the latest allocation information.
[0078] Through the above technical solution, updating the allocation relationship table ensures the consistency of the management and allocation of heterogeneous acceleration card resources within the system. It helps to avoid resource conflicts, incorrect allocations or omissions, thereby improving the stability and reliability of the system. By accurately recording the allocation situation of each heterogeneous acceleration card, the system can manage and utilize resources more effectively. When an acceleration card fails, the faulty heterogeneous acceleration card can be quickly located and replaced according to the updated allocation relationship table, reducing the fault recovery time and improving the availability and business continuity of the system.
[0079] Optionally, in some embodiments, after receiving the power-down instruction of the faulty heterogeneous acceleration card, it further includes: controlling the heterogeneous acceleration card to be removed to power down when the heterogeneous acceleration card to be removed is in an unallocated state.
[0080] It can be understood that if the heterogeneous acceleration card to be removed is in an unallocated state, the fusion architecture system controls the heterogeneous acceleration card to be removed to power down according to the power-down instruction of the faulty heterogeneous acceleration card.
[0081] Through the above technical solution, it is ensured that the heterogeneous acceleration card to be removed is powered off when it is not used by any computing resource pool, which can avoid data loss or service interruption caused by sudden power failure.
[0082] Optionally, in some embodiments, when controlling the power-off of the heterogeneous acceleration card to be removed, it further includes: determining whether the heterogeneous acceleration card to be removed is successfully powered off; if the heterogeneous acceleration card to be removed is successfully powered off, generating a first power-off success signal to the pooling management controller of the integrated cabinet of the fusion architecture, and using the pooling management controller to send the power-off success information to a preset terminal to remind the operation and maintenance personnel to replace the new heterogeneous acceleration card.
[0083] Specifically, the fusion architecture system monitors the power status of the heterogeneous acceleration card to be removed to detect whether it is successfully powered off. Once it is confirmed that the heterogeneous acceleration card to be removed is successfully powered off, the system generates a first power-off success signal to the pooling management controller of the integrated cabinet of the fusion architecture. After receiving the first power-off success signal, the pooling management controller generates the corresponding power-off success information.
[0084] Then, the pooling management controller sends the power-off success information to a preset terminal (such as the mobile device or workstation of the operation and maintenance personnel) through a preset communication channel (such as email, text message, operation and maintenance management system, etc.). Among them, the power-off success information includes the number of the heterogeneous acceleration card to be removed, the number of the heterogeneous acceleration card device where the heterogeneous acceleration card to be removed is located, the power-off time, and a reminder for the operation and maintenance personnel to replace the new heterogeneous acceleration card.
[0085] After receiving the power-off success information, the operation and maintenance personnel learn that the heterogeneous acceleration card to be removed has been safely powered off, take out the heterogeneous acceleration card to be removed from the integrated cabinet, and install a new heterogeneous acceleration card.
[0086] If the power-off fails, it may be necessary to retry the power-off command, record the error information, notify the administrator, etc.
[0087] Through the above technical solution, it can be ensured that the acceleration card has indeed been safely powered off by determining whether the heterogeneous acceleration card to be removed is successfully powered off. Once it is confirmed that the heterogeneous acceleration card to be removed is successfully powered off, the power-off status can be quickly conveyed to the operation and maintenance personnel without manual inspection or manual reporting. After receiving the power-off success information, the operation and maintenance personnel can immediately know when they can safely remove the faulty acceleration card and replace it with a new one, reducing the waiting time and improving the efficiency and response speed of the operation and maintenance operation.
[0088] Optionally, in some embodiments, after using the pooling management controller to send the power-off success information to a preset terminal, it includes: determining whether the pooling management controller receives a power-on command; if the pooling management controller receives a power-on command, controlling the new heterogeneous acceleration card to be powered on based on the first number and the second number of the new heterogeneous acceleration card.
[0089] It is understandable that once the pooling management controller receives a power-on instruction, which includes the first number of the new heterogeneous acceleration card and the second number of the heterogeneous acceleration card device where it is located, the pooling management controller prepares to execute the power-on operation according to the first number of the new heterogeneous acceleration card and the second number of the heterogeneous acceleration card device where it is located. It should be noted that since the new heterogeneous acceleration card replaces the heterogeneous acceleration card to be removed, therefore, the first number and the second number of the heterogeneous acceleration card to be removed are the same as the first number and the second number of the new heterogeneous acceleration card.
[0090] The pooling management controller identifies the location of the new heterogeneous acceleration card that needs to be powered on according to the first number and the second number in the power-on instruction, and sends a power-on instruction to the new heterogeneous acceleration card. The system monitors the power status of the new heterogeneous acceleration card in real time to ensure that the power-on operation is successfully executed.
[0091] Through the above technical solution, by judging whether the pooling management controller receives a power-on instruction, the system can automatically trigger the power-on process of the new heterogeneous acceleration card, reduce manual intervention, and improve the degree of automation. Once it is confirmed that the new heterogeneous acceleration card needs to be powered on, the system can immediately control its power-on based on the first number and the second number, ensuring that the resources can be quickly restored and reallocated after the faulty acceleration card is removed.
[0092] Optionally, in some embodiments, after controlling the new heterogeneous acceleration card to power on based on the first number and the second number, it includes: detecting whether the new heterogeneous acceleration card has completed power-on; if the new heterogeneous acceleration card has completed power-on, then reset the new heterogeneous acceleration card, and after the reset of the new heterogeneous acceleration card is successful, allocate it to the corresponding computing resource pool.
[0093] It is understandable that the fusion architecture system continuously monitors the power status of the new heterogeneous acceleration card to judge whether its power-on is completed.
[0094] Once it is detected that the new heterogeneous acceleration card has completed power-on, the system prepares to execute the next reset operation, that is, the system sends a reset instruction to the new heterogeneous acceleration card, the system monitors the reset process of the new heterogeneous acceleration card, and detects whether the reset of the new heterogeneous acceleration card is successful.
[0095] Once the reset of the new heterogeneous acceleration card is successful, the system allocates the new heterogeneous acceleration card to the computing resource pool of the fusion architecture whole cabinet.
[0096] Through the above technical solution, by automating the detection of the power-on, reset, and allocation processes, the need for manual intervention is reduced, and the operation and maintenance personnel can rely on the automated processing of the system to ensure the correct deployment and configuration of the new heterogeneous acceleration card, thereby reducing the operation complexity and error rate.
[0097] Optionally, in some embodiments, after allocating the target heterogeneous acceleration card to the computing resource pool corresponding to the heterogeneous acceleration card to be removed, the following steps are further included: controlling the heterogeneous acceleration card to be removed to power off, and determining whether the heterogeneous acceleration card to be removed has powered off successfully; if the heterogeneous acceleration card to be removed has powered off successfully, generating a second power-off success signal to the pooling management controller of the integrated cabinet of the fusion architecture, and using the pooling management controller to send the second power-off success information to a preset terminal to remind the operation and maintenance personnel to replace the new heterogeneous acceleration card.
[0098] It should be understood that after the fusion architecture system has completed the operations of judging the faulty heterogeneous acceleration card, selecting and replacing the target heterogeneous acceleration card, and allocating the target heterogeneous acceleration card to the corresponding resource pool, it sends a power-off instruction to the heterogeneous acceleration card to be removed to control the heterogeneous acceleration card to power off. The integrated cabinet of the fusion architecture judges whether the heterogeneous acceleration card to be removed has powered off successfully by monitoring the power state of the heterogeneous acceleration card to be removed.
[0099] Once it is confirmed that the heterogeneous acceleration card to be removed has powered off successfully, the system generates a second power-off success signal to notify the pooling management controller of the integrated cabinet of the fusion architecture. After receiving the second power-off success signal, the pooling management controller generates the corresponding second power-off success information and sends it to the preset terminal (such as the mobile device or workstation of the operation and maintenance personnel) through a preset communication channel (such as email, SMS, operation and maintenance management system, etc.).
[0100] The second power-off success information includes the number of the heterogeneous acceleration card to be removed, the power-off time, and a reminder of the new heterogeneous acceleration card that the operation and maintenance personnel may need to replace.
[0101] After receiving the second power-off success information, the operation and maintenance personnel learn that the heterogeneous acceleration card to be removed has powered off safely and the system is ready for replacement. The operation and maintenance personnel perform the physical replacement operation, remove the faulty acceleration card from the integrated cabinet, install the new heterogeneous acceleration card, and verify the new acceleration card to ensure that it works properly and meets the system requirements, and may add it to the corresponding heterogeneous acceleration card device.
[0102] If the power-off fails, the system may need to execute an error handling process, such as retrying the power-off instruction, recording the error information, notifying the administrator, etc.
[0103] Through the above technical solution, after allocating a new heterogeneous acceleration card to the computing resource pool, the system immediately controls the power-off of the acceleration card to be removed, ensuring that the computing resources will not be interrupted due to the replacement of the acceleration card, thereby improving the availability and continuity of the system. By determining whether the heterogeneous acceleration card to be removed has been powered off successfully, the system can ensure that the acceleration card has indeed been safely powered off, preventing potential security risks caused by incomplete power-off. Once it is confirmed that the acceleration card to be removed has been powered off successfully, the system will automatically generate a second power-off success signal and send it to a preset terminal through the pooling management controller. This automated notification mechanism can quickly convey the power-off status to the operation and maintenance personnel, reminding them to replace the heterogeneous acceleration card to be removed. After receiving the power-off success information, the operation and maintenance personnel can immediately know when they can safely remove the heterogeneous acceleration card to be removed and replace it, reducing the waiting time and improving the efficiency and response speed of the operation and maintenance operations.
[0104] Optionally, in some embodiments, controlling the power-off of the heterogeneous acceleration card to be removed includes: generating a power-off instruction and determining the transmission mode of the power-off instruction; according to the transmission mode, sending the power-off instruction to the controller corresponding to the heterogeneous acceleration card to be removed, so as to control the power-off of the heterogeneous acceleration card to be removed based on the power-off instruction through the controller corresponding to the heterogeneous acceleration card to be removed.
[0105] Among them, in some embodiments, the transmission mode includes cable transmission and / or network transmission.
[0106] Specifically, the power-off instruction of the heterogeneous acceleration card sent by the PSMC can be transmitted through the CDFP cable, or the power-off instruction can be directly sent to the BMC of the heterogeneous acceleration card Box through the network. The BMC of the heterogeneous acceleration card Box sends the power-off command to the CPLD of the heterogeneous acceleration card Box to control the power-off of the heterogeneous acceleration card. In this way, it is not dependent on the hardware design and is more flexible.
[0107] It can be understood that if the transmission mode is the cable transmission mode, the PSMC sends the power-off instruction to the BMC corresponding to the heterogeneous acceleration card to be removed through the CDFP cable. After receiving the power-off instruction, the BMC corresponding to the heterogeneous acceleration card to be removed sends the power-off instruction to the corresponding CPLD to control the power-off of the heterogeneous acceleration card.
[0108] If the transmission mode is the network transmission mode, the PSMC sends the power-off instruction to the BMC corresponding to the heterogeneous acceleration card to be removed through a network protocol (such as TCP / IP). The BMC corresponding to the heterogeneous acceleration card to be removed sends the power-off instruction to the corresponding CPLD, thereby controlling the power-off of the heterogeneous acceleration card to be removed.
[0109] It should be noted that not only the power-off instruction can be transmitted through the cable or the network, but also the power-on instruction can be transmitted through the cable or the network.
[0110] Through the above technical solution, through a series of operations including generating a power-off instruction, determining a transmission mode, and sending the power-off instruction to the BMC of the heterogeneous acceleration card Box, the system can control the power-off of the acceleration card more flexibly and accurately, reduce the complexity of manual operations, improve the overall operation and maintenance efficiency, provide two or more optional transmission modes such as cable transmission and network transmission, enhance the flexibility and adaptability of the system, help ensure that the heterogeneous acceleration card to be removed can be powered off safely and efficiently, and at the same time improve the efficiency and convenience of operation and maintenance operations.
[0111] Optionally, in some embodiments, the above method for handling faults of heterogeneous acceleration cards further includes: separately powering all heterogeneous acceleration cards in the integrated architecture whole cabinet.
[0112] It can be understood that 16 heterogeneous acceleration cards among the multiple heterogeneous acceleration card devices in the embodiments of the present invention are separately powered, and the power-on and power-off of different heterogeneous acceleration cards do not interfere with each other.
[0113] The power-on and power-off of each heterogeneous acceleration card are controlled by the CPLD in the heterogeneous acceleration card Box. As Figure 3 shown, the CPLD receives three groups of control signal inputs from the IO Box CPLD (Host_CARDx_PWR_EN), the heterogeneous acceleration card Box BMC (BMC_CARDx_BTN_N), and the power button on the left ear of the heterogeneous acceleration card Box (FP_PWR_BTN_N) respectively, and outputs the power-on enable signals (CPLD_P12V_CARDx_EN, CPLD_P3V3_CARDx_EN) of the heterogeneous acceleration card through internal logic processing. The CPLD judges the power supply status of the heterogeneous acceleration card according to the powerenable (CPLD_P12V_CARDx_EN) and powergood (P3V3_CARDx_PWRGD) signals of the VR chip.
[0114] After receiving the power-on instruction of the heterogeneous acceleration card, the heterogeneous acceleration card Box CPLD enables the MP5991GLU to output P12V_CARDx, and at the same time receives the P12V_CARDx_PWRGD signal and enables the MPQ8626VR to output P3V3_CARDx. The P12V and P3V3 required for the operation of the heterogeneous acceleration card are output through the above chips to realize the power supply of the heterogeneous acceleration card.
[0115] When the enable signal (Host_CARDx_PWR_EN) at the IO Box CPLD end is valid, the two signals of the IO Box CPLD (Host_CARDx_PWR_EN) and the heterogeneous acceleration card Box BMC (BMC_CARDx_BTN_N) will be ignored by the CPLD to prevent abnormal power-off of the heterogeneous acceleration card.
[0116] The 40 CDFP ports of the IO Box and the CDFP ports of the device Box are connected by CDFP cables. The CDFP cables can transmit the power-on and power-off instructions (Host_CARDx_PWR_EN) for controlling the heterogeneous acceleration card by the CPLD, and transmit the CPLD signals in the IO Box to the heterogeneous acceleration card Box CPLD.
[0117] Through the above technical solution, separate power supply means that each heterogeneous acceleration card has an independent power supply, reducing the risk of the entire system or multiple acceleration cards failing due to a single power failure. When a power failure occurs, only the affected acceleration card will lose power, and other acceleration cards can still continue to work, thus improving the overall reliability and stability of the system. Separate power supply enables the system to more flexibly manage the power states of each acceleration card, and separate power supply allows the system to precisely control the power of each acceleration card according to actual needs.
[0118] To enable those skilled in the art to further understand the fault handling method of the heterogeneous acceleration card in the embodiments of the present invention, the following will be elaborated in detail with specific embodiments, as Figure 2 and 4 shown.
[0119] In step S401, when a fault occurs in the heterogeneous acceleration card device or in certain specific necessary scenarios where the heterogeneous acceleration card needs to be replaced, the operation and maintenance personnel determine the number of the heterogeneous acceleration card to be replaced and the number of the heterogeneous acceleration card Box where it is located.
[0120] In step S402, the operation and maintenance personnel send a power-off instruction for the heterogeneous acceleration card to the PSMC through the IPMI command, Redfish interface or web page. The instruction parameters include the number of the heterogeneous acceleration card and the number of the heterogeneous acceleration card Box where it is located.
[0121] In step S403, the PSMC determines whether the heterogeneous acceleration card has been assigned to a certain Host BOX according to the allocation relationship table between the heterogeneous acceleration card to be replaced and the computing resource pool.
[0122] In step S404, if it is in the assigned state, the PSMC removes the heterogeneous acceleration card to be replaced from under the Host BOX, and updates the allocation relationship table and the device asset information at the Host BMC end.
[0123] In step S405, the PSMC selects an unallocated heterogeneous acceleration card from the computing resource pool to replace the heterogeneous acceleration card that needs to be replaced, and allocates an unallocated heterogeneous acceleration card under the HostBOX corresponding to the original heterogeneous acceleration card that needs to be replaced to ensure its computing power requirement, and updates the allocation relationship table and the device asset information at the Host BMC end. If the heterogeneous acceleration card is in an unallocated state, directly execute steps S406 - S412. After allocating an unallocated heterogeneous acceleration card under the Host BOX corresponding to the original heterogeneous acceleration card that needs to be replaced, also power down the original heterogeneous acceleration card that needs to be replaced, that is, execute steps S406 - S412.
[0124] In step S406, the PSMC sends the heterogeneous acceleration card power - down instruction to the IO Box CPLD through the I2C link according to the number of the heterogeneous acceleration card that needs to be replaced and the Box number where it is located.
[0125] In step S407, the IO Box CPLD sends the power - down instruction to the heterogeneous acceleration card BoxCPLD through the CDFP cable to control the corresponding heterogeneous acceleration card to power down.
[0126] In step S408, after the heterogeneous acceleration card Box BMC detects that the heterogeneous acceleration card has been powered down successfully through the CPLD, it reports the power - down success message to the PSMC, and the PSMC updates the power supply status of the heterogeneous acceleration card.
[0127] In step S409, the operation and maintenance personnel replace the new heterogeneous acceleration card.
[0128] In step S410, after the operation and maintenance personnel complete replacing the new heterogeneous acceleration card, they send a heterogeneous acceleration card power - on instruction to the PSMC.
[0129] In step S411, after the PSMC detects that the new heterogeneous acceleration card has been powered on, it sends a reset signal to the new heterogeneous acceleration card to reset the new heterogeneous acceleration card.
[0130] In step S412, after the new heterogeneous acceleration card is reset, it can be continued to be allocated to the computing resource pool as a resource.
[0131] In summary, based on the analysis of the above - mentioned specific embodiments, the present invention can achieve the following beneficial effects:
[0132] (1) By receiving the power-down instruction for the faulty heterogeneous acceleration card and determining the status of the acceleration card to be removed, the rapid identification and handling of the faulty acceleration card are achieved. When it is confirmed that the acceleration card to be removed is in the allocated state, an unallocated target acceleration card can be automatically selected from the resource pool for replacement and allocated to the original computing resource pool, thus ensuring the continuity and stability of the computing resources. This not only improves the efficiency of fault handling but also realizes the dynamic optimal allocation of resources, enhancing the flexibility and reliability of the overall system;
[0133] (2) When controlling the power-down of the heterogeneous acceleration card to be removed, by generating a power-down instruction and determining the transmission mode (such as cable transmission or network transmission), the accurate transmission and execution of the power-down instruction are ensured. At the same time, this method also includes the detection and confirmation links for successful power-down, as well as the mechanism of using the pooling management controller to send the power-down success information to the preset terminal, which helps the operation and maintenance personnel to timely understand the status of the acceleration card and perform safe physical replacement operations. In addition, the design of separate power supply for all heterogeneous acceleration cards in the integrated cabinet of the fusion architecture further enhances the power supply stability and security of the system.
[0134] (3) After allocating the target heterogeneous acceleration card to the computing resource pool corresponding to the acceleration card to be removed, not only the seamless replacement of resources is achieved, but also the accuracy and consistency of the internal resource information of the system are ensured by updating the allocation relationship table, etc., which helps the operation and maintenance personnel to better monitor and manage the acceleration card resources in the whole cabinet. At the same time, the fault handling process and information feedback mechanism provided by this method also greatly facilitate the daily maintenance and fault troubleshooting work of the operation and maintenance personnel, reducing the operation and maintenance cost and complexity.
[0135] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.
[0136] According to the fault handling method of the heterogeneous acceleration card proposed by the embodiments of the present invention, it is judged whether a power-down instruction for the faulty heterogeneous acceleration card is received, where the power-down instruction for the faulty heterogeneous acceleration card includes the heterogeneous acceleration card to be removed. If the power-down instruction for the faulty heterogeneous acceleration card is received, when the heterogeneous acceleration card to be removed is in an allocated state, the heterogeneous acceleration card to be removed is removed, and a target heterogeneous acceleration card in an unallocated state is determined from at least one heterogeneous acceleration card device in the integrated cabinet of the fusion architecture. The target heterogeneous acceleration card is used to replace the heterogeneous acceleration card to be removed, and the target heterogeneous acceleration card is allocated to the computing resource pool corresponding to the heterogeneous acceleration card to be removed. Thus, it can solve the technical problem that when a heterogeneous acceleration card device in the integrated cabinet of the fusion architecture fails, the existing processing method requires powering down the entire heterogeneous acceleration card device, affecting the business operations of all computing resource pools in the integrated cabinet and bringing a great business impact to users. It achieves the technical effect of being able to complete the replacement operation of the faulty acceleration card while minimizing the impact on the business functions of the entire system, and promptly restoring the normal operation of the business, reducing the maintenance pressure on the operation and maintenance personnel.
[0137] An embodiment of the present invention also provides a fault handling device for a heterogeneous acceleration card, as Figure 5 shown. The device includes: a judgment module 100, a fault handling module 200, and a replacement module 300.
[0138] Among them, the judgment module 100 is used to judge whether a power-down instruction for the faulty heterogeneous acceleration card is received, where the power-down instruction for the faulty heterogeneous acceleration card includes the heterogeneous acceleration card to be removed; the fault handling module 200 is used to, if the power-down instruction for the faulty heterogeneous acceleration card is received, when the heterogeneous acceleration card to be removed is in an allocated state, remove the heterogeneous acceleration card to be removed, and determine a target heterogeneous acceleration card in an unallocated state from at least one heterogeneous acceleration card device in the integrated cabinet of the fusion architecture; the replacement module 300 is used to use the target heterogeneous acceleration card to replace the heterogeneous acceleration card to be removed, and allocate the target heterogeneous acceleration card to the computing resource pool corresponding to the heterogeneous acceleration card to be removed.
[0139] Optionally, in some embodiments, the power-down instruction for the faulty heterogeneous acceleration card further includes the first number of the heterogeneous acceleration card to be removed and the second number of the heterogeneous acceleration card device to which the heterogeneous acceleration card to be removed belongs.
[0140] Optionally, in some embodiments, after receiving the power-down instruction for the faulty heterogeneous acceleration card, the fault handling module 200 is further used to: based on the first number and the second number, judge whether the heterogeneous acceleration card to be removed has a corresponding allocation relationship with at least one computing resource pool in the integrated cabinet of the fusion architecture; if the heterogeneous acceleration card to be removed has a corresponding allocation relationship with at least one computing resource pool in the integrated cabinet of the fusion architecture, it is determined that the heterogeneous acceleration card to be removed is in an allocated state, otherwise, it is determined that the heterogeneous acceleration card to be removed is in an unallocated state.
[0141] Optionally, in some embodiments, after receiving the power-down instruction for the faulty heterogeneous acceleration card, the fault handling module 200 is further configured to: when the heterogeneous acceleration card to be removed is in an unallocated state, control the heterogeneous acceleration card to be removed to power down.
[0142] Optionally, in some embodiments, when controlling the heterogeneous acceleration card to be removed to power down, the fault handling module 200 is further configured to: determine whether the heterogeneous acceleration card to be removed has powered down successfully; if the heterogeneous acceleration card to be removed has powered down successfully, generate a first power-down success signal to the pooling management controller of the fusion architecture whole cabinet, and use the pooling management controller to send a power-down success message to a preset terminal to remind the operation and maintenance personnel to replace the new heterogeneous acceleration card.
[0143] Optionally, in some embodiments, after using the pooling management controller to send a power-down success message to a preset terminal, the fault handling module 200 is further configured to: determine whether the pooling management controller has received a power-on instruction; if the pooling management controller has received a power-on instruction, based on the first number and the second number of the new heterogeneous acceleration card, control the new heterogeneous acceleration card to power on.
[0144] Optionally, in some embodiments, after controlling the new heterogeneous acceleration card to power on based on the first number and the second number, the fault handling module 200 is further configured to: detect whether the new heterogeneous acceleration card has completed power-on; if the new heterogeneous acceleration card has completed power-on, reset the new heterogeneous acceleration card, and after the new heterogeneous acceleration card is successfully reset, allocate it to the corresponding computing resource pool.
[0145] Optionally, in some embodiments, after allocating the target heterogeneous acceleration card to the computing resource pool corresponding to the heterogeneous acceleration card to be removed, the allocation module 300 is further configured to: control the heterogeneous acceleration card to be removed to power down, and determine whether the heterogeneous acceleration card to be removed has powered down successfully; if the heterogeneous acceleration card to be removed has powered down successfully, generate a second power-down success signal to the pooling management controller of the fusion architecture whole cabinet, and use the pooling management controller to send a second power-down success message to a preset terminal to remind the operation and maintenance personnel to replace the new heterogeneous acceleration card.
[0146] Optionally, in some embodiments, the fault handling module 200 is further configured to: generate a power-down instruction and determine the transmission mode of the power-down instruction; according to the transmission mode, send the power-down instruction to the controller corresponding to the heterogeneous acceleration card to be removed, so as to control the heterogeneous acceleration card to be removed to power down based on the power-down instruction through the controller corresponding to the heterogeneous acceleration card to be removed.
[0147] Optionally, in some embodiments, the transmission mode includes cable transmission and / or network transmission.
[0148] Optionally, in some embodiments, the fault handling device 10 of the above heterogeneous acceleration card further includes: a power supply module for separately powering all the heterogeneous acceleration cards in the integrated cabinet of the fusion architecture.
[0149] Optionally, in some embodiments, after allocating the target heterogeneous acceleration card to the computing resource pool corresponding to the heterogeneous acceleration card to be removed, the allocation module 300 is further configured to: update the allocation relationship table of the target heterogeneous acceleration card in the integrated cabinet of the fusion architecture.
[0150] It should be noted that for the description of the features in the embodiments corresponding to the fault handling device of the heterogeneous acceleration card, reference can be made to the relevant descriptions in the embodiments corresponding to the fault handling method of the heterogeneous acceleration card, which will not be elaborated here one by one.
[0151] An embodiment of the present invention further provides an electronic device. Figure 6 The following is a schematic structural diagram of the electronic device provided by the embodiment of the present invention. The electronic device may include:
[0152] A memory 601, a processor 602, and a computer program stored on the memory 601 and executable on the processor 602.
[0153] When the processor 602 executes the program, it implements the fault handling method of the heterogeneous acceleration card provided in the above embodiments.
[0154] Furthermore, the electronic device further includes:
[0155] A communication interface 603 for communication between the memory 601 and the processor 602.
[0156] The memory 601 is used to store a computer program executable on the processor 602.
[0157] The memory 601 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0158] If the memory 601, the processor 602, and the communication interface 603 are implemented independently, the communication interface 603, the memory 601, and the processor 602 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only a thick line is used to represent it in Figure 6 , but it does not mean that there is only one bus or one type of bus.
[0159] Optionally, in a specific implementation, if the memory 601, the processor 602, and the communication interface 603 are integrated on a single chip, the memory 601, the processor 602, and the communication interface 603 can communicate with each other through an internal interface.
[0160] The processor 602 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0161] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the fault handling method for a heterogeneous acceleration card when running.
[0162] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a Read-Only Memory (ROM), a Random Access Memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.
[0163] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered as exceeding the scope of the present invention.
[0164] The above has introduced in detail a method for handling faults of a heterogeneous acceleration card provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A method for handling a fault of a heterogeneous accelerator card, characterized in that: The following steps are involved: Determining whether a power-off instruction of a faulty heterogeneous accelerator card is received, wherein the power-off instruction of the faulty heterogeneous accelerator card includes the heterogeneous accelerator card to be removed; If a power-off instruction of the faulty heterogeneous acceleration card is received, then, if the heterogeneous acceleration card to be removed is in an allocated state, the heterogeneous acceleration card to be removed is removed, and a target heterogeneous acceleration card in an unallocated state is determined from at least one heterogeneous acceleration card device of the converged architecture whole cabinet, wherein the converged architecture whole cabinet includes an IO resource pool, a heterogeneous acceleration card resource pool, and a computing resource pool; The target heterogeneous accelerator card is used to replace the heterogeneous accelerator card to be removed, and the target heterogeneous accelerator card is allocated to the computing resource pool corresponding to the heterogeneous accelerator card to be removed. The failed heterogeneous accelerator card power-off instruction also includes the first serial number of the to-be-removed heterogeneous accelerator card and the second serial number of the heterogeneous accelerator card device to which the to-be-removed heterogeneous accelerator card belongs. After allocating the target heterogeneous acceleration card to the computing resource pool corresponding to the heterogeneous acceleration card to be removed, the method includes: updating an allocation relationship table of the target heterogeneous acceleration card in the entire cabinet of the converged architecture, wherein the allocation relationship table is an allocation relationship table between the target heterogeneous acceleration card and the computing resource pool.
2. The method for handling a fault of a heterogeneous accelerator card according to claim 1, characterized in that: After receiving the power-off instruction of the faulty heterogeneous acceleration card, the method further includes: Based on the first number and the second number, determining whether the heterogeneous acceleration card to be removed has a corresponding allocation relationship with at least one computing resource pool of the converged architecture whole cabinet; If there is a corresponding allocation relationship between the heterogeneous acceleration card to be removed and at least one computing resource pool of the converged architecture whole cabinet, it is determined that the heterogeneous acceleration card to be removed is in the allocated state; otherwise, it is determined that the heterogeneous acceleration card to be removed is in the unallocated state.
3. The method for handling a fault of a heterogeneous accelerator card according to claim 1, characterized in that: After receiving the power-off instruction of the faulty heterogeneous acceleration card, the method further includes: When the heterogeneous acceleration card to be removed is in an unallocated state, the heterogeneous acceleration card to be removed is controlled to be powered off.
4. The method for handling a fault of a heterogeneous accelerator card according to claim 3, characterized in that: When controlling the power-off of the heterogeneous acceleration card to be removed, the method further includes: Determine whether the heterogeneous acceleration card to be removed is powered off successfully; If the heterogeneous acceleration card to be removed is powered off successfully, a first power-off success signal is generated to the pooled management controller of the converged architecture whole cabinet, and the pooled management controller is used to send power-off success information to a preset terminal to remind the operation and maintenance personnel to replace the new heterogeneous acceleration card.
5. The method for handling a fault of a heterogeneous accelerator card according to claim 4, characterized in that: After the pooling management controller is used to send the power-off success information to the preset terminal, the method includes: Determining whether the pooling management controller has received a power-on instruction; If the pooling management controller receives the power-on instruction, it controls the new heterogeneous accelerator card to power on based on the first number and the second number of the new heterogeneous accelerator card.
6. The method for handling a fault of a heterogeneous accelerator card according to claim 5, characterized in that: After controlling the new heterogeneous accelerator card to be powered on based on the first number and the second number, the method includes: Detect whether the new heterogeneous acceleration card is powered on; If the new heterogeneous acceleration card is powered on, the new heterogeneous acceleration card is reset, and is allocated to the corresponding computing resource pool after the new heterogeneous acceleration card is successfully reset.
7. The method for handling a fault of a heterogeneous accelerator card according to claim 1, characterized in that: After allocating the target heterogeneous accelerator card to the computing resource pool corresponding to the heterogeneous accelerator card to be removed, the method further includes: Controlling the heterogeneous accelerator card to be removed to power off, and determining whether the heterogeneous accelerator card to be removed is successfully powered off; If the heterogeneous acceleration card to be removed is powered off successfully, a second power-off success signal is generated to the pooled management controller of the converged architecture whole cabinet, and the pooled management controller is used to send the second power-off success information to a preset terminal to remind the operation and maintenance personnel to replace the new heterogeneous acceleration card.
8. The method for handling a fault of a heterogeneous accelerator card according to claim 3 or 7, characterized in that: The controlling the heterogeneous accelerator card to be removed to power off includes: Generate a power-off instruction and determine a transmission mode of the power-off instruction; According to the transmission mode, the power-off instruction is sent to a controller corresponding to the heterogeneous acceleration card to be removed, so that the controller corresponding to the heterogeneous acceleration card to be removed controls the power-off of the heterogeneous acceleration card to be removed based on the power-off instruction.
9. The method for handling a fault of a heterogeneous accelerator card according to claim 8, characterized in that: The transmission mode includes cable transmission and / or network transmission.
10. The method for handling a fault of a heterogeneous accelerator card according to claim 1, characterized in that: Also includes: Provide discrete power supply for all heterogeneous acceleration cards in the entire cabinet of the converged architecture.
11. A fault handling device for a heterogeneous accelerator card, characterized in that: include: A judgment module, used to judge whether a power-off instruction of a faulty heterogeneous accelerator card is received, wherein the power-off instruction of the faulty heterogeneous accelerator card includes a heterogeneous accelerator card to be removed; a fault processing module, configured to, if receiving a power-off instruction for the faulty heterogeneous acceleration card, remove the heterogeneous acceleration card to be removed if the heterogeneous acceleration card to be removed is in an allocated state, and determine a target heterogeneous acceleration card in an unallocated state from at least one heterogeneous acceleration card device of a converged architecture whole cabinet, wherein the converged architecture whole cabinet includes an IO resource pool, a heterogeneous acceleration card resource pool, and a computing resource pool; a replacement module, configured to replace the heterogeneous acceleration card to be removed with the target heterogeneous acceleration card, and allocate the target heterogeneous acceleration card to the computing resource pool corresponding to the heterogeneous acceleration card to be removed, The failed heterogeneous accelerator card power-off instruction also includes a first serial number of the heterogeneous accelerator card to be removed and a second serial number of the heterogeneous accelerator card device to which the heterogeneous accelerator card to be removed belongs. After allocating the target heterogeneous acceleration card to the computing resource pool corresponding to the heterogeneous acceleration card to be removed, the replacement module is further used to: update the allocation relationship table of the target heterogeneous acceleration card in the entire cabinet of the converged architecture, wherein the allocation relationship table is an allocation relationship table between the target heterogeneous acceleration card and the computing resource pool.
12. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the fault handling method for a heterogeneous acceleration card as described in any one of claims 1 to 10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the fault handling method for a heterogeneous acceleration card as described in any one of claims 1 to 10.
Citation Information
Patent Citations
FPGA heterogeneous acceleration card management system
CN109614293A
Accelerator card fault determination method and device, equipment and medium
CN118897747A