Processing method, device, equipment, storage medium and program product of open accelerator module cluster

By acquiring real-time and historical hardware status data of the Open Accelerator module cluster and combining it with multiple preset health monitoring items, the health status of each module is assessed, which solves the problem of accuracy in task allocation and operation and maintenance of the Open Accelerator module cluster, and realizes optimized resource utilization and improved operation and maintenance efficiency.

CN121277720BActive Publication Date: 2026-02-24SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511845938.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-02-24
Estimated Expiration
2045-12-09

AI Technical Summary

Technical Problem

In existing technologies, the task allocation and operation and maintenance of open accelerator module clusters lack accuracy, resulting in uneven resource utilization and low operation and maintenance efficiency. It is also impossible to effectively distinguish the health status of open accelerator modules, causing operation and maintenance to become a passive "firefighting" process.

Method used

By acquiring real-time and historical hardware status data of each module in the open accelerator module cluster, and combining it with health monitoring data from multiple preset health monitoring projects, the real-time health status of each module is assessed, enabling precise task allocation and operation and maintenance.

Benefits of technology

It improves the accuracy of task allocation and operation and maintenance of the open accelerator module cluster, realizes optimized resource utilization and proactive prediction of problematic modules, and enhances operation and maintenance efficiency and service level agreement compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121277720B_ABST
    Figure CN121277720B_ABST
Patent Text Reader

Abstract

The application relates to a processing method and device for an open accelerator module cluster, equipment, a storage medium and a program product, and relates to the technical field of artificial intelligence chips. The application can improve the accuracy of task allocation processing and operation and maintenance processing of the open accelerator module cluster. The method comprises the following steps: obtaining health monitoring data of each open accelerator module in the open accelerator module cluster in a plurality of preset health monitoring items, the health monitoring data being obtained based on real-time hardware state data and historical hardware state data of the open accelerator module; obtaining real-time health state evaluation information of each open accelerator module according to the health monitoring data of the plurality of preset health monitoring items; and performing task allocation processing and operation and maintenance processing of each open accelerator module based on the real-time health state evaluation information of each open accelerator module, the real-time hardware state data, the historical hardware state data and historical health state evaluation information of each open accelerator module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence chip technology, and in particular to a processing method, apparatus, device, storage medium and program product for an open accelerator module cluster. Background Technology

[0002] The rapid development of artificial intelligence (AI) technology is profoundly impacting all industries, with large-scale AI models achieving significant breakthroughs in intelligence, operational efficiency, and application costs. Among these advancements, computing clusters are evolving towards multi-GPU clusters, which can contain tens of thousands of Open Accelerator Modules (OAM, OCP). This places higher demands on their performance and maintenance. These OAM modules contain AI chips, which are hardware chips specifically designed and optimized for AI tasks, including but not limited to GPUs (Graphics Processing Units), NPUs (Neural Network Processing Units), and GPGPUs (General-Purpose Computing on Graphics Processing Units).

[0003] In current technology, the current power consumption and temperature signals of the AI ​​chips in the open accelerator module are usually obtained and task allocation and maintenance are performed accordingly. The accuracy of task allocation and maintenance of the open accelerator module cluster still needs to be improved. Summary of the Invention

[0004] Therefore, it is necessary to provide a processing method, apparatus, equipment, storage medium, and program product for an open accelerator module cluster to address the aforementioned technical problems.

[0005] Firstly, this application provides a method for processing open accelerator module clusters, including:

[0006] The system acquires health monitoring data for each open accelerator module in the open accelerator module cluster across multiple preset health monitoring items; the health monitoring data is obtained based on the real-time hardware status data and historical hardware status data of the open accelerator module.

[0007] Based on the health monitoring data of the multiple preset health monitoring items, the real-time health status assessment information of each of the open accelerator modules is obtained;

[0008] Based on the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information of each open accelerator module, task allocation and operation and maintenance processing are performed for each open accelerator module.

[0009] Secondly, this application also provides a processing apparatus for an open accelerator module cluster, comprising:

[0010] The data acquisition module is used to acquire health monitoring data of each open accelerator module in the open accelerator module cluster for multiple preset health monitoring items; the health monitoring data is obtained based on the real-time hardware status data and historical hardware status data of the open accelerator module.

[0011] The health assessment module is used to obtain real-time health status assessment information for each of the open accelerator modules based on the health monitoring data of the multiple preset health monitoring items.

[0012] The processing and execution module is used to perform task allocation and operation and maintenance processing for each open accelerator module based on the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information of each open accelerator module.

[0013] Thirdly, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method.

[0014] Fourthly, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method.

[0015] Fifthly, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method.

[0016] The aforementioned processing method, apparatus, equipment, storage medium, and program product for the open accelerator module cluster acquires health monitoring data for each open accelerator module in the cluster across multiple preset health monitoring items. This health monitoring data is derived from the real-time and historical hardware status data of the open accelerator modules. Then, based on the health monitoring data from the multiple preset health monitoring items, real-time health status assessment information for each open accelerator module is obtained. Based on this real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information, task allocation and maintenance processing for each open accelerator module are executed. This solution, by accurately sensing the real-time health status assessment information of each open accelerator module in the cluster, and combining this with the real-time hardware status data, historical hardware status data, and historical health status assessment information, can comprehensively assess the health status and performance of each open accelerator module, thereby accurately executing task allocation and maintenance processing for each open accelerator module in the cluster and improving the accuracy of task allocation and maintenance processing for the open accelerator module cluster. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is an application environment diagram of a processing method for an open accelerator module cluster in one embodiment;

[0019] Figure 2 This is a schematic diagram of an open accelerator module cluster in one embodiment;

[0020] Figure 3 This is a schematic diagram of an open accelerator module server in one embodiment;

[0021] Figure 4 This is a flowchart illustrating a method for processing an open accelerator module cluster in one embodiment;

[0022] Figure 5 This is a flowchart illustrating the processing method for an open accelerator module cluster in another embodiment;

[0023] Figure 6 This is a schematic diagram of task allocation and operation and maintenance processing in one embodiment;

[0024] Figure 7 This is a structural block diagram of a processing device for an open accelerator module cluster in one embodiment;

[0025] Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0027] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various objects, but these objects are not limited by these terms. These terms are only used to distinguish the first object from the second object. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the solutions, or any combination of multiple solutions.

[0028] With the rapid development of artificial intelligence technology, computing clusters are evolving towards multi-GPU clusters. A computing cluster can contain multiple open accelerator module servers, and each open accelerator module server can contain multiple open accelerator modules. Open accelerator modules can include various hardware components such as AI chips, microprocessors (or other logic devices), discrete voltage devices, and high-bandwidth memory. How the cluster scheduler can accurately allocate and manage tasks within the open accelerator module cluster is a crucial issue.

[0029] Current technology typically only acquires the current power consumption and temperature signals of the AI ​​chips in open accelerator modules. However, this makes it difficult for the cluster scheduler to accurately perceive the status of the open accelerator modules (such as aging, runtime, and wear and tear). Consequently, the cluster scheduler cannot accurately allocate tasks and perform maintenance on the open accelerator module cluster. The cluster scheduler can only assign tasks to "idle" AI chips, unable to distinguish whether the AI ​​chip is brand new or about to be scrapped. Moreover, the cluster scheduler treats all open accelerator modules equally in resource utilization, without differentiation. If problems occur, it can only react passively, stopping training or replacing the open accelerator module after training ends. This significantly reduces training efficiency and turns maintenance into "firefighting" maintenance—replacing whatever is broken, reducing maintenance efficiency and increasing maintenance costs.

[0030] To address this, this application provides a method for processing open accelerator module clusters. By accurately sensing the real-time health status assessment information of each open accelerator module in the cluster, and combining this with real-time hardware status data, historical hardware status data, and historical health status assessment information, the health status and performance of each open accelerator module can be comprehensively evaluated. This allows for accurate execution of task allocation and maintenance for each open accelerator module in the cluster, improving the accuracy of task allocation and maintenance. It maximizes the utilization of open accelerator modules within the cluster, enabling the cluster to be self-aware, make intelligent decisions, and remain stable and reliable. From an operational perspective, it allows for planned and more precise location of problematic open accelerator modules, and proactive prediction and avoidance of the impact of problematic open accelerator modules on performance. This is significantly helpful in ensuring the service level agreements of cloud service providers and maximizing the return on investment of data centers.

[0031] The processing method for open accelerator module clusters provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, the application environment can be an open accelerator module cluster. The open accelerator module cluster may include multiple open accelerator module servers (open accelerator module server 1, open accelerator module server 2, ..., open accelerator module server N), and may also include a cluster scheduler. The cluster scheduler can interact with each open accelerator module server through cluster management software at the interface layer. The cluster scheduler can aggregate data and information sent by each open accelerator module server through the cluster management software at the interface layer. The cluster scheduler can also send relevant scheduling instructions (including instructions for task allocation and maintenance) to each open accelerator module server through the cluster management software at the interface layer. The processing method of the open accelerator module cluster in this application embodiment can be applied to... Figure 1 The cluster scheduler in the system.

[0032] like Figure 2The diagram illustrates an open accelerator module cluster using a single open accelerator module server as an example. Each open accelerator module server can contain multiple open accelerator modules (e.g., open accelerator module 0 to open accelerator module 7), and may also include a baseboard management controller. The baseboard management controller can connect to each open accelerator module via management buses such as I2C (Inter-Integrated Circuit) or SMBUS (System Management Bus). The baseboard management controller can interact with the cluster management software at the interface layer via APIs (Application Programming Interfaces) such as Redfish.

[0033] like Figure 3 The diagram illustrates an open accelerator module server, using a single open accelerator module as an example. Each open accelerator module may include a microcontroller unit and various hardware components connected to it, such as a GPU, power supply circuitry, a temperature sensor (for monitoring the open accelerator module's temperature), and onboard storage units. The GPU may contain a temperature sensor (for monitoring the GPU's temperature) and high-bandwidth memory (HBM). Alternatively, the temperature sensor can be integrated into the high-bandwidth memory or the microcontroller unit within the open accelerator module to obtain the temperature of the hardware to be monitored.

[0034] The following describes the processing method of the open accelerator module cluster of this application in conjunction with various embodiments.

[0035] In one exemplary embodiment, such as Figure 4 As shown, a method for processing open accelerator module clusters is provided, which can be applied to... Figure 1 In a cluster scheduler, the method may include the following steps:

[0036] Step S401: Obtain health monitoring data for each Open Accelerator module in the Open Accelerator module cluster across multiple preset health monitoring items.

[0037] The preset health monitoring items are pre-set health monitoring items used to monitor the health status of the open accelerator modules. The data used to monitor the health status can be referred to as health monitoring data. Health monitoring data can be obtained based on the real-time and historical hardware status data of the open accelerator modules. The real-time and historical hardware status data are hardware status data, which can include the status data of various hardware components contained in the open accelerator modules. The hardware status data can be monitored by, for example, the microcontroller units in the open accelerator modules at monitoring time intervals. The microcontroller units can obtain health monitoring data based on the hardware status data, and the health monitoring data can be stored, for example, in the onboard storage units of the open accelerator modules. The health monitoring data and hardware status data can be reported to the cluster scheduler by, for example, the microcontroller units in the open accelerator modules at reporting time intervals. Thus, the cluster scheduler can obtain the health monitoring data of each open accelerator module in the open accelerator module cluster for multiple preset health monitoring items. The microcontroller continuously monitors hardware status data (in real time), and can upload health monitoring data and hardware status data to the cluster scheduler after a period of time (reporting interval). The hardware status data obtained from the continued monitoring at this time can be recorded as real-time hardware status data, and the previously reported hardware status data can be recorded as historical hardware status data. The real-time hardware status data, historical hardware status data, and the aforementioned health monitoring data can all be stored by the microcontroller in, for example, the onboard storage unit of the open accelerator module, or can be read by the microcontroller from, for example, the onboard storage unit of the open accelerator module and reported to the cluster scheduler.

[0038] For example, hardware status data may include GPU temperature, high-bandwidth memory temperature, microcontroller temperature, voltage regulator module (VRM) temperature, GPU power consumption, high-bandwidth memory power consumption, voltage regulator (VR) power consumption, voltage discrete device power consumption, ECC (Error-Correcting Code) memory error correction count, PCIe (Peripheral Component Interconnect Express) link errors, internal errors, etc. This hardware status data can be real-time or historical.

[0039] For example, the health monitoring data for multiple preset health monitoring items may include cumulative working time, number of thermal cycles, duration of high temperature, duration of high power consumption, error count, real-time temperature, real-time power consumption, etc.

[0040] The cumulative runtime item can be used to monitor the runtime of the GPU and HBM under high load, such as a board with a power consumption of 500W running continuously at 450W for 3 hours or 5 hours. Therefore, as an example, the cumulative runtime of the GPU and HBM under high load can be obtained from real-time and historical hardware status data of the GPU and HBM as the health monitoring data for this item.

[0041] The thermal cycle count can be used to monitor the number of times the GPU reaches a certain temperature threshold (e.g., 80 degrees Celsius) and a certain heating rate threshold (e.g., 20 degrees Celsius per minute) during GPU heating (this can be recorded as heating count). Similarly, the number of times the GPU reaches a certain initial temperature threshold (e.g., 80 degrees Celsius) and a certain cooling rate threshold (e.g., 20 degrees Celsius per minute) during GPU cooling (this can be recorded as cooling count). Frequent thermal expansion and contraction is a major cause of hardware aging. Therefore, as an example, the number of heating and cooling cycles the GPU has experienced can be obtained from real-time and historical hardware status data. Furthermore, the thermal cycle count can be calculated from the heating and cooling counts as health monitoring data for this item.

[0042] The high-temperature duration item can be used to monitor the duration of an open accelerator module at high temperatures (e.g., above 90 degrees Celsius). For example, the duration of the open accelerator module at high temperatures can be obtained from real-time and historical hardware status data of the module's temperature sensor as health monitoring data for this project. Furthermore, data such as the highest temperature of the open accelerator module can also be recorded in this project.

[0043] For projects involving high-power duration, this can be used to monitor power consumption events and their durations that exceed the standard TDP (Thermal Design Power) of the open accelerator module. Therefore, as an example, the cumulative duration of the open accelerator module exceeding the standard TDP can be obtained from the real-time and historical hardware status data of the open accelerator module as the health monitoring data for this project.

[0044] The error count item can be used to monitor error counts such as ECC memory error correction count, PCIe link error, and internal errors in the open accelerator module. Therefore, as an example, the total error count of the open accelerator module, including the aforementioned ECC memory error correction count, PCIe link error, and internal error counts, can be obtained from real-time and historical hardware status data and used as the health monitoring data for this item. A continuous increase in the ECC memory error correction count is an important indicator of memory aging.

[0045] For the real-time temperature project, it can be used to monitor the real-time temperatures of the GPU, HBM, and VRMs. Therefore, as an example, the highest real-time temperature among the GPU, HBM, and VRMs of the open accelerator module can be obtained from the real-time hardware status data of the open accelerator module and used as the health monitoring data for this project.

[0046] For real-time power consumption, this can be used to monitor the power consumption of the entire open accelerator module and its various hardware components (such as GPUs, VRMs, voltage discrete devices, etc.). Therefore, as an example, the power consumption of the entire open accelerator module can be obtained from the real-time hardware status data of the open accelerator module as health monitoring data for the project.

[0047] In this step, the microcontroller unit or other logic device of the open accelerator module can monitor the hardware status data of the open accelerator module and store it in the onboard storage unit. The microcontroller unit or other logic device of the open accelerator module can obtain the health monitoring data of the open accelerator module for multiple preset health monitoring items based on the real-time hardware status data and historical hardware status data. The microcontroller unit or other logic device of the open accelerator module can store the health monitoring data of the open accelerator module for multiple preset health monitoring items in the onboard storage unit. The microcontroller unit or other logic device of the open accelerator module can report the real-time hardware status data, historical hardware status data and health monitoring data of multiple preset health monitoring items from the onboard storage unit to the cluster scheduler. Thus, the cluster scheduler can obtain the health monitoring data of each open accelerator module in the open accelerator module cluster for multiple preset health monitoring items.

[0048] Step S402: Based on the health monitoring data of multiple preset health monitoring items, obtain the real-time health status assessment information of each open accelerator module.

[0049] In this step, for each open accelerator module, the cluster scheduler can obtain its real-time health status assessment information based on its health monitoring data across multiple preset health monitoring items. Each open accelerator module can report its health monitoring data across these items to the cluster scheduler at regular intervals. Upon receiving the reported health monitoring data, the cluster scheduler can obtain the module's real-time health status assessment information. As described above, data monitoring and reporting can be continuous; therefore, each report generates a real-time health status assessment, while health status assessments prior to this report can be recorded as historical health status assessments.

[0050] For health status assessment information (including real-time and historical health status assessment information), as an example, a preset mapping relationship can be constructed to map health monitoring data from multiple preset health monitoring items to health status assessment information. This health status assessment information can include health status assessment values ​​or health status assessment levels. The health status assessment value can be a numerical value within, for example, the range of 0 to 100, with a higher value indicating a better health status. The health status assessment level can include, for example, three levels: highest, second-highest, and lowest, with higher levels indicating a better health status. Health status assessment information can also be mapped to health status assessment levels; for example, a health status assessment value greater than 90 maps to the highest level, a value greater than 70 and less than or equal to 90 maps to the second-highest level, and a value less than or equal to 70 maps to the lowest level, and so on. This allows relatively complex health monitoring data to be integrated into a relatively simple health status assessment, facilitating decision-making for task allocation and operation and maintenance of the open accelerator module.

[0051] Step S403: Based on the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information of each open accelerator module, perform task allocation and operation and maintenance processing for each open accelerator module.

[0052] In this step, the cluster scheduler obtains real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information for each open accelerator module. Based on this information, it performs task allocation and maintenance for each open accelerator module. For example, for a task to be assigned, the cluster scheduler can identify several open accelerator modules whose health status has consistently remained good (e.g., health status assessment values ​​consistently greater than 90) based on real-time and historical health status assessment information. Then, it can further identify, for example, open accelerator modules with relatively low error counts from these modules to handle the task, based on real-time and historical hardware status data. As an example, for operation and maintenance, the cluster scheduler can identify several open accelerator modules whose health status is consistently declining based on real-time and historical health status assessment information. Then, it can further identify open accelerator modules with deteriorating heat dissipation efficiency, such as GPUs, based on real-time and historical hardware status data, and send the corresponding operation and maintenance information to the operation and maintenance system of the open accelerator module cluster, etc.

[0053] The processing method for the open accelerator module cluster in this embodiment acquires health monitoring data for each open accelerator module in the cluster across multiple preset health monitoring items. This health monitoring data is derived from the real-time and historical hardware status data of the open accelerator modules. Then, based on the health monitoring data from the multiple preset health monitoring items, real-time health status assessment information for each open accelerator module is obtained. Based on this real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information, task allocation and maintenance processing for each open accelerator module are executed. This solution, by accurately sensing the real-time health status assessment information of each open accelerator module in the cluster and combining it with the real-time hardware status data, historical hardware status data, and historical health status assessment information, can comprehensively assess the health status and performance of each open accelerator module, thereby accurately executing task allocation and maintenance processing for each open accelerator module in the cluster and improving the accuracy of task allocation and maintenance processing for the open accelerator module cluster.

[0054] In an exemplary embodiment, step S401, obtaining health monitoring data for each open accelerator module in the open accelerator module cluster across multiple preset health monitoring items, may include:

[0055] The system receives health monitoring data of open accelerator modules in multiple preset health monitoring items from the baseboard management controller of the open accelerator module server in the open accelerator module cluster; based on the health monitoring data of each open accelerator module in multiple preset health monitoring items sent by the baseboard management controller of each open accelerator module server, the system obtains health monitoring data of each open accelerator module in the open accelerator module cluster in multiple preset health monitoring items.

[0056] In this embodiment, combined with Figures 1 to 3 An open accelerator module cluster can include several open accelerator module servers, and each open accelerator module server can contain several open accelerator modules. The microcontroller units of the open accelerator modules can send health monitoring data for multiple preset health monitoring items to the baseboard management controller of the open accelerator module server via management buses such as I2C. The cluster scheduler can retrieve the health monitoring data for multiple preset health monitoring items from the baseboard management controller via standard APIs such as Redfish. Thus, the cluster scheduler can receive the health monitoring data for multiple preset health monitoring items sent by the baseboard management controller of each open accelerator module server in the open accelerator module cluster. Based on the health monitoring data for multiple preset health monitoring items sent by the baseboard management controllers of each open accelerator module server, the cluster scheduler can obtain the health monitoring data for each open accelerator module in the open accelerator module cluster for multiple preset health monitoring items. Therefore, the solution of this embodiment can realize the reporting path for the health monitoring data of the open accelerator modules.

[0057] In an exemplary embodiment, step S402, which involves obtaining real-time health status assessment information for each open accelerator module based on health monitoring data from multiple preset health monitoring items, may include:

[0058] For each open accelerator module, based on the health monitoring data of multiple preset health monitoring items of the open accelerator module and the preset mapping relationship corresponding to each preset health monitoring item, a real-time health status reference value corresponding to the health monitoring data of each preset health monitoring item of the open accelerator module is obtained; based on the real-time health status reference value corresponding to each preset health monitoring item of the open accelerator module, a real-time health status assessment value of the open accelerator module is obtained; based on the real-time health status assessment value of the open accelerator module, the real-time health status assessment information of the open accelerator module is determined.

[0059] In this embodiment, for each open accelerator module, the cluster scheduler obtains the real-time health status reference value corresponding to the health monitoring data of each preset health monitoring item of the open accelerator module based on the health monitoring data of multiple preset health monitoring items of the open accelerator module and the preset mapping relationship corresponding to each preset health monitoring item. The preset mapping relationship is the mapping relationship between health monitoring data and real-time health status reference values. Each preset health monitoring item can correspond to a different preset mapping relationship. Through the preset mapping relationship, the health monitoring data of the corresponding item can be mapped to a real-time health status reference value.

[0060] For example, preset mapping relationships may include preset mapping relationships for cumulative working time, thermal cycle count, high temperature duration, high power consumption duration, error count, real-time temperature, and real-time power consumption.

[0061] As an example, the preset mapping relationship for the cumulative working time item can be represented as follows:

[0062] ;

[0063] P time T represents the real-time health status reference value corresponding to the project's cumulative working time. accum This indicates the runtime of the GPU and HBM under high load. High load can be defined as above 90% of rated power consumption. max This indicates the maximum allowed cumulative working hours, with an example value of 10,000 hours.

[0064] As an example, the preset mapping relationship for the thermal cycle number item can be represented as follows:

[0065] ;

[0066] P thermal C represents the real-time health status reference value corresponding to the thermal cycle count item. thermal C represents the number of thermal cycles of the GPU (which can be calculated from the number of heating and cooling cycles). max This indicates the maximum permissible number of thermal cycles.

[0067] As an example, the preset mapping relationship for the high temperature duration item can be represented as follows:

[0068] ;

[0069] P temp hist D represents the real-time health status reference value corresponding to the duration of high temperature. high tempD represents the cumulative duration (in hours) at high temperatures (e.g., above 90 degrees Celsius). maxtemp This indicates the maximum permissible duration of high temperature; example value D. maxtemp =1000 hours.

[0070] As an example, the preset mapping relationship for high-power duration items can be represented as follows:

[0071] ;

[0072] P power hist D represents the real-time health status reference value corresponding to the high-power duration item. high power This indicates the cumulative duration (in hours) of power exceeding the standard TDP (e.g., greater than 500W). maxpower Indicates the maximum permissible high power duration, example value D. maxpower =100 hours.

[0073] As an example, the default mapping relationship for error counting items can be represented as follows:

[0074] ;

[0075] P error E represents the real-time health status reference value corresponding to the error count item. total This represents the total error count, which can include ECC memory error correction count, PCIe link errors, internal errors, etc. max This represents the maximum allowed error count, example value E. max =10000 times.

[0076] As an example, the preset mapping relationship for real-time temperature items can be represented as follows:

[0077] ;

[0078] P current temp This represents the real-time health status reference value corresponding to the real-time temperature item, T. current This represents the real-time temperature (in degrees Celsius), and can be the highest real-time temperature among the GPU, HBM, and VRMs. T nomal This represents the normal temperature threshold; example value T. nomal =80 degrees Celsius. T critical This represents the critical temperature threshold, with an example value T. critical =105 degrees Celsius.

[0079] As an example, the preset mapping relationship for real-time power consumption items can be represented as follows:

[0080] ;

[0081] P current power This represents the real-time health status reference value corresponding to the real-time power consumption item, P. current P represents the real-time power consumption of the board (W). TDP This represents the standard TDP power consumption, example value P. TDP =500W, P critical This represents the critical power consumption threshold, with an example value P. critical =600W.

[0082] The value 10 or 20 in the above preset mapping relationship can represent the weight of each preset health monitoring item. These weights can be set based on hardware aging factors and can be adjusted according to actual importance.

[0083] The parameters in the above-mentioned preset mapping relationship (such as T) max (etc.) are example values. In actual applications, they can be adjusted according to the design specifications, warranty policy or historical data of the open accelerator module.

[0084] Therefore, the cluster scheduler can obtain the real-time health status reference value corresponding to each preset health monitoring item of the open accelerator module, and then calculate the real-time health status evaluation value of the open accelerator module based on the real-time health status reference value corresponding to each preset health monitoring item of the open accelerator module.

[0085] As an example, the cluster scheduler can calculate the real-time health status assessment value H of the open accelerator module using the following formula:

[0086] .

[0087] The above formula integrates various data points into a real-time health status assessment value ranging from 0 to 100. A higher real-time health status assessment value indicates a better health status for the open accelerator module. This formula can be based on a penalty numerical model: starting from a value of 100, a penalty value is deducted based on the real-time health status reference value corresponding to each preset health monitoring item, with the total penalty value not exceeding 100, ensuring that the real-time health status assessment value is between 0 and 100.

[0088] The cluster scheduler can calculate the real-time health status assessment value of the open accelerator module according to the reporting time interval to reflect its real-time health status.

[0089] Therefore, the cluster scheduler can obtain the real-time health status assessment information of the open accelerator modules based on their real-time health status assessment values. This can be achieved by using the real-time health status assessment value as the primary information, or by determining the corresponding real-time health status assessment level for each open accelerator module based on the value, and then using both the real-time health status assessment value and / or the real-time health status assessment level as the primary information. For example, an open accelerator module with a real-time health status assessment value greater than 90 can be designated as the highest level, suitable for new cards or cards used under light loads, where temperature, power consumption, and error counts are all within excellent ranges. An open accelerator module with a real-time health status assessment value greater than 70 and less than or equal to 90 can be designated as the second highest level, suitable for cards that have undergone a period of high-load operation, where ECC errors have slightly increased and thermal cycling counts are high, but still within specifications, allowing for stable operation of the vast majority of tasks. The real-time health status assessment level of open accelerator modules with a real-time health status assessment value of less than or equal to 70 can be determined as the lowest level. This level can be used to characterize long-term high-load operation, frequent memory ECC error correction, and long-term operation at high temperatures. It is recommended to run the module at a lower frequency or assign it to non-critical tasks.

[0090] In an exemplary embodiment, step S403, which involves performing task allocation processing for each open accelerator module based on its real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information, may include:

[0091] Based on the task type of the task to be assigned, determine the task allocation basis information in the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information; based on the task allocation basis information, assign the task to be assigned to the target open accelerator module among multiple open accelerator modules.

[0092] In this embodiment, after determining the task to be assigned, the cluster scheduler can obtain the task type of the task to be assigned. Based on the task type, it determines the task allocation basis information for the task to be assigned from real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information. This task allocation basis information can be used as the allocation basis for the task to be assigned. Specifically, one or more of the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information can be selected as the task allocation basis information. The selection of which information(s) is used can be determined based on the task type. This allows for the pre-setting of corresponding task allocation basis information selection strategies. Different task allocation basis information selection strategies are used for different task types to select the appropriate task allocation basis information, and the task to be assigned is then allocated to a target open accelerator module among multiple open accelerator modules. The target open accelerator module refers to the open accelerator module selected from multiple open accelerator modules based on the task allocation basis information. The solution in this embodiment can handle the task allocation and processing of tasks of different task types, achieve refined decision-making, and realize load balancing based on health status, so as to maximize the rational utilization of open accelerator modules in the open accelerator module cluster.

[0093] In some exemplary embodiments, further, the process of determining the task allocation basis information in the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information based on the task type of the task to be assigned may include: determining the real-time health status assessment information as the task allocation basis information when the task type is a first task type; the process of assigning the task to be assigned to a target open accelerator module among multiple open accelerator modules based on the task allocation basis information may include: assigning the task to be assigned to the target open accelerator module among the multiple open accelerator modules corresponding to the one with the highest real-time health status assessment value. The real-time health status assessment information includes a real-time health status assessment value.

[0094] In this embodiment, the task type can be a first task type, which can represent a very important AI training task that requires long-term execution (days or even weeks). When the task type is the first task type, the cluster scheduler can determine real-time health status assessment information as the basis for task allocation. Based on the real-time health status assessment information, it obtains the real-time health status assessment value of the open accelerator module. This real-time health status assessment information may include the real-time health status assessment value, thereby identifying the target open accelerator module with the highest real-time health status assessment value among multiple open accelerator modules. The task to be allocated is then assigned to this target open accelerator module. In this way, the cluster scheduler can select the target open accelerator module with the largest and most reliable real-time health status assessment value to execute the task. Because the cost of failure due to hardware failure midway through this task type is extremely high (wasting computing resources and time), it can effectively improve the success rate of critical tasks.

[0095] In some exemplary embodiments, further, the process of determining the task allocation basis information in the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information based on the task type of the task to be assigned may include: determining the real-time health status assessment information as the task allocation basis information when the task type is a second task type; the process of assigning the task to be assigned to a target open accelerator module among multiple open accelerator modules based on the task allocation basis information may include: assigning the task to be assigned to a target open accelerator module among the multiple open accelerator modules whose real-time health status assessment level is not the highest level. The real-time health status assessment information includes a real-time health status assessment level.

[0096] In this embodiment, the task type can be a second task type, which can represent a large inference task composed of multiple parallel instances, where the failure of one or two instances does not affect the overall task. When the task type is the second task type, the cluster scheduler can determine the real-time health status assessment information as the basis for task allocation. Based on the real-time health status assessment information, it obtains the real-time health status assessment level of the open accelerator module. This real-time health status assessment information may include the real-time health status assessment level, thereby identifying the target open accelerator module among multiple open accelerator modules whose real-time health status assessment level is not the highest (this non-highest level can be the second highest level). The task to be allocated is then assigned to this target open accelerator module. In this way, the cluster scheduler can allocate these task instances to open accelerator modules with lower health status. Even if one open accelerator module fails midway through the task, it will not cause catastrophic consequences; the instance can simply be restarted on other healthy open accelerator modules. This fully utilizes all hardware resources and achieves tiered resource utilization.

[0097] In some exemplary embodiments, further, the above-mentioned determination of the task allocation basis information in the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information based on the task type of the task to be assigned may include: when the task type is the third task type, determining the real-time hardware status data and historical hardware status data as the task allocation basis information; the above-mentioned allocation of the task to be assigned to the target open accelerator module among multiple open accelerator modules based on the task allocation basis information may include: for several candidate open accelerator modules of the same model among multiple open accelerator modules, determining the target open accelerator module among the several candidate open accelerator modules that meets the operating temperature conditions and operating frequency conditions based on the real-time hardware status data and historical hardware status data.

[0098] In this embodiment, the task type can be a third task type, which can represent a latency-sensitive high-performance computing task. When the task type is a third task type, the cluster scheduler can use real-time hardware status data and historical hardware status data as the basis for task allocation. For several candidate open accelerator modules of the same model among multiple open accelerator modules, the target open accelerator module that meets the operating temperature and operating frequency conditions is determined based on the real-time hardware status data and historical hardware status data. The operating temperature condition can be that the temperature of the open accelerator module is relatively low, and the operating frequency condition can be that the operating frequency is relatively high. For example, for two open accelerator modules of the same model, if the core temperature of one open accelerator module is 10 degrees Celsius higher than that of the other open accelerator module due to dust accumulation on the heatsink or aging of the thermal paste, in order to prevent overheating, the GPU Boost mechanism of the one open accelerator module will automatically reduce the frequency, resulting in its actual operating frequency being lower than that of the other open accelerator module. The cluster scheduler can identify this phenomenon based on the historical temperature and frequency data reported by the MCU. When there is a high-performance computing task that is sensitive to latency, the cluster scheduler can prioritize the other open accelerator module with a lower operating temperature and higher performance, thereby ensuring that the task obtains the optimal single-card performance.

[0099] In an exemplary embodiment, step S403, which involves performing maintenance processing on each open accelerator module based on its real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information, may include:

[0100] Based on the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information of each open accelerator module, health status prediction information for each open accelerator module is obtained; based on the health status prediction information, the open accelerator modules to be maintained among multiple open accelerator modules are identified; and the operation and maintenance information corresponding to the open accelerator modules to be maintained is sent to the operation and maintenance system.

[0101] In this embodiment, the cluster scheduler can predict the health status of each open accelerator module based on its real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information. For example, it can determine whether the health status is continuously declining based on real-time and historical health status assessment information, and whether the HBM temperature curve shows deteriorating heat dissipation efficiency based on real-time and historical hardware status data, thus obtaining the health status prediction information for each open accelerator module. Therefore, the cluster scheduler can further determine the open accelerator modules requiring maintenance based on the health status prediction information. For example, it can identify open accelerator modules whose health status is continuously declining and whose HBM temperature curve shows deteriorating heat dissipation efficiency, and send the corresponding maintenance information to the maintenance system. The cluster scheduler can stop assigning new tasks to this open accelerator module and mark it as "requiring maintenance." Meanwhile, the cluster scheduler can send maintenance information to the operations and maintenance system, such as: "The card in OAM Slot 5 is expected to need replacement; please prepare spare parts." This enables predictive maintenance, avoiding sudden downtime during peak business periods, allowing operations and maintenance work to be carried out calmly and systematically.

[0102] In one exemplary embodiment, such as Figure 5 As shown, a method for processing open accelerator module clusters is also provided, which can be applied to a cluster scheduler and may include the following steps:

[0103] Step S501: Receive health monitoring data of the open accelerator module in multiple preset health monitoring items sent by the baseboard management controller of the open accelerator module server of the open accelerator module cluster.

[0104] The open accelerator module cluster includes several open accelerator module servers, each containing several open accelerator modules. The health monitoring data is obtained based on the real-time and historical hardware status data of the open accelerator modules.

[0105] Step S502: Based on the health monitoring data of each open accelerator module in the open accelerator module cluster in multiple preset health monitoring items sent by the baseboard management controller of each open accelerator module server, obtain the health monitoring data of each open accelerator module in multiple preset health monitoring items.

[0106] Step S503: For each open accelerator module, based on the health monitoring data of multiple preset health monitoring items of the open accelerator module and the preset mapping relationship corresponding to each preset health monitoring item, obtain the real-time health status reference value corresponding to the health monitoring data of each preset health monitoring item of the open accelerator module.

[0107] The preset mapping relationship is the mapping relationship between health monitoring data and real-time health status reference values.

[0108] Step S504: Obtain the real-time health status assessment value of the Open Accelerator Module based on the real-time health status reference value corresponding to each preset health monitoring item of the Open Accelerator Module.

[0109] Step S505: Determine the real-time health status assessment information of the open accelerator module based on the real-time health status assessment value of the open accelerator module.

[0110] Step S506: When the task type is the first task type, determine the real-time health status assessment information as the task allocation basis information; assign the task to be assigned to the target open accelerator module with the maximum real-time health status assessment value among multiple open accelerator modules.

[0111] The real-time health status assessment information includes real-time health status assessment values.

[0112] Step S507: When the task type is the second task type, determine the real-time health status assessment information as the task allocation basis information; assign the task to be assigned to the target open accelerator module that corresponds to the real-time health status assessment level that is not the highest level among multiple open accelerator modules.

[0113] The real-time health status assessment information includes the real-time health status assessment level.

[0114] Step S508: When the task type is the third task type, determine the real-time hardware status data and historical hardware status data as the task allocation basis information; for several candidate open accelerator modules of the same model among multiple open accelerator modules, determine the target open accelerator module that meets the operating temperature and operating frequency conditions among the candidate open accelerator modules based on the real-time hardware status data and historical hardware status data.

[0115] Step S509: Based on the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information of each open accelerator module, obtain the health status prediction information of each open accelerator module; based on the health status prediction information, identify the open accelerator modules to be maintained among multiple open accelerator modules; send the operation and maintenance information corresponding to the open accelerator modules to be maintained to the operation and maintenance system.

[0116] In this embodiment, the task allocation and maintenance of steps S506 to S509 are combined with... Figure 6 In the case of the first task type, the cluster scheduler can select the target open accelerator module with the highest real-time health status assessment value (highest health score) and the greatest reliability to execute the task. This is because the cost of failure due to hardware failure midway through a task of this type is extremely high (wasting computing resources and time), effectively improving the success rate of critical tasks. In the case of the second task type, the cluster scheduler can assign task instances to open accelerator modules with lower health status and prepare a failover mechanism. Even if one open accelerator module fails midway through the task, it will not cause catastrophic consequences; the instance can simply be restarted on other healthy open accelerator modules. This fully utilizes all hardware resources and achieves tiered resource utilization. In the case of the third task type, the cluster scheduler can compare the historical temperature and frequency data of open accelerator modules of the same model and prioritize open accelerator modules with lower operating temperatures, better performance, and greater stability, thereby ensuring that the task achieves optimal single-card performance. For operations and maintenance (O&M), the cluster scheduler can continuously monitor the health status of all open accelerator modules and perform predictive maintenance and avoidance. For example, based on health status prediction information, it can identify open accelerator modules whose health status is continuously declining (reaching the warning threshold) and whose HBM temperature curve shows deteriorating heat dissipation efficiency (reaching the warning threshold) that require maintenance. The cluster scheduler can then send the corresponding O&M information for these modules to the O&M system. If the warning threshold is not reached, the cluster scheduler can continue monitoring their health status. Specifically, for open accelerator modules requiring maintenance, the cluster scheduler can stop assigning new tasks to them and mark them as "requiring maintenance." Simultaneously, the cluster scheduler can send O&M information to the O&M system, such as: "The card for OAM Slot 5 is expected to need replacement; please prepare spare parts." This enables predictive maintenance, avoiding sudden downtime during peak business periods, allowing O&M work to be carried out calmly and systematically.

[0117] The solution in this embodiment, compared to traditional technologies or modes, belongs to an intelligent collaborative mode (microcontroller + cluster scheduler) in terms of monitoring, task allocation, and operation and maintenance. At the open accelerator module level, in the traditional mode, the open accelerator module is a black box when the cluster scheduler allocates and manages tasks, unable to accurately perceive the status of each open accelerator module (such as aging, runtime, wear and tear), resulting in relatively static monitoring. The solution in this embodiment, however, can monitor the health status and various hardware status data of the open accelerator modules. When the cluster scheduler allocates and manages tasks, the open accelerator module is a white box, and real-time data and historical data can be monitored. At the target level, the cluster scheduler in the traditional mode can usually only allocate tasks to "idle" open accelerator modules, filling idle computing power, without distinguishing between brand new and soon-to-be-obsolete open accelerator modules. The solution in this embodiment maximizes reliability and total throughput. At the fault handling level, in the traditional mode, if a problem occurs, the cluster scheduler can only react passively, restarting after a task fails. The solution in this embodiment can proactively predict and avoid problems. At the resource utilization level, the traditional model treats all open accelerator modules equally, without differentiation in usage. The solution in this embodiment, however, enables tiered utilization of open accelerator modules. At the operation and maintenance level, in the traditional model, if a problem occurs, the cluster scheduler can only react passively, stopping training or replacing the open accelerator module after training ends. This significantly reduces training efficiency, turning operation and maintenance into "firefighting"—replacing whatever is broken, reducing efficiency and increasing costs. The solution in this embodiment enables planned and refined operation and maintenance. This solution combines the health status reporting of open accelerator modules with cluster load balancing, bringing about a fundamental change. It not only provides load balancing but also a comprehensive evaluation of the health status and performance of open accelerator modules, making them a white box to the cluster. By combining real-time health scores and historical values, it maximizes the utilization of open accelerator modules within the cluster. The solution in this embodiment enables the cluster to be self-aware, make intelligent decisions, and remain stable and reliable. For operations and maintenance (O&M), this allows for more planned and precise identification of problematic open accelerator modules, enabling proactive prediction and mitigation of their impact on performance. This is crucial for ensuring compliance with cloud service provider service level agreements (SLPs) and maximizing the return on investment (ROI) of data centers.

[0118] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0119] Based on the same inventive concept, this application also provides a processing apparatus for an open accelerator module cluster to implement the processing method for the open accelerator module cluster described above. The solution provided by this apparatus is similar to the implementation described in the above method. Therefore, the specific limitations in one or more embodiments of the processing apparatus for an open accelerator module cluster provided below can be found in the limitations of the processing method for the open accelerator module cluster described above, and will not be repeated here.

[0120] In one exemplary embodiment, such as Figure 7 As shown, a processing apparatus for an open accelerator module cluster is provided. The processing apparatus 700 for the open accelerator module cluster may include:

[0121] The data acquisition module 701 is used to acquire health monitoring data of each open accelerator module in the open accelerator module cluster for multiple preset health monitoring items; the health monitoring data is obtained based on the real-time hardware status data and historical hardware status data of the open accelerator module.

[0122] The health assessment module 702 is used to obtain real-time health status assessment information for each of the open accelerator modules based on the health monitoring data of the multiple preset health monitoring items.

[0123] The processing execution module 703 is used to perform task allocation processing and operation and maintenance processing for each open accelerator module based on the real-time health status assessment information, real-time hardware status data, historical hardware status data and historical health status assessment information of each open accelerator module.

[0124] In an exemplary embodiment, the data acquisition module 701 is configured to receive health monitoring data of the open accelerator modules in multiple preset health monitoring items sent by the baseboard management controller of the open accelerator module server of the open accelerator module cluster; wherein, the open accelerator module cluster includes a plurality of open accelerator module servers; the open accelerator module servers include a plurality of open accelerator modules; and based on the health monitoring data of each open accelerator module in the open accelerator module cluster in multiple preset health monitoring items sent by the baseboard management controller of each open accelerator module server, the health monitoring data of each open accelerator module in the open accelerator module cluster in multiple preset health monitoring items is obtained.

[0125] In an exemplary embodiment, the health assessment module 702 is configured to, for each of the open accelerator modules, obtain a real-time health status reference value corresponding to the health monitoring data of each of the multiple preset health monitoring items of the open accelerator module and a preset mapping relationship corresponding to each preset health monitoring item; the preset mapping relationship is a mapping relationship between the health monitoring data and the real-time health status reference value; obtain a real-time health status assessment value of the open accelerator module based on the real-time health status reference value corresponding to each of the preset health monitoring items of the open accelerator module; and determine the real-time health status assessment information of the open accelerator module based on the real-time health status assessment value of the open accelerator module.

[0126] In an exemplary embodiment, the processing execution module 703 is configured to determine the task allocation basis information in the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information based on the task type of the task to be assigned; and to allocate the task to be assigned to a target open accelerator module among multiple open accelerator modules based on the task allocation basis information.

[0127] In an exemplary embodiment, the processing execution module 703 is configured to: determine the real-time health status assessment information as the task allocation basis information when the task type is a first task type; allocate the task to be allocated to a target open accelerator module among multiple open accelerator modules corresponding to the highest real-time health status assessment value; the real-time health status assessment information includes the real-time health status assessment value; or, when the task type is a second task type, determine the real-time health status assessment information as the task allocation basis information; allocate the task to be allocated to a target open accelerator module among multiple open accelerator modules corresponding to a real-time health status assessment level other than the highest level; the real-time health status assessment information includes the real-time health status assessment level; or, when the task type is a third task type, determine the real-time hardware status data and historical hardware status data as the task allocation basis information; and, for a plurality of candidate open accelerator modules of the same model among multiple open accelerator modules, determine the target open accelerator module among the plurality of candidate open accelerator modules that meets the operating temperature and operating frequency conditions based on the real-time hardware status data and historical hardware status data.

[0128] In an exemplary embodiment, the processing execution module 703 is configured to obtain health status prediction information for each open accelerator module based on real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information for each open accelerator module; determine the open accelerator modules to be maintained among the multiple open accelerator modules based on the health status prediction information; and send the operation and maintenance information corresponding to the open accelerator modules to be maintained to the operation and maintenance system.

[0129] Each module in the processing device of the aforementioned open accelerator module cluster can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.

[0130] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external devices via a network connection. When the computer program is executed by the processor, it implements a processing method for an open accelerator module cluster.

[0131] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0132] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0133] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.

[0134] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0135] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0136] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0137] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0138] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for processing open accelerator module clusters, characterized in that, The method includes: The system acquires health monitoring data for each open accelerator module in the open accelerator module cluster across multiple preset health monitoring items; the health monitoring data is obtained based on the real-time hardware status data and historical hardware status data of the open accelerator module. Based on the health monitoring data of the multiple preset health monitoring items, the real-time health status assessment information of each of the open accelerator modules is obtained; Based on the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information of each open accelerator module, task allocation and operation and maintenance processing are performed for each open accelerator module.

2. The method according to claim 1, characterized in that, The acquisition of health monitoring data for each open accelerator module in the open accelerator module cluster across multiple preset health monitoring items includes: The baseboard management controller receives health monitoring data of the open accelerator modules in multiple preset health monitoring items sent by the open accelerator module server of the open accelerator module cluster; wherein, the open accelerator module cluster includes a plurality of the open accelerator module servers; and the open accelerator module server contains a plurality of the open accelerator modules. Based on the health monitoring data of each open accelerator module in the open accelerator module cluster in multiple preset health monitoring items sent by the baseboard management controller of each open accelerator module server, the health monitoring data of each open accelerator module in the open accelerator module cluster in multiple preset health monitoring items is obtained.

3. The method according to claim 1, characterized in that, The step of obtaining real-time health status assessment information for each open accelerator module based on health monitoring data from the plurality of preset health monitoring items includes: For each of the open accelerator modules, based on the health monitoring data of the multiple preset health monitoring items of the open accelerator module and the preset mapping relationship corresponding to each preset health monitoring item, a real-time health status reference value corresponding to the health monitoring data of each preset health monitoring item of the open accelerator module is obtained; the preset mapping relationship is the mapping relationship between the health monitoring data and the real-time health status reference value. Based on the real-time health status reference value corresponding to each preset health monitoring item of the open accelerator module, the real-time health status assessment value of the open accelerator module is obtained. Based on the real-time health status assessment value of the open accelerator module, the real-time health status assessment information of the open accelerator module is determined.

4. The method according to any one of claims 1 to 3, characterized in that, Based on the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information of each open accelerator module, task allocation processing for each open accelerator module is performed, including: Based on the task type of the task to be assigned, determine the task allocation basis information in the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information. Based on the task allocation information, the tasks to be assigned are allocated to target open accelerator modules among multiple open accelerator modules.

5. The method according to claim 4, characterized in that, The step of determining the task allocation basis information from the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information based on the task type of the task to be assigned includes: When the task type is the first task type, the real-time health status assessment information is determined as the task allocation basis information; The step of allocating the task to be assigned to a target open accelerator module among multiple open accelerator modules according to the task allocation basis information includes: The task to be assigned is assigned to the target open accelerator module that corresponds to the maximum real-time health status assessment value among multiple open accelerator modules; the real-time health status assessment information includes the real-time health status assessment value. or, The step of determining the task allocation basis information from the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information based on the task type of the task to be assigned includes: When the task type is the second task type, the real-time health status assessment information is determined as the task allocation basis information; The step of allocating the task to be assigned to a target open accelerator module among multiple open accelerator modules according to the task allocation basis information includes: The tasks to be assigned are distributed to target open accelerator modules among multiple open accelerator modules, corresponding to open accelerator modules with a real-time health status assessment level that is not the highest level; the real-time health status assessment information includes the real-time health status assessment level. or, The step of determining the task allocation basis information from the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information based on the task type of the task to be assigned includes: When the task type is the third task type, the real-time hardware status data and historical hardware status data are determined as the basis information for task allocation; The step of allocating the task to be assigned to a target open accelerator module among multiple open accelerator modules according to the task allocation basis information includes: For several candidate open accelerator modules of the same model among multiple open accelerator modules, the target open accelerator module that meets the operating temperature and operating frequency conditions is determined based on the real-time hardware status data and historical hardware status data.

6. The method according to any one of claims 1 to 3, characterized in that, Based on the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information of each open accelerator module, the operation and maintenance processing of each open accelerator module is performed, including: Based on the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information of each open accelerator module, the health status prediction information of each open accelerator module is obtained. Based on the health status prediction information, identify the open accelerator modules to be maintained among multiple open accelerator modules; Send the maintenance information corresponding to the open accelerator module to be maintained to the maintenance system.

7. A processing apparatus for an open accelerator module cluster, characterized in that, The device includes: The data acquisition module is used to acquire health monitoring data of each open accelerator module in the open accelerator module cluster for multiple preset health monitoring items; the health monitoring data is obtained based on the real-time hardware status data and historical hardware status data of the open accelerator module. The health assessment module is used to obtain real-time health status assessment information for each of the open accelerator modules based on the health monitoring data of the multiple preset health monitoring items. The processing and execution module is used to perform task allocation and operation and maintenance processing for each open accelerator module based on the real-time health status assessment information, real-time hardware status data, historical hardware status data, and historical health status assessment information of each open accelerator module.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Mechanical hard disk monitoring method and device, computer equipment and storage medium

    CN113778797A

  • Task processing method of artificial intelligence processor, storage medium and electronic equipment

    CN120066806A