An abnormality processing method and device for heterogeneous acceleration resources, a storage medium and an electronic device

By monitoring the hardware and device health of heterogeneous acceleration resources on the cloud computing platform, identifying and handling abnormal resources, the problem of inconsistency between the registration and actual use of virtualized heterogeneous acceleration resources is solved, ensuring the reliability and stability of the cloud platform.

CN117149474BActive Publication Date: 2025-12-09ZTE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210563855.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-23
Publication Date
2025-12-09
Estimated Expiration
2042-05-23

AI Technical Summary

Technical Problem

Existing technologies cannot effectively identify and handle the inconsistency between the registration and actual use of virtualized heterogeneous acceleration resources managed by cloud computing platforms, leading to abnormal resource allocation and losses to cloud computing platforms and users.

Method used

By performing hardware health monitoring and device usage health monitoring on heterogeneous acceleration resources, healthy or unhealthy hardware resources can be identified, and faulty resources can be handled to ensure the reliability and stability of resources.

Benefits of technology

It enables rapid identification and timely processing of heterogeneous acceleration resources, avoiding losses caused by resource anomalies and improving the reliability and stability of the cloud platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117149474B_ABST
    Figure CN117149474B_ABST
Patent Text Reader

Abstract

The application provides an isomeric acceleration resource exception handling method and device, a storage medium and an electronic device. The method comprises the following steps: determining whether the isomeric acceleration resource of a cloud computing platform is a hardware healthy resource or a hardware unhealthy resource by means of hardware health monitoring of the isomeric acceleration resource; determining whether the isomeric acceleration resource is a use healthy resource or an allocation failure resource by means of equipment use health monitoring of the isomeric acceleration resource; performing hardware exception handling on the hardware unhealthy resource; and performing allocation exception handling on the allocation failure resource. By means of the method, the problem that, in the related art, only traditional server ordinary hardware resource detection is focused on, the inconsistency between registration and actual use of virtualized isomeric acceleration resources managed by the cloud computing platform cannot be identified, and thus losses are brought to the cloud computing platform and users can be solved, and the reliability, stability, timeliness, etc. of the cloud platform managed isomeric acceleration resources can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of cloud computing, in particular, to a heterogeneous acceleration resource exception processing method and device, a storage medium and an electronic device. BACKGROUND

[0002] With the development of AI technologies such as deep learning, users' demand for computing power and performance is becoming more and more urgent, and more and more users hope to obtain heterogeneous computing capabilities through a cloud computing platform to achieve performance acceleration of business, and the heterogeneous computing service provided by the cloud computing platform has become an indispensable function.

[0003] Heterogeneous acceleration resources of a cloud computing platform usually include a graphics processing unit (GPU), an AI acceleration card (NPU), a field programmable gate array (FPGA), and a smart NIC. Compared with traditional hardware, the heterogeneous acceleration resources of the cloud computing platform have the characteristics of multiple types of acceleration resources, easy pluggability, multiple virtualization modes, unified allocation and recovery, frequent use, special service bearing, etc.

[0004] When an exception occurs in the heterogeneous acceleration hardware, if it cannot be identified, reported, and recovered in time, it will cause serious loss to the customer business carried on the cloud computing platform. Especially for the heterogeneous acceleration resources allocated in a virtualization mode, such as GPU, NPU, and FPGA, in the process of frequent allocation and frequent recovery of resources, the problem of loss of recovery information or untimely recovery of resources may occur due to communication exceptions, which is prone to cause inconsistency between the registration and actual use of the heterogeneous acceleration resources, thereby causing an exception in the allocation of cloud platform resources and causing loss to the cloud computing platform and customers.

[0005] At present, most of the traditional hardware detection methods are detected and judged by the server's own system. On the one hand, the judgment is not accurate, and on the other hand, with the increase in the number of types, it cannot be well managed, and most importantly, it cannot identify the inconsistency between the registration and actual use of the virtualized heterogeneous acceleration resources managed by the cloud computing platform.

[0006] Since there is no exception detection and exception handling method for the heterogeneous acceleration resources of the cloud computing platform in the related art, especially when the registration exception of the virtualized acceleration hardware (GPU, NPU), the allocation exception of the virtualized acceleration hardware during the maintenance of the administrator, the health status exception of the device itself, and the misoperation of the device occur, it cannot be sensed and handled in time, thereby affecting the normal use of the cloud computing platform and causing loss to the cloud computing platform and users.

[0007] The related art only focuses on detection of traditional server ordinary hardware resources, and cannot identify inconsistency between registration and actual use of virtualized heterogeneous acceleration resources managed by a cloud computing platform, thereby causing losses to the cloud computing platform and users. No solution has been proposed. SUMMARY

[0008] Embodiments of the present application provide a heterogeneous acceleration resource exception processing method and device, a storage medium and an electronic device to at least solve the problem that the related art only focuses on detection of traditional server ordinary hardware resources, and cannot identify inconsistency between registration and actual use of virtualized heterogeneous acceleration resources managed by a cloud computing platform, thereby causing losses to the cloud computing platform and users. When a heterogeneous acceleration resource is abnormal, the non-healthy state of the heterogeneous acceleration resource can be quickly perceived and timely alarm and recovery are performed, thereby ensuring reliability, stability, timeliness and the like of the cloud platform in managing the heterogeneous acceleration resource.

[0009] According to an embodiment of the present application, a heterogeneous acceleration resource exception processing method is provided, and the method comprises:

[0010] The hardware health of the heterogeneous acceleration resource of the cloud computing platform is monitored to determine whether the heterogeneous acceleration resource is a hardware healthy resource or a hardware non-healthy resource;

[0011] The device usage health of the heterogeneous acceleration resource is monitored to determine whether the heterogeneous acceleration resource is a usage healthy resource or an allocation failure resource;

[0012] The hardware non-healthy resource is subjected to hardware exception processing;

[0013] The allocation failure resource is subjected to allocation exception processing.

[0014] According to another embodiment of the present application, a heterogeneous acceleration resource exception processing device is further provided, and the device comprises:

[0015] A first monitoring module is configured to monitor the hardware health of the heterogeneous acceleration resource of the cloud computing platform to determine whether the heterogeneous acceleration resource is a hardware healthy resource or a hardware non-healthy resource;

[0016] A second monitoring module is configured to monitor the device usage health of the heterogeneous acceleration resource to determine whether the heterogeneous acceleration resource is a usage healthy resource or an allocation failure resource;

[0017] A first response module is configured to perform hardware exception processing on the hardware non-healthy resource;

[0018] A second response module is configured to perform allocation exception processing on the allocation failure resource.

[0019] According to still another embodiment of the present application, a computer readable storage medium is also provided, in which a computer program is stored, wherein the computer program is configured to perform the steps of any of the above method embodiments when executed.

[0020] According to still another embodiment of the present application, an electronic device is also provided, comprising a memory in which a computer program is stored, and a processor configured to execute the computer program to perform the steps of any of the above method embodiments.

[0021] The embodiments of the present application can solve the problem in the related art that only the detection of the traditional server ordinary hardware resources is concerned, the inconsistency between the registration and actual use of the virtualized heterogeneous acceleration resources managed by the cloud computing platform cannot be identified, and thus the cloud computing platform and the user are brought losses. When the heterogeneous acceleration resources are abnormal, the unhealthy state of the heterogeneous acceleration resources can be quickly perceived and timely alarm and recovery are performed, and the reliability, stability, timeliness and the like of the cloud platform management of the heterogeneous acceleration resources are ensured. BRIEF DESCRIPTION OF DRAWINGS

[0022] Figure 1 is a hardware structure block diagram of a computer terminal of the heterogeneous acceleration resource abnormality processing method of the embodiments of the present application;

[0023] Figure 2 is a flowchart of the heterogeneous acceleration resource abnormality processing method of the embodiments of the present application;

[0024] Figure 3 is a flowchart of the heterogeneous acceleration resource hardware health monitoring method of the embodiments of the present application;

[0025] Figure 4 is a flowchart of the heterogeneous acceleration resource device use health monitoring method of the embodiments of the present application;

[0026] Figure 5 is a timing diagram of the device use health monitoring and processing of the optional embodiments of the present application;

[0027] Figure 6 is a timing diagram of the heterogeneous acceleration resource abnormality recovery processing of the optional embodiments of the present application;

[0028] Figure 7 is a block diagram of the heterogeneous acceleration resource abnormality processing apparatus of the embodiments of the present application;

[0029] Figure 8 is a heterogeneous acceleration resource health monitoring and abnormality processing architecture of the embodiments of the present application. Detailed Implementation

[0030] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.

[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0032] The methods and embodiments provided in this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure block diagram of the computer terminal for the heterogeneous acceleration resource anomaly handling method according to an embodiment of this application, as shown below. Figure 1 As shown, a computer terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the computer terminal described above. For example, the computer terminal may also include components that are more complex than those described above. Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0033] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the heterogeneous acceleration resource exception handling method in this embodiment. The processor 102 executes various functional applications and business chain address pool slicing processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0034] The transmission device 106 is configured to receive or send data via a network. The network can include a wireless network provided by a communication provider of the computer terminal. In an example, the transmission device 106 includes a network adapter (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In an example, the transmission device 106 can be a radio frequency (RF) module configured to communicate with the Internet in a wireless manner.

[0035] In the embodiment, a heterogeneous acceleration resource exception handling method running on the computer terminal or the network architecture is provided, Figure 2 is a flowchart of the heterogeneous acceleration resource exception handling method of the embodiment, as shown in the figure, the flow includes the following steps: Figure 2

[0036] In step S202, the heterogeneous acceleration resource is determined to be a hardware healthy resource or a hardware unhealthy resource by performing hardware health monitoring on the heterogeneous acceleration resource of the cloud computing platform.

[0037] In step S204, the heterogeneous acceleration resource is determined to be a usage healthy resource or an allocation failure resource by performing device usage health monitoring on the heterogeneous acceleration resource.

[0038] In step S206, hardware exception handling is performed on the hardware unhealthy resource.

[0039] In step S208, allocation exception handling is performed on the allocation failure resource.

[0040] In an embodiment, before step S202, it is determined whether the heterogeneous acceleration resource exists by scanning the PCI slot. If the heterogeneous acceleration resource exists, the resource information of the heterogeneous acceleration resource is obtained. Specifically, the resource information of the heterogeneous acceleration resource can be identified in combination with the configuration of the cloud computing platform. The heterogeneous acceleration resource includes GPU, NPU, FPGA, and SmartNIC. The resource information of the heterogeneous acceleration resource can include PCI address, vendor information, device model, device ID, etc. The PCI address includes slot number.

[0041] In the embodiment, step S202 can specifically include: calling a corresponding hardware health detection interface according to the resource information of the heterogeneous acceleration resource; judging the hardware state of the heterogeneous acceleration resource through the hardware health detection interface; if the hardware state is healthy, the heterogeneous acceleration resource is determined to be a hardware healthy resource; if the hardware state is unhealthy, the heterogeneous acceleration resource is determined to be a hardware unhealthy resource.

[0042] ​Specifically, the hardware health detection interface of the heterogeneous acceleration resource that has passed the security authentication in the cloud computing platform can be called in cycles according to the type, vendor information and device model of the heterogeneous acceleration resource, and the hardware state of the heterogeneous acceleration resource is determined by the hardware health detection interface.

[0043] In another embodiment, the hardware health monitoring method in step S202 can be performed on each heterogeneous acceleration resource according to a preset hardware health detection period.

[0044] Figure 3 is a flowchart of the heterogeneous acceleration resource hardware health monitoring method of the embodiment of the present application, as shown in Figure 3 The heterogeneous acceleration resource hardware health monitoring method specifically includes the following steps:

[0045] Step S302: Scan each heterogeneous acceleration resource on the PCI slot on the computing node to obtain the PCI address of the acceleration resource;

[0046] Step S304: Identify the vendor and model of the specific acceleration resource (GPU, NPU, FPGA, SmartNIC) in combination with the cloud platform configuration;

[0047] Step S306: Take the PCI address, vendor and model as core identification parameters, and call the hardware health detection interface recognized by the cloud platform in cycles to determine the hardware state of each heterogeneous acceleration resource;

[0048] Step S308: Determine whether the hardware state of the heterogeneous acceleration resource is healthy; if the determination result is yes, execute step S310a, and if the determination result is no, execute step S310b;

[0049] Step S310a: Determine that the heterogeneous acceleration resource is a hardware healthy resource;

[0050] Step S310b: Determine that the heterogeneous acceleration resource is a hardware unhealthy resource;

[0051] Step S312: Determine whether there are still heterogeneous acceleration resources on the current node that have not been subjected to hardware health detection;

[0052] Step S314: Output the hardware healthy resource and the hardware unhealthy resource.

[0053] In the embodiment, step S302 can specifically include: scanning each PCI slot on the PCI slot to determine whether the slot is installed with an entity acceleration resource, and if there is an entity acceleration resource, obtaining the PCI address corresponding to the acceleration resource. Specifically, only one entity acceleration resource can be installed in each PCI slot, and the PCI address includes the slot number. The types of entity acceleration resources can include GPU, NPU, FPGA, SmartNIC, etc.

[0054] By the method in this embodiment, the problem that in the related art, only the system itself can detect the traditional hardware, and the detection result of the various heterogeneous acceleration resources is inaccurate and inconvenient to manage, can be solved. By detecting the manufacturer information and the device model of the heterogeneous acceleration resources to call the corresponding interface, not only the accuracy of the detection result is improved, but also the unified management of the various heterogeneous acceleration resources is realized.

[0055] In another embodiment, the step S204 can specifically include: obtaining allocation data of the heterogeneous acceleration resources; and determining the use healthy resources and the allocation failure resources according to the allocation data.

[0056] In this embodiment, the determination of the use healthy resources or the allocation failure resources according to the allocation data includes: determining actual use data of the heterogeneous acceleration resources; and sequentially performing data comparison on the allocation data and the actual use data of each heterogeneous acceleration resource. If the allocation data and the actual use data are consistent, the heterogeneous acceleration resource is determined as the use healthy resource, otherwise, the heterogeneous acceleration resource is determined as the allocation failure resource.

[0057] In an embodiment, the device use health monitoring method in the step S204 can be performed according to a preset device use health detection period.

[0058] Specifically, each heterogeneous acceleration resource can be allocated to multiple customers by virtualization, and the allocation data includes an allocation customer and an allocation quantity, and the actual use data includes a use customer and a use quantity.

[0059] Further, the allocation customer and the use customer of each heterogeneous acceleration resource are compared, and the allocation quantity and the use quantity are compared. If all the data are consistent, the heterogeneous acceleration resource is determined as the use healthy resource, otherwise, the heterogeneous acceleration resource is determined as the allocation failure resource.

[0060] Figure 4 is a flowchart of the heterogeneous acceleration resource device use health monitoring method of the embodiment of the application, as shown in Figure 4 The heterogeneous acceleration resource device use health monitoring method specifically includes the following steps:

[0061] Step S402: calling a cloud platform heterogeneous acceleration resource interface to obtain allocation data details (including an allocation customer, an allocation quantity, etc.) of the heterogeneous acceleration resources;

[0062] Step S404: detecting each allocated acceleration resource;

[0063] Step S406: judging whether the corresponding customer exists. If the judgment result is yes, directly executing step S410, if the judgment result is no, executing step S408;

[0064] Step S408: add the heterogeneous acceleration resource to the allocation failure list, and record the customer of the abnormal allocation;

[0065] Step S410: determine whether there is a heterogeneous acceleration resource that has not been determined, if the result is yes, return to step S404, if the result is no, execute step S412;

[0066] Step S412: output the heterogeneous acceleration resource of the allocation failure

[0067] In the embodiment, each heterogeneous acceleration resource can be virtualized and allocated to multiple customers, and the customer types usually include virtual machines, bare machines, containers, etc.

[0068] The step S406 can specifically include determining whether the virtual machines, bare machines, and containers allocated to the heterogeneous acceleration resource exist, if all exist, determining that the customers allocated to the heterogeneous acceleration resource are normal.

[0069] By the method in the embodiment, the problem that the resource allocation registration and the actual use of the customers are inconsistent when the heterogeneous acceleration resource is virtualized and allocated in the related art can be solved, the heterogeneous acceleration resource of the allocation failure can be identified in time, and the virtualized and allocated heterogeneous acceleration resource is prevented from being allocated to multiple customers repeatedly, thereby ensuring the security and stability of the cloud computing platform.

[0070] In an embodiment, the allocation abnormality processing of the allocation failure resource includes: updating the allocation data of the allocation failure resource according to the actual use data, specifically, updating the allocation customer in the allocation data by using the use customer in the actual use data, and updating the allocation quantity in the allocation data by using the use quantity in the actual use data.

[0071] Figure 5 is a time sequence diagram of device use health monitoring and processing of an optional embodiment of the application, as shown in Figure 5 The heterogeneous acceleration resource device use health monitoring and processing method specifically includes the following steps:

[0072] Step S502: output the allocation failure resource according to the device use health monitoring method;

[0073] Step S504: call the response module to perform allocation abnormality processing on the allocation failure resource;

[0074] Step S506: update the heterogeneous acceleration resource information of the allocation failure resource;

[0075] Step S508: return the update result;

[0076] Step S510: return.

[0077] In another embodiment, the method for handling exceptions of heterogeneous acceleration resources further comprises performing exception alarm on the hardware non-healthy resources and the allocation failure resources.

[0078] In an embodiment, the hardware exception handling of the hardware non-healthy resources specifically comprises the following steps:

[0079] determining whether the usage state of the hardware non-healthy resource is unavailable, and if the result of the determination is no, setting the usage state of the hardware non-healthy resource as unavailable and setting the recovery state of the hardware non-healthy resource as recoverable;

[0080] determining whether the hardware non-healthy resource has been allocated to a customer, and if the result of the determination is yes, notifying the cloud computing platform to migrate the customer to which the hardware non-healthy resource has been allocated, and / or setting the recovery state of the hardware non-healthy resource as unrecoverable.

[0081] Specifically, the usage state of the heterogeneous acceleration resource is divided into available and unavailable, and the recovery state of the heterogeneous acceleration resource is divided into recoverable and unrecoverable. When the usage state of the heterogeneous acceleration resource is set, the system automatically records the source of the setting of the usage state. If the usage state is set by an administrator, it is marked as an administrator, and the corresponding recovery state is unrecoverable. If the usage state is automatically set by the exception response module, it is marked as a response module, and the corresponding recovery state is recoverable.

[0082] In this embodiment, notifying the cloud computing platform to migrate the customer to which the hardware non-healthy resource has been allocated specifically can include notifying the administrator associated with the cloud computing platform to timely determine the usage of the hardware non-healthy resource and perform a hot migration action (reallocate a normal heterogeneous acceleration resource to the customer) or other actions on all virtual machines, bare machines, containers and other customers that have used the hardware non-healthy resource.

[0083] In an embodiment, the exception resource information corresponding to the hardware non-healthy resource and the allocation failure resource can be obtained; the exception resource information is standardized to obtain standardized exception information; the standardized exception information is reported to the cloud computing platform, so as to timely notify relevant personnel to handle the exception information, and the standardized exception information can be stored in the cloud computing platform for subsequent searching.

[0084] In another embodiment, the standardized abnormal information can also be obtained from the cloud computing platform; the hardware health resource and the health resource information corresponding to the usage health resource are obtained; the health resource information is standardized to obtain standardized health information; the recoverable resource is determined from the standardized abnormal information according to the standardized health information; and if the recovery state of the recoverable resource is recoverable, the recoverable resource is recovered. Specifically, the standardized abnormal information and the standardized health information at least include the PCI address, the manufacturer information, the device model, the device ID, etc. of the heterogeneous acceleration resource, wherein the PCI address includes the slot number.

[0085] In the embodiment, the recoverable resource is determined from the standardized abnormal information according to the standardized health information, which includes: the standardized health information and the standardized abnormal information are matched according to a preset matching rule, wherein the preset matching rule includes matching at least one of the following resource information: the PCI address, the manufacturer information, and the model; and the heterogeneous acceleration resource corresponding to the matched standardized abnormal information is determined as the recoverable resource.

[0086] In the embodiment, the recovery processing of the recoverable resource specifically can include: if there is an abnormal alarm corresponding to the recoverable resource, the abnormal alarm is canceled; and the usage state of the recoverable resource is set to available.

[0087] Figure 6 is a timing diagram of the heterogeneous acceleration resource abnormal recovery processing according to an optional embodiment of the application, as shown in Figure 6 The heterogeneous acceleration resource abnormal recovery processing method specifically includes the following steps:

[0088] Step S601: the healthy heterogeneous acceleration resource is output according to the hardware health monitoring method;

[0089] Step S602: the healthy heterogeneous acceleration resource information is sent;

[0090] Step S603: the reported unhealthy heterogeneous acceleration resource information is obtained;

[0091] Step S604: the reported unhealthy heterogeneous acceleration resource information is returned;

[0092] Step S605: the recoverable heterogeneous acceleration resource is identified by a specific method;

[0093] Step S606: the heterogeneous acceleration resource information is standardized, and the alarm recovery interface of the cloud computing platform is called;

[0094] Step S607: return;

[0095] Step S608: it is judged whether the heterogeneous acceleration resource needs to be recovered to available;

[0096] Step S609: return;

[0097] In this embodiment, the specific method in step S605 can include: comparing data according to PCI address, manufacturer information, device model, device ID, official interface, or identifying through a specific algorithm.

[0098] In this embodiment, the step S608 of determining whether the heterogeneous acceleration resource needs to be restored to available can specifically include: determining according to the recovery state of the heterogeneous acceleration resource, and if the recovery state is recoverable, restoring the heterogeneous acceleration resource to available.

[0099] In another embodiment, the heterogeneous acceleration resource abnormal recovery processing method in steps S601 to S609 can be performed according to a preset recovery period.

[0100] According to the method for processing the abnormal recovery of the heterogeneous acceleration resource in this embodiment, when the abnormal condition of the heterogeneous acceleration resource is detected, an alarm prompt can be sent to the customer and the cloud computing platform administrator in time, so as to avoid causing serious loss. In addition, through human intervention or system automatic processing, the abnormal heterogeneous acceleration resource can be restored to a healthy state. For this case, this embodiment can automatically restore the heterogeneous acceleration resource to an available state, respond in time, process quickly, reduce the adverse effects on the using customer, and improve the reliability of the cloud computing platform.

[0101] According to another aspect of the embodiment of the application, a heterogeneous acceleration resource abnormal processing device is also provided, Figure 7 is a block diagram of the heterogeneous acceleration resource abnormal processing device of the embodiment of the application, as Figure 7 shown, the device includes:

[0102] The first monitoring module 702 is configured to determine the heterogeneous acceleration resource as a hardware healthy resource or a hardware unhealthy resource by monitoring the hardware health of the heterogeneous acceleration resource of the cloud computing platform.

[0103] The second monitoring module 704 is configured to determine the heterogeneous acceleration resource as a use healthy resource or an allocation failure resource by monitoring the device use health of the heterogeneous acceleration resource.

[0104] The first response module 706 is configured to perform hardware abnormal processing on the hardware unhealthy resource.

[0105] The second response module 708 is configured to perform allocation abnormal processing on the allocation failure resource.

[0106] In an embodiment, the device further includes:

[0107] The scanning module is configured to determine whether the heterogeneous acceleration resource exists by scanning a PCI slot.

[0108] The first obtaining module is configured to obtain resource information of the heterogeneous acceleration resource if the heterogeneous acceleration resource exists.

[0109] In an embodiment, the first monitoring module 702 further includes:

[0110] The calling unit is configured to call a corresponding hardware health detection interface according to the resource information of the heterogeneous acceleration resource.

[0111] The detection unit is configured to determine a hardware state of the heterogeneous acceleration resource through the hardware health detection interface.

[0112] The first judging unit is configured to determine that the heterogeneous acceleration resource is the hardware healthy resource if the hardware state is healthy, and determine that the heterogeneous acceleration resource is the hardware unhealthy resource if the hardware state is unhealthy.

[0113] In an embodiment, the apparatus further includes:

[0114] The exception alarm module is configured to perform exception alarm on the hardware unhealthy resource and the allocation failure resource.

[0115] In an embodiment, the second monitoring module 704 further includes:

[0116] The first obtaining unit is configured to obtain allocation data of the heterogeneous acceleration resource.

[0117] The second judging unit is configured to determine the usage healthy resource and the allocation failure resource according to the allocation data.

[0118] In an embodiment, the second judging unit further includes:

[0119] The second obtaining unit is configured to determine actual usage data of the heterogeneous acceleration resource.

[0120] The data comparison unit is configured to sequentially compare the allocation data and the actual usage data of each heterogeneous acceleration resource, and determine that the heterogeneous acceleration resource is the usage healthy resource if the allocation data and the actual usage data are consistent, or determine that the heterogeneous acceleration resource is the allocation failure resource otherwise.

[0121] In an embodiment, the second response module 708 is further configured to:

[0122] Update the allocation data of the allocation failure resource according to the actual usage data.

[0123] In an embodiment, the first response module 706 further includes:

[0124] a setting unit, configured to determine whether the usage state of the hardware unhealthy resource is unavailable, and if the determination result is no, set the usage state of the hardware unhealthy resource as unavailable, and set the recovery state of the hardware unhealthy resource as recoverable;

[0125] a processing unit, configured to determine whether the hardware unhealthy resource has been allocated to a customer, and if the determination result is yes, notify the cloud computing platform to migrate the customer to which the hardware unhealthy resource has been allocated, and / or set the recovery state of the hardware unhealthy resource as unrecoverable.

[0126] In an embodiment, the apparatus further includes:

[0127] a second acquisition module, configured to acquire abnormal resource information corresponding to the hardware unhealthy resource and the allocation failure resource;

[0128] a first standardization module, configured to perform standardization processing on the abnormal resource information to obtain standardized abnormal information;

[0129] a reporting module, configured to report the standardized abnormal information to the cloud computing platform.

[0130] In an embodiment, the apparatus further includes:

[0131] a third acquisition module, configured to acquire the standardized abnormal information from the cloud computing platform;

[0132] a fourth acquisition module, configured to acquire healthy resource information corresponding to the hardware healthy resource and the usage healthy resource;

[0133] a second standardization module, configured to perform standardization processing on the healthy resource information to obtain standardized healthy information;

[0134] a recovery determination module, configured to determine recoverable resources from the standardized abnormal information according to the standardized healthy information;

[0135] a recovery processing module, configured to perform recovery processing on the recoverable resources if the recovery state of the recoverable resources is recoverable.

[0136] In an embodiment, the recovery determination module includes:

[0137] a matching unit, configured to match the standardized healthy information and the standardized abnormal information according to a preset matching rule, wherein the preset matching rule includes matching at least one of the following resource information: PCI address, manufacturer information, model number;

[0138] The recovery judging unit is configured to determine the heterogeneous acceleration resource corresponding to the standardized abnormal information with the successful matching as the recoverable resource.

[0139] In an embodiment, the recovery processing module comprises:

[0140] The canceling unit is configured to cancel the abnormal alarm if the abnormal alarm corresponding to the recoverable resource exists.

[0141] The recovery unit is configured to set the usage state of the recoverable resource as available.

[0142] According to another aspect of the embodiments of the present application, a heterogeneous acceleration resource health monitoring and exception handling architecture is further provided.

[0143] Figure 8 The heterogeneous acceleration resource health monitoring and exception handling architecture of the embodiments of the present application is shown in FIG. 8, which comprises: Figure 8

[0144] The health identification module 81 comprises a hardware health monitoring module 811, a device usage health monitoring module 812, and a cloud platform heterogeneous resource usage interface 813.

[0145] The exception handling module 82 comprises an abnormal alarm module 821, an abnormal response module 822, an abnormal recovery module 823, a cloud platform alarm interface 824, and a cloud platform heterogeneous resource management interface 825.

[0146] In the embodiments, the hardware health monitoring module 811 is configured to realize part or all of the functions of the first monitoring module 702; the device usage health monitoring module 812 is configured to realize part or all of the functions of the second monitoring module 704; and the cloud platform heterogeneous resource usage interface 813 is configured to realize part or all of the functions of the second obtaining unit.

[0147] Specifically, the hardware health monitoring module 811 is configured to determine the heterogeneous acceleration resource as a hardware healthy resource or a hardware unhealthy resource by means of hardware health monitoring of the heterogeneous acceleration resource of the cloud computing platform; the device usage health monitoring module 812 is configured to determine the heterogeneous acceleration resource as a usage healthy resource or an allocation failure resource by means of device usage health monitoring of the heterogeneous acceleration resource; and the cloud platform heterogeneous resource usage interface 813 is configured to determine the actual usage data of the heterogeneous acceleration resource.

[0148] ​In another embodiment, the anomaly alarm module 821 is configured to alarm for the hardware non-healthy resource and the allocation failure resource; the anomaly response module 822 is configured to implement part or all of the functions of the first response module 706 and the second response module 708, including being configured to perform anomaly processing on the hardware non-healthy resource and the allocation failure resource; the cloud platform alarm interface 824 is configured to inform the cloud computing platform of the anomaly alarm information; and the cloud platform heterogeneous resource management interface 825 is configured to manage the heterogeneous acceleration resource, including being configured to set the usage state of the heterogeneous acceleration resource.

[0149] By the embodiments of the present application, the problem that only the traditional server ordinary hardware resource detection is focused on in the related art, the inconsistency between the registration and actual use of the virtualized heterogeneous acceleration resource managed by the cloud computing platform cannot be identified, and thus the cloud computing platform and the user are brought losses can be solved. When the heterogeneous acceleration resource is abnormal, the non-healthy state of the heterogeneous acceleration resource can be quickly perceived and timely alarmed and recovered, and the reliability, stability, timeliness, and the like of the cloud platform in managing the heterogeneous acceleration resource can be ensured.

[0150] The embodiments of the present application further provide a computer readable storage medium, which stores a computer program, and the computer program is configured to execute the steps in any one of the method embodiments when running.

[0151] In an example embodiment, the computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0152] The embodiments of the present application further provide an electronic device, which includes a memory storing a computer program and a processor configured to execute the computer program to perform the steps in any one of the method embodiments.

[0153] In an example embodiment, the electronic device can further include a transmission device connected to the processor and an input and output device connected to the processor.

[0154] The specific examples in the embodiments can refer to the examples described in the above embodiments and example embodiments, and the embodiments will not be described here again.

[0155] It should be apparent to those skilled in the art that the modules or steps of the application described above can be implemented with general computing devices, which can be centralized on a single computing device or distributed on a network of multiple computing devices, and which can be implemented with program codes executable by the computing devices, so that they can be stored in storage devices and executed by the computing devices, and in some cases, the steps shown or described can be executed in different orders than shown, or made into individual integrated circuit modules, or made into a single integrated circuit module. Thus, the present application is not limited to any particular combination of hardware and software.

[0156] The preferred embodiments of the present application described above are only used to explain the principles of the present application and not limit the present application. Any modification, equivalent replacement, improvement, etc. within the principles of the present application should be included in the protection scope of the present application.

Claims

1. A method for handling exceptions of heterogeneous acceleration resources, the method comprising: The method comprises: determining the heterogeneous acceleration resources as hardware healthy resources or hardware unhealthy resources by means of hardware health monitoring of the heterogeneous acceleration resources of the cloud computing platform; determining the heterogeneous acceleration resources as usage healthy resources or allocation failure resources by means of device usage health monitoring of the heterogeneous acceleration resources, comprising: obtaining allocation data of the heterogeneous acceleration resources; determining the usage healthy resources and the allocation failure resources according to the allocation data; wherein determining the usage healthy resources or the allocation failure resources according to the allocation data comprises: determining actual usage data of the heterogeneous acceleration resources; sequentially comparing the allocation data and the actual usage data of each of the heterogeneous acceleration resources, and if the allocation data and the actual usage data are consistent, determining the heterogeneous acceleration resources as usage healthy resources, otherwise, determining the heterogeneous acceleration resources as allocation failure resources; performing hardware exception processing on the hardware unhealthy resources; performing allocation exception processing on the allocation failure resources, comprising: updating the allocation data of the allocation failure resources according to the actual usage data; before determining the heterogeneous acceleration resources as hardware healthy resources or hardware unhealthy resources by means of hardware health monitoring of the heterogeneous acceleration resources of the cloud computing platform, the method further comprises: determining whether the heterogeneous acceleration resources exist by scanning PCI slots; if the heterogeneous acceleration resources exist, obtaining resource information of the heterogeneous acceleration resources, comprising: identifying the resource information of the heterogeneous acceleration resources in combination with the configuration of the cloud computing platform, the heterogeneous acceleration resources comprising: GPU, NPU, FPGA, and Smart NIC, and the resource information of the heterogeneous acceleration resources comprising: PCI address, vendor information, device model, and device ID, wherein the PCI address comprises slot number.

2. The method of claim 1, wherein, determining the heterogeneous acceleration resources as hardware healthy resources or hardware unhealthy resources by means of hardware health monitoring of the heterogeneous acceleration resources of the cloud computing platform, comprising: calling a corresponding hardware health detection interface according to the resource information of the heterogeneous acceleration resources; judging the hardware state of the heterogeneous acceleration resources through the hardware health detection interface; if the hardware state is healthy, determining the heterogeneous acceleration resources as the hardware healthy resources; if the hardware state is unhealthy, determining the heterogeneous acceleration resources as the hardware unhealthy resources.

3. The method of claim 1, wherein, The method further comprises: performing exception alarm on the hardware unhealthy resources and the allocation failure resources.

4. The method of claim 1, wherein, performing hardware exception processing on the hardware unhealthy resources, comprising: judging whether the usage state of the hardware unhealthy resources is unavailable, and if the result of the judgment is no, setting the usage state of the hardware unhealthy resources as unavailable and setting the recovery state of the hardware unhealthy resources as recoverable; judging whether the hardware unhealthy resources have been allocated to a customer, and if the result of the judgment is yes, notifying the cloud computing platform to migrate the customer to which the hardware unhealthy resources have been allocated, and / or setting the recovery state of the hardware unhealthy resources as unrecoverable.

5. The method of claim 1, wherein, The method further comprises: Obtain abnormal resource information corresponding to the hardware unhealthy resource and the allocation failure resource; Standardize the abnormal resource information to obtain standardized abnormal information; Report the standardized abnormal information to a cloud computing platform.

6. The method of claim 5, wherein, The method further includes: Obtain the standardized abnormal information from the cloud computing platform; Obtain health resource information corresponding to the hardware healthy resource and the use healthy resource; Standardize the health resource information to obtain standardized health information; Determine a recoverable resource from the standardized abnormal information according to the standardized health information; If a recovery state of the recoverable resource is recoverable, perform recovery processing on the recoverable resource.

7. The method of claim 6, wherein, Determine a recoverable resource from the standardized abnormal information according to the standardized health information, including: Match the standardized health information and the standardized abnormal information according to a preset matching rule, wherein the preset matching rule includes matching at least one of the following resource information: PCI address, manufacturer information, model number; Determine a heterogeneous acceleration resource corresponding to the standardized abnormal information that matches successfully as the recoverable resource.

8. The method of claim 6, wherein, If a recovery state of the recoverable resource is recoverable, perform recovery processing on the recoverable resource, including: If there is an abnormal alarm corresponding to the recoverable resource, cancel the abnormal alarm; Set a use state of the recoverable resource to available.

9. An apparatus for handling exceptions of heterogeneous acceleration resources, the apparatus comprising: The device includes: A first monitoring module configured to determine, by hardware health monitoring of heterogeneous acceleration resources of a cloud computing platform, whether the heterogeneous acceleration resources are hardware healthy resources or hardware unhealthy resources; A second monitoring module configured to determine, by device use health monitoring of the heterogeneous acceleration resources, whether the heterogeneous acceleration resources are use healthy resources or allocation failure resources, including: obtaining allocation data of the heterogeneous acceleration resources; determining the use healthy resources and the allocation failure resources according to the allocation data; wherein determining the use healthy resources or the allocation failure resources according to the allocation data includes: determining actual use data of the heterogeneous acceleration resources; sequentially performing data comparison on the allocation data and the actual use data of each of the heterogeneous acceleration resources, and if the allocation data and the actual use data are consistent, determining that the heterogeneous acceleration resources are use healthy resources, otherwise, determining that the heterogeneous acceleration resources are allocation failure resources; A first response module configured to perform hardware abnormality processing on the hardware unhealthy resources; A second response module configured to perform allocation abnormality processing on the allocation failure resources, including: performing data update on allocation data of the allocation failure resources according to the actual use data; The device further includes: A scanning module configured to determine whether the heterogeneous acceleration resources exist by scanning PCI slot positions; The first obtaining module is configured to, if the heterogeneous acceleration resource exists, obtain resource information of the heterogeneous acceleration resource, including: identifying the resource information of the heterogeneous acceleration resource in combination with a configuration of a cloud computing platform, the heterogeneous acceleration resource including: a GPU, an NPU, an FPGA, and a Smart NIC, and the resource information of the heterogeneous acceleration resource including: a PCI address, vendor information, a device model, and a device ID, the PCI address including a slot number.

10. A computer-readable storage medium having stored therein a computer program, wherein, The computer program is configured to execute the method of any one of claims 1-8 when running. 11.An electronic device comprising a memory and a processor, the memory having stored therein a computer program, the processor being configured to execute the computer program to perform the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Optimal allocation of dynamic cloud computing platform resources

    CN107743611A

  • Embedded reconfigurable heterogeneous measurement method and system, storage medium and processor

    CN111694789A