Equipment fault processing method, electronic equipment, storage medium and program product

By isolating and resetting the failed GPU, the problem of low GPU fault repair efficiency is solved, and rapid automated repair is achieved, which improves the repair success rate and avoids data loss.

CN120256187AActive Publication Date: 2025-07-04ALIBABA CLOUD COMPUTING CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510736973.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

In the prior art, GPU fault repair efficiency is low, especially the fault repair caused by dual-bit ECC errors, and data loss may occur. The existing method of restarting the server takes 10 to 20 minutes and has a low success rate.

Method used

By isolating the failed GPU, blocking the access task, and returning to normal working state after the reset is successful, the entire process is automated and completed within 10 to 15 seconds.

Benefits of technology

It realizes automatic and rapid repair of faulty GPUs, improves repair efficiency, avoids data loss, and does not require human participation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256187A_ABST
    Figure CN120256187A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an equipment fault processing method, electronic equipment, a storage medium and a program product, and relates to the technical field of computers.The method comprises the steps that under the condition that any graphics processor in a graphics processor set has a target fault, access isolation is conducted on a fault graphics processor with the target fault, the access isolation is used for blocking an access task aiming at the fault graphics processor; the fault graphics processor is reset, under the condition that resetting succeeds, the fault graphics processor is recovered to the target state, and the target state is the normal work state. According to the technical scheme provided by the embodiment of the invention, on the basis of avoiding data loss, automatic and rapid fault repair is realized, and the fault repair efficiency of the graphics processor is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method for processing device failures, an electronic device, a storage medium, and a program product, which can be applied to the field of fault processing technology. Background Art

[0002] Graphics Processing Unit (GPU) is widely used in the field of artificial intelligence due to its excellent parallel processing ability. For example, during the model training process, a single server is usually configured with 8 to 16 GPUs. However, GPUs may often fail during operation, and the failure caused by double-bit Error-Correcting Code (ECC) error is one of the typical failures. Since double-bit ECC error is an uncorrectable error, that is, the GPU cannot self-correct, after a double-bit ECC error occurs in the GPU, the server where the GPU is located will be removed from the scheduling domain of the training cluster. This means that when a single GPU has a double-bit ECC error, all the GPUs deployed on the server will be idle. In this regard, in the related art, the server is manually restarted to repair the GPU failure. However, restarting the server takes 10 to 20 minutes, the repair efficiency of the failure is low, and there may be problems such as data loss. Summary of the Invention

[0003] Embodiments of this application provide a method for processing device failures, an electronic device, a storage medium, and a program product to alleviate or solve one or more technical problems existing in the prior art.

[0004] In a first aspect, embodiments of this application provide a method for processing device failures. The method includes: when any graphics processor in a set of graphics processors has a target failure, performing access isolation on the failed graphics processor with the target failure, where the access isolation is used to block access tasks for the failed graphics processor; resetting the failed graphics processor, where, when the reset is successful, the failed graphics processor returns to a target state, and the target state is a normal working state.

[0005] In a second aspect, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored on the memory. The processor implements the method according to any one of the embodiments of this application when executing the computer program.

[0006] In a third aspect, embodiments of this application provide a computer-readable storage medium, in which a computer program is stored. The computer program implements the method according to any one of the embodiments of this application when executed by a processor.

[0007] Fourthly, an embodiment of the present application provides a computer program product, including a computer program which, when executed by a processor, implements the method of any one of the embodiments of the present application.

[0008] In the device fault handling method provided by the embodiment of the present application, when any graphics processor in the set of graphics processors has a target fault, access isolation is performed on the faulty graphics processor with the target fault to block the access task for the faulty graphics processor, and the faulty graphics processor is reset so that the faulty graphics processor returns to the target state. This can not only improve the success rate of resetting, but also does not require human participation. The whole process takes about 10 to 15 seconds. On the basis of avoiding data loss, it realizes the automatic and rapid repair of faults, that is, improves the fault repair efficiency of the graphics processor.

[0009] For an overview of the technical solution, in order to be able to more clearly understand the technical means of the present application, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically illustrates the specific embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In the drawings, unless otherwise specified, the same reference numerals throughout the several views denote the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be regarded as limiting the scope of the present application.

[0011] Figure 1 Shows a schematic diagram of the fault distribution of the graphics processor; Figure 2 Shows a schematic diagram of the application scenario of the device fault handling method according to the embodiment of the present application; Figure 3 Shows a flowchart of the device fault handling method according to the embodiment of the present application; Figure 4 Shows a schematic flow diagram of the device fault handling method according to the embodiment of the present application; Figure 5 Shows a schematic diagram of the modules of the device fault handling device according to the embodiment of the present application; Figure 6 Shows a block diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0012] In the following, only some exemplary embodiments are simply described. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the concept or scope of the present application. Therefore, the drawings and the description are considered to be exemplary in nature and not restrictive.

[0013] To facilitate the understanding of the technical solutions of the embodiments of the present application, the related technologies of the embodiments of the present application are described below. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application.

[0014] The following terms will be used hereinafter: GPU: A processor for processing graphics and parallel computing tasks, which can be applied to accelerate computing tasks that require high parallelism, such as graphics rendering, deep learning, scientific computing, and data analysis.

[0015] ECC: A technology for detecting and correcting errors, which can detect and correct single-bit errors (Single-Bit Error, SBE) in data during storage and transmission, and can detect but not repair multi-bit errors (Multi-Bit Error, MBE) in data during storage and transmission. Among them, single-bit errors can also be called single-bit ECC errors, and multi-bit errors can also be called multi-bit ECC errors, such as double-bit ECC errors, etc.

[0016] Reset refers to the reset operation of a system or device, which can be a hardware restart (such as pressing a restart button) or a software reset (such as restarting the system through a command), and is used to restore the system to the target state and clear errors.

[0017] Reattach: Generally refers to the process of reconnecting or attaching to a certain device or resource. In the field of computing, it may involve reconnecting storage devices, network connections, or other external devices.

[0018] ECC Page Retirement: The video memory page is marked as "retired" due to excessive ECC errors.

[0019] With the continuous development of artificial intelligence technology, the demand for high-performance computing resources has also been continuously increasing. Due to its better parallel processing ability, GPU has become the main device to meet this demand. For example, during model training, usually 8 to 16 GPUs are configured in a server. However, this dependence on GPUs also brings new challenges, that is, various types of failures may often occur during the operation of GPUs.

[0020] Such as Figure 1As shown, GPU failures can include failures caused by double-bit ECC errors accounting for 64.4%, single-bit ECC errors accounting for 9.2%, and other errors accounting for 26.4%. It can be seen that double-bit ECC errors are common failures of GPUs. Since double-bit ECC errors are uncorrectable errors, that is, the GPU cannot correct itself, after a double-bit ECC error occurs in the GPU, the server where the GPU is located will be removed from the scheduling domain of the training cluster. And this means that when a single GPU has a double-bit ECC error, all GPUs deployed on the server will be idle.

[0021] In response to this, several coping methods have been proposed in related technologies. One coping method is to manually restart the server. By restarting the server, all hardware devices such as GPUs can be re-initialized, thus solving problems such as double-bit ECC errors, video memory leaks, and card drops. However, restarting the server takes 10 to 20 minutes, the fault repair efficiency is low, and there may be problems such as data loss.

[0022] Another coping method is to directly reset the GPU when a double-bit ECC error occurs in the GPU. However, as long as there is a process other than the reset process using the GPU, it will cause the GPU reset to fail. In an actual computing environment, various monitoring components are usually deployed to monitor the performance of the GPU. These monitoring components will continuously occupy the GPU. Even if all these monitoring components are stopped, new processes will soon access the GPU, so the success rate of resetting is relatively low. Currently, the success rate of resetting is often less than 10%.

[0023] There is also a coping method called Reattach, that is, after all processes occupying the GPU exit, the GPU driver is uninstalled; and when a new process accesses the GPU, the GPU driver is reloaded to achieve the reconnection of the GPU. However, there are the following deficiencies in this method. On the one hand, for GPU architectures after Ampere and later, the GPU supports the Row-Remapping function, and each memory bank in the HBM (High Bandwidth Memory) is equipped with a SpareRow in hardware. Different from the traditional PageRetirement, Row-Remapping uses the spare row to replace the memory unit with an ECC error to avoid holes when managing video memory in software. However, Row-Remapping needs to be reset to take effect. On the other hand, after the reset is successful, the relevant processes of the GPU need to be manually started by the user, that is, user participation is required, which increases the complexity.

[0024] Based on this, the embodiments of the present application provide a device fault handling method, an electronic device, a storage medium, and a program product, aiming to improve the fault repair efficiency of the graphics processor.Figure 2 FIG. 0 shows a schematic diagram of an application scenario of a device fault handling method provided by an embodiment of the present application. As Figure 2 shown, this scenario includes an electronic device, in which a graphics processing unit (GPU) cluster is configured, and the GPU cluster includes at least one graphics processing unit (GPU).

[0025] Exemplarily, the electronic device can be a terminal device or a server. The terminal device can be a mobile phone, a tablet computer, a desktop computer, a portable notebook, a vehicle-mounted terminal, etc. The server can be a physical server or a cloud server for cloud computing, etc. Figure 2 In this example, the electronic device is taken as a physical server for illustrative purposes.

[0026] The electronic device maintains a fault handling policy, which is used to indicate that in the case where any graphics processing unit in the set of graphics processing units has a target fault, access isolation is performed on the faulty graphics processing unit with the target fault, and then the faulty graphics processing unit is reset. Among them, access isolation is used to block access tasks for the faulty graphics processing unit. And in the case where the faulty image processor is successfully reset, the graphics processing unit returns to the target state, that is, the state before the target fault occurs. The target fault includes but is not limited to the double-bit ECC error, single-bit ECC error described above and other types of errors described below.

[0027] Thus, by performing access isolation on the faulty graphics processing unit with the target fault to block access tasks for the faulty graphics processing unit and resetting the faulty graphics processing unit, the faulty graphics processing unit can be restored to the target state. This not only improves the success rate of resetting, but also does not require human participation. The entire process takes about 10 to 15 seconds, achieving automatic and rapid repair of the fault on the basis of avoiding data loss, that is, improving the fault repair efficiency of the graphics processing unit.

[0028] It should be noted that the above application scenarios or application examples provided in the embodiments of the present application are for easy understanding, and the embodiments of the present application do not specifically limit the application of the technical solutions. In addition, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0029] The following uses specific embodiments to elaborate in detail on the technical solution of this application and how the technical solution of this application solves the foregoing technical problems. The several specific embodiments listed may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of this application in detail with reference to the accompanying drawings.

[0030] Figure 3 The flowchart of the device fault handling method according to the embodiment of this application is shown. As Figure 3 shown, this method may include step S301 and step S302.

[0031] Step S301: When any graphics processor in the set of graphics processors has a target fault, perform access isolation on the faulty graphics processor with the target fault, where the access isolation is used to block access tasks for the faulty graphics processor.

[0032] In order to be able to repair the faulty graphics processor with the target fault in a timely manner, in some embodiments, a repair tool with fault detection and fault repair functions may be configured in the electronic device. When the electronic device starts successfully, the electronic device can start this repair tool and perform fault detection on the set of graphics processors deployed in the electronic device through this repair tool.

[0033] In other embodiments, when the electronic device starts successfully, the electronic device can start a repair process with fault detection and fault repair functions and perform fault detection on the set of graphics processors deployed in the electronic device through this repair process.

[0034] Considering that when the graphics processor is not occupied, performing a reset operation on the graphics processor will increase the success rate of the reset. Based on this, when the detection result of the fault detection indicates that any graphics processor in the set of graphics processors has a target fault, access isolation is performed on the faulty graphics processor with the target fault through the above repair tool or repair process to block access tasks for the faulty graphics processor. The access tasks may include access processes, access threads, etc.

[0035] Among them, the target fault may be a fault caused by double-bit ECC error, ECC Page Retirement, Remapped Rows failure, etc. The following uses the double-bit ECC error as an example for illustration.

[0036] Step S302: Reset the faulty graphics processor. Among them, when the reset is successful, the faulty graphics processor returns to the target state, and the target state is a state where it can work normally.

[0037] In some embodiments, the electronic device may execute a reset instruction according to the device identifier of the faulty graphics processing unit with the target fault through the above repair tool or repair process to reset the faulty graphics processing unit and obtain a reset result. When the reset result indicates successful reset, the target fault is repaired and the corresponding faulty graphics processing unit is restored to the target state. When the reset result indicates failed reset, the electronic device may generate a reset failure record and report it to the designated operation and maintenance personnel. The reset failure record may include information about the target fault, detailed information about the failed reset, and so on.

[0038] It should be noted that the reset instruction may vary depending on the manufacturer of the graphics processing unit, and no specific limitation is made in this application. Accordingly, the operations performed on the graphics processing unit by different reset instructions and the corresponding target states may also be different. Exemplarily, for the graphics processing unit produced by manufacturer A, the operations performed on the corresponding faulty graphics processing unit after executing the reset instruction may include terminating all running CUDA (Compute Unified Device Architecture) processes, clearing the video memory, re-initializing the driver of the faulty graphics processing unit, etc.; accordingly, the target state may be a normal working state where there are no running CUDA tasks, the video memory is in a cleared state, and the driver of the graphics processing unit is set to default parameters. For the graphics processing unit produced by manufacturer B, the operations performed on the corresponding faulty graphics processing unit after executing the reset instruction may include terminating the computing task, clearing the video memory, resetting the error counter to zero, etc.; accordingly, the target state may be a normal working state where there are no running computing tasks, the video memory is in a cleared state, and the error counter is in a zeroed state.

[0039] It should be further noted that the target state in the embodiments of this application is not the factory state, that is, it will not delete the driver or physical hardware of the graphics processing unit, but only restore the configuration of the graphics processing unit to the basic configuration that can work normally, so that the graphics processing unit is in a state where it can work normally.

[0040] Thus, when any graphics processing unit in the set of graphics processing units has a target fault, by isolating the access to the graphics processing unit with the target fault to block the access task for this graphics processing unit and resetting this graphics processing unit, this graphics processing unit can be restored to the target state. This not only improves the success rate of reset; but also does not require human participation. The entire process takes about 10 to 15 seconds. On the basis of avoiding data loss, it realizes the automated and rapid repair of faults, that is, it improves the fault repair efficiency of the graphics processing unit.

[0041] In order to timely detect the graphics processing unit with the target fault, in some embodiments, such asFigure 4 As shown, fault detection can be performed on each graphics processor in the set of graphics processors. That is to say, before step S301, it can also include: for any graphics processor in the set of graphics processors, obtaining the operating status data of the graphics processor according to a preset detection period; when the preset fault identifier is included in the operating status data, determining that the graphics processor is a faulty graphics processor with a target fault.

[0042] In some implementation manners, the operating status data can be log data, and the log data can include fault logs, service logs, security logs, etc. of each graphics processor in the set of graphics processors, and each log is arranged in the order of the generation time. Correspondingly, the electronic device can query the candidate log data within the corresponding detection period according to a preset query path according to the preset detection period; or execute a second query instruction to obtain the candidate log data within the corresponding detection period. For any graphics processor in the set of graphics processors, query the target log data including the device identifier of the graphics processor from the candidate log data. Determine whether the preset fault identifier is included in the target log data. When the preset fault identifier is included in the target log data, determine that the graphics processor is a faulty graphics processor with a target fault. Wherein, the second query instruction is used to query the candidate log data of each graphics processor in the set of graphics processors within the corresponding detection period.

[0043] In other implementation manners, the operating status data can be log data, and the log data is divided into at least one log subset, and the at least one log subset corresponds to at least one graphics processor in the set of graphics processors one by one, and the logs in any log subset are arranged in the order of the generation time. Correspondingly, for any graphics processor in the set of graphics processors, the electronic device can obtain the log subset corresponding to the graphics processor from the log data according to the device identifier of the graphics processor according to the detection period. Obtain the target log data within the corresponding detection period from the log subset, and determine whether the preset fault identifier is included in the target log data. When the preset fault identifier is included in the target log data, determine that the graphics processor is a faulty graphics processor with a target fault.

[0044] In still other implementation manners, the operating status data can be fault data. Correspondingly, for any graphics processor in the set of graphics processors, the electronic device can execute a third query instruction according to the detection period to query the fault data of the graphics processor within the corresponding detection period. Determine whether the preset fault identifier is included in the fault data. When the preset fault identifier is included in the fault data, determine that the graphics processor is a faulty graphics processor with a target fault. Wherein, the third query instruction is used to query the fault data of any graphics processor within the detection period.

[0045] Exemplarily, the target fault may be a fault caused by a double-bit ECC error. Correspondingly, when the running state data is log data, the preset fault identifier may be a specific type identifier used to indicate the log data, such as "XID48"; when the running state data is fault data, the preset fault identifier may be an identifier used to indicate the error type, such as "ECC Errors" (used to indicate ECC errors), "Uncorrectable Errors" (used to indicate uncorrectable errors), etc.

[0046] Thus, by periodically obtaining the running state data of each graphics processor in the graphics processor set during the corresponding detection period and determining whether the preset fault identifier is included in the running state data, the effective detection of the target fault is achieved, which can provide guarantee for timely fault repair.

[0047] Considering that in practical applications, when a task occupies a graphics processor, the success rate of resetting the graphics processor is often relatively low. Based on this, in some embodiments, as Figure 4 shown, the access isolation of the graphics processor with the target fault includes both preventing new access tasks and terminating the ongoing access tasks. That is, the access isolation of the graphics processor in the foregoing step S301 may include: configuring the access state of the faulty graphics processor to an inaccessible state, and this inaccessible state is used to prevent new access tasks for the faulty graphics processor; and, when the faulty graphics processor is currently in an accessed state, terminating the access task that is accessing the faulty graphics processor.

[0048] To avoid having new access tasks to access the faulty graphics processor with the target fault when the graphics processor has a target fault, in some implementation manners, when it is determined that any graphics processor has a target fault, the access state of the faulty graphics processor may be first configured to an inaccessible state. Then, it is detected whether the faulty graphics processor is currently in an accessed state. When the faulty graphics processor is currently in an accessed state, it indicates that there is an access task accessing the faulty graphics processor, and a termination signal is sent to the access task that is accessing the faulty graphics processor to terminate the access task that is accessing the faulty graphics processor.

[0049] Exemplarily, the electronic device may call the kill command to send a termination signal to the access task that is accessing the faulty graphics processor, where the kill command is a command for managing processes in some operating systems and is used to send a signal to a process.

[0050] It should be noted that the execution order of the operation of configuring the access status of the faulty graphics processor to an inaccessible state and the operation of detecting whether the faulty graphics processor is currently in an accessed state can be interchanged, or they can be executed simultaneously.

[0051] Thus, for the graphics processor with a target fault, by preventing its new access tasks and terminating the ongoing access tasks, it is ensured that the graphics processor with a target fault is in an unoccupied state, providing favorable conditions for subsequent reset operations and improving the accuracy of the reset operations.

[0052] To ensure that the graphics processors without target faults in the graphics processor set can be accessed normally, in some embodiments, configuring the access status of the faulty graphics processor to an inaccessible state may include: querying a pointer to an access function according to the function name of the access function, where the access function is used to access any graphics processor in the graphics processor set. Modifying the pointed object of the pointer to a proxy function, where the proxy function includes the device identifier of the faulty graphics processor with a target fault to indicate that the access status of the faulty graphics processor with a target fault is inaccessible.

[0053] Among them, the proxy function can include the device identifier of the graphics processor with a target fault in different ways. In some implementation manners, a proxy function and a target array corresponding to the proxy function may be predefined, where the proxy function includes the processing logic of access requests, and the target array may include the device identifiers and corresponding status identifiers of each graphics processor in the graphics processor set. Correspondingly, when it is determined that any graphics processor in the graphics processor set has a target fault, an update function interface can be called to update the status identifier corresponding to the device identifier of the faulty graphics processor with a target fault in the target array to preset data to indicate that the access status of the corresponding faulty graphics processor is inaccessible. And, according to the function name of the preset access function, a query interface is called to obtain a pointer to the access function. Modifying the currently pointed object of the pointer (i.e., the access function) to a proxy function. Wherein, the query interface is used to query the corresponding function according to the function name. Exemplarily, the preset data is 1, and the function name "gpu_open" of the access function currently pointed to by the pointer is modified to the function name "gpu_dlopen" of the proxy function.

[0054] In some other implementations, a proxy function and a target array corresponding to the proxy function can be predefined. The proxy function includes the processing logic of the access request, and the target array can include the device identifiers of the faulty graphics processors with target faults. Correspondingly, when it is determined that any one of the graphics processors in the graphics processor set has a target fault, the add function interface can be called to add the device identifier of the faulty graphics processor with the target fault to the target array, so as to represent that the access status of the faulty graphics processor with the target fault is an inaccessible state. And, according to the function name of the preset access function, the query interface is called to obtain a pointer to the access function. The currently pointed object of this pointer (i.e., the access function) is modified to the proxy function.

[0055] Thus, by querying the pointer to the access function and modifying its pointed object, that is, hijacking the access function and redirecting it to the proxy function, and injecting the device identifier of the graphics processor with the target fault into the proxy function, it is possible to block new access tasks for the graphics processor with the target fault based on the proxy function subsequently, and ensure that other graphics processors in the graphics processor set (i.e., graphics processors other than the graphics processor with the target fault) can be accessed normally.

[0056] Further, in some embodiments, the method of the embodiments of the present application may further include: returning an access failure message when the call information of the proxy function is obtained and the device identifier in the call information is included in the proxy function; or, calling the access function when the call information of the proxy function is obtained and the device identifier in the call information is not included in the proxy function.

[0057] Specifically, after the pointer to the access function is modified to point to the proxy function, any access task (such as an access process, an access thread, etc.) needs to call the proxy function when accessing the graphics processor. Correspondingly, when the electronic device obtains the call information of the proxy function, the device identifier of the graphics processor to be accessed is obtained from the call information. And, when the target array corresponding to the access function is preset and the target array includes the device identifiers of each graphics processor and the corresponding status identifiers, if the status identifier corresponding to the device identifier of the graphics processor to be accessed in the target array is preset data, it is determined that the device identifier in the call information is included in the proxy function, that is, it is represented that the image processor to be accessed is a faulty graphics processor with a target fault, so an access failure message is returned. If the status identifier corresponding to the device identifier of the graphics processor to be accessed in the target array is not preset data, it is determined that the device identifier in the call information is not included in the proxy function, that is, it is represented that the image processor to be accessed is not a faulty graphics processor with a target fault, so the access function is called to access the corresponding graphics processor without a target fault.

[0058] When a target array corresponding to an access function is preset and the device identifier of a graphics processing unit with a target fault is included in the target array, if the device identifier of the to-be-accessed graphics processing unit is included in the target array, it is determined that the device identifier in the call information is included in the proxy function, that is, it is characterized that the to-be-accessed image processor is a faulty graphics processing unit with a target fault. Therefore, an access failure message is returned. If the device identifier of the to-be-accessed graphics processing unit is not included in the target array, it is determined that the device identifier in the call information is not included in the proxy function, that is, it is characterized that the to-be-accessed image processor is not a faulty graphics processing unit with a target fault. Therefore, the access function is called to access the corresponding graphics processing unit without a target fault.

[0059] It can be seen that when the call information of the proxy function is obtained, by determining whether the device identifier in the call information is included in the proxy function, it is possible to both prevent new access tasks for the graphics processing unit with a target fault and ensure that other graphics processing units in the graphics processing unit set can be accessed normally.

[0060] In order to accurately determine whether the faulty graphics processing unit with a target fault is currently in an accessed state, in some embodiments, before terminating the access task of accessing the faulty graphics processing unit, the method may further include: executing a first query instruction to obtain a query result; in the case where the query result indicates that there is an access task of accessing the faulty graphics processing unit, determining that the graphics processing unit is currently in an accessed state.

[0061] Specifically, after determining that any graphics processing unit in the graphics processing unit set has a target fault, a first query instruction may be executed according to the device identifier of the faulty graphics processing unit with a target fault to obtain a query result. In some embodiments, the query result may include a task list. Correspondingly, in the case where the task list in the query result is non-empty, it indicates that there is an access task of accessing the faulty graphics processing unit, that is, it is determined that the faulty graphics processing unit is currently in an accessed state. In other embodiments, the query result may include a task identifier. Correspondingly, in the case where the task identifier is non-empty, it indicates that there is an access task of accessing the faulty graphics processing unit, that is, it is determined that the faulty graphics processing unit is currently in an accessed state. The specific form of the query result is not limited to the above description and can be set as needed in actual applications. The first query instruction is used to query the access task for the corresponding graphics processing unit according to the device identifier.

[0062] Thus, by executing the first query instruction, it is possible to accurately determine whether the graphics processing unit is currently in an accessed state based on the query result, thereby providing an accurate basis for effective access isolation.

[0063] In some embodiments, the set of graphics processors configured in the electronic device may be part or all of the graphics processors corresponding to high-performance computing tasks. Exemplarily, the high-performance computing tasks are tasks such as neural network training and neural network inference. In order to accurately schedule each graphics processor, in some embodiments, as Figure 4 shown, the electronic device may also send a fault message to the scheduling device when a target fault is detected; and send a recovery message to the scheduling device when the reset is successful. That is to say, on the basis of any of the above embodiments, the method may further include: Send the fault message of the faulty graphics processor with the target fault to the scheduling device of the set of graphics processors. The fault message is used to indicate removing the faulty graphics processor with the target fault from the scheduling domain, and the scheduling domain includes the set of graphics processors. And, when the reset is successful, send the recovery message of the faulty graphics processor to the scheduling device, and the recovery message is used to indicate adding the faulty graphics processor to the scheduling domain.

[0064] Specifically, when it is determined that any graphics processor in the set of graphics processors has a target fault, the fault message may be sent to the scheduling device according to the device identifier of the faulty graphics processor. When the scheduling device receives the fault message, it obtains the device identifier from the fault message and marks the scheduling status of the faulty graphics processor corresponding to the device identifier as an invalid status to remove the faulty graphics processor from the scheduling domain. When the electronic device successfully resets the faulty graphics processor with the target fault, according to the device identifier of the faulty graphics processor, send a recovery message to the scheduling device. The scheduling device obtains the device identifier from the recovery message and marks the scheduling status of the faulty graphics processor corresponding to the device identifier as a valid status to add the faulty graphics processor to the scheduling domain. When the electronic device fails to reset the faulty graphics processor with the target fault, a reset failure record may be generated for subsequent analysis.

[0065] Among them, the scheduling device may be a device in a distributed cluster. The distributed cluster may include at least one electronic device, and any electronic device may be configured with a set of graphics processors. The scheduling device is used to manage the scheduling domain including each graphics processor.

[0066] It should be noted that Figure 4 for illustration only and not for limitation, the execution order of some operation steps may be interchanged, and some operation steps may also be executed simultaneously. Exemplarily, the operation of sending the fault message to the scheduling device may also be executed after terminating the access task being accessed; preventing new access tasks may also be executed simultaneously with terminating the access task being accessed, etc.

[0067] Therefore, when a target fault occurs in the graphics processing unit, a fault message is sent to the scheduling device, and when the reset of the graphics processing unit with the target fault is successful, a recovery message is sent to the scheduling device. This can ensure that the scheduling device effectively manages the scheduling domain, thereby improving the accuracy of graphics processing unit scheduling.

[0068] To further demonstrate the advantages of the device fault handling method provided in the embodiments of the present application (hereinafter referred to as the innovative solution), the following is a comparison with the centralized response method adopted in the related technology described above, as shown in the following table:

[0069] Based on the above comparison, in terms of the comprehensive elapsed time, success rate, whether Row Remapping is supported, and whether Page Retirement is supported, etc., the innovative solution provided in the embodiments of the present application is superior to each solution in the related technology, and has the advantages of low elapsed time, high success rate, and support for Row Remapping and Page Retirement.

[0070] It can be understood that the physical structure of the video memory has multiple combinations of storage units (rows / columns). When a certain row (row) frequently makes errors due to hardware aging or other defects, the memory controller or firmware of the graphics processing unit will redirect the address of this row to a reserved spare row, and this process is Row Remapping. Since after the reset of the faulty graphics processing unit is successful, it will trigger the error checking mechanism of the faulty graphics processing unit at the hardware or firmware level, and when the result of the error check indicates that an error occurs in a certain video memory row, the address of this row will be redirected to the reserved spare row through the memory controller or firmware of the faulty graphics processing unit. And, when the error check result indicates that an irreparable damage occurs in a certain area of the video memory (such as a video memory page), the area is marked as retired through the driver or firmware of the faulty graphics processing unit so as not to be allocated for use anymore. Therefore, Row Remapping and Page Retirement are supported in the embodiments provided in the present application.

[0071] Corresponding to the application scenario and method of the method provided in the embodiments of the present application, the embodiments of the present application also provide a device fault handling device, and this device can be applied to Figure 2 the electronic device shown in Figure 5As shown in the figure, the device includes: an isolation module 501, configured to isolate access to a faulty graphics processor with a target fault among any graphics processors in a set of graphics processors, where the access isolation is used to block access tasks for the faulty graphics processor. A reset module 502, configured to reset the faulty graphics processor, where, in the case of successful reset, the faulty graphics processor resumes to a target state, and the target state is a normal working state.

[0072] Optionally, the isolation module 501 is specifically configured to configure the access state of the faulty graphics processor to an inaccessible state, where the inaccessible state is used to prevent new access tasks for the faulty graphics processor; and, in the case where the faulty graphics processor is currently in an accessed state, terminate the access task that is accessing the faulty graphics processor.

[0073] Optionally, the isolation module 501 is further specifically configured to query a pointer to the access function according to the function name of the access function, where the access function is used to access any graphics processor in the set of graphics processors; modify the pointed object of the pointer to a proxy function, where the proxy function includes the device identifier of the faulty graphics processor to indicate that the access state of the faulty graphics processor is the inaccessible state.

[0074] Optionally, the device further includes a processing module, configured to return an access failure message in the case of obtaining call information of the proxy function and the proxy function includes the device identifier in the call information; or, in the case of obtaining call information of the proxy function and the proxy function does not include the device identifier in the call information, call the access function.

[0075] Optionally, the isolation module 501 is further configured to execute a first query instruction before terminating the access task that is accessing the faulty graphics processor to obtain a query result; in the case where the query result indicates that there is an access task accessing the faulty graphics processor, determine that the faulty graphics processor is currently in the accessed state.

[0076] Optionally, the device further includes a detection module, configured to, before the isolation module 501 isolates access to the faulty graphics processor, obtain the operation state data of any graphics processor in the set of graphics processors according to a preset detection period; in the case where the operation state data includes a preset fault identifier, determine that the graphics processor is a faulty graphics processor with the target fault.

[0077] Optionally, the apparatus further includes a sending module, configured to send fault information of the faulty graphics processor to a scheduling device of the set of graphics processors, where the fault information is used to indicate removing the faulty graphics processor from a scheduling domain, and the scheduling domain includes the set of graphics processors; and send recovery information of the faulty graphics processor to the scheduling device when the reset is successful, where the recovery information is used to indicate adding the faulty graphics processor to the scheduling domain.

[0078] For the functions of the modules in each apparatus of the embodiments of the present application, reference may be made to the corresponding descriptions in the above methods, and the corresponding beneficial effects are achieved, which will not be elaborated here. In addition, the apparatus embodiments described above are merely illustrative, where the modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of the present application.

[0079] Figure 6 It is a block diagram of an electronic device for implementing the embodiments of the present application. As Figure 6 shown, the electronic device includes: a memory 601 and a processor 602, where a computer program that can run on the processor 602 is stored in the memory 601. When the processor 602 executes the computer program, the method in the above embodiments is implemented. The number of the memory 601 and the processor 602 may be one or more. In a specific implementation, the electronic device may further include a communication interface 603, configured to communicate with external devices and perform data interaction and transmission.

[0080] In a specific implementation, if the memory 601, the processor 602, and the communication interface 603 are implemented independently, the memory 601, the processor 602, and the communication interface 603 may be connected to each other through a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0081] Optionally, in a specific implementation, if the memory 601, the processor 602, and the communication interface 603 are integrated on a single chip, the memory 601, the processor 602, and the communication interface 603 can communicate with each other through an internal interface.

[0082] An embodiment of the present application provides a computer-readable storage medium that stores a computer program, and when the program is executed by a processor, the method provided in the embodiment of the present application is implemented.

[0083] An embodiment of the present application provides a computer program product that includes a computer program, and when the program is executed by a processor, the method provided in the embodiment of the present application is implemented.

[0084] An embodiment of the present application further provides a chip that includes a processor for calling and running instructions stored in a memory from the memory, so that a communication device installed with the chip executes the method provided in the embodiment of the present application.

[0085] An embodiment of the present application further provides a chip that includes an input interface, an output interface, a processor, and a memory. The input interface, the output interface, the processor, and the memory are connected through an internal connection path. The processor is used to execute the code in the memory, and when the code is executed, the processor is used to execute the method provided in the embodiment of the application.

[0086] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the advanced reduced instruction set machine (ARM) architecture.

[0087] Further, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).

[0088] In the above embodiments, it may be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium.

[0089] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0090] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means two or more unless otherwise specifically defined.

[0091] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed.

[0092] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with such instruction execution systems, apparatuses, or devices.

[0093] It should be understood that each part of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by a program instructing relevant hardware. This program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiment.

[0094] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, an optical disc, or the like.

[0095] As mentioned above, the above is only an exemplary embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope recorded in the present application can easily think of various changes or substitutions thereof, and these should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A method for processing device failures, the method comprising: In the case that any graphics processor in a set of graphics processors has a target failure, performing access isolation on the failed graphics processor with the target failure, wherein the access isolation is used to block access tasks for the failed graphics processor; Resetting the failed graphics processor, wherein in the case of successful reset, the failed graphics processor is restored to a target state, and the target state is a state of normal operation.

2. The method according to claim 1, wherein, The performing access isolation on the graphics processor with the target failure includes: Configuring the access state of the failed graphics processor to an inaccessible state, where the inaccessible state is used to prevent new access tasks for the failed graphics processor; and, In the case that the failed graphics processor is currently in an accessed state, terminating the access task that is accessing the failed graphics processor.

3. The method according to claim 2, wherein The configuring the access state of the failed graphics processor to an inaccessible state includes: Querying for a pointer to an access function according to the function name of the access function, where the access function is used to access any graphics processor in the set of graphics processors; Modifying the pointed object of the pointer to a proxy function, where the proxy function includes the device identifier of the failed graphics processor to indicate that the access state of the failed graphics processor is the inaccessible state.

4. The method according to claim 3, wherein, The method further includes: In the case of obtaining call information of the proxy function and the proxy function contains the device identifier in the call information, returning an access failure message; or, In the case of obtaining call information of the proxy function and the proxy function does not contain the device identifier in the call information, calling the access function.

5. The method according to claim 2, wherein Before the terminating the access task that is accessing the failed graphics processor, the method further includes: Executing a first query instruction to obtain a query result; In the case that the query result indicates that there is an access task accessing the failed graphics processor, determining that the failed graphics processor is currently in the accessed state.

6. The method according to any one of claims 1-5, wherein, Before the performing access isolation on the failed graphics processor, the method further includes: For any graphics processor in the set of graphics processors, obtaining the operation status data of the graphics processor according to a preset detection period; In the case that the operation status data contains a preset failure identifier, determining that the graphics processor is a failed graphics processor with the target failure.

7. The method according to any one of claims 1-5, wherein The method further includes: Sending the failure information of the failed graphics processor to a scheduling device of the set of graphics processors, where the failure information is used to indicate removing the failed graphics processor from a scheduling domain, and the scheduling domain includes the set of graphics processors; In the case of successful reset, sending the recovery information of the failed graphics processor to the scheduling device, where the recovery information is used to indicate adding the failed graphics processor to the scheduling domain.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, wherein the processor implements the method according to any one of claims 1 to 7 when executing the computer program.

9. A computer-readable storage medium storing a computer program therein, wherein the computer program implements the method according to any one of claims 1 to 7 when executed by a processor.

10. A computer program product comprising a computer program, wherein the computer program implements the method according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Data access method, chip, electronic equipment and storage medium

    CN116795731A

  • Fault maintenance method and device for graphics processor

    CN118567892A

  • GPU video memory isolation method and device based on Kubernetes, medium and product

    CN118796352A

  • GPU exception processing method and device, equipment and storage medium

    CN119311464A

  • GPU video memory channel allocation method, apparatus and device, and storage medium

    CN119938331A