MCM-GPU-based error recovery method, apparatus and device, and storage medium

By selectively recovering the target error area affected by errors based on dynamic analysis and idempotent area technology in MCM-GPU, the problem of large-scale soft error impact under the MCM-GPU architecture is solved, and efficient error recovery and low storage overhead are achieved.

CN119988097APending Publication Date: 2025-05-13JILIN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510123463.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Under the MCM-GPU architecture, as the scale of parallel computing tasks increases, the impact of soft errors becomes more significant, resulting in inaccurate calculation results, data corruption, model training failure, task interruption or crash. The existing GPU full-thread rollback technology leads to waste of computing resources and performance losses.

Method used

By determining the error area of ​​the target MCM-GPU based on the preset dynamic analysis method, identifying the target GPM, and resuming the error area using the target registers to avoid the high overhead of rollback of the entire thread.

Benefits of technology

Improves the error recovery efficiency of MCM-GPU programs, reduces storage overhead and recovery time, and optimizes energy consumption performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988097A_ABST
    Figure CN119988097A_ABST
Patent Text Reader

Abstract

The invention discloses an MCM-GPU-based error recovery method and device, equipment and a storage medium, and relates to the technical field of error recovery, and the method comprises the steps: determining a target error region of a target program corresponding to a target MCM-GPU during current operation based on a preset dynamic analysis method; if a preset target error is detected in the target error area, determining a corresponding target GPM in the target error area; the target GPM is any GPU module in a plurality of GPU modules obtained after the GPU is divided; and determining a target register in the target MCM-GPU so as to recover the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register. It can be seen from the above that through the multi-level efficient fine-grained recovery technology and the lightweight check point technology, on the premise that the correctness of the program is guaranteed, the error recovery efficiency of the MCM-GPU program can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of error recovery, and in particular to an error recovery method, device, equipment and storage medium based on MCM-GPU. Background Art

[0002] MCM-GPU (Multi-chip-module GPUs, MCM-GPU) adopts a multi-chip module architecture, which integrates multiple GPU modules (GPU Chip Modules, GPMs) in one package to achieve modular expansion of performance. Under the MCM-GPU architecture, as the scale of parallel computing tasks increases and the number of computing units increases, the impact of soft errors becomes more significant, leading to inaccurate computing results, data corruption, model training failure, task interruption or crash, and other problems.

[0003] At present, in order to recover MCM-GPU applications from error states, GPU full-thread rollback technology is usually used, that is, during program execution, the state of the GPU (Graphics Processing Unit) is periodically saved as a checkpoint. Once an error is detected, the GPU full-thread rollback technology will load the last saved checkpoint to restore the program state. However, even if the calculation results of many threads are correct, it is necessary to roll back the previous checkpoint and re-execute, resulting in a waste of computing resources and performance loss. In addition, there are multiple GPU kernel function calls or a large number of computing tasks between two checkpoints, including mixed calculations of CPU and GPU. Therefore, saving system-level checkpoints will take up a lot of storage resources. When an error occurs, all tasks between checkpoints need to be re-executed, including multiple kernel function calls or complex computing tasks, which significantly increases the recovery time and system overhead.

[0004] In summary, how to improve the error recovery efficiency of MCM-GPU programs is a technical problem that needs to be solved urgently. Summary of the invention

[0005] In view of this, the purpose of the present invention is to provide an error recovery method, device, equipment and storage medium based on MCM-GPU, which can improve the error recovery efficiency of MCM-GPU programs. The specific scheme is as follows:

[0006] In a first aspect, the present application provides an error recovery method based on MCM-GPU, comprising:

[0007] Determine the target error area of ​​the target program corresponding to the target MCM-GPU during the current runtime based on a preset dynamic analysis method;

[0008] If a preset target error is detected from the target error area, a corresponding target GPM in the target error area is determined; the target GPM is any GPU module among a plurality of GPU modules obtained after dividing the GPU;

[0009] A target register in the target MCM-GPU is determined so as to restore the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register.

[0010] Optionally, determining the target error area when the target program corresponding to the target MCM-GPU is currently running based on a preset dynamic analysis method includes:

[0011] Determine the global memory access mode of the target program corresponding to the target MCM-GPU when it is currently running based on a preset dynamic analysis method;

[0012] According to the global memory access pattern, a preset read-after-write dependency relationship between threads in the target program is identified to obtain a corresponding identification result, so as to determine the target error area based on the identification result.

[0013] Optionally, if a preset target error is detected from the target error area, determining a corresponding target GPM in the target error area includes:

[0014] If soft errors are detected from the target error region, and it is determined based on the soft errors that the number of corresponding GPMs in the target error region meets a preset number condition, the GPM is determined to be a target GPM.

[0015] Optionally, before restoring the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register, the method further includes:

[0016] If it is determined that the target program meets the preset time-sensitive partitioning condition, the target program is partitioned based on the boundaries corresponding to the preset synchronization primitives and preset basic blocks to obtain the idempotent region corresponding to each thread in the target GPM.

[0017] Optionally, before restoring the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register, the method further includes:

[0018] If it is determined that the target program satisfies the preset storage-limited partitioning condition, the target program is partitioned based on preset load instructions and preset storage instructions to obtain the idempotent region corresponding to each thread in the target GPM.

[0019] Optionally, determining a target register in the target MCM-GPU includes:

[0020] Identify state changes of each register in the target MCM-GPU in the idempotent region;

[0021] If it is determined based on the identification result that the state of the register satisfies a preset active state, and the value of the register is modified in different idempotent areas, then the register is determined to be the target register in the target MCM-GPU.

[0022] Optionally, the recovering the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register includes:

[0023] Determine the target checkpoint data corresponding to the target register, and perform rollback processing on each thread in the target GPM based on the target checkpoint data, a preset rollback mechanism, and the idempotent region corresponding to each thread in the target GPM to restore the target error region;

[0024] If it is determined that the target error area is restored to the preset normal operating state, jump to the step of determining the target error area when the target program corresponding to the target MCM-GPU is currently running based on the preset dynamic analysis method, so as to perform the next round of error recovery operation.

[0025] In a second aspect, the present application provides an error recovery device based on MCM-GPU, comprising:

[0026] A target error region determination module, used to determine the target error region of the target program corresponding to the target MCM-GPU when it is currently running based on a preset dynamic analysis method;

[0027] a target GPM determination module, configured to determine a corresponding target GPM in the target error region if a preset target error is detected from the target error region; the target GPM is any GPU module among a plurality of GPU modules obtained by dividing the GPU;

[0028] The target error region recovery module is used to determine the target register in the target MCM-GPU so as to recover the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register.

[0029] In a third aspect, the present application provides an electronic device, including:

[0030] Memory, used to store computer programs;

[0031] The processor is used to execute the computer program to implement the aforementioned MCM-GPU based error recovery method.

[0032] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned MCM-GPU-based error recovery method is implemented.

[0033] In this application, firstly, the target error area of ​​the target program corresponding to the target MCM-GPU during the current operation is determined based on the preset dynamic analysis method; then, if a preset target error is detected from the target error area, the corresponding target GPM in the target error area is determined; the target GPM is any GPU module among the multiple GPU modules obtained after dividing the GPU; finally, the target register in the target MCM-GPU is determined, so as to restore the target error area based on the idempotent area corresponding to each thread in the target GPM and the target register. As can be seen from the above, in this application, based on dynamic analysis and idempotent areas, the target error area affected by the error in the target program corresponding to the target MCM-GPU is selectively restored to avoid the high overhead generated by the GPU full thread rollback. By determining the target register and combining the idempotent area, the high efficiency and low storage overhead of the error recovery of the MCM-GPU program can be ensured. In this way, through multi-level efficient fine-grained recovery technology, the storage overhead and recovery time are balanced, and the energy consumption performance is optimized under the premise of ensuring program correctness and fast error recovery. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0035] Figure 1 A flow chart of an error recovery method based on MCM-GPU provided in this application;

[0036] Figure 2 A specific MCM-GPU-based error recovery method flow chart provided in this application;

[0037] Figure 3 A schematic diagram of the structure of an error recovery device based on MCM-GPU provided in this application;

[0038] Figure 4 A structural diagram of an electronic device provided for this application. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] At present, in order to recover the MCM-GPU application from the error state, the GPU full-thread rollback technology is usually adopted, that is, during the program execution, the GPU state is periodically saved as a checkpoint. Once an error is detected, the GPU full-thread rollback technology will load the last saved checkpoint to restore the program state. However, even if the calculation results of many threads are correct, it is necessary to roll back the previous checkpoint and re-execute it, resulting in a waste of computing resources and performance loss. In addition, there are multiple GPU kernel function calls or a large number of computing tasks between the two checkpoints, including mixed calculations of the CPU and GPU. Therefore, saving system-level checkpoints will take up a lot of storage resources. When an error occurs, it is necessary to re-execute all tasks between the checkpoints, including multiple kernel function calls or complex computing tasks, which significantly increases the recovery time and system overhead. To this end, the present application provides an error recovery scheme based on MCM-GPU, which can improve the error recovery efficiency of MCM-GPU programs.

[0041] See also Figure 1 As shown, an embodiment of the present invention discloses an error recovery method based on MCM-GPU, which may include:

[0042] Step S11: determining a target error area when a target program corresponding to the target MCM-GPU is currently running based on a preset dynamic analysis method.

[0043] In this embodiment, in order to determine the error area in the program corresponding to the MCM-GPU, the above-mentioned determination of the target error area when the target program corresponding to the target MCM-GPU is currently running based on the preset dynamic analysis method may include: determining the global memory access mode when the target program corresponding to the target MCM-GPU is currently running based on the preset dynamic analysis method; identifying the preset read-after-write dependency between the threads in the target program according to the global memory access mode to obtain a corresponding identification result, so as to determine the target error area based on the identification result. Specifically, when the target program corresponding to the target MCM-GPU is running, the specific address of the read and write operations of each thread on the memory is recorded. When the storage operation of a thread modifies the data of a memory address, and the data of the memory address is subsequently read by the load operation of another thread, it is determined that there is a read-after-write dependency between the two threads. Dependencies include: intra-thread dependency, that is, the dependency only exists in the same thread, and inter-thread dependency, that is, the dependency involves data exchange between different threads. By analyzing the global memory access pattern of the target program corresponding to the target MCM-GPU during runtime, and due to the execution model of GPU hardware design, such as the independence of thread blocks, and the memory access pattern of most applications, it can be determined that there is usually no read-after-write dependency between threads between thread blocks within the GPU kernel, and the GPU hardware limits the data transfer path between thread blocks, thereby avoiding the propagation of errors between thread blocks. Therefore, when performing error recovery on the target program corresponding to the target MCM-GPU, only the GPU module containing the error, that is, the GPM, needs to be rolled back without re-executing all threads. The GPU module containing the error is determined as the target error area in the target program corresponding to the target MCM-GPU.

[0044] Step S12: If a preset target error is detected from the target error area, a corresponding target GPM in the target error area is determined; the target GPM is any GPU module among a plurality of GPU modules obtained after dividing the GPU.

[0045] In this embodiment, in order to avoid resetting the state of the entire thread and reduce the waste of computing resources, if a preset target error is detected from the target error region, the corresponding target GPM in the target error region is determined, including: if a soft error is detected from the target error region, and based on the soft error, the number of corresponding GPMs in the target error region is determined to meet the preset number condition, then the GPM is determined to be the target GPM. It can be understood that when the target program corresponding to the MCM-GPU is in an error running state, a variety of errors can be detected from the target error region corresponding to the target program, including but not limited to: soft errors, hard errors, manufacturing defects, data transmission errors, etc. The soft errors existing in the MCM-GPU can be resolved by rewriting or resetting. Specifically, the source and nature of the soft errors in the target error region can be determined by analyzing the system state, monitoring logs, electromagnetic environment, etc. when the error occurs, and the target GPM corresponding to the target error region can be determined based on the source and nature of the soft errors. If it is determined that the soft error only affects one GPU module, that is, one GPM, a local rollback method can be used for error recovery, that is, only the affected GPM is rolled back, and other GPMs are not rolled back. After that, based on the error recovery at the GPM level, the granularity of error recovery can be further reduced, that is, each thread in the target GPM can be determined for recovery. Through fine-grained recovery technology, the calculation results that have been correctly executed in the MCM-GPU can be retained to the maximum extent.

[0046] Step S13: determine a target register in the target MCM-GPU so as to restore the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register.

[0047] In this embodiment, the recovery range of GPM can be accurately determined by optimizing the partitioning technology of idempotent regions, thereby reducing the time and resources required for the full-thread rollback technology and improving the error recovery efficiency of the program of the MCM-GPU system. It is understandable that in order to enhance memory throughput by hiding memory latency, the GPU architecture usually relies on register storage variables instead of cache performance improvement. However, storing variables through registers means that during the error recovery process, the preservation of register states will directly affect the storage and operation overhead of the MCM-GPU. At the same time, the checkpoint rollback mechanism can ensure the reliability and continuity of the computing task by saving and restoring the previous state when an error or failure occurs in the system. In the MCM-GPU, the computing tasks of the GPU and the CPU are included between the two checkpoints, so the checkpoint needs to save the state of the entire MCM-GPU system, including the memory data and intermediate calculation results of the CPU and GPU. And multiple GPU kernel function calls or a large number of computing tasks are included between the two checkpoints. When an error occurs during the operation of the target program corresponding to the MCM-GPU, the computing tasks between the checkpoints must be re-executed. Traditional checkpoint methods often save all active registers, including registers that have not been modified but still survived in the subsequent idempotent region. Therefore, the traditional checkpoint method brings unnecessary storage redundancy, especially in GPUs with a large number of thread copies, the overhead of this storage redundancy will be multiplied.

[0048] It can be understood that in order to further improve the error recovery efficiency on the basis of fine-grained recovery, the above-mentioned determination of the target register in the target MCM-GPU may include: identifying the state changes of each register in the target MCM-GPU in the idempotent region; if it is determined based on the identification result that the state of the register satisfies the preset active state, and the value of the register is modified in different idempotent regions, then the register is determined to be the target register in the target MCM-GPU. Specifically, it is first necessary to accurately identify the state changes of each register in the target MCM-GPU, and for registers that survive across idempotent regions but whose register values ​​are not modified in subsequent idempotent regions, the corresponding register values ​​can be directly used without repeated saving; and registers that survive across idempotent regions and whose register values ​​are modified in subsequent idempotent regions can be determined as target registers in the target MCM-GPU, and only the checkpoint data corresponding to the target register needs to be saved.

[0049] In this embodiment, the above-mentioned recovery of the target error area based on the idempotent area corresponding to each thread in the target GPM and the target register includes: determining the target checkpoint data corresponding to the target register, and rolling back each thread in the target GPM based on the target checkpoint data, the preset rollback mechanism and the idempotent area corresponding to each thread in the target GPM to restore the target error area; if it is determined that the target error area is restored to the preset normal operating state, then jump to the step of determining the target error area when the target program corresponding to the target MCM-GPU is currently running based on the preset dynamic analysis method, so as to perform the next round of error recovery operations. Specifically, it is necessary to first determine the variables stored in the target register, that is, the target checkpoint data, and then roll back each thread in the target GPM based on the target checkpoint data to roll back to the starting position of the idempotent area corresponding to each thread. If the target error area is restored to the preset normal operating state, the target error area when the target program corresponding to the target MCM-GPU is currently running can be determined based on the preset dynamic analysis method to perform the next round of error recovery operations.

[0050] As can be seen from the above, in this embodiment, the target error area of ​​the target program corresponding to the target MCM-GPU is first determined based on the preset dynamic analysis method when it is currently running; if a preset target error is detected from the target error area, the corresponding target GPM in the target error area is determined; the target GPM is any GPU module among the multiple GPU modules obtained after dividing the GPU; finally, the target register in the target MCM-GPU is determined, so as to restore the target error area based on the idempotent area corresponding to each thread in the target GPM and the target register. As can be seen from the above, in this embodiment, based on dynamic analysis and idempotent areas, the target error area affected by the error in the target program corresponding to the target MCM-GPU is selectively restored to avoid the high overhead generated by the GPU full thread rollback. By determining the target register and combining the idempotent area, the high efficiency and low storage overhead of the error recovery of the MCM-GPU program can be ensured. In this way, through multi-level efficient fine-grained recovery technology, the storage overhead and recovery time are balanced, and the energy consumption performance is optimized under the premise of ensuring program correctness and fast error recovery.

[0051] Based on the previous embodiment, it can be seen that the present application can determine the corresponding target GPM in the target error area through multi-level efficient fine-grained recovery technology, and use lightweight checkpoint technology to roll back each thread in the target GPM, thereby improving the error recovery efficiency of the MCM-GPU program. Next, this embodiment will explain in detail how to determine the idempotent area corresponding to each thread in the target GPM. Figure 2As shown, the embodiment of the present invention further discloses an error recovery method based on MCM-GPU, which may include:

[0052] Step S21: If it is determined that the target program meets the preset time-sensitive partitioning condition, the target program is partitioned based on the boundaries corresponding to the preset synchronization primitives and preset basic blocks to obtain the idempotent region corresponding to each thread in the target GPM.

[0053] In this embodiment, if it is determined that the target program meets the preset time-sensitive partitioning conditions, the target program can be partitioned by the fast recovery strategy (i.e., Fast Recovery, FRC) to obtain the idempotent region corresponding to each thread in the target GPM, that is, the idempotent region boundary can be set at the synchronization primitive. Synchronization primitives are an important mechanism to ensure orderly operation between threads. By setting the boundary at the synchronization primitive, the error propagation between threads can be isolated; wherein, the synchronization primitives include but are not limited to memory barriers, atomic operations, etc. At the same time, the idempotent region boundary can be set at the boundary of the basic block. The basic block is a section of continuous instruction code without branches and is the basic execution unit of the program. Setting the idempotent region boundary at the basic block can minimize the amount of code that needs to be re-executed during error recovery. It can be understood that since the idempotent region is small, only a small amount of affected code needs to be re-executed during error recovery, which significantly reduces the time for error recovery, and there are no branches and jumps inside the basic block, which is convenient for control flow analysis and recovery point positioning.

[0054] Step S22: If it is determined that the target program meets the preset storage-restricted partitioning condition, the target program is partitioned based on preset load instructions and preset storage instructions to obtain the idempotent region corresponding to each thread in the target GPM.

[0055] In this embodiment, if it is determined that the target program meets the preset storage-restricted partitioning conditions, the target program can be partitioned by a low storage overhead strategy (i.e., Low Storage Overhead, LSO) to obtain the idempotent region corresponding to each thread in the target GPM, that is, the idempotent region boundary can be set at the load instruction and the store instruction. Load instructions usually appear at the program startup stage to obtain initial data; store instructions are usually located at the program end stage to save calculation results. The stability of global memory data, such as ECC protection (Error Correction Code, i.e., error correction code), is used to reduce the storage requirements of register checkpoints.

[0056] It can be understood that an idempotent region refers to a program code fragment that can obtain consistent output after being re-executed from the beginning at any point, and will not cause abnormal results due to repeated execution. The division of idempotent regions needs to be based on the following basic principles:

[0057] (1) Eliminate read-after-write dependencies: Ensure that variables in idempotent regions will not be modified during re-execution due to read-after-write dependencies. For example, if the input data of a region changes during execution, the region cannot meet the idempotency requirement.

[0058] (2) Maintain independence: The context input between regions is clear, such as register status, and ensures that re-execution only needs to restore the affected region without involving other regions.

[0059] It is understandable that the target thread may be divided into idempotent regions based on different application scenarios. Specifically, the difference between the two idempotent region division strategies of the target thread in this embodiment is shown in Table 1.

[0060] Table 1

[0061]

[0062] Step S23: determine the target register in the target MCM-GPU so as to restore the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register.

[0063] For a more specific processing procedure of the above step S23, reference may be made to the corresponding contents disclosed in the above embodiments, which will not be described in detail here.

[0064] As can be seen from the above, in this embodiment, the target program can be divided based on preset time-sensitive partitioning conditions or preset storage-limited partitioning conditions to determine the idempotent region corresponding to each thread in the target GPM, and the target error region can be restored based on the target register and the idempotent region corresponding to each thread in the target GPM. In this way, by optimizing the partitioning technology of the idempotent region, the error recovery efficiency of the MCM-GPU program can be further improved. In addition, the idempotent region partitioning strategy in this embodiment supports the flexible application of fast recovery and low storage overhead, is suitable for high-concurrency computing scenarios, and significantly improves the fault tolerance and performance efficiency of the GPU system.

[0065] Accordingly, see Figure 3 As shown, the embodiment of the present application also provides an error recovery device based on MCM-GPU, which may include:

[0066] A target error region determination module 11 is used to determine a target error region when a target program corresponding to a target MCM-GPU is currently running based on a preset dynamic analysis method;

[0067] A target GPM determination module 12 is used to determine a corresponding target GPM in the target error area if a preset target error is detected from the target error area; the target GPM is any GPU module among a plurality of GPU modules obtained after dividing the GPU;

[0068] The target error region recovery module 13 is used to determine the target register in the target MCM-GPU so as to recover the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register.

[0069] As can be seen from the above, in this application, the target error area of ​​the target program corresponding to the target MCM-GPU is first determined based on the preset dynamic analysis method when it is currently running; if a preset target error is detected from the target error area, the corresponding target GPM in the target error area is determined; the target GPM is any GPU module among the multiple GPU modules obtained after dividing the GPU; finally, the target register in the target MCM-GPU is determined, so as to restore the target error area based on the idempotent area corresponding to each thread in the target GPM and the target register. As can be seen from the above, in this application, based on dynamic analysis and idempotent areas, the target error area affected by the error in the target program corresponding to the target MCM-GPU is selectively restored to avoid the high overhead generated by the GPU full thread rollback. By determining the target register and combining the idempotent area, the high efficiency and low storage overhead of the error recovery of the MCM-GPU program can be ensured. In this way, through multi-level efficient fine-grained recovery technology, the storage overhead and recovery time are balanced, and the energy consumption performance is optimized under the premise of ensuring program correctness and fast error recovery.

[0070] In some specific implementations, the target error area determination module 11 may include:

[0071] A global memory access mode determination unit, configured to determine, based on a preset dynamic analysis method, a global memory access mode of the target program corresponding to the target MCM-GPU when it is currently running;

[0072] A target error area determination unit is used to identify a preset read-after-write dependency relationship between threads in the target program according to the global memory access pattern to obtain a corresponding identification result, so as to determine the target error area based on the identification result.

[0073] In some specific implementations, the target GPM determination module 12 may include:

[0074] The target GPM determining unit determines the GPM as a target GPM if soft errors are detected from the target error region and the number of corresponding GPMs in the target error region determined based on the soft errors satisfies a preset number condition.

[0075] In some specific implementations, the MCM-GPU based error recovery device may further include:

[0076] The first idempotent region determination module is used to divide the target program based on the boundaries corresponding to the preset synchronization primitives and preset basic blocks if it is determined that the target program meets the preset time-sensitive division conditions, so as to obtain the idempotent region corresponding to each thread in the target GPM.

[0077] In some specific implementations, the MCM-GPU based error recovery device may further include:

[0078] The second idempotent region determination module is used to divide the target program based on preset load instructions and preset storage instructions to obtain the idempotent region corresponding to each thread in the target GPM if it is determined that the target program meets the preset storage-restricted partitioning condition.

[0079] In some specific implementations, the target error region recovery module 13 may include:

[0080] A state change identification unit, used to identify state changes of each register in the target MCM-GPU in the idempotent region;

[0081] The target register determining unit is used to determine that the state of the register satisfies a preset active state based on the identification result, and the value of the register is modified in different idempotent areas, and then determine that the register is the target register in the target MCM-GPU.

[0082] In some specific implementations, the target error region recovery module 13 may include:

[0083] a first target error region recovery unit, configured to determine target checkpoint data corresponding to the target register, and perform rollback processing on each thread in the target GPM based on the target checkpoint data, a preset rollback mechanism, and the idempotent region corresponding to each thread in the target GPM, so as to recover the target error region;

[0084] The second target error area recovery unit is used to jump to the step of determining the target error area of ​​the target program corresponding to the target MCM-GPU during the current runtime based on the preset dynamic analysis method if it is determined that the target error area has recovered to a preset normal operating state, so as to perform the next round of error recovery operations.

[0085] Furthermore, the present application also discloses an electronic device. Figure 4 It is a structural diagram of an electronic device 20 shown according to an exemplary embodiment, and the content in the figure cannot be regarded as any limitation on the scope of use of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input and output interface 25 and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the error recovery method based on MCM-GPU disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0086] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.

[0087] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0088] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the MCM-GPU-based error recovery method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.

[0089] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the error recovery method based on MCM-GPU disclosed above is implemented. The specific steps of the method can refer to the corresponding contents disclosed in the above embodiments, and will not be repeated here.

[0090] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0091] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0092] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0093] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0094] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. An error recovery method based on MCM-GPU, characterized in that: include: Determine the target error area of ​​the target program corresponding to the target MCM-GPU during the current runtime based on a preset dynamic analysis method; If a preset target error is detected from the target error area, a corresponding target GPM in the target error area is determined; the target GPM is any GPU module among a plurality of GPU modules obtained after dividing the GPU; A target register in the target MCM-GPU is determined so as to restore the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register.

2. The MCM-GPU-based error recovery method according to claim 1, characterized in that: The method of determining the target error area when the target program corresponding to the target MCM-GPU is currently running based on a preset dynamic analysis method includes: Determine the global memory access mode of the target program corresponding to the target MCM-GPU when it is currently running based on a preset dynamic analysis method; According to the global memory access pattern, a preset read-after-write dependency relationship between threads in the target program is identified to obtain a corresponding identification result, so as to determine the target error area based on the identification result.

3. The MCM-GPU-based error recovery method according to claim 1, characterized in that: If a preset target error is detected from the target error area, determining a corresponding target GPM in the target error area includes: If soft errors are detected from the target error region, and it is determined based on the soft errors that the number of corresponding GPMs in the target error region meets a preset number condition, the GPM is determined to be a target GPM.

4. The MCM-GPU-based error recovery method according to claim 1, characterized in that: Before restoring the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register, the method further includes: If it is determined that the target program meets the preset time-sensitive partitioning condition, the target program is partitioned based on the boundaries corresponding to the preset synchronization primitives and preset basic blocks to obtain the idempotent region corresponding to each thread in the target GPM.

5. The MCM-GPU-based error recovery method according to claim 1, characterized in that: Before restoring the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register, the method further includes: If it is determined that the target program satisfies the preset storage-limited partitioning condition, the target program is partitioned based on preset load instructions and preset storage instructions to obtain the idempotent region corresponding to each thread in the target GPM.

6. The MCM-GPU based error recovery method according to claim 1, characterized in that: The determining of the target register in the target MCM-GPU includes: Identify state changes of each register in the target MCM-GPU in the idempotent region; If it is determined based on the identification result that the state of the register satisfies a preset active state, and the value of the register is modified in different idempotent areas, then the register is determined to be the target register in the target MCM-GPU.

7. The MCM-GPU based error recovery method according to any one of claims 1 to 6, characterized in that: The recovering the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register includes: Determine the target checkpoint data corresponding to the target register, and perform rollback processing on each thread in the target GPM based on the target checkpoint data, a preset rollback mechanism, and the idempotent region corresponding to each thread in the target GPM to restore the target error region; If it is determined that the target error area is restored to the preset normal operating state, jump to the step of determining the target error area when the target program corresponding to the target MCM-GPU is currently running based on the preset dynamic analysis method, so as to perform the next round of error recovery operation.

8. An error recovery device based on MCM-GPU, characterized in that: include: A target error region determination module, used to determine the target error region of the target program corresponding to the target MCM-GPU when it is currently running based on a preset dynamic analysis method; a target GPM determination module, configured to determine a corresponding target GPM in the target error region if a preset target error is detected from the target error region; the target GPM is any GPU module among a plurality of GPU modules obtained by dividing the GPU; The target error region recovery module is used to determine the target register in the target MCM-GPU so as to recover the target error region based on the idempotent region corresponding to each thread in the target GPM and the target register.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory; wherein the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the MCM-GPU based error recovery method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the MCM-GPU-based error recovery method according to any one of claims 1 to 7.