Method, apparatus and related device for detecting silent data corruption
By using multiple detection modules to inspect hardware before the machine learning task begins, the problem of detecting silent data corruption is solved, enabling accurate location of faulty hardware and improved task reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-08
- Publication Date
- 2026-07-07
AI Technical Summary
Existing technologies struggle to efficiently detect and locate silent data corruption in machine learning tasks, leading to contaminated computation results and task failures, and making it difficult to trace the source.
Before the machine learning task begins, multiple different detection modules are used to detect the hardware associated with the task, obtain the detection results, and accurately locate the abnormal hardware that is corrupted by silent data based on the abnormal detection results.
Accurately locating faulty hardware before task execution avoids wasting computing resources and improves the reliability and efficiency of task execution.
Smart Images

Figure CN122346403A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a method, apparatus and related equipment for detecting silent data corruption. Background Technology
[0002] As machine learning, especially deep learning models, continues to increase in scale and complexity, their training and inference tasks place extremely high demands on the computing power and reliability of the underlying computing hardware. These tasks typically require long-term operation on high-performance computing clusters, consuming enormous amounts of electricity and computing resources.
[0003] In this process, the reliability of computing hardware is fundamental to the success of the task. Hardware may suffer from silent data corruption (SDC) for various reasons, which refers to errors that occur during data computation, storage, or transmission that are not detected by the system in real time. These errors are insidious; once they occur, they can contaminate the computational results of machine learning tasks, potentially leading to serious consequences such as model training failure and decreased accuracy. Furthermore, they are difficult to debug and trace, severely impacting the reliability of task execution. Therefore, how to detect silent data corruption has become a pressing technical challenge. Summary of the Invention
[0004] This disclosure provides a method, apparatus, and related equipment for detecting silent data corruption.
[0005] Firstly, this disclosure provides a method for detecting silent data corruption, the method comprising:
[0006] Before running the machine learning task, multiple different detection modules are invoked to detect multiple different hardware associated with the running process of the machine learning task, so as to obtain the detection results corresponding to the multiple different detection modules and obtain multiple detection results.
[0007] If any of the multiple detection results contain anomaly detection results, then based on the anomaly detection results, identify the anomalous hardware that has silently corrupted data from among the multiple different hardware components.
[0008] Secondly, this disclosure provides a method for processing machine learning tasks, wherein the machine learning tasks are executed in parallel by scheduling nodes and multiple worker nodes from a preset set of worker nodes, the method comprising:
[0009] The scheduling node controls multiple work nodes in the preset work node set, so that each work node executes the above-mentioned silent data corruption detection method.
[0010] If any working node reports an abnormal detection result, the first working node with the abnormal detection result is removed from the preset working node set by the scheduling node, and the second working node is rescheduled from the resource pool so that the second working node replaces the first working node.
[0011] If the detection results of multiple working nodes included in the preset working node set are all normal, the machine learning task is executed in parallel through the multiple working nodes in the preset working node set.
[0012] Thirdly, this disclosure provides a device for detecting silent data corruption, the device comprising:
[0013] The detection module is used to call multiple different detection modules before running the machine learning task, and to detect multiple different hardware associated with the running process of the machine learning task, so as to obtain the detection results corresponding to the multiple different detection modules and obtain multiple detection results.
[0014] The determination module is used to determine, in the case that an anomaly detection result is included among the multiple detection results, an anomaly hardware with silent data corruption from among the multiple different hardware based on the anomaly detection result.
[0015] Fourthly, this disclosure provides an electronic device, including:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The processor is configured to execute the above-described method.
[0019] Fifthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described method.
[0020] In a sixth aspect, this disclosure provides a computer program product comprising computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the method described above.
[0021] In the embodiments provided in this disclosure, on the one hand, by setting hardware detection before running the machine learning task, problems such as detection being disconnected from task execution and faults being discovered only in the later stages of the task are avoided. This allows hardware with potential silent data corruption to be intercepted from participating in actual computation at the source, effectively preventing the risk of the entire machine learning task failing due to underlying hardware errors and avoiding a huge waste of computing resources. On the other hand, by calling multiple different detection modules to detect multiple different hardware, the correspondence between the detection modules and the hardware to be detected can be used to accurately locate abnormal hardware with silent data corruption from among multiple hardware. Therefore, this method can accurately locate faulty hardware before task execution, improving the reliability of task execution.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0023] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0024] Figure 1 A flowchart illustrating a method for detecting silent data corruption provided in this embodiment of the disclosure;
[0025] Figure 2 A flowchart illustrating a method for processing a machine learning task according to yet another embodiment of this disclosure is shown.
[0026] Figure 3 A flowchart illustrating a machine learning task processing method based on an example implementation of this disclosure is shown.
[0027] Figure 4 A structural diagram of a silent data corruption detection device provided in an embodiment of this disclosure;
[0028] Figure 5 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0029] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0030] Unless otherwise specified, the various embodiments and features of this disclosure may be combined with each other. As used herein, the term "and / or" includes any and all combinations of one or more of the associated enumerated entries.
[0031] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0032] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0033] Figure 1 This is a flowchart illustrating a method for detecting silent data corruption according to an embodiment of the present disclosure. This method can be applied to systems running machine learning tasks. (Refer to...) Figure 1 The method includes the following steps:
[0034] Step S110: Before running the machine learning task, call multiple different detection modules to detect multiple different hardware associated with the machine learning task's operation process, so as to obtain the detection results corresponding to the multiple different detection modules and obtain multiple detection results.
[0035] The core of this step is to proactively schedule multiple detection modules to detect multiple key hardware components on which the task depends before the resource-intensive machine learning task begins computation, so that when an anomaly is detected, it can be accurately traced back to the specific faulty hardware unit.
[0036] The detection module can be an independent software program unit encapsulated to verify the correctness and integrity of specific hardware functions. Each detection module can be used to execute a specific test algorithm for a type of hardware. The specific implementation of the detection module can be an independent executable file, dynamic link library, or script, and can be managed in the form of plugins, etc., to facilitate dynamic loading and configuration by the system.
[0037] In this disclosure, hardware refers to a physical computing device or a specific functional unit within a device provided for performing machine learning tasks. Typically, multiple hardware components are functional parts within the device associated with the task execution process. For example, in a graphics processing unit (GPU), this may include streaming multiprocessors, tensor cores, a global memory controller, caches, and other hardware. Correspondingly, abnormal hardware refers to physical hardware or hardware sub-units identified as having functional defects based on anomaly detection results. By performing multi-module detection at this critical point before task execution, the discovery of potential hardware failures can be advanced from the task computation process to the task initialization phase.
[0038] Step S120: In the case where multiple detection results include abnormal detection results, based on the abnormal detection results, identify the abnormal hardware with silent data corruption from multiple different hardware.
[0039] Specifically, when configuring each detection module, the target object to be detected corresponding to that module can be pre-determined; that is, the hardware target object detected by that module. When a detection module reports an anomaly, the anomaly result can be directly associated with the corresponding specific hardware unit according to the preset mapping relationship. Furthermore, when multiple interconnected detection modules malfunction simultaneously, information such as the anomaly pattern, the hardware dependencies detected by each module, and the cluster topology can be comprehensively analyzed to identify the faulty hardware. This method enables fault location capabilities, thereby accurately pinpointing the specific faulty hardware.
[0040] Therefore, this disclosure demonstrates two key advantages. First, by timing hardware detection before the machine learning task runs, it avoids issues such as detection being disconnected from task execution and faults being discovered only late in the task. This approach intercepts hardware with potential silent data corruption from participating in actual computation, effectively preventing the risk of the entire machine learning task failing due to underlying hardware errors and avoiding significant waste of computing resources. Second, by calling multiple different detection modules to detect multiple different hardware components, the correspondence between the detection modules and the hardware under test can be used to accurately locate the abnormal hardware with silent data corruption. Thus, this method can accurately locate faulty hardware before task execution, improving the reliability of task execution.
[0041] In one optional implementation, to achieve flexibility and targeting in the detection process, before invoking multiple different detection modules to detect multiple different hardware components associated with the machine learning task's execution, the following steps are performed: Based on the task type of the machine learning task, multiple different detection modules corresponding to that task type are selected from a pre-configured pool of candidate detection modules, and these modules are loaded. The pre-configured candidate detection modules are plug-in modules that support dynamic loading or unloading. The task type can be categorized based on factors such as the machine learning task's computational objectives, resource requirements, and business functions. For example, it can include model training tasks and model inference tasks. Furthermore, task types can be further subdivided into: large-scale distributed pre-training tasks, lightweight fine-tuning tasks, high-throughput online inference tasks, and high-precision offline batch inference tasks. Different types of tasks exhibit significant differences in hardware stress patterns and the degree of dependence on critical hardware paths. A candidate detection module refers to the set of all pre-configured, selectable detection modules. The types and number of candidate detection modules can be dynamically added or removed as hardware iterations or detection requirements change. Each candidate detection module encapsulates the verification logic for a specific type or function of hardware (such as general matrix computation, tensor computation, special function computation, memory consistency, etc.). Through a filtering operation, a subset can be selected from the candidate detection module set to construct a set of modules that are relevant to the current task and have good detection efficiency. This avoids unnecessary detections and prevents the omission of critical checks. Through dynamic loading, the program code, configuration, and dependent resources of the selected detection modules can be imported into the current runtime environment or the memory of a specific computing node, making the detection modules executable and callable. Plug-in modules refer to each detection module being configured as an independent plugin conforming to a unified interface specification. Each plugin can be independently developed, compiled, deployed, added, or removed, thus achieving strong scalability and maintainability.
[0042] The following sections will introduce each of the detection modules separately:
[0043] Optionally, to achieve high reliability verification of the GPU core computing unit, multiple different detection modules may include a computing core verification module for verifying the numerical computing cores in the graphics processing unit. Correspondingly, when obtaining multiple detection results corresponding to multiple different detection modules, the hardware calculation result obtained by the computing core verification module after executing a calculation task based on a first preset input data is obtained. The detection result corresponding to the computing core verification module is determined based on the degree of difference between the hardware calculation result and a reference calculation result. The reference calculation result can be obtained by the central processing unit after executing the calculation task based on the first preset input data. Furthermore, if the degree of difference meets a first preset difference condition, the detection result corresponding to the computing core verification module is determined to be an abnormal detection result. The first preset difference condition can be: the degree of difference is greater than a first difference threshold. Of course, the first preset difference condition can also be implemented in various ways, such as directly comparing whether the hardware calculation result and the reference calculation result are the same.
[0044] In machine learning tasks, the numerical computing cores of a GPU (such as CUDA Cores and Tensor Cores) are fundamental for performing massive arithmetic operations. If the internal circuitry of these cores suffers silent data corruption, it can lead to subtle deviations in the computational results. Current technologies lack efficient and reliable verification of the computational correctness of these cores. For example, relying solely on the GPU's own repetitive computations or simple checksums is insufficient to detect systemic hardware logic errors; while comparing with another GPU of the same model is costly and cannot guarantee the absolute correctness of the comparison source. To address these issues, the aforementioned approach introduces a comparison and verification mechanism that uses CPU computation results as a reference standard, effectively resolving these problems. This method provides a correctness reference for the GPU computing core with relatively low overhead. Any deviations exceeding a reasonable error range can be directly attributed to the GPU hardware's SDC (Software-Defined Core), thus achieving high-confidence detection and precise location of computing core faults. This method is clearly defined, easily automated, and can be seamlessly integrated into the pre-task detection process.
[0045] The computational core verification module can be a software component used to verify the correctness of the numerical computation core function in the graphics processing unit (GPU). Its core function is to drive the GPU to perform specific calculations and determine whether the calculation results are correct. The numerical computation core can be a hardware unit in the GPU responsible for performing basic arithmetic operations (such as floating-point addition, multiplication, multiplication-addition, etc.), for example, it can include general-purpose stream processors (i.e., CUDA Cores) and tensor computation cores. The hardware computation result refers to the output data generated by the GPU's numerical computation core after performing a computation task based on a first preset input data. The reference computation result refers to the standard result used for comparison with the hardware computation result; that is, the reference computation result is the result generated by the CPU independently performing a computation based on exactly the same first preset input data and computation task description. Because the CPU architecture differs from the GPU and typically uses high-precision or verified mathematical libraries for computation, the computation result can be considered a reliable reference benchmark. The difference is used to quantify the magnitude of the deviation between the hardware computation result and the reference computation result; for example, it can be the maximum absolute value of the difference between corresponding elements between two sets of data, the root mean square error, or the norm of the relative error. The specific value of the first dissimilarity threshold can be determined based on the numerical characteristics, accuracy requirements, and fault tolerance needs of the computational task. For example, different task types can correspond to different first dissimilarity thresholds.
[0046] Typically, the numerical computation cores in a graphics processing unit (GPU) can include a general-purpose computation core and a tensor computation core. Furthermore, the computation core verification module includes a general-purpose matrix verification module for performing general-purpose matrix calculations on the general-purpose computation core, and a mixed-precision matrix verification module for performing mixed-precision matrix calculations on the tensor computation core. Specifically, the general-purpose computation core's arithmetic fundamentals are verified through general-purpose matrix calculations, while the tensor computation core's dedicated acceleration correctness is verified through mixed-precision matrix calculations. This approach allows for in-depth coverage of all critical computational paths, from basic computing power to dedicated acceleration, ensuring comprehensiveness and efficiency in the detection process.
[0047] In one optional implementation, the multiple different detection modules may include: a function calculation verification module for verifying function units in a graphics processor; correspondingly, multiple detection results corresponding to the multiple different detection modules can be obtained by: obtaining at least two function calculation results obtained by the function calculation verification module after performing at least two related function operations on preset function data; performing self-verification on the at least two function calculation results according to a preset mathematical identity, and obtaining the detection result of the function calculation verification module based on the self-verification result; wherein, if the self-verification result does not conform to the mathematical identity, the detection result of the function calculation verification module is determined to be an abnormal detection result.
[0048] The above method enables unique and efficient verification of specific function units within the GPU. It abandons the traditional comparison model that relies on external reference values (such as CPU calculation results), instead utilizing mathematical principles to achieve self-verification within the GPU. Specifically, by driving the GPU to perform two or more function operations on the same input data that have a mathematically defined identity, and verifying whether the calculation results conform to the identity, independent, closed-loop verification of the correctness of the function unit's calculations is achieved. This method is particularly suitable for verifying transcendental functions implemented through complex approximation circuits.
[0049] Special function units in graphics processing units (GPUs) are used to compute transcendental functions (such as sin, cos, exp, log, etc.). The hardware implementation of these functions is typically based on lookup tables combined with polynomial approximation or iterative algorithms, resulting in complex circuitry. If the same result comparison method as general-purpose computing cores is used for verification, the following challenges arise: the algorithms and precision used to implement the same transcendental function on CPUs and GPUs may differ, leading to systematic differences in the comparison reference standards themselves, making it difficult to set precise fault tolerance thresholds. The self-verification method in this disclosure enables independent verification without relying on CPUs or other external computing resources. Verification can be completed solely within the GPU under test, reducing system dependence and complexity. Furthermore, a single run can generate multiple cross-verifiable results, which are then judged using a unified mathematical relationship, significantly improving detection efficiency. The related function operation refers to the computation of two or more functions that have a definite mathematical identity. For example, the computation of sin(x) and cos(x) is related because they satisfy sin²(x) + cos²(x) = 1.
[0050] In one optional implementation, the multiple different detection modules may include a memory consistency verification module for detecting the memory consistency verification unit. Accordingly, multiple detection results corresponding to the multiple different detection modules can be obtained in the following way: The memory consistency verification module obtains the hardware test results obtained after the memory consistency verification unit in the graphics processor performs a memory consistency test on the second preset input data; the hardware test results are compared with reference test results for consistency; wherein, if the hardware test results are inconsistent with the reference test results, the detection result corresponding to the memory consistency verification module is determined to be an abnormal detection result.
[0051] Memory consistency testing can be implemented in various ways, such as through atomic operation correctness verification testing and memory barrier and synchronization primitive effectiveness testing. In one optional implementation, memory consistency testing can be achieved through a global reduction operation. Accordingly, the hardware test result is obtained by the memory consistency verification unit performing a global reduction operation on the second preset input data; and the reference test result is obtained by the central processing unit performing a global reduction operation on the second preset input data, i.e., the reference test result can also be called the reference reduction result. Specifically, the memory consistency verification module obtains the hardware reduction result obtained by the memory consistency verification unit in the graphics processor performing a global reduction operation on the second preset input data; the hardware reduction result is compared with the reference reduction result; if the hardware reduction result and the reference reduction result are inconsistent, the detection result corresponding to the memory consistency verification module is determined to be an anomaly detection result. The memory consistency verification unit can include at least one of the following: a global memory controller, a cache consistency protocol unit, and an atomic operation unit.
[0052] The memory consistency verification unit ensures the correctness, order, and eventual consistency of data during multi-threaded concurrent access. Specifically, it can be implemented collaboratively by a global memory controller, a cache consistency protocol unit, and atomic operation units. The memory consistency verification module can be a software component used to detect whether the memory consistency verification unit is functioning correctly. The global reduction operation combines a data set scattered across numerous parallel threads into a single final value using binary operators (such as addition, finding the maximum value, and finding the minimum value). The key aspect of this type of operation is that all threads ultimately need to accumulate a portion of the results into a shared global variable, which can easily lead to read / write resource contention. The hardware reduction result refers to the single result value obtained after multiple threads of the graphics processor execute the global reduction operation in parallel. The reference reduction result refers to the correct reduction result used for comparison. In this embodiment, it can be the result obtained by the central processing unit performing the same reduction operation on identical input data in a single thread, serial manner. Since serial execution avoids concurrency conflicts, the result can be considered a deterministic truth value.
[0053] The above methods can verify the reliability of the GPU memory subsystem. Specifically, by designing a high-conflict, high-concurrency global reduction operation as a stress test load, potential defects in the memory subsystem's handling of data contention and synchronization can be exposed. The result of the GPU executing this operation is compared with a deterministic reference result executed serially on the CPU. Inconsistency indicates silent data corruption during data writing from the computation unit to memory or during data synchronization between multiple threads. In parallel computing tasks such as distributed machine learning training, the correctness of the computation unit (such as CUDA Cores) itself does not guarantee the final success of the task. After computation, data needs to be written to global memory through a complex path and maintain consistency among a large number of parallel threads. Silent data corruption in the memory subsystem (including the global memory controller, cache coherency protocol, and atomic operation units) can lead to problems such as correct computation but incorrect data write-back or lost data synchronization. Because the computation unit does not report errors, these types of errors are often subtle, and conventional detection methods usually cannot cover them. The above method, through global reduction operations, enables a large number of threads to perform high-frequency, concurrent operations on shared memory addresses, thereby simulating data contention patterns in critical scenarios such as gradient synchronization in distributed training. This efficiently triggers issues such as scheduling conflicts in the memory controller, cache consistency failures, and atomic operation errors. Furthermore, assuming the computing unit has been verified, inconsistencies between the hardware reduction result and the reference reduction result indicate that the root cause of the fault lies in the memory subsystem, thus narrowing the scope of the fault to specific components such as the global memory controller, cache consistency protocol, or atomic operation unit, providing crucial information for subsequent maintenance.
[0054] In one optional implementation, the multiple different detection modules may include an end-to-end detection module. Accordingly, multiple detection results corresponding to the multiple different detection modules can be obtained as follows: The end-to-end detection module obtains the hardware kernel output result obtained after the graphics processor runs a preset composite computing task kernel; wherein the composite computing task kernel is configured to call multiple hardware functions of the graphics processor; the hardware kernel output result is compared with the reference kernel output result to obtain the detection result of the end-to-end detection module; wherein, if the difference between the hardware kernel output result and the reference kernel output result meets a second preset difference condition, the detection result corresponding to the end-to-end detection module is determined to be an abnormal detection result.
[0055] This test simulates the processing of a real workload by driving the graphics processing unit (GPU) to run a predefined composite computing task kernel, and verifies the correctness of the final output. This kernel is designed to intensively call various hardware functional units of the GPU. By quantitatively comparing the obtained hardware output with independent and reliable reference output, and verifying it according to preset difference judgment conditions, the final acceptance test is achieved to determine whether the GPU exhibits silent data corruption in a near-real-world, system-level, integrated operating state. The composite computing task kernel can be a complete GPU program that encapsulates various computational operations and data flows. Such a kernel can call different hardware functional units of the GPU, such as general-purpose computing cores, tensor computing cores, and special function units, and involves frequent global memory access and synchronization operations.
[0056] For example, the composite computing task kernel can be an attention mechanism computing kernel. Correspondingly, the hardware kernel output result obtained by the graphics processor after running the attention mechanism computing kernel can be obtained through the end-to-end detection module. The hardware kernel output result is compared with the reference kernel output result to obtain the detection result of the end-to-end detection module. Among them, if the difference between the hardware kernel output result and the reference kernel output result exceeds the second difference threshold, the detection result corresponding to the end-to-end detection module is determined to be an abnormal detection result.
[0057] The above method is used to implement end-to-end system testing. Unlike testing individual hardware units in isolation, the end-to-end testing module drives the graphics processor to run a complete, realistic, and complex AI computing kernel, and compares the overall output of this kernel on the GPU with the output of an independent and reliable reference kernel. This method aims to verify the overall correctness and stability of individual hardware units (such as computing cores, function units, memory subsystems, etc.) when they actually work together and process complex data streams, thereby detecting silent data corruption that may only occur during system-level integration and interaction.
[0058] Even if all unit and integration tests for the computing core, function units, and memory subsystem pass, it still cannot completely guarantee that the GPU will not encounter errors when running complex machine learning models in actual operation. The reason is that complex computing kernels (such as attention mechanisms) frequently and interleavedly call different hardware units, which may lead to timing errors, state conflicts, or data corruption that only occur under specific scheduling orders, data dependencies, or resource contention conditions. These problems cannot be effectively simulated through unit tests. This embodiment introduces an end-to-end detection module that can run realistic and complex attention mechanism kernels, thereby reproducing the actual working state of the GPU in the target application (such as large language model training) to the greatest extent possible. This places all underlying hardware units in a collaborative, high-pressure real working environment, thus discovering silent data corruption caused by complex interactions between units. Furthermore, when this module fails to detect the problem while other unit detection modules pass, it can determine that the root cause of the fault lies in the collaboration, scheduling, or complex data paths between hardware units, thus providing guidance for in-depth fault diagnosis.
[0059] The end-to-end testing module can be a software component used to ultimately verify the graphics processing unit's (GPU) ability to execute complete and complex computational tasks. The input to the end-to-end testing module can be a complete description of the computational task, and the output can be a comparison between the final result of the task executed on the GPU and a reference result. The attention mechanism computation kernel can be a highly optimized GPU computation program that implements the attention mechanism algorithm. The hardware kernel output refers to the output data generated after the GPU runs the attention mechanism computation kernel. The reference kernel output refers to the standard output considered correct and used for comparison with the hardware kernel output. The source of the reference kernel output can be various, such as the result of a reference implementation running the same logic on a CPU without GPU-level optimization; or a small batch of reference implementations running the same kernel on a GPU with a smaller batch size. The second difference threshold can be set comprehensively based on the numerical characteristics of the attention mechanism computation kernel, possible cumulative rounding errors, and the precision loss that the application can tolerate. It can be more lenient than the threshold for unit testing, and different second difference thresholds can be configured for different task types.
[0060] Figure 2 This diagram illustrates a flowchart of a machine learning task processing method according to another embodiment of this disclosure, wherein the machine learning task can be executed in parallel by a scheduling node and multiple worker nodes from a preset set of worker nodes. Figure 2 As shown, the processing method includes the following steps:
[0061] Step S210: Control multiple work nodes in a preset set of work nodes through a scheduling node, so that each work node executes a silent data corruption detection method.
[0062] In a distributed computing system, the scheduling node can be the central control node responsible for task reception, decomposition, resource scheduling, worker node management, status monitoring, and global coordination. For example, in a Kubernetes-based training task, the Master node can act as the scheduling node. The pre-defined worker node set refers to a group of computing nodes pre-allocated or designated by the scheduling node to execute the current machine learning task. This set defines the scope of computing resources for the task, and the multiple parallel computing loads of the task will be distributed to multiple nodes in this set for execution. A worker node refers to the physical or virtual computing node that actually undertakes the computing task load. Each worker node is typically equipped with accelerated hardware such as a GPU for running SDC detection and machine learning computations. The silent data corruption detection method executed by each worker node can be implemented as described in the previous embodiment.
[0063] Step S220: If the detection result reported by any working node is an abnormal detection result, the first working node with the abnormal detection result is removed from the preset working node set by the scheduling node, and the second working node is rescheduled from the resource pool so that the second working node replaces the first working node.
[0064] The resource pool refers to the set of available computing resources maintained by the cluster management system and dynamically allocated by scheduling nodes. The first working node is the faulty working node that reported an anomaly detection result after performing SDC checks. The second working node is a new working node scheduled from the resource pool to replace the first working node.
[0065] Step S230: If the detection results of multiple working nodes in the preset working node set are all normal, the machine learning task is executed in parallel through multiple working nodes in the preset working node set.
[0066] This embodiment ensures that each working node participating in the computation passes a hardware health check before the task begins. Furthermore, it establishes a dynamic, self-healing fault handling mechanism to ensure that the distributed machine learning task can start under the premise that all nodes are healthy, thereby ensuring reliability.
[0067] In distributed machine learning scenarios, tasks are typically executed in parallel by multiple worker nodes (such as GPU servers). If SDC detection is performed only locally on the worker nodes, the lack of centralized coordination and decision-making leads to the following problems: First, the hardware SDC of a single node may affect the calculated gradients or model shards, and through inter-node communication, impact the entire training task, resulting in global failure. Second, after a faulty node is detected, manual intervention is required to remove it from the task and find a replacement node, causing significant interruptions and time consumption, which cannot meet the needs of automated operation and maintenance for large-scale clusters. In this embodiment, once an abnormal node is detected, the scheduling node can remove it from the preset set of worker nodes, ensuring that the faulty node will not participate in any subsequent task steps (such as gradient synchronization or model synchronization), thereby cutting off the error propagation path and controlling the impact of the fault within a single node. Furthermore, by automatically rescheduling healthy nodes from the resource pool to replace the faulty node, the system can restore the full scale of computing resources without manual intervention, ensuring the sustainability of the parallelism required for distributed tasks.
[0068] Optionally, the machine learning task includes at least one of the following: model training task and model inference task; and after the second worker node is rescheduled from the resource pool to replace the first worker node, multiple worker nodes in the preset worker node set can repeatedly execute the above-mentioned silent data corruption detection method until the detection termination condition is met. The detection termination condition can be one or a set of logical conditions used to determine whether the preset worker node set has reached a state where it is safe to start the task, or whether the attempt should be abandoned and an error reported. For example, the detection termination condition can be set to all worker nodes in the preset worker node set reporting normal detection results in a complete detection loop. Alternatively, the detection termination condition can be that the number of retry scheduling and detection loops reaches a preset maximum threshold N. The maximum retry threshold N is used to prevent infinite looping in the event of a large-scale cluster hardware failure.
[0069] To facilitate understanding, an example is provided below to illustrate the specific implementation details of the silent data corruption detection method provided in this application. This example relates to a method for detecting silent data corruption (SDC) on a graphics processing unit (GPU) during large model training, specifically involving technical fields such as large model training, reliability, and fault detection. As the scale of large language model (LLM) training increases dramatically, the size, complexity, and continuous runtime of the computing hardware (such as GPUs) required for model training and application also increase dramatically. Correspondingly, SDC occurring at the hardware level becomes an extremely serious and insidious problem. SDC refers to errors that occur during data computation, transmission, or storage without being detected by the underlying hardware or system, thus "silently" contaminating the results of upper-layer applications (such as model training). Because large model training is highly sensitive to computational accuracy, SDC on GPUs can lead to unpredictable gradual shifts in model weights during training, ultimately resulting in model convergence failure, decreased accuracy, or absurd outputs, which are difficult to trace. Due to the sporadic and covert nature of SDC, it will result in huge losses (such as wasting a lot of GPU computing power) and is difficult to debug.
[0070] The related technology provides a method for detecting silent data corruption, specifically implemented as follows: A target machine learning model is loaded onto the target computing chip under test; a preset dataset is input into the target machine learning model to perform calculations, and the calculation process ends when the storage stress test conditions of the target computing chip are met; the model calculation data generated during the calculation process is compared with standard calculation data, and if the comparison results are inconsistent, it is determined that the target computing chip has silent data corruption at the data storage level. This method determines whether silent data corruption exists in the hardware by comparing the differences in the results of the target machine learning model running on a standard chip and the chip under test.
[0071] The above-mentioned methods in related technologies have at least the following problems: In the above scheme, the health status of the computing chip is detected by running a target machine learning model. The detection effect may be different when using different machine learning models. Moreover, it can only detect the existence of problems, but cannot accurately locate the specific faulty hardware unit. For example, it cannot locate which GPU is faulty, or whether it is a CUDA core or a Tensor Core. Such information is crucial for maintenance personnel. At the same time, the above detection method is difficult to automate. To achieve fully automated detection, it is necessary to build complex judgment logic and knowledge base.
[0072] To address the aforementioned issues, this example proposes a systematic and high-coverage SDC (Software-Defined Computation) comprehensive detection method for large model training hardware. This method aims to rapidly and comprehensively perform SDC detection on multiple key hardware paths, including GPU computing cores, dedicated acceleration units, special function units, and memory consistency units, before the training task begins (i.e., the "takeoff check" phase). This mitigates the risk of large model training failure due to hardware SDC at the source. This example employs a layered, progressive fault-localization detection strategy, focusing on converging the detection scope from the entire training resource cluster to the hardware computing modules on specific cards, facilitating post-fault handling by maintenance personnel. Furthermore, the test setup in this example can be integrated into the takeoff check of large model training, eliminating the need for manual checks in the test environment. It can be executed during training tasks, and detection modules can be dynamically added and removed, allowing for flexible adjustment of detection items and achieving a highly customizable effect.
[0073] This example provides a comprehensive method for silent data corruption detection on GPUs for training large models. It can be executed before the training task begins and includes multiple cooperating detection modules, each designed for a specific hardware unit and SDC mode. The following sections provide a detailed explanation of each detection module in this example:
[0074] (1) General matrix verification module:
[0075] In this example, the General Matrix Verification (GCM) module can also be called the General Matrix Multiplication (GEMM) computation core detection module. Correspondingly, the target hardware to be detected can be the arithmetic logic units (such as FP32 CUDACores) in the streaming multiprocessor of a GPU. The detection principle is to perform deterministic, large-scale FP32 precision matrix multiplication operations. GEMM consists of massive addition (FMA) operations, which are the most fundamental operations of the computation core. The SDC determination method is to compare the GPU's computation result with a reference result obtained on the CPU using a high-precision or verifiable algorithm element-by-element. Any difference exceeding a predetermined floating-point error tolerance (significantly greater than normal rounding error) is considered evidence of SDC in the computation core. This module can be used to expose fundamental logical errors in the computation core.
[0076] (2) Mixed precision matrix verification module:
[0077] In this example, the Mixed Precision Matrix Verification (SDC) module can also be called the Tensor Core Dedicated Acceleration Unit (SDC) Detection Module. Correspondingly, the target hardware to be detected can be hardware units in the GPU specifically designed for mixed precision matrix computation (such as Tensor Cores). The SDC detection principle is to perform mixed precision (e.g., FP16 input / FP32 accumulation) matrix multiplication operations to simulate the actual computation path of large model training. The SDC determination method is to diagnose problems such as Tensor Core circuit faults, precision overflow, or underflow by analyzing whether abnormal values (such as NaN, Inf) appear in the calculation results, or by observing the difference patterns from the CPU reference results. This module can specifically detect SDCs of mixed precision dedicated hardware that may be missed by standard CUDA Core testing.
[0078] (3) Function calculation verification module:
[0079] In this example, the function computation verification module can also be called the Special Function Units (SFU) detection module. Correspondingly, the target hardware to be detected can be the special function units in the GPU, used to compute transcendental functions such as sin, cos, exp, and log. Detection principle: Self-verification is achieved on the GPU using mathematical identities, without relying on CPU reference values. For example, for the same input vector x, sin(x) and cos(x) are calculated simultaneously, and the deviation of sin(x)^2 + cos(x)^2 from 1 is verified to be within a reasonable tolerance. The SDC determination method is: when the deviation of the mathematical identity exceeds a preset scientific computing tolerance, the special function unit is determined to have experienced SDC. This method is particularly suitable for detecting function computation errors implemented through complex circuits such as lookup tables and iterative algorithms.
[0080] (4) Memory consistency verification module:
[0081] In this example, the memory consistency verification module can also be called the global memory and consistency detection module for reduction operations. Accordingly, the target hardware to be tested can include the global memory controller, cache consistency protocol unit, and atomic operation unit. The detection principle is as follows: A global reduction operation (such as sum, max, min) is performed, which requires data synchronization and coordination across multiple thread blocks / grids on the GPU. The SDC determination method is to compare the GPU's reduction result with the CPU's reduction result on the same dataset. Any inconsistency indicates that SDC has occurred from the computation unit to memory write, or during global data synchronization and atomic operations. This module is used to detect the "last mile" problem where the computation is correct but the data is written back or synchronization fails.
[0082] (5) End-to-end detection module:
[0083] In this example, the end-to-end detection module can also be called the Flash Attention kernel end-to-end function detection module. Accordingly, the target hardware to be detected is: to integrate and verify the collaborative working state of the hardware units involved in the above (1)-(4) modules in the actual, complex AI kernel. The detection principle is: to run the standard Flash Attention algorithm kernel, which intensively calls GEMM, element-level operations (may involve SFU) and reduction operations, and belongs to the core of the Transformer architecture. The SDC determination method is: to compare the Flash Attention output on the GPU with the verified CPU output or a small batch reference implementation. This module is used to implement integrated system testing and can expose SDC caused by complex data flow and hardware unit interaction that may not be apparent in single module testing.
[0084] Figure 3 This diagram illustrates a flowchart of a machine learning task processing method implemented in this example, using a machine learning task as the training task for illustration. Figure 3 As shown, the task processing method includes the following steps:
[0085] Step S301: Start the training task and create the training load.
[0086] The training workload includes a master scheduling node and worker nodes. The master scheduling node is mainly responsible for the overall management and coordination of the training task lifecycle, while the worker nodes are responsible for the actual operation and computation of the training task.
[0087] Step S302: Initiate takeoff check to check whether the training environment of the working node meets the training requirements.
[0088] Step S303: Each working node performs silent data corruption detection.
[0089] The silent data corruption detection performed by each working node includes: detection of the general matrix multiplication calculation core, detection of the tensor core dedicated acceleration unit, detection of the transcendental function special function unit, detection of the global memory and consistency of the reduction operation, and detection of the end-to-end function of the attention mechanism kernel.
[0090] Step S304: The worker node collects the detection results and reports them to the scheduling node.
[0091] Step S305: The scheduling node determines whether the detection result of each working node is passed.
[0092] Specifically, after the scheduling node receives an event from the worker node reporting the detection result through the SDC detection module, it checks the attributes of the currently reported event and determines whether the detection result is passed based on the event attributes.
[0093] Step S306: If the detection results of multiple detection modules in the working node are all passed, then start the training task.
[0094] Step S307: If the detection result of at least one detection module in at least one working node is a failure, a new working node is rescheduled for a new round of takeoff checks.
[0095] Specifically, when a worker node fails the SDC check, that worker node is removed from the training worker node list, and a new worker node is rescheduled to replace the deleted worker node. This initiates a new round of takeoff checks to ensure that any subsequent training steps (such as gradient synchronization) do not include the deleted faulty node, thus preventing error propagation. The checks are repeated until the node passes or the maximum number of restarts is reached.
[0096] This example combines different types of SDC detection methods (GEMM alignment, TensorCore outlier analysis, SFU mathematical identity verification, reduction result consistency, and Flash Attention end-to-end verification) for different hardware subsystems into an organic, holistic detection scheme specifically designed to ensure the computational reliability of large model training hardware. Furthermore, by employing multiple detection modules and a hierarchical strategy, the fault range is narrowed down, enabling precise localization to the smallest fault unit. Moreover, SDC detection can be integrated with the training process, achieving automated environment detection and fault node handling, thereby reducing manual costs. Additionally, since multiple detection modules can be designed as plug-ins, detection items can be flexibly adjusted, supporting the dynamic addition and deletion of detection items, facilitating flexible formulation of detection tasks based on different hardware platforms and detection objectives.
[0097] Therefore, this example can achieve at least the following technical effects: (1) High systematic coverage: It can cover the complete SDC detection chain from basic computing units, dedicated acceleration units, special function units to data path consistency and end-to-end AI kernel, eliminating the blind spots of single tests. (2) Strong targeting: Each detection module is designed for specific computing modes and hardware paths of large model training workloads, which can effectively trigger and expose typical SDC failure modes related to them, and the detection efficiency is high. (3) Advance risk avoidance: Placing the detection before the start of the training task ("take-off check") can discover hardware with potential SDC risks in advance, avoiding its use in large model training tasks that last for weeks or even months, thereby saving huge computing and time costs. (4) Rich diagnostic information: Failure of different modules can point to different types of suspected hardware failure units (such as GEMM failure pointing to general computing core, Tensor Core test failure pointing to mixed precision unit), providing clear fault location clues for operation and maintenance personnel, and helping to quickly repair or replace hardware. (5) Good practicality: The detection scheme is based on the standard GPU programming model and AI operators, which can be easily integrated into the existing cluster management, task scheduling and health check process to achieve automation.
[0098] In summary, compared with general hardware stress testing or functional testing, this example has at least the following significant technical advantages: (1) Deep customization: It is designed entirely around the core computational patterns (matrix operations, transcendental functions, reduction, attention mechanisms) of large model training, and the testing scenarios are highly consistent with real loads. (2) Diverse testing dimensions: It not only tests the correctness of computation, but also tests numerical stability (NaN / Inf), mathematical consistency (identity) and system consistency (GPU-CPU consistency), which is more comprehensive and easy to expand. (3) Easy to automate: The testing scheme is based on standard GPU programming models and AI operators, which can be easily integrated into existing cluster management, task scheduling and health check processes to achieve automation.
[0099] Figure 4 A schematic diagram of a silent data corruption detection device according to another embodiment of this disclosure is shown, as follows: Figure 4 As shown, this silent data corruption detection device is applied to a system running machine learning tasks, and the device includes:
[0100] The detection module 41 is used to call multiple different detection modules before running the machine learning task to detect multiple different hardware associated with the running process of the machine learning task, so as to obtain the detection results corresponding to the multiple different detection modules and obtain multiple detection results.
[0101] The determination module 42 is used to determine, in the case that there are abnormal detection results among the multiple detection results, abnormal hardware with silent data corruption from the multiple different hardware according to the abnormal detection results.
[0102] The specific working principles of each of the above modules can be found in the descriptions of the corresponding parts in the method embodiments, and will not be repeated here.
[0103] This disclosure also provides an electronic device, including:
[0104] At least one processor; and
[0105] A memory communicatively connected to the at least one processor; wherein,
[0106] The processor is configured as described above.
[0107] It is understood that the various method embodiments mentioned above in this disclosure can be combined with each other to form combined embodiments without violating the principle and logic. Due to space limitations, this disclosure will not elaborate further. Those skilled in the art will understand that in the above methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0108] Figure 5 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.
[0109] Reference Figure 5 This disclosure provides an electronic device comprising: at least one processor 701; at least one memory 702; and one or more I / O interfaces 703 connected between the processor 701 and the memory 702; wherein the memory 702 stores one or more computer programs executable by the at least one processor 701, the one or more computer programs being executed by the at least one processor 701 to enable the at least one processor 701 to perform the aforementioned silent data corruption detection method.
[0110] This disclosure also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the above-described method. The computer-readable storage medium may be volatile or non-volatile.
[0111] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device performs the above-described method.
[0112] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0113] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable program instructions, data structures, program modules, or other data. Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0114] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0115] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0116] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0117] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0118] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0119] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0120] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0121] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. A method for detecting silent data corruption, characterized in that, The method includes: Before running the machine learning task, multiple different detection modules are invoked to detect multiple different hardware associated with the running process of the machine learning task, so as to obtain the detection results corresponding to the multiple different detection modules and obtain multiple detection results. If any of the multiple detection results contain anomaly detection results, then based on the anomaly detection results, identify the anomalous hardware that has silently corrupted data from among the multiple different hardware components.
2. The method according to claim 1, characterized in that, Before invoking multiple different detection modules to detect multiple different hardware components associated with the execution of the machine learning task, the method further includes: Based on the task type of the machine learning task, multiple different detection modules corresponding to the task type are selected from a pre-configured pool of candidate detection modules, and the multiple different detection modules are loaded; wherein, the pre-configured pool of candidate detection modules are plug-in modules that support dynamic loading or unloading.
3. The method according to claim 1 or 2, characterized in that, The various detection modules include: a computational core verification module for verifying the numerical computation cores in the graphics processor; The step of obtaining the detection results corresponding to the multiple different detection modules includes: Obtain the hardware calculation result obtained by the computing core verification module after performing a calculation task based on the first preset input data, and determine the detection result corresponding to the computing core verification module based on the degree of difference between the hardware calculation result and the reference calculation result; The reference calculation result is obtained by the central processing unit after executing the calculation task based on the first preset input data; and, if the difference meets the first preset difference condition, the detection result corresponding to the calculation core verification module is determined to be an abnormal detection result.
4. The method according to claim 3, characterized in that, The numerical computation core in the graphics processor includes: a general-purpose computation core and a tensor computation core; and the computation core verification module includes: A general matrix verification module for performing general matrix calculations on the general computing core, and a mixed precision matrix verification module for performing mixed precision matrix calculations on the tensor computing core.
5. The method according to claim 2, characterized in that, The various detection modules include: a function calculation verification module for verifying function units in a graphics processor; The step of obtaining the detection results corresponding to the multiple different detection modules includes: Obtain at least two function calculation results obtained by the function calculation verification module after performing at least two related function operations on the preset function data; The calculation results of the at least two functions are self-verified according to a preset mathematical identity, and the detection result of the function calculation verification module is obtained based on the self-verification result. In cases where the self-verification result does not conform to the mathematical identity, the detection result of the function calculation verification module is determined to be an abnormal detection result.
6. The method according to claim 2, characterized in that, The plurality of different detection modules also include: a memory consistency verification module for detecting the memory consistency verification unit; The step of obtaining the detection results corresponding to the multiple different detection modules includes: The memory consistency verification module obtains the hardware test results after the memory consistency verification unit in the graphics processor performs a memory consistency test on the second preset input data. The hardware test results are compared with the reference test results for consistency. In cases where the hardware test results are inconsistent with the reference test results, the detection result corresponding to the memory consistency verification module is determined to be an abnormal detection result.
7. The method according to claim 6, characterized in that, When the memory consistency test is implemented through a global reduction operation, the hardware test result is the hardware reduction result obtained by the memory consistency verification unit performing a global reduction operation on the second preset input data; and the reference test result is obtained by the central processing unit performing a global reduction operation on the second preset input data.
8. The method according to claim 2, characterized in that, The plurality of different detection modules also include: an end-to-end detection module; The step of obtaining the detection results corresponding to the multiple different detection modules includes: The end-to-end detection module obtains the hardware kernel output result obtained after the graphics processor runs a preset composite computing task kernel; wherein, the composite computing task kernel is configured to call multiple hardware functions of the graphics processor. The output of the hardware kernel is compared with the output of the reference kernel to obtain the detection result of the end-to-end detection module; Specifically, if the difference between the output of the hardware kernel and the output of the reference kernel satisfies the second preset difference condition, the detection result corresponding to the end-to-end detection module is determined to be an abnormal detection result.
9. A method for processing machine learning tasks, characterized in that, The machine learning task is executed in parallel by a scheduling node and multiple worker nodes from a preset set of worker nodes. The method includes: The scheduling node controls multiple working nodes in the preset set of working nodes, so that each working node executes the silent data corruption detection method according to any one of claims 1-8. If any working node reports an abnormal detection result, the first working node with the abnormal detection result is removed from the preset working node set by the scheduling node, and the second working node is rescheduled from the resource pool so that the second working node replaces the first working node. If the detection results of multiple working nodes included in the preset working node set are all normal, the machine learning task is executed in parallel through the multiple working nodes in the preset working node set.
10. The method according to claim 9, characterized in that, The machine learning task includes at least one of the following: a model training task and a model inference task; and after rescheduling the second worker node from the resource pool to replace the first worker node, the process further includes: Multiple working nodes in the preset set of working nodes repeatedly execute the silent data corruption detection method according to any one of claims 1-8 until the detection termination condition is met.
11. A detection device for silent data corruption, characterized in that, The device includes: The detection module is used to call multiple different detection modules before running the machine learning task, and to detect multiple different hardware associated with the running process of the machine learning task, so as to obtain the detection results corresponding to the multiple different detection modules and obtain multiple detection results. The determination module is used to determine, in the case that an anomaly detection result is included among the multiple detection results, an anomaly hardware with silent data corruption from among the multiple different hardware based on the anomaly detection result.
12. An electronic device, characterized in that, include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in, The processor is configured to perform the method according to any one of claims 1-10.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-10.
14. A computer program product comprising computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, characterized in that, When the computer-readable code is run in an electronic device, the processor in the electronic device performs the method of any one of claims 1-10.