Fault-tolerant rescheduling method and apparatus for distributed training scenario

By listening to kernel logs and training logs to identify hardware failures and mark failed devices, the uncertainty problem of fault-tolerant scheduling in hardware failures in distributed training is solved, and the success rate and stability of scheduling are improved.

WO2025124179A1PCT designated stage expired Publication Date: 2025-06-19CHINA TELECOM CLOUD TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/135829
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-11-29
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

In distributed training scenarios, existing fault-tolerant rescheduling methods rely on the error return code of the training program in case of hardware failure, there is uncertainty in encoding normativeness and error recognition, and may be rescheduled to the fault node.

Method used

By listening to the kernel log of the hardware device and the log of the training program instance, identify hardware failures and add error marks to the failed device, putting it out of the scheduling range, ensuring rescheduling to a healthy node.

Benefits of technology

The dependence of fault tolerance function on the encoding standardization and error recognition of training programs is reduced, the success rate of fault tolerance scheduling is improved, and the possibility of rescheduling to faulty devices is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135829_19062025_PF_FP_ABST
    Figure CN2024135829_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of rescheduling, and provides a fault-tolerant rescheduling method and apparatus for a distributed training scenario. The method comprises: presetting a training number and scheduling training program instances for distributed training; monitoring kernel logs of resource devices and training logs of program instances; when information about a faulty hardware device appears in the kernel logs, adding an error tag to the faulty hardware device; checking the training log of the program instance corresponding to the faulty hardware device, and adding an error tag to the program instance having reported an error in the training log; placing the hardware device having the error tag and the program instance having the error tag outside a scheduling scope; and supplementing the scheduling scope with pre-trained program instances until the training number is met; and scheduling the program instances within the supplemented scheduling scope to the resource devices within the scheduling scope for distributed training. The present invention can improve the rescheduling efficiency and success rate.
Need to check novelty before this filing date? Find Prior Art

Description

A fault-tolerant rescheduling method and device for distributed training scenarios Technical Field

[0001] The present invention relates to the field of rescheduling technology, and in particular to a fault-tolerant rescheduling method and device for distributed training scenarios. Background Art

[0002] In model training scenarios, especially those involving large models, hardware failures are more likely. Automatic training recovery is necessary after hardware failures, and fault-tolerant rescheduling is one method for automatically resuming training. Fault-tolerant rescheduling relies on the return value of training instances. If one or more of a group of distributed training instances exit with a non-zero error code, the scheduler will reschedule the training instances that exited with a non-zero error code. During rescheduling, new hardware resources are allocated, and distributed training resumes.

[0003] However, in the fault-tolerant rescheduling method, the training program's exit with a non-zero error code is merely a specification, not a hard constraint. If the training program does not return a non-zero error according to the specification, the platform cannot trigger rescheduling, resulting in a loss of fault tolerance. This makes the fault-tolerant function highly dependent on the standardization of the training program coding. If the hard constraint is intrusive to the training program, the training program instance may exit due to hardware or software issues, requiring the correct distinction between these issues and the return of different error codes. This makes the fault-tolerant rescheduling function highly dependent on the accuracy of the training program's error classification. Even if the training program instance returns an error code and the scheduler process reschedules, there is a chance that the training program instance will be rescheduled back to the node with the hardware failure because the hardware-faulty device is still within the scheduling range. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the related art. To this end, the present invention provides a fault-tolerant rescheduling method and apparatus for distributed training scenarios.

[0005] The present invention provides a fault-tolerant rescheduling method for a distributed training scenario, comprising:

[0006] S1: Preset the number of training programs and schedule training program instances, and schedule the program instances to resource devices for distributed training;

[0007] S2: monitoring the kernel log of the resource device performing distributed training and the training log of the program instance;

[0008] S3: When information about a hardware device with a fault error appears in the kernel log, an error mark is added to the hardware device with the fault error; a training log of a program instance corresponding to the hardware device with the fault error is checked, and an error mark is added to the program instance with the error reported in the training log;

[0009] S4: Set the scheduling range and place the hardware devices with error marks and the program instances with error marks outside the scheduling range;

[0010] S5: When the number of program instances remaining in the scheduling range is less than the training quantity, add pre-trained program instances that meet the training quantity to the scheduling range;

[0011] S6: Dispatching the supplemented program instances within the scheduling range to the resource devices within the scheduling range for distributed training to complete fault-tolerant rescheduling.

[0012] According to a fault-tolerant rescheduling method for a distributed training scenario provided by the present invention, the hardware device in step S3 includes a GPU processor and an IB card.

[0013] According to a fault-tolerant rescheduling method for a distributed training scenario provided by the present invention, the error mark in step S3 includes an error keyword and an error code value.

[0014] According to a fault-tolerant rescheduling method for a distributed training scenario provided by the present invention, the error keyword and the error code value can both be configured by regular filtering.

[0015] According to a fault-tolerant rescheduling method for a distributed training scenario provided by the present invention, in step S3, when performing hardware device failure error judgment and program instance error judgment, the kernel log has a higher priority than the training log.

[0016] According to a fault-tolerant rescheduling method for a distributed training scenario provided by the present invention, step S5 further includes:

[0017] S51: When the number of program instances remaining in the scheduling range is greater than or equal to the training quantity, the program instances in the scheduling range are not supplemented.

[0018] The present invention also provides a fault-tolerant rescheduling device for a distributed training scenario, comprising:

[0019] Pre-scheduling module: used to preset the number of training sessions and schedule training program instances, and dispatch the program instances to resource devices for distributed training;

[0020] Monitoring module: used for monitoring and obtaining the kernel log of the resource device for distributed training and the training log of the program instance;

[0021] The error monitoring module is used to add an error mark to a hardware device when a hardware device with a fault error appears in the kernel log; it is also used to check the training log of the program instance corresponding to the hardware device with the fault error and add an error mark to the program instance with an error reported in the training log;

[0022] Range supplement module: used to set the scheduling range, place hardware devices with error marks and program instances with error marks outside the scheduling range; also used to supplement the pre-trained program instances that meet the training quantity to the scheduling range when the program instances remaining in the scheduling range are less than the training quantity;

[0023] Rescheduling module: used to schedule the supplemented program instances within the scheduling range to the resource devices within the scheduling range for distributed training to complete fault-tolerant rescheduling.

[0024] According to a fault-tolerant rescheduling device for a distributed training scenario provided by the present invention, the error monitoring module further includes:

[0025] Log screening unit: used to screen the kernel log for hardware device information with fault errors, and also used to screen the training log of the program instance corresponding to the hardware device with fault errors for errors;

[0026] Error marking unit: used for marking the error devices and error-reporting program instances obtained by the log screening unit.

[0027] The present invention also provides a fault-tolerant rescheduling device for a distributed training scenario, comprising:

[0028] a memory and at least one processor, wherein instructions are stored in the memory;

[0029] At least one of the processors calls the instructions in the memory to enable the document database audit device to execute a fault-tolerant rescheduling method for a distributed training scenario as described in any one of the above items.

[0030] The present invention also provides a computer-readable storage medium having instructions stored thereon, and when the instructions are executed by a processor, a fault-tolerant rescheduling method for a distributed training scenario as described in any one of the above items is implemented.

[0031] The present invention also provides a fault-tolerant rescheduling method, apparatus, equipment and storage medium for distributed training scenarios. In a distributed training scenario, when a hardware failure occurs, the hardware failure is identified by comparing the kernel log and the training instance log, and the training instance is automatically rescheduled for fault tolerance, which greatly reduces the dependence of the fault-tolerant function on the training program coding standardization and the effectiveness of error identification. At the same time, by labeling the faulty device with a prohibition scheduling label, repeated scheduling to the faulty device is avoided, effectively improving the success rate of fault-tolerant scheduling.

[0032] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0034] FIG1 is a flow chart of a fault-tolerant rescheduling method for a distributed training scenario provided by the present invention;

[0035] FIG2 is a schematic structural diagram of a fault-tolerant rescheduling device for a distributed training scenario provided by the present invention.

[0036] Reference numerals:

[0037] 100, pre-scheduling module; 200, monitoring module; 300, error monitoring module; 310, log screening unit; 320, error marking unit; 400, range supplement module; 500, rescheduling module. DETAILED DESCRIPTION

[0038] To make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the embodiments described are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.

[0039] In the description of the embodiments of the present invention, it should be noted that the terms "center", "longitudinal", "lateral", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the embodiments of the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore should not be understood as limiting the embodiments of the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only and should not be understood as indicating or implying relative importance.

[0040] In the description of the embodiments of the present invention, it should be noted that, unless otherwise specified or limited, the terms "connected" and "connection" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; mechanical connections, electrical connections; and direct connections or indirect connections through an intermediary. Those skilled in the art will understand the specific meanings of the above terms in the embodiments of the present invention based on the specific circumstances.

[0041] In the embodiments of the present invention, unless otherwise expressly specified or limited, a first feature being "above" or "below" a second feature may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, a first feature being "above," "above," or "above" a second feature may mean that the first feature is directly above or diagonally above the second feature, or simply means that the first feature is at a higher level than the second feature. A first feature being "below," "below," or "below" a second feature may mean that the first feature is directly below or diagonally below the second feature, or simply means that the first feature is at a lower level than the second feature.

[0042] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0043] In order to better understand the present invention, the technical terms mentioned in the embodiments of the present invention are explained below.

[0044] Distributed training is essentially distributed computing, which uses a cluster of multiple machines to break down a large, complex problem into multiple small, simple problems that are solved in parallel, and the results of the small problems are combined into the final result.

[0045] The embodiment of the present invention will be described below with reference to FIG1 and FIG2 .

[0046] The present invention provides a fault-tolerant rescheduling method for a distributed training scenario, comprising:

[0047] S1: Preset the number of training programs and schedule training program instances, and schedule the program instances to resource devices for distributed training;

[0048] S2: monitoring the kernel log of the resource device performing distributed training and the training log of the program instance;

[0049] Furthermore, the training scheduler is first trained to schedule the set number of training program instances to a device with resources, and start the training program in the instance, monitoring the kernel log of the training device and the training program instance log.

[0050] S3: When information about a hardware device with a fault error appears in the kernel log, an error mark is added to the hardware device with the fault error; a training log of a program instance corresponding to the hardware device with the fault error is checked, and an error mark is added to the program instance with the error reported in the training log;

[0051] Furthermore, when hardware failure errors such as GPU errors and communication card errors appear in the kernel log, an error mark is added to the device, and it is checked whether the training program instance on the device also reports an error. If an error is reported, an error mark is added to the training program instance.

[0052] The hardware devices in step S3 include a GPU processor and an IB card.

[0053] Furthermore, the content monitored by the present invention includes but is not limited to hardware fault errors such as GPU errors and IB card communication errors appearing in the kernel log.

[0054] The error mark in step S3 includes an error keyword and an error code value.

[0055] The error keywords and the error code values ​​can both be configured by regular filtering.

[0056] Among them, in step S3, when performing hardware device failure error judgment and program instance error judgment, the priority of the kernel log is higher than the training log.

[0057] Furthermore, the kernel log has a higher priority than the training log because the kernel log standard specification facilitates inspection, so the kernel log is mainly filtered, supplemented by the training program instance log.

[0058] Furthermore, it is necessary to monitor the device status. When an error mark is found, the device where the instance is located is marked with a scheduling prohibition label to exclude the node from the scheduling range, traverse the training instances on the error device, and delete the training instances with error marks.

[0059] S4: Set the scheduling range and place the hardware devices with error marks and the program instances with error marks outside the scheduling range;

[0060] S5: When the number of program instances remaining in the scheduling range is less than the training quantity, add pre-trained program instances that meet the training quantity to the scheduling range;

[0061] Wherein, step S5 further includes:

[0062] S51: When the number of program instances remaining in the scheduling range is greater than or equal to the training quantity, the program instances in the scheduling range are not supplemented.

[0063] S6: Dispatching the supplemented program instances within the scheduling range to the resource devices within the scheduling range for distributed training to complete fault-tolerant rescheduling.

[0064] Furthermore, when the problem training program instance is deleted, the training scheduler checks that the number of training program instances is less than the set value, and schedules the new training program instance to the device with resources, while excluding the device that is prohibited from scheduling. After the scheduling is successful, the training program of the instance is started until the number of instances reaches the expected value, completing the fault-tolerant rescheduling.

[0065] The present invention also provides a fault-tolerant rescheduling device for a distributed training scenario, comprising:

[0066] Pre-scheduling module 100: used to preset the number of trainings and schedule training program instances, and dispatch the program instances to resource devices for distributed training;

[0067] Monitoring module 200: used for monitoring and obtaining the kernel log of the resource device performing distributed training and the training log of the program instance;

[0068] The error monitoring module 300 is configured to add an error mark to a hardware device when a hardware device with a fault error appears in the kernel log; and is further configured to check the training log of the program instance corresponding to the hardware device with the fault error and add an error mark to the program instance with an error reported in the training log;

[0069] Range supplement module 400: used to set a scheduling range, place hardware devices with error marks and program instances with error marks outside the scheduling range; and further used to supplement pre-trained program instances that meet the training quantity to the scheduling range when the number of program instances remaining in the scheduling range is less than the training quantity;

[0070] The rescheduling module 500 is used to schedule the supplemented program instances within the scheduling range to the resource devices within the scheduling range for distributed training, so as to complete fault-tolerant rescheduling.

[0071] Wherein, the error monitoring module further includes:

[0072] Log screening unit 310: used to screen the kernel log for hardware device information with fault errors, and also to screen the training log of the program instance corresponding to the hardware device with fault errors for errors;

[0073] The error marking unit 320 is used to mark the error devices and error-reporting program instances obtained by the log screening unit 310 as errors.

[0074] When applying a fault-tolerant rescheduling device for a distributed training scenario provided by the present invention, the training quantity is first preset, and the training program instance is scheduled through the pre-scheduling module, and the program instance is scheduled to the resource device for distributed training; when performing distributed training, the kernel log of the resource device and the training log of the program instance are monitored through the monitoring module; then the log screening unit in the error monitoring module is used to screen whether there is any hardware device information with fault errors in the kernel log, and screen whether the training log of the program instance corresponding to the hardware device with fault errors reports an error; the error marking unit in the error monitoring module is used to mark the faulty device and the program instance with error obtained by the log screening unit; a scheduling range is set, and the hardware device with error mark and the program instance with error mark are placed outside the scheduling range through the range supplement module; and when the program instances retained in the scheduling range are less than the training quantity, the pre-trained program instances that meet the training quantity are supplemented to the scheduling range; finally, the rescheduling module is used to schedule the supplemented program instances in the scheduling range to the resource devices in the scheduling range for distributed training to complete fault-tolerant rescheduling.

[0075] The present invention also provides a fault-tolerant rescheduling device for a distributed training scenario, comprising:

[0076] a memory and at least one processor, wherein instructions are stored in the memory;

[0077] At least one of the processors calls the instructions in the memory to enable the document database audit device to execute a fault-tolerant rescheduling method for a distributed training scenario as described in any one of the above items.

[0078] In some embodiments, the device may vary significantly due to different configurations or performance, and may include one or more processors, for example, one or more processors and memories, and one or more storage media for storing applications or data, such as one or more mass storage devices. The memories and storage media may be either transient or persistent storage. The program stored on the storage medium may include one or more modules, each of which may include a series of instruction operations on the device. Furthermore, the processor may be configured to communicate with the storage medium and execute the series of instruction operations on the storage medium on the device.

[0079] The device may also include one or more power supplies, one or more wired or wireless network interfaces, one or more input and output interfaces, and / or one or more operating systems, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will appreciate that the device structure shown above does not limit the device, and the device may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0080] The present invention also provides a computer-readable storage medium having instructions stored thereon, and when the instructions are executed by a processor, a fault-tolerant rescheduling method for a distributed training scenario as described in any one of the above items is implemented.

[0081] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0082] If the integrated unit is implemented in the form of a software functional unit and sold or passed as an independent product, it can be stored in a computer-readable storage medium. Based on the understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk and other media that can store program code.

[0083] The present invention provides a fault-tolerant rescheduling method, apparatus, device and storage medium for distributed training scenarios. By comparing kernel logs and training instance logs to identify hardware failures, the fault-tolerant rescheduling of training instances is automatically performed, greatly reducing the dependence of the fault-tolerant function on uncertain factors such as the standardization of training program coding and the effectiveness of error identification. In addition, repeated scheduling to the faulty device is avoided by labeling the faulty device with a prohibition scheduling label, thereby improving the efficiency and success rate of rescheduling.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A fault-tolerant rescheduling method for distributed training scenarios, characterized in that: include: S1: Preset the number of trainings and schedule training program instances, and schedule the program instances to resource devices for distributed training; S2: monitoring the kernel log of the resource device for distributed training and the training log of the program instance; S3: when information about a hardware device with a fault error appears in the kernel log, an error mark is added to the hardware device with the fault error; Check the training log of the program instance corresponding to the hardware device with the fault error, and add an error mark to the program instance with the error reported in the training log; S4: Setting the scheduling scope, placing the hardware devices with error marks and the program instances with error marks outside the scheduling scope; S5: When the number of program instances retained in the scheduling range is less than the training quantity, adding pre-trained program instances that meet the training quantity to the scheduling range; S6: Schedule the supplemented program instances within the scheduling range to the resource devices within the scheduling range for distributed training to complete fault-tolerant rescheduling.

2. The fault-tolerant rescheduling method for a distributed training scenario according to claim 1, characterized in that: The hardware device in step S3 includes a GPU processor and an IB card.

3. The fault-tolerant rescheduling method for a distributed training scenario according to claim 1, characterized in that: The error mark in step S3 includes an error keyword and an error code value.

4. The fault-tolerant rescheduling method for a distributed training scenario according to claim 3, characterized in that: Both the error keyword and the error code value can be configured by regular filtering.

5. The fault-tolerant rescheduling method for a distributed training scenario according to claim 1, characterized in that: In step S3, when performing hardware device failure error determination and program instance error determination, the kernel log has a higher priority than the training log.

6. The fault-tolerant rescheduling method for a distributed training scenario according to claim 1, characterized in that: Step S5 also includes: S51: When the number of program instances remaining in the scheduling range is greater than or equal to the training quantity, the program instances in the scheduling range are not supplemented.

7. A fault-tolerant rescheduling device for a distributed training scenario, characterized in that: include: Pre-scheduling module: used to preset the number of trainings and schedule training program instances, and schedule the program instances to resource devices for distributed training; Monitoring module: used for monitoring and obtaining the kernel log of the resource device for distributed training and the training log of the program instance; Error monitoring module: used for adding an error mark to the hardware device when a hardware device with a fault error appears in the kernel log; It is also used to check the training log of the program instance corresponding to the hardware device that has a fault error, and add an error mark to the program instance that reports an error in the training log; Range supplement module: used to set the scheduling range and put the hardware devices with error marks and program instances with error marks outside the scheduling range; Also used for supplementing the pre-trained program instances satisfying the training quantity to the scheduling scope when the program instances retained in the scheduling scope are less than the training quantity; Rescheduling module: used to schedule the supplemented program instances within the scheduling range to the resource devices within the scheduling range for distributed training to complete fault-tolerant rescheduling.

8. The fault-tolerant rescheduling device for a distributed training scenario according to claim 7, characterized in that: The error monitoring module further comprises: Log screening unit: used to screen the kernel log for hardware device information that has a fault error, and also used to screen the training log of the program instance corresponding to the hardware device that has a fault error for an error report; Error marking unit: used for marking the error devices and error-reporting program instances obtained by the log screening unit.

9. A fault-tolerant rescheduling device for distributed training scenarios, characterized in that: include: A memory and at least one processor, wherein instructions are stored in the memory; At least one of the processors calls the instructions in the memory so that the document database audit device executes the fault-tolerant rescheduling method for a distributed training scenario as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, which, when executed by the processor, implement a fault-tolerant rescheduling method for a distributed training scenario as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Task scheduling method, device and system

    CN112988340A

  • Distributed training task scheduling method and device, equipment and storage medium

    CN116483546A

  • Fault-tolerant rescheduling method and device for distributed training scene

    CN118245257A

  • System for training optimisation

    US20100076278A1