Multi-accelerator card scheduling method and system for model acceleration training

By dynamically adjusting the allocation of accelerator cards, generating computing power containers and accelerator card sequences according to model training requirements, evaluating computing power utilization efficiency in real time, and automatically detecting hardware and software faults, the problem of static and inflexible allocation of accelerator card resources in existing technologies is solved, achieving efficient matching of computing power resources and rapid fault location.

CN119902892BActive Publication Date: 2025-12-16四川并济科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411985346.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-12-16
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

In existing technologies, the allocation of accelerator card resources is static and inflexible, resulting in wasted computing power, inflexible scheduling, and low efficiency in fault handling, making it difficult to meet the dynamic needs of model training.

Method used

By dynamically adjusting the allocation of accelerator cards, generating computing power containers and accelerator card sequences according to model training requirements, evaluating computing power utilization efficiency in real time, and automatically detecting hardware and software faults, the system achieves flexible scheduling of accelerator cards and rapid fault location.

Benefits of technology

Ensuring that computing resources match model training needs improves system flexibility and fault handling efficiency, and shortens troubleshooting time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119902892B_ABST
    Figure CN119902892B_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence, and discloses a multi-acceleration card scheduling method and system for model acceleration training, which comprises the following steps: obtaining a first required computing power according to the data processing speed of model training requirements and the set computing power use efficiency, obtaining the first computing power use efficiency of the current model according to the data processing speed of the current model training and the data processing speed corresponding to the first required computing power; obtaining a second required computing power according to the first computing power use efficiency of the current model and the data processing speed of model training requirements, and forming a second acceleration card sequence; and completing acceleration card scheduling according to the second acceleration card sequence. The application provides data support for the optimized allocation of computing power resources by evaluating the computing power use efficiency in real time.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, in particular to a multi-accelerator card scheduling method and system for model acceleration training. BACKGROUND

[0002] In the field of deep learning, as the complexity of models continues to increase, the demand for computing resources also grows. To accelerate the model training process, multiple accelerator cards (such as GPUs, FPGAs, etc.) are often used for parallel computing. However, how to efficiently and dynamically schedule these accelerator card resources to maximize the efficiency of computing power usage while ensuring the stability and reliability of the training process is a major challenge currently faced.

[0003] Traditional accelerator card scheduling methods are often based on static configuration, that is, before the training begins, a certain number of accelerator cards are manually allocated according to the estimated demand of the model. This method has several obvious problems:

[0004] Waste of computing power: Since the actual computing power demand during the model training process may vary with the training phase, static configuration often leads to an excess of computing power resources in some phases and a shortage in others, resulting in waste of computing power.

[0005] Inflexible scheduling: Once the accelerator cards are allocated, it is difficult to make dynamic adjustments according to real-time demand during the training process, lacking flexibility.

[0006] Low efficiency in handling faults: When accelerator cards fail, the traditional scheduling method usually requires human intervention for fault diagnosis and repair, affecting training efficiency.

[0007] Difficulty in evaluating computing power usage efficiency: There is a lack of effective mechanisms to evaluate and adjust computing power usage efficiency in real time, leading to increased training costs. SUMMARY

[0008] The present application aims to overcome the shortcomings of the prior art and provide a multi-accelerator card scheduling method for model acceleration training, comprising the following steps:

[0009] Step one, according to the data processing speed of the model training demand and the set computing power usage efficiency, the first demand computing power is obtained, the accelerator card scheduling module generates the computing power container according to the first demand computing power, the first computing power container schedules the corresponding accelerator card according to the first demand computing power, and generates the first accelerator card sequence, the accelerator card control module performs linkage test according to the first accelerator card sequence and the target model, obtains the data processing speed of the current model training, and according to the data processing speed of the current model training and the data processing speed corresponding to the first demand computing power, obtains the first computing power usage efficiency of the current model;

[0010] Step two, according to the first computing power utilization efficiency of the current model and the data processing speed of the model training requirement, a second demand computing power is obtained, according to the computing power difference between the second demand computing power and the first demand computing power, the acceleration card with the corresponding computing power difference is scheduled, and the first acceleration card sequence is formed into a second acceleration card sequence;

[0011] Step three, the acceleration card control module performs linkage test according to the second acceleration card sequence and the target model, obtains the data processing speed of the current model training, and obtains the second computing power utilization efficiency of the current model training according to the data processing speed of the current model training and the second demand computing power. If the first computing power utilization efficiency is consistent with the second computing power utilization efficiency, step seven is entered, otherwise, step four is entered.

[0012] Step four, according to the ratio of the total power of the acceleration card of the first acceleration card sequence to the total power of the acceleration card of the second acceleration card sequence, the acceleration card fault type is judged. If it is a hardware fault, step five is entered, if it is a software fault, step six is entered.

[0013] Step five, the power of each acceleration card in the second acceleration card sequence is obtained, and the difference between the power of each acceleration card and the designed power is obtained respectively. The acceleration card with a difference greater than a set difference threshold is a fault acceleration card. According to the fault acceleration card information, the acceleration card is replaced, and step three is returned.

[0014] Step six, according to the second acceleration card sequence and the first computing power utilization efficiency, the target data processing rate of each acceleration card is obtained respectively, and the data processing rate difference between the data processing rate of each acceleration card and the corresponding target data processing rate is obtained respectively. Wherein, the data processing rate difference greater than the set data processing rate difference threshold is the fault acceleration card. According to the fault acceleration card information, the software repair is carried out on the fault acceleration, and step three is returned.

[0015] Step seven, according to the second acceleration card sequence, the acceleration card scheduling is completed.

[0016] Further, the first demand computing power is obtained according to the data processing speed of the model training requirement, which includes:

[0017] According to the corresponding computing power of the model training requirement data processing speed, the first demand computing power is obtained according to the ratio of the corresponding computing power and the set computing power utilization efficiency.

[0018] Further, the first computing power utilization efficiency of the current model is obtained according to the data processing speed of the current model training and the first demand computing power, which includes:

[0019] According to the ratio of the data processing speed of the current model training to the data processing speed corresponding to the first demand computing power, the first computing power utilization efficiency of the current model is obtained, and the following formula is used:

[0020]

[0021] Further, the first demand algorithm power use efficiency according to the current model and the set algorithm power use efficiency are used to obtain the second demand algorithm power, including:

[0022]

[0023] According to the data processing speed corresponding to the second demand algorithm power, the second demand algorithm power is obtained.

[0024] Further, the second algorithm power use efficiency of the current model training is obtained according to the data processing speed of the current model training and the second demand algorithm power, including:

[0025]

[0026] Further, the proportion of the total power of the acceleration card in the first acceleration card sequence and the total power of the acceleration card in the second acceleration card sequence is used to judge the acceleration card fault type, including:

[0027] If the proportion of the total power of the acceleration card in the first acceleration card sequence and the total power of the acceleration card in the second acceleration card sequence is consistent with the proportion of the number of acceleration cards in the first acceleration card sequence and the number of acceleration cards in the second acceleration card sequence, it is a software fault, otherwise, it is a hardware fault.

[0028] Further, the target data processing rate of each acceleration card is obtained according to the second acceleration card sequence and the first algorithm power use efficiency, including:

[0029] According to the data processing speed corresponding to the algorithm power of each acceleration card in the second acceleration card sequence and the first algorithm power use efficiency, the target data processing rate of each acceleration card is obtained.

[0030] The multi-acceleration card scheduling system for model acceleration training applies the multi-acceleration card scheduling method for model acceleration training, and includes a cloud data server, an acceleration card scheduling module, an acceleration card control module, a data acquisition module, a communication module and a data processing module.

[0031] The beneficial effects of the application are: by dynamically adjusting the distribution of acceleration cards, it is ensured that the algorithm power resource always matches the actual demand of model training, and the waste of algorithm power is avoided.

[0032] The number and configuration of the acceleration cards can be dynamically adjusted according to real-time requirements during the model training process, improving the flexibility of the system. The acceleration cards can be added or removed at any time during the training process, adapting to the model training requirements of different scales.

[0033] Hardware and software faults of the acceleration cards are automatically detected, reducing the need for manual intervention and improving fault handling efficiency. Through power ratio analysis and data processing rate comparison, the faulty acceleration card is quickly located, shortening the troubleshooting time. BRIEF DESCRIPTION OF DRAWINGS

[0034] Figure 1 A flowchart of a multi-acceleration card scheduling method for model acceleration training is shown.

[0035] Figure 2 A schematic diagram of a multi-acceleration card scheduling system for model acceleration training is shown. DETAILED DESCRIPTION

[0036] The technical solutions of the present application will be described in further detail below in conjunction with the accompanying drawings, but the scope of protection of the present application is not limited to the following description.

[0037] The features and performance of the present application will be described in further detail below in conjunction with the embodiments.

[0038] As shown in Figure 1 The multi-acceleration card scheduling method for model acceleration training includes the following steps:

[0039] Step one, according to the data processing speed of the model training requirement and the set computing power utilization efficiency, the first demand computing power is obtained, the acceleration card scheduling module generates the computing power container according to the first demand computing power, the first computing power container schedules the corresponding acceleration card according to the first demand computing power, and generates the first acceleration card sequence, the acceleration card control module performs linkage test according to the first acceleration card sequence and the target model, obtains the data processing speed of the current model training, and according to the data processing speed of the current model training and the data processing speed corresponding to the first demand computing power, obtains the first computing power utilization efficiency of the current model;

[0040] Step two, according to the first computing power utilization efficiency of the current model and the data processing speed of the model training requirement, the second demand computing power is obtained, according to the computing power difference value between the second demand computing power and the first demand computing power, the acceleration card corresponding to the computing power difference value is scheduled, and the first acceleration card sequence is constructed to form the second acceleration card sequence;

[0041] Step three, the acceleration card control module performs linkage test according to the second acceleration card sequence and the target model, obtains a data processing speed of the current model training, and obtains a second algorithm power use efficiency of the current model training according to the data processing speed of the current model training and the second demand algorithm power. If the first algorithm power use efficiency is consistent with the second algorithm power use efficiency, step seven is entered. Otherwise, step four is entered.

[0042] Step four, the acceleration card fault type is judged according to the ratio of the total power of the acceleration cards in the first acceleration card sequence to the total power of the acceleration cards in the second acceleration card sequence. If it is a hardware fault, step five is entered. If it is a software fault, step six is entered.

[0043] Step five, the power of each acceleration card in the second acceleration card sequence is obtained, and the difference between the power of each acceleration card and the designed power is obtained. If the difference is greater than a set difference threshold value, the acceleration card is a fault acceleration card. The acceleration card is replaced according to the fault acceleration card information, and step three is returned.

[0044] Step six, the target data processing rate of each acceleration card is obtained according to the second acceleration card sequence and the first algorithm power use efficiency. The data processing rate difference between the data processing rate of each acceleration card and the corresponding target data processing rate is obtained. If the data processing rate difference is greater than a set data processing rate difference threshold value, the acceleration card is a fault acceleration card. The fault acceleration is repaired by software according to the fault acceleration card information, and step three is returned.

[0045] Step seven, the acceleration card scheduling is completed according to the second acceleration card sequence.

[0046] The data processing speed according to the model training requirement is obtained, and the first demand algorithm power is obtained.

[0047] The corresponding algorithm power is obtained according to the model training requirement of the data processing speed. The first demand algorithm power is obtained according to the ratio of the corresponding algorithm power to the set algorithm power use efficiency.

[0048] The first algorithm power use efficiency of the current model is obtained according to the data processing speed of the current model training and the first demand algorithm power.

[0049] The first algorithm power use efficiency of the current model is obtained according to the ratio of the data processing speed of the current model training to the corresponding data processing speed of the first demand algorithm power. The following formula is used:

[0050]

[0051] The second demand algorithm power is obtained according to the first demand algorithm power use efficiency of the current model and the set algorithm power use efficiency.

[0052]

[0053] According to the speed of the data corresponding to the second demand computing power, the second demand computing power is obtained.

[0054] According to the data processing speed of the current model training and the second demand computing power, the second computing power utilization efficiency of the current model training is obtained.

[0055]

[0056] According to the ratio of the total power of the acceleration cards in the first acceleration card sequence to the total power of the acceleration cards in the second acceleration card sequence, the acceleration card fault type is judged.

[0057] If the ratio of the total power of the acceleration cards in the first acceleration card sequence to the total power of the acceleration cards in the second acceleration card sequence is consistent with the number ratio of the number of acceleration cards in the first acceleration card sequence to the number of acceleration cards in the second acceleration card sequence, it is a software fault, otherwise, it is a hardware fault.

[0058] According to the second acceleration card sequence and the first computing power utilization efficiency, the target data processing rate of each acceleration card is obtained.

[0059] According to the data processing speed corresponding to the computing power of each acceleration card in the second acceleration card sequence and the first computing power utilization efficiency, the target data processing rate of each acceleration card is obtained.

[0060] According to the second acceleration card sequence, the acceleration card scheduling is completed, including: according to the second acceleration card sequence

[0061] As shown in Figure 2 The multi-acceleration card scheduling system for model acceleration training applies the multi-acceleration card scheduling method for model acceleration training, and includes a cloud data server, an acceleration card scheduling module, an acceleration card control module, a data acquisition module, a communication module and a data processing module. The acceleration card scheduling module, the acceleration card control module, the data acquisition module and the communication module are connected with the data processing module. The cloud data server is in communication connection with the communication module.

[0062] Specifically, the application provides a multi-acceleration card scheduling method for model acceleration training, which comprises the following steps:

[0063] Step 1: initial computing power allocation and efficiency evaluation

[0064] The first demand computing power is calculated as follows:

[0065] According to the data processing speed of the model training demand, the required computing power is calculated.

[0066] The first demand computing power is obtained by the formula "first demand computing power = required computing power / set computing power use efficiency" in combination with the set computing power use efficiency (e.g. desired computing power utilization rate).

[0067] Generating a computing power container and an accelerator card sequence:

[0068] The accelerator card scheduling module generates a computing power container according to the first demand computing power, and the container is responsible for scheduling a corresponding number of accelerator cards. The scheduled accelerator cards form a first accelerator card sequence.

[0069] Linkage test and efficiency calculation: the accelerator card control module performs linkage test according to the first accelerator card sequence and the target model, and measures the actual data processing speed of the current model training.

[0070] According to the actual data processing speed and the data processing speed corresponding to the first demand computing power, the first computing power use efficiency of the current model is calculated, and the formula is: first demand computing power use efficiency = actual data processing speed / data processing speed corresponding to first demand computing power.

[0071] Step 2: Computing power adjustment and redistribution

[0072] Calculate the second demand computing power:

[0073] According to the first computing power use efficiency of the current model and the data processing speed required by the model training, the second demand computing power is recalculated.

[0074] Specifically, the actual achieved computing power (i.e. the computing power corresponding to the actual data processing speed) is first back calculated according to the first computing power use efficiency, and then the second demand computing power corresponding to the data processing speed is calculated in combination with the set computing power use efficiency and the data processing speed required by the model training, and then the second demand computing power is obtained.

[0075] Adjusting the accelerator card sequence: according to the difference between the second demand computing power and the first demand computing power, a corresponding number of accelerator cards are scheduled, combined with the first accelerator card sequence to form a second accelerator card sequence.

[0076] Step 3: Efficiency evaluation again

[0077] Linkage test and efficiency calculation: the accelerator card control module performs linkage test according to the second accelerator card sequence and the target model, and measures the actual data processing speed of the current model training.

[0078] According to the actual data processing speed and the second demand computing power, the second computing power use efficiency of the current model training is calculated, and the formula is: second computing power use efficiency = actual data processing speed / data processing speed corresponding to second demand computing power.

[0079] Efficiency consistency: If the first computing power usage efficiency is consistent with the second computing power usage efficiency, it means that the computing power adjustment is effective, and step seven is entered; otherwise, step four is entered for troubleshooting.

[0080] Step four: fault type judgment

[0081] Power ratio analysis: by comparing the total power ratio of the first and second accelerator card sequences and the ratio of the number of accelerator cards, the fault type is determined.

[0082] If the power ratio and the number ratio are consistent, it may be a software failure; otherwise, it may be a hardware failure.

[0083] Step five: hardware failure handling

[0084] Identify faulty accelerator cards: compare the power of each accelerator card in the second accelerator card sequence with the designed power, and the accelerator card with a difference greater than the set threshold is considered a faulty accelerator card.

[0085] Replace the accelerator card: replace the faulty accelerator card according to the faulty accelerator card information, and then return to step three for re-evaluation.

[0086] Step six: software failure handling

[0087] Identify faulty accelerator cards: calculate the target data processing rate of each accelerator card according to the second accelerator card sequence and the first computing power usage efficiency. Compare the actual data processing rate of each accelerator card with the target data processing rate, and the accelerator card with a difference greater than the set threshold is considered a faulty accelerator card.

[0088] Software repair: repair the faulty accelerator card according to the faulty accelerator card information, and then return to step three for re-evaluation.

[0089] Step seven: complete accelerator card scheduling

[0090] According to the second accelerator card sequence, complete the final scheduling of the accelerator card to ensure that the model training is carried out under the optimal computing power configuration.

[0091] Specifically:

[0092] According to the data processing speed of the model training requirement, the first required computing power is obtained:

[0093] In specific implementation, the data processing speed required by the model training can be estimated through historical data or empirical formula, and then the required computing power is calculated.

[0094] The set computing power usage efficiency is usually determined according to the actual situation and performance target of the system, such as 90%, 95%, etc.

[0095] The calculation formula of the first computing power usage efficiency is:

[0096] The first computing power use efficiency = actual data processing speed / (first demand computing power * the reciprocal of the set computing power use efficiency corresponding to the data processing speed).

[0097] Calculation of the second demand computing power:

[0098] First, according to the first computing power use efficiency, the actual achieved computing power is deduced, that is, the actual data processing speed divided by the first computing power use efficiency.

[0099] Then, combined with the set computing power use efficiency and the data processing speed of model training demand, the data processing speed corresponding to the second demand computing power is calculated by adjusting, and then the second demand computing power is deduced.

[0100] The logic of fault type judgment:

[0101] If the total power ratio is consistent with the number ratio, but the computing power use efficiency is still not up to standard, it is likely to be a software problem, such as driver incompatibility, low algorithm implementation efficiency, etc.

[0102] If the total power ratio and the number ratio are inconsistent, and there is a significant deviation, it is likely to be a hardware problem, such as damaged accelerator card, poor heat dissipation, etc.

[0103] Detailed steps of hardware fault handling:

[0104] Get the real-time power data of each accelerator card and compare it with the designed power.

[0105] For accelerator cards with abnormally low power, further hardware testing such as temperature detection, signal integrity testing, etc. is performed to confirm the fault.

[0106] After replacing the faulty accelerator card, retest the linkage to ensure the system returns to normal.

[0107] Detailed steps of software fault handling:

[0108] According to the second accelerator card sequence and the first computing power use efficiency, the data processing rate of each accelerator card should be calculated.

[0109] For accelerator cards with actual data processing rate significantly lower than the target rate, check its software configuration, such as driver version, algorithm parameters, etc.

[0110] After repairing the software problem, retest the linkage to verify the repair effect.

[0111] The application also provides a multi-acceleration card scheduling system for model acceleration training, which applies the multi-acceleration card scheduling method for model acceleration training.

[0112] The cloud data server is responsible for storing model training data, historical computing power usage data, etc., and provides data support for the system.

[0113] The acceleration card scheduling module generates computing power containers according to computing power requirements, schedules acceleration card resources, and forms an acceleration card sequence.

[0114] The acceleration card control module is responsible for communication with the acceleration card, performs linkage testing, and obtains the state information of the acceleration card.

[0115] The data acquisition module is responsible for collecting real-time power and temperature data of the acceleration card, and provides a basis for troubleshooting.

[0116] The communication module is responsible for data communication between the cloud data server and other modules of the system.

[0117] The data processing module is responsible for processing collected data, calculating computing power usage efficiency, determining fault types, and issuing corresponding control instructions.

[0118] Example 1: Multi-acceleration card scheduling in deep learning model training

[0119] An AI research company is developing a complex deep learning model for image recognition tasks. In order to improve training efficiency, they decide to use the multi-acceleration card scheduling method provided by the application to accelerate the model training process using multiple GPU acceleration cards.

[0120] Specific steps:

[0121] Initial computing power allocation and efficiency evaluation:

[0122] Calculate the first demand computing power:

[0123] Through historical data and empirical formula, the data processing speed required for model training is estimated to be 1000 GFLOPS (billion floating point operations per second).

[0124] Set the computing power usage efficiency to 90%.

[0125] According to the formula "first demand computing power = required computing power / set computing power usage efficiency", the first demand computing power is calculated to be 1111 GFLOPS (1000 GFLOPS / 0.9).

[0126] Generate computing power containers and acceleration card sequence:

[0127] The acceleration card scheduling module generates an algorithm container according to the first demand algorithm, and schedules four GPUs with similar performance to meet the demand. The first acceleration card sequence is formed.

[0128] Linkage test and efficiency calculation:

[0129] The acceleration card control module performs linkage test with the target model, and measures the actual data processing speed as 950 GFLOPS.

[0130] According to the formula, the first algorithm usage efficiency is calculated as 85.5% (950 GFLOPS / 1111 GFLOPS corresponding data processing speed, that is, the ratio of actual to theoretical).

[0131] Algorithm adjustment and redistribution:

[0132] Calculate the second demand algorithm: According to the first algorithm usage efficiency of 85.5%, it is deduced that the actual algorithm is close to the theoretical value, but it does not reach the set efficiency.

[0133] Combined with the set 90% algorithm usage efficiency and the model training demand of 1000 GFLOPS, the second demand algorithm corresponding to the data processing speed needs to be adjusted to a more efficient state, and then the new second demand algorithm is obtained.

[0134] Adjust the acceleration card sequence: according to the difference between the second demand algorithm and the first demand algorithm, it is decided to add one more GPU to improve the overall algorithm. The second acceleration card sequence is formed, which includes five GPUs.

[0135] Efficiency evaluation again:

[0136] Linkage test and efficiency calculation: new linkage test is performed, and the actual data processing speed is measured to be increased to 1080 GFLOPS. The second algorithm usage efficiency is calculated as 97.2% (1080 GFLOPS / second demand algorithm corresponding to adjusted data processing speed).

[0137] Efficiency consistency:

[0138] Since the second algorithm usage efficiency of 97.2% is close to the set 90% and the system is stable, it is considered that the algorithm adjustment is effective.

[0139] Complete acceleration card scheduling: according to the second acceleration card sequence, the final scheduling of the acceleration card is completed to ensure that the model training is carried out under the optimal algorithm configuration.

[0140] Results: Through the multi-acceleration card scheduling method of the present application, the training efficiency of the deep learning model is successfully improved, the expected algorithm usage efficiency is achieved, and the training time is shortened.

[0141] Example Two: Multi-GPU Card Scheduling Optimization for Natural Language Processing Model A natural language processing (NLP) company is training a large language model and decides to use the multi-GPU card scheduling system of the present invention to improve training speed and efficiency.

[0142] Specific steps:

[0143] Initial computing power allocation and efficiency evaluation:

[0144] Calculate the first demand computing power: By analyzing the model structure and historical data, it is estimated that the data processing speed required for model training is 1500TFLOPS (trillion floating-point operations per second).

[0145] Set the computing power usage efficiency to 95%.

[0146] According to the formula, the first demand computing power is calculated as 1579TFLOPS (1500TFLOPS / 0.95).

[0147] Generate computing power container and acceleration card sequence: The acceleration card scheduling module generates a computing power container and schedules 8 high-performance GPU acceleration cards to meet the demand. Form the first acceleration card sequence.

[0148] Linkage test and efficiency calculation: The acceleration card control module performs linkage test with the target model, and measures the actual data processing speed as 1400TFLOPS. Calculate the first computing power usage efficiency as 88.6% (1400TFLOPS / 1579TFLOPS corresponding data processing speed).

[0149] Computing power adjustment and redistribution:

[0150] Calculate the second demand computing power: According to the first computing power usage efficiency of 88.6%, it is deduced that the actual computing power has not reached the set efficiency.

[0151] Re-calculate the second demand computing power based on the set 95% computing power usage efficiency and the model training demand of 1500TFLOPS.

[0152] Adjust the acceleration card sequence: Decide to add 2 more GPU acceleration cards to optimize computing power allocation. Form the second acceleration card sequence, a total of 10 GPU acceleration cards.

[0153] Re-evaluate efficiency:

[0154] Linkage test and efficiency calculation: Perform new linkage test, measure the actual data processing speed to 1550TFLOPS. Calculate the second computing power usage efficiency as 98.1% (1550TFLOPS / second demand computing power corresponding to the adjusted data processing speed).

[0155] Efficiency consistency: The second computing power efficiency of 98.1% is close to the set 95%, and the system is stable, so the computing power adjustment is effective.

[0156] Troubleshooting (assuming):

[0157] Fault type judgment: In a certain adjustment, it was found that the computing power efficiency did not significantly improve, so power proportion analysis was performed. The total power proportion and the number proportion were inconsistent, indicating a possible hardware failure.

[0158] Hardware failure handling: Through the data acquisition module, the real-time power of each acceleration card was obtained, and it was found that one of the acceleration cards had abnormally low power. After replacing the faulty acceleration card, the system returned to normal after retesting.

[0159] Complete acceleration card scheduling: Based on the final second acceleration card sequence, the acceleration card scheduling was completed to ensure that the model training was conducted under the optimal computing power configuration.

Claims

1. A multi-accelerator card scheduling method for model acceleration training, characterized in that, The method comprises the following steps: Step one, according to the data processing speed of model training requirements and the set algorithm power utilization efficiency, the first demand algorithm power is obtained, the acceleration card scheduling module generates the algorithm power container according to the first demand algorithm power, the first algorithm power container schedules the corresponding acceleration card according to the first demand algorithm power, and the first acceleration card sequence is generated, the acceleration card control module carries out linkage test according to the first acceleration card sequence and the target model, the data processing speed of the current model training is obtained, and the first algorithm power utilization efficiency of the current model is obtained according to the data processing speed of the current model training and the data processing speed corresponding to the first demand algorithm power; Step two, according to the first algorithm power utilization efficiency of the current model and the data processing speed of model training requirements, the second demand algorithm power is obtained, the acceleration card corresponding to the algorithm power difference value is scheduled according to the algorithm power difference value between the second demand algorithm power and the first demand algorithm power, and the second acceleration card sequence is formed with the first acceleration card sequence; Step three, the acceleration card control module carries out linkage test according to the second acceleration card sequence and the target model, the data processing speed of the current model training is obtained, the second algorithm power utilization efficiency of the current model training is obtained according to the data processing speed of the current model training and the second demand algorithm power, if the first algorithm power utilization efficiency is consistent with the second algorithm power utilization efficiency, step seven is entered, otherwise, step four is entered; Step four, the proportion of the total power of the acceleration card in the first acceleration card sequence and the total power of the acceleration card in the second acceleration card sequence is used to judge the acceleration card fault type, if it is a hardware fault, step five is entered, if it is a software fault, step six is entered; Step five, the power of each acceleration card in the second acceleration card sequence is obtained, the difference value between the power of each acceleration card and the designed power is obtained respectively, the acceleration card with the difference value greater than the set difference value threshold is the fault acceleration card, the acceleration card is replaced according to the fault acceleration card information, and step three is returned; Step six, the target data processing rate of each acceleration card is obtained according to the second acceleration card sequence and the first algorithm power utilization efficiency, the data processing rate difference value between the data processing rate of each acceleration card and the corresponding target data processing rate is obtained respectively, wherein the data processing rate difference value greater than the set data processing rate difference threshold is the fault acceleration card, the software of the fault acceleration is repaired according to the fault acceleration card information, and step three is returned; Step seven, the acceleration card scheduling is completed according to the second acceleration card sequence.

2. The multi-accelerator card scheduling method for model acceleration training of claim 1, wherein, The first demand algorithm power is obtained according to the data processing speed of model training requirements. The first demand algorithm power is obtained according to the corresponding algorithm power and the ratio of the set algorithm power utilization efficiency.

3. The multi-accelerator card scheduling method for model acceleration training of claim 2, wherein, The first algorithm power utilization efficiency of the current model is obtained according to the ratio of the data processing speed of the current model training and the data processing speed corresponding to the first demand algorithm power, and the following formula is used: The second demand algorithm power is obtained according to the first demand algorithm power utilization efficiency of the current model and the set algorithm power utilization efficiency.

4. The multi-accelerator card scheduling method for model acceleration training of claim 3, wherein, ​ According to the data processing speed corresponding to the second demand computing power, the second demand computing power is obtained.

5. The multi-accelerator card scheduling method for model acceleration training of claim 4, wherein, The second computing power usage efficiency of the current model training is obtained according to the data processing speed of the current model training and the second demand computing power.

6. The multi-accelerator card scheduling method for model acceleration training of claim 5, wherein, The proportion of the total power of the acceleration cards in the first acceleration card sequence to the total power of the acceleration cards in the second acceleration card sequence is determined. If the proportion of the total power of the acceleration cards in the first acceleration card sequence to the total power of the acceleration cards in the second acceleration card sequence is consistent with the proportion of the number of acceleration cards in the first acceleration card sequence to the number of acceleration cards in the second acceleration card sequence, it is a software failure, otherwise, it is a hardware failure.

7. The multi-accelerator card scheduling method for model acceleration training of claim 6, wherein, The target data processing rate of each acceleration card is obtained according to the second acceleration card sequence and the first computing power usage efficiency. The target data processing rate of each acceleration card is obtained according to the data processing speed corresponding to the computing power of each acceleration card in the second acceleration card sequence and the first computing power usage efficiency.

8. A multi-accelerator card scheduling system for model acceleration training, characterized in that, The multi-acceleration card scheduling method for model acceleration training according to any one of claims 1-7 comprises a cloud data server, an acceleration card scheduling module, an acceleration card control module, a data acquisition module, a communication module and a data processing module; the acceleration card scheduling module, the acceleration card control module, the data acquisition module and the communication module are connected with the data processing module; the cloud data server is in communication connection with the communication module.

Citation Information

Patent Citations

  • Computing power scheduling method and device based on Kubernetes

    CN112241321A

  • Accelerator card load balancing scheduling method and device, communication equipment and storage medium

    CN117234734A