Model Training Method, Device, Storage Medium and Computer Equipment
By adaptively adjusting the frequency regulation delay of the neural network processor, adaptive frequency reduction is carried out at different stages of AI large model training in intelligent computing clusters, solving the problem of inefficient energy caused by the setting of a single NPU frequency and achieving energy efficiency improvement.
Patent Information
- Application Number
- CN202510572168.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-06
AI Technical Summary
In AI large-scale AI clusters, due to dynamic changes in computational loads, the frequency and voltage settings of single NPUs in the prior art lead to inefficient energy efficiency.
By obtaining the frequency regulation delay of the neural network processor, comparing with the threshold, adaptively adjusting the processing method of the NPU, and down frequency at different training stages, including sequential processing, interval processing and phased down frequency.
It improves the energy efficiency of model training, reduces energy consumption, and improves the utilization rate of computing resources.
Smart Images

Figure CN120086025B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model training method, device, storage medium, and computer device. Background Art
[0002] Today, with the rapid development of artificial intelligence (AI) technology, intelligent computing clusters (referred to as intelligent computing clusters for short) have become the key pillars for promoting AI scientific research and industrial applications. Intelligent computing clusters are built by adopting advanced AI processors, including graphics processing units (GPUs), neural network processing units (NPUs), etc., and can efficiently process complex AI tasks and the computing requirements of large-scale data. However, the popularization and wide application of intelligent computing clusters have also brought severe energy consumption challenges, which pose a huge pressure on global energy use and environmental protection.
[0003] When training large AI models on intelligent computing clusters, due to the huge model scale, the computing resources of a single computing node are difficult to meet the requirements. Therefore, parallel training methods need to be adopted to accelerate the training process and make full use of the computing power resources of intelligent computing clusters. Parallel training mainly includes several common methods such as data parallelism, pipeline parallelism, and tensor parallelism. These methods usually set a single NPU frequency and voltage for a computing task according to the severity of the computing task. Although this method reduces the implementation complexity, it ignores the dynamic changes of the computing load, resulting in low energy efficiency. Therefore, related technologies urgently need to propose a model training method to solve the above technical problems. Summary of the Invention
[0004] The main purpose of this application is to provide a model training method, device, storage medium, and computer device, which can adaptively reduce the frequency at different stages of model training according to the regulation delay of the frequency regulation of the neural network processor, and improve energy efficiency.
[0005] In a first aspect, an embodiment of this application provides a model training method, including:
[0006] Obtain the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple said neural network processors sequentially process multiple batches of tasks;
[0007] Compare the regulation delay with a regulation delay threshold to obtain a comparison result;
[0008] When the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each said neural network processor to sequentially process multiple batches of tasks, and reduce the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage;
[0009] When the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, during the forward propagation stage of the model training process, control each of the neural network processors to sequentially process multiple batches of tasks, during the backward propagation stage, control each of the neural network processors to process multiple batches of tasks at intervals, and during the forward propagation stage and the backward propagation stage, downscale the frequency when at least one neural network processor processes at least one batch of tasks.
[0010] In a second aspect, an embodiment of the present application provides a model training device, including:
[0011] A first acquisition unit, configured to acquire the regulation delay of frequency regulation of any neural network processor required for model training, and multiple batches of tasks are sequentially processed among the multiple neural network processors;
[0012] A comparison unit, configured to compare the regulation delay with the regulation delay threshold to obtain a comparison result;
[0013] A first downscaling unit, configured to, when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each of the neural network processors to sequentially process multiple batches of tasks, and downscale the frequency when at least one neural network processor processes at least one batch of tasks during the forward propagation stage;
[0014] A second downscaling unit, configured to, when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, during the forward propagation stage of the model training process, control each of the neural network processors to sequentially process multiple batches of tasks, during the backward propagation stage, control each of the neural network processors to process multiple batches of tasks at intervals, and during the forward propagation stage and the backward propagation stage, downscale the frequency when at least one neural network processor processes at least one batch of tasks.
[0015] In a third aspect, an embodiment of the present application provides a storage medium. The computer-readable storage medium stores multiple instructions, and these instructions are suitable for being loaded by a processor to execute the model training method as described in any one of the above.
[0016] In a fourth aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the model training method as described in any one of the above is implemented.
[0017] In the embodiments of the present application, by obtaining the regulation delay of the frequency regulation of any neural network processor required for model training, multiple neural network processors sequentially process multiple batches of tasks; comparing the regulation delay with a regulation delay threshold to obtain a comparison result; when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, controlling each neural network processor to sequentially process multiple batches of tasks, and reducing the frequency of at least one neural network processor when processing at least one batch of tasks in the forward propagation stage; when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, controlling each neural network processor to sequentially process multiple batches of tasks in the forward propagation stage of the model training process, controlling each neural network processor to process multiple batches of tasks at intervals in the backward propagation stage, and reducing the frequency of at least one neural network processor when processing at least one batch of tasks in the forward propagation stage and the backward propagation stage. Compared with the related art where a single NPU frequency and voltage are set for one computing task, the embodiments of the present application adaptively reduce the frequency in different stages of model training according to the regulation delay of the frequency regulation of the neural network processor, thereby improving energy efficiency.
[0018] Other features and advantages of the present disclosure will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for describing the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 It is a schematic diagram of the scenario of the model training system provided by the embodiments of the present application.
[0021] Figure 2 It is a schematic flowchart of the model training method provided by the embodiments of the present application.
[0022] Figure 3 It is a schematic diagram of the pipeline parallel training of the AI large model provided by the embodiments of the present application.
[0023] Figure 4 It is a schematic diagram of the bubble merging scheduling at high regulation delay and the training before frequency regulation provided by the embodiments of the present application.
[0024] Figure 5The bubble merging scheduling under high regulation delay provided by the embodiments of the present application, and the training schematic diagram after frequency regulation.
[0025] Figure 6 The bubble merging scheduling under low regulation delay provided by the embodiments of the present application, and the training schematic diagram before frequency regulation.
[0026] Figure 7 The bubble merging scheduling under low regulation delay provided by the embodiments of the present application, and the training schematic diagram after frequency regulation.
[0027] Figure 8 The structural schematic diagram of the model training device provided by the embodiments of the present application.
[0028] Figure 9 The structural schematic diagram of the computer device provided by the embodiments of the present application. Detailed implementation manners
[0029] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope of protection of the present application.
[0030] It should be noted that in some processes described in the specification, claims and the above-mentioned drawings, there are multiple steps that appear in a specific order. However, it should be clearly understood that these steps may not be executed in the order in which they appear in this document or may be executed in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, descriptions such as "first", "second" or "target" in this document are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0031] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are described. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:
[0032] Neural - Processing Unit (NPU): A processor designed specifically for neural network operations, especially for a large number of matrix and vector operations in deep learning. With the rapid development of deep learning technology in the field of artificial intelligence, traditional CPUs and GPUs gradually expose problems of low efficiency and high energy consumption when processing deep learning tasks. The NPU emerges as the times require and has the following characteristics:
[0033] Highly parallel computing: It integrates a large number of computing units internally, can process numerous data simultaneously, and significantly accelerates the training and inference speeds of deep learning models. For example, in image recognition tasks, it can quickly complete the extraction and analysis of image features.
[0034] Specialized architecture: Designed based on the characteristics of neural network algorithms, it has special optimizations for common operations such as convolution and activation, and can efficiently execute various operations of neural networks.
[0035] Low power consumption: Compared with general-purpose processors, it can operate with lower power consumption when performing deep learning tasks, and is suitable for mobile devices, smart home devices, etc. with strict power consumption requirements.
[0036] Today, NPUs are widely used in fields such as intelligent security, autonomous driving, and intelligent voice assistants, promoting the popularization and development of artificial intelligence applications.
[0037] Forward propagation: It refers to inputting the input data into the input layer of the neural network. The data passes through the calculations of each layer in turn according to the network hierarchical structure. For each layer, first calculate the weighted sum of the input data and add the bias, and then apply the activation function to obtain the output of that layer. Finally, obtain the predicted value of the model from the output layer. Then use a loss function (such as mean squared error, cross entropy, etc.) to calculate the gap between the predicted value and the actual label, that is, the loss value, and also calculate the gradient of the loss with respect to the output. Simply put, it is the process from inputting data to generating a prediction result and calculating the loss.
[0038] Backward propagation: It is the core algorithm for training deep neural networks, aiming to optimize model parameters by calculating and propagating gradients. Its core is the chain rule. Using the chain rule, the gradient of the loss function with respect to the model output is propagated backward layer by layer to each parameter in the network, and the gradients of the loss function with respect to the weights and biases of each layer are calculated. Then, based on the calculated gradients, use the gradient descent algorithm or other optimization algorithms to update the weights and biases of each layer, thereby reducing the value of the loss function and making the prediction result of the model closer to the actual label.
[0039] Forward propagation provides the intermediate results required for calculating gradients for backward propagation, and backward propagation optimizes the model by adjusting parameters according to the loss obtained from forward propagation. The two complement each other and jointly complete the model training process.
[0040] When training large AI models on an intelligent computing cluster, due to the huge scale of the model, the computing resources of a single computing node are difficult to meet the requirements. Therefore, parallel training methods need to be adopted to accelerate the training process and make full use of the computing power resources of the intelligent computing cluster. Parallel training mainly includes several common methods such as data parallelism, pipeline parallelism, and tensor parallelism. Each method has its applicable scenarios and advantages, which will be introduced one by one below.
[0041] (1)Data Parallelism
[0042] Data parallelism is one of the most commonly used parallel training methods. Its core idea is to keep the model parameters consistent across all computing devices, and then divide the training data into multiple batches, which are separately allocated to different computing nodes for independent calculation. Each node performs the same forward and backward propagations but uses different subsets of data. After each batch of training is completed, the nodes will merge the gradients they calculated respectively through a gradient synchronization mechanism and update the global parameters of the model. Its advantages are that data parallelism can be extended to multiple nodes without changing the model structure, is suitable for training large-scale datasets, and is easy to implement. Frameworks such as PyTorch and TensorFlow provide support for data parallelism.
[0043] (2)Tensor Parallelism
[0044] Tensor parallelism is a fine-grained implementation method of model parallelism. In large models, many computational operations (such as matrix multiplication) involve operations on large tensors. Tensor parallelism splits these large tensors into smaller parts and distributes them to different computing nodes for parallel calculation. For example, a large matrix can be split by rows or columns among multiple NPUs, and each NPU is responsible for processing a part of the matrix calculation and finally merging the results. This method is suitable for extremely large models, especially those cases where a single computing node cannot accommodate the complete model parameters, and can make full use of hardware resources by finely decomposing the computing tasks to multiple devices for execution.
[0045] (3)Pipeline Parallelism
[0046] Pipeline parallelism is another model parallelism method. Its core idea is to split the model hierarchically and allocate different layers to different NPUs. During training, the input data passes through each part of the model in a pipeline manner. The first batch of data is immediately passed to the next NPU after passing through the first NPU, and this NPU can start processing the second batch of data. This is applicable to cases where the model has an obvious hierarchical structure and can be split by layer, such as deep neural networks, and reduces the memory footprint because each node only needs to store a part of the model parameters.
[0047] (4)Hybrid Parallelism
[0048] The hybrid parallel method combines data parallelism, tensor parallelism, and pipeline parallelism to perform parallel computing at different levels. For very large-scale models, a single parallel method may not be able to fully utilize the efficiency of computing resources, so a hybrid parallel method is usually adopted. For example, tensor parallelism can be first adopted within a computing node to split the tensors of each layer among multiple devices for processing; at the same time, different layers of the entire model can be distributed to multiple nodes through pipeline parallelism; finally, data parallelism is used to split large datasets among different computing nodes for processing.
[0049] The above method usually sets a single NPU frequency and voltage for a computing task according to the severity of the computing task. Although this method reduces the implementation complexity, it ignores the dynamic changes in the computing load, resulting in low energy efficiency.
[0050] To solve the above problems, embodiments of this application obtain the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple said neural network processors sequentially process multiple batches of tasks; compare the regulation delay with a regulation delay threshold to obtain a comparison result; when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each said neural network processor to sequentially process multiple batches of tasks, and reduce the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage; when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, control each said neural network processor to sequentially process multiple batches of tasks during the forward propagation stage of the model training process, control each said neural network processor to process multiple batches of tasks at intervals during the backward propagation stage, and reduce the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage and the backward propagation stage. Compared with the related art of setting a single NPU frequency and voltage for a computing task, embodiments of this application adaptively reduce the frequency according to the regulation delay of the frequency regulation of the neural network processor at different stages of model training, improving energy efficiency. Please continue to refer to the following specific embodiments for details.
[0051] Please refer to Figure 1 , Figure 1 which is a scenario schematic diagram of the model training system provided by embodiments of this application. It includes a terminal 140, the Internet 130, a gateway 120, a server 110, etc.
[0052] The terminal 140 includes, but is not limited to, pre-configured laptop computers, or tablet computers, desktop computers, and other electronic devices with the ability to submit data. In addition, it can be a single device or a collection of multiple devices. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.
[0053] The terminal 140 refers to a computer system that can report data to the server 110. Compared with ordinary terminals, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, etc.
[0054] The gateway 120 is also called an internetwork connector and a protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. Between two systems using different communication protocols, data formats, or languages, and even with completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. The messages sent from the terminal 140 to the server 110 need to be sent to the corresponding server 110 through the gateway 120. The messages sent from the server 110 to the terminal 140 also need to be sent to the corresponding terminal 140 through the gateway 120.
[0055] The model training method of the embodiments of the present disclosure can be implemented on the server 110.
[0056] It should be noted that Figure 1 The scenario schematic diagram of the model training system shown is only an example. The model training system and scenario described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of image processing technology and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0057] In this embodiment, the description will be made from the perspective of the model training device, which can be specifically integrated in a computer device with a storage unit and installed with a microprocessor and having computing capabilities.
[0058] Please refer to Figure 2 , Figure 2 , which is a schematic flowchart of the model training method provided by the embodiments of the present application. The model training method includes:
[0059] In step 201, obtain the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple neural network processors sequentially process multiple batches of tasks.
[0060] Among them, in the training process of the embodiments of this application, multiple neural network processors (NPUs) with the same regulation delay are used for training. Therefore, only by obtaining the regulation delay of frequency regulation of any one neural network processor required for model training can we know the regulation delay of each NPU used for model training. In view of the differences in different NPU hardware conditions, in order to obtain the accurate frequency-voltage regulation delay value required for the energy efficiency optimization of the AI large model pipeline parallel training, the following specific measurement methods for frequency-voltage regulation delay are proposed:
[0061] Turn on the NPU frequency reading module. First, ensure that the frequency reading module of the NPU is in the on state to monitor the NPU frequency change in real time.
[0062] Maintain the low-frequency state. Use a script to set the frequency of the NPU to a low frequency and maintain it at the low frequency for 5 seconds. This step ensures that the NPU is in a stable low-frequency state, providing a benchmark for subsequent operations.
[0063] Frequency switching operation. Immediately adjust the frequency of the NPU from the low frequency to the high frequency through a script: Record the time when the frequency adjustment instruction is issued by the script, denoted as t_1. At this time, the frequency-voltage regulation mechanism inside the NPU starts to operate to adjust to the set high frequency.
[0064] Monitor the high-frequency arrival time. Through the NPU frequency reading module, monitor the time point when the NPU reaches the set high frequency, denoted as t_2.
[0065] Calculate the frequency-voltage regulation delay. Calculate the NPU frequency-voltage regulation delay according to the following formula: (Frequency regulation delay) Delay = t_2 - t_1.
[0066] Delay range and analysis. Through multiple experiments, verify and record the regulation delay data in different scenarios. Generally, the NPU frequency-voltage regulation delay is between 10 milliseconds and 1000 milliseconds. The specific value of the delay depends on the model, design, and operating environment of the NPU hardware.
[0067] Through the above method, the accurate NPU frequency-voltage regulation delay can be obtained, providing key data support for the subsequent energy efficiency optimization of AI large model training. Based on obtaining the regulation delay of each NPU before model training, a group of NPUs with the same or regulation delays differing by no more than a preset difference can be used as the NPUs required for training the model.
[0068] Specifically, please refer to Figure 3 , Figure 3 which is the schematic diagram of the AI large model pipeline parallel training provided by the embodiments of this application. As shown in Figure 3As shown, the figure shows the process of training a micro-batch in parallel through a pipeline. The micro-batch needs to run on NPU1 - NPU4 in sequence. The forward propagation is represented in light color, and the backward propagation is represented in dark color. The tasks of a micro-batch are carried out in the order of NPU1 - NPU2 - NPU3 - NPU4 in the forward propagation stage (for example Figure 3 for the forward propagation of micro-batch 1 in Figure 3 is in the order of NPU1 - NPU2 - NPU3 - NPU4), and the tasks of a micro-batch are carried out in the order of NPU4 - NPU3 - NPU2 - NPU1 in the backward propagation stage (for example Figure 3 for the backward propagation 1B of micro-batch 1 in Figure 3 is in the order of NPU4 - NPU3 - NPU2 - NPU1). Since the amount of calculation in the backward propagation is about 2 times that in the forward propagation, the calculation duration is also about 2 times the relationship. Due to the logical dependencies between these 8 calculations, the calculation arrangement cannot be further compressed and must be carried out in sequence.
[0069] In step 202, the regulation delay is compared with the regulation delay threshold to obtain a comparison result.
[0070] Among them, in the embodiments of the present application, different model training schemes are set for the high or low regulation delay of the NPU used in training. Therefore, a regulation delay threshold needs to be set. By comparing the regulation delay with this regulation delay threshold, it is determined whether the regulation delay of the NPU used in this training belongs to high regulation delay or low regulation delay.
[0071] Specifically, if it belongs to low regulation delay, the comparison result is that the regulation delay is less than or equal to the regulation delay threshold; if it belongs to high regulation delay, the comparison result is that the regulation delay is greater than the regulation delay threshold.
[0072] In step 203, when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each of the neural network processors to process multiple batches of tasks in sequence, and reduce the frequency of at least one neural network processor processing at least one batch of tasks in the forward propagation stage.
[0073] Among them, please refer to Figure 4 , Figure 4 which is the training schematic diagram before frequency regulation of the bubble merging scheduling under high regulation delay provided by the embodiments of the present application. In order to further improve the training efficiency of large models, the embodiments of the present application generally use not only 1 micro-batch, but a group composed of multiple micro-batches (for example Figure 4 the micro-batch tasks 1, 2, 3, and 4 in Figure 4 )) to improve the utilization rate of the NPU. There is no logical dependency between different micro-batches and they can be executed concurrently. The specific arrangement method of multiple micro-batches needs to be determined according to the frequency voltage regulation delay. For the high regulation delay scenario, the embodiments of the present application propose to arrange multiple micro-batches inFigure 4 The bubble merging scheduling method in Figure 4 (for example, the uncalculated part in the central white area) are merged together, reducing the number of times of execution frequency regulation, thus avoiding the impact brought by high regulation delay.
[0074] Specifically, as Figure 5 shown, Figure 5 This is a training schematic diagram of bubble merging scheduling under high regulation delay provided by an embodiment of the present application after frequency regulation. In the scenario of high regulation delay, the calculation part in the backpropagation stage is not downscaled because the calculation time in the backpropagation stage is long, and downscaling will significantly increase the total calculation duration and affect the overall calculation duration. Therefore, downscaling is only performed for at least one neural network processor to process at least one batch of tasks. For example Figure 5 shown, the forward propagation calculation of batch task 4 on NPU2, the forward propagation calculations of batch tasks 3 and 4 on NPU3, and the forward propagation calculations of batch tasks 2, 3, and 4 on NPU4 can be downscaled, and their operation speed and power consumption are appropriately reduced by reducing the frequency. Through this strategy, energy consumption can be effectively reduced and energy efficiency can be improved.
[0075] In some embodiments, downscaling at least one neural network processor to process at least one batch of tasks in the forward propagation stage includes:
[0076] (1) Determining the first batch quantity of batch tasks that need to be downscaled in the forward propagation stage;
[0077] (2) Obtaining multiple adjusted frequencies and corresponding processor voltages, and obtaining the processing rate ratio corresponding to each of the adjusted frequencies;
[0078] (3) Based on the first batch quantity, each of the adjusted frequencies and the corresponding processor voltages, and the processing rate ratio corresponding to each of the adjusted frequencies, determining the first post-downscaling power consumption value corresponding to each of the adjusted frequencies in the forward propagation stage;
[0079] (4) Screening out the first adjusted frequency corresponding to the smallest first post-downscaling power consumption value from the multiple adjusted frequencies;
[0080] (5) Before each neural network processor processes the batch tasks that need to be downscaled, adjusting the frequency to the first adjusted frequency.
[0081] Among them, although it is known that in the scenario of high regulation delay, downscaling is performed on the batch tasks in the forward propagation stage, but specifically to what frequency can minimize the power consumption value in the training process needs to be determined through specific calculations. The determination method is:
[0082] Determine the first batch quantity of the batch tasks that need to be downclocked during the forward propagation stage of model training. Here, the first batch quantity is the sum of the task quantities of the batch tasks that need to be downclocked in the batch tasks corresponding to each NPU. For example, for the forward propagation calculation of batch task 4 on NPU2, the forward propagation calculations of batch tasks 3 and 4 on NPU3, and the forward propagation calculations of batch tasks 2, 3, and 4 on NPU4 are downclocked. Then the first batch quantity is the sum of the task quantity 1 corresponding to NPU2, the task quantity 2 corresponding to NPU3, and the task quantity 3 corresponding to NPU4, that is, 1 + 2 + 3 = 6.
[0083] Specifically, for an NPU chip, it is necessary to determine what the corresponding voltage, processing rate, and power are under different frequency settings. For this purpose, a frequency-voltage-processing rate-power relationship mapping table needs to be measured.
[0084] To perform the measurement, first find the forward propagation operators in the training load of the AI large model. For each adjustable frequency f_0, f_1, f_2, f_3,... of the NPU chip, measure its corresponding processor voltage V_0, V_1, V_2, V_3,... and the corresponding operator processing rate s_0, s_1, s_2, s_3,... and the power P_0, P_1, P_2, P_3,... of the NPU. Assume that f_0 is the highest frequency of the NPU chip. Then normalize the operator processing rate to obtain the processing rate ratios corresponding to f_0, f_1, f_2, f_3,... as 1, s_1 / s_0, s_2 / s_0, s_3 / s_0,....
[0085] Using this rate table, the frequency-voltage setting for a specific processing rate can be looked up. In the following text, assume that for a processing rate ratio x, the required frequency setting is f(x) and the voltage is V(x). For example, for a certain NPU chip, f_3 = 1500 MHz, V_3 = 810 mV, s_3 / s_0 = 0.75, and P_3 = 280 W. Then when it is necessary to reduce the NPU processing rate to 0.75, according to this table, it can be found that the frequency needs to be set to f(0.75) = 1500 MHz and the voltage to V(0.75) = 810 mV.
[0086] Based on this, different adjustment frequencies f(x) and corresponding processor voltages V(x) can be obtained, and the processing rate ratio x = s(x) / s_0 corresponding to each adjustment frequency f. Based on the above information, the first power consumption value E1 after frequency reduction corresponding to each adjustment frequency in the forward propagation stage can be calculated. The first adjustment frequency P1 with the minimum first power consumption value E1min after frequency reduction is selected from multiple adjustment frequencies. It shows that after reducing the frequency according to the first adjustment frequency P1, the processing power consumption value is the lowest. Therefore, before each neural network processor processes a batch of tasks that require frequency reduction, the frequency is adjusted to the first adjustment frequency P1.
[0087] For example, there are adjustment frequencies P1, P2, P3, and P4. The first power consumption value after frequency reduction corresponding to P1 is 150W, the first power consumption value after frequency reduction corresponding to P2 is 120W, the first power consumption value after frequency reduction corresponding to P3 is 130W, and the first power consumption value after frequency reduction corresponding to P4 is 140W. Then, the adjustment frequency P2 corresponding to the minimum first power consumption value after frequency reduction of 120W is determined as the first adjustment frequency. Before each neural network processor processes a batch of tasks that require frequency reduction, the frequency is adjusted to P2.
[0088] In some embodiments, determining the first power consumption value after frequency reduction corresponding to each adjustment frequency in the forward propagation stage based on the first batch quantity, each adjustment frequency and the corresponding processor voltage, and the processing rate ratio corresponding to each adjustment frequency includes:
[0089] (1.1) Obtain the first original processing duration for processing a batch of tasks in the forward propagation stage and the capacitance load constant;
[0090] (1.2) Calculate the ratio of the first original processing duration to the processing rate ratio corresponding to each adjustment frequency to obtain the first processing duration after frequency reduction corresponding to each adjustment frequency;
[0091] (1.3) Based on the first batch quantity, the capacitance load constant, the first processing duration after frequency reduction, each adjustment frequency, the first processing duration after frequency reduction corresponding to each adjustment frequency, and the processor voltage corresponding to each adjustment frequency, determine the first power consumption value after frequency reduction corresponding to each adjustment frequency in the forward propagation stage.
[0092] Among them, the specific method for determining the first power consumption value after frequency reduction corresponding to each adjustment frequency in the forward propagation stage is: obtain the first original processing duration T for processing a batch of tasks in the forward propagation stage and the capacitance load constant C, calculate the ratio of the first original processing duration T to the processing rate corresponding to each adjustment frequency x to obtain the first processing duration after frequency reduction T / x corresponding to each adjustment frequency. According to the power consumption value calculation formula It can be known that the power consumption value after frequency reduction corresponding to each adjustment frequency f(x) for processing each batch of tasks is , where t is the processing duration, and in combination with the number of the first batch, the first power consumption value after frequency reduction corresponding to each adjustment frequency f(x) can be calculated, that is, the number of the first batch P(x).
[0093] Taking Figure 5 as an example, Figure 5 in which the number of the first batch is 6, then the first power consumption value after frequency reduction corresponding to each adjustment frequency f(x) is 6 .
[0094] In some embodiments, the first total power consumption value after frequency reduction corresponding to each adjustment frequency can also be calculated to determine the first adjustment frequency. For example in, the calculation duration of a single forward propagation is T, and the reverse calculation duration is generally 2 times the calculation duration of the forward propagation, so the reverse calculation duration is 2T. The processing rate before frequency reduction is 1, the corresponding frequency is f(1), the regulation frequencies of 3 NPUs regulated by the voltage V(1) are f(x), the voltage is V(x), and the calculation duration of each block becomes T / x. Then the total first power consumption value after frequency reduction Figure 5 is , 6 is the number of the first batch, the number of tasks in the remaining non-frequency-reduced forward propagation stage is 10, and the number of tasks in the remaining non-frequency-reduced backward propagation stage is 16. Among them, the calculation duration corresponding to 16 is 2T, and the calculation duration corresponding to 10 is T. Then 42 in the formula is obtained by 16 2 + 10. By calculating the first total power consumption value after frequency reduction corresponding to each adjustment frequency to determine the first adjustment frequency with the smallest first total power consumption value after frequency reduction .
[0095] In some embodiments, obtaining the ratio of the processing rate corresponding to each adjustment frequency includes:
[0096] (1.1) Obtaining the processing rate corresponding to each adjustment frequency of any one of the neural network processors and the processing rate corresponding to the highest frequency;
[0097] (1.2) Determining the ratio of the processing rate corresponding to different adjustment frequencies to the processing rate corresponding to the highest frequency to obtain the ratio of the processing rate corresponding to each adjustment frequency.
[0098] Among them, the determination method of the processing rate ratio corresponding to each adjustment frequency is to obtain the processing rate s(x) corresponding to each adjustment frequency of any one of the neural network processors, and the processing rate s(0) corresponding to the highest frequency, and determine the ratio of the processing rate corresponding to different adjustment frequencies to the processing rate corresponding to the highest frequency, so as to obtain s(x) / s(0) corresponding to each adjustment frequency.
[0099] In step 204, when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, during the forward propagation stage of the model training process, control each neural network processor to process multiple batches of tasks in sequence, and during the backward propagation stage, control each neural network processor to process multiple batches of tasks at intervals, and reduce the frequency when at least one neural network processor processes at least one batch of tasks during the forward propagation stage and the backward propagation stage.
[0100] Among them, as Figure 6 shown, Figure 6 This is a training schematic diagram of bubble merging scheduling under low regulation delay provided by the embodiments of the present application before frequency regulation. If the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, it means that the current is in a low regulation delay scenario. During the forward propagation stage of the model training process, control each neural network processor to process multiple batches of tasks in sequence. For example, control NPU1 to process batch tasks 1, 2, 3, and 4 in sequence; during the backward propagation stage, control each neural network processor to process multiple batches of tasks at intervals. For example, control NPU3 to process batch task 2 at an interval after processing batch task 1. This interval is because NPU4 immediately processes the backward propagation calculation of batch task 1 after processing batch task 1, and waits until the backward propagation calculation of batch task 1 is processed before processing batch task 2, resulting in an interval between the backward propagation calculations of batch task 1 and batch task 2 of NPU3.
[0101] The bubble reduction scheduling in the low regulation delay scenario is different from the Figure 4 bubble merging scheduling in terms of the scheduling method. Its purpose is to reduce the size of each bubble, so that more computing units can perform frequency regulation. Figure 4 All backward propagations in the bubble merging scheduling in Figure 6 , including 1B, 2B, 3B, and 4B calculations, are difficult to perform frequency regulation, otherwise it will significantly extend the total calculation duration; on the contrary,
[0102] Please refer to Figure 7 , Figure 7The bubble merging scheduling under low regulation delay provided by the embodiments of this application, and the training schematic diagram after frequency regulation. During the forward propagation stage and the reverse propagation stage, when at least one neural network processor processes at least one batch of tasks, its NPU frequency and voltage are lowered, so that the calculation duration is appropriately extended, and the remaining calculation durations remain unchanged. In this way, the power consumption of the underlined calculation can be reduced, but the overall calculation progress is not affected at the same time.
[0103] For example, Figure 7 For the 3B calculations on NPU1, NPU2, and NPU3, by setting the frequency to f(0.86) and the voltage to V(0.86), their operation speed ratio is reduced to 0.86 of the original; the remaining calculations are also set with frequencies in a similar way. Through this strategy, the energy consumption can be effectively reduced without affecting the overall training time, and the energy efficiency can be improved.
[0104] In some embodiments, the frequency reduction during the forward propagation stage and the reverse propagation stage when at least one neural network processor processes at least one batch of tasks includes:
[0105] (1) Determine the second batch quantity of the batch tasks that need to be frequency-reduced during the forward propagation stage, and determine the third batch quantity of the batch tasks that need to be frequency-reduced during the reverse propagation stage;
[0106] (2) Obtain the total batch quantity of multiple said batch tasks, the second original processing duration for processing one batch of tasks during the forward propagation stage, and the third original processing duration for processing one batch of tasks during the reverse propagation stage;
[0107] (3) Screen out the target batch quantity with the largest batch quantity from the batch quantities of the batch tasks that need to be frequency-reduced by each said neural network processor during the reverse propagation stage;
[0108] (4) Based on the total batch quantity, the target batch quantity, the second original processing duration, the third original processing duration, and the processing rate ratio corresponding to each adjusted frequency, screen out the candidate adjusted frequencies from multiple said adjusted frequencies;
[0109] (5) Based on the second batch quantity, the third batch quantity, each said candidate adjusted frequency and the corresponding processor voltage, and the processing rate ratio corresponding to each said candidate adjusted frequency, determine the second power consumption value after frequency reduction corresponding to each said candidate adjusted frequency;
[0110] (6) Screen out the second adjusted frequency with the smallest second total power consumption value after frequency reduction from multiple said candidate adjusted frequencies;
[0111] (7) Before each of the neural network processors processes a batch task that needs to be downclocked, adjust the frequency to the second adjusted frequency.
[0112] Among them, to avoid the total processing duration of the entire processing process after downclocking exceeding the total processing duration before downclocking after adjusting the frequency of the processor according to certain adjusted frequencies, it is necessary to screen out candidate adjusted frequencies that do not affect the total processing duration from the adjusted frequencies, and screen out the second adjusted frequency with the smallest power consumption value according to the second power consumption value after downclocking corresponding to each candidate adjusted frequency, so as to adjust the frequency to the second adjusted frequency before each neural network processor processes a batch task that needs to be downclocked. Thus, the lowest power consumption can be achieved without affecting the total processing duration.
[0113] Specifically, the screening process of candidate adjusted frequencies needs to determine the second batch quantity of the batch tasks that need to be downclocked in the forward propagation stage, and the third batch quantity of the batch tasks that need to be downclocked in the backward propagation stage; obtain the total batch quantity of multiple batch tasks, the second original processing duration for processing one batch task in the forward propagation stage, and the third original processing duration for processing one batch task in the backward propagation stage; screen out the target batch quantity with the largest batch quantity from the batch quantities of the batch tasks that need to be downclocked in the backward propagation stage of each neural network processor; based on the total batch quantity, the target batch quantity, the second original processing duration, the third original processing duration, and the processing rate ratio corresponding to each adjusted frequency, screen out candidate adjusted frequencies from multiple adjusted frequencies.
[0114] For example Figure 7 among them, the multiple batch tasks are batch task 1, batch task 2, batch task 3, and batch task 4 respectively, so the total batch quantity is 4. In the forward propagation stage, it is determined that the batch task that needs to be downclocked is the batch task in NPU3 4 , so the second batch quantity is 1; in the backward propagation stage, the batch tasks that need to be downclocked are respectively in NPU1, NPU2, and NPU3 1B , 2B , 3B , so the third batch quantity is 9; the second original processing duration for processing one batch task in the forward propagation stage is T, and the third original processing duration for processing one batch task in the backward propagation stage is 2T. The batch quantities that need to be downclocked in NPU1, NPU2, and NPU3 are the largest, all being 3, so the target batch quantity is 3; finally, according to the total batch quantity 4, the target batch quantity 3, the second original processing duration T, the third original processing duration 2T, and the processing rate ratio x corresponding to each adjusted frequency, candidate adjusted frequencies are screened out from multiple adjusted frequencies.
[0115] The calculation method of the second power consumption value after downscaling corresponding to each candidate adjustment frequency is to determine the second power consumption value after downscaling corresponding to each candidate adjustment frequency based on the second batch quantity, the third batch quantity, each candidate adjustment frequency and the corresponding processor voltage, and the processing rate ratio corresponding to each candidate adjustment frequency.
[0116] In some embodiments, screening candidate adjustment frequencies from multiple adjustment frequencies based on the total batch quantity, the target batch quantity, the second original processing duration, the third original processing duration, and the processing rate ratio corresponding to each adjustment frequency includes:
[0117] (1.1) Obtain the product of the total batch quantity and the third original processing duration to get a first calculation result;
[0118] (1.2) Calculate the sum of the first calculation result and the second original processing duration to get a reference processing duration;
[0119] (1.3) Calculate the ratio of the third original processing duration to the processing rate ratio corresponding to each adjustment frequency to get the second processing duration after downscaling corresponding to each adjustment frequency;
[0120] (1.4) Determine the product of the second processing duration after downscaling corresponding to each adjustment frequency and the target batch quantity to get a second calculation result corresponding to each adjustment frequency;
[0121] (1.5) Obtain the difference between the total batch quantity and the target batch quantity to get the number of non-downscaled batches;
[0122] (1.6) Obtain the product of the number of non-downscaled batches and the third original processing duration to get a third calculation result;
[0123] (1.7) Determine the sum of the second calculation result corresponding to each adjustment frequency and the third calculation result to get the total processing duration corresponding to each adjustment frequency in the backward propagation stage;
[0124] (1.8) Determine the adjustment frequencies with the total processing duration not exceeding the reference processing duration as candidate adjustment frequencies.
[0125] Among them, the specific method for screening candidate adjustment frequencies from multiple adjustment frequencies according to the total batch quantity, the target batch quantity, the second original processing duration, the third original processing duration, and the processing rate ratio corresponding to each adjustment frequency is as follows: First, determine a reference processing duration, which is mainly the reference duration in the backpropagation stage. Therefore, obtain the processing duration of the backpropagation stage before frequency reduction, that is, the product of the total batch quantity and the third original processing duration, to obtain the first calculation result. Calculate the sum value of the first calculation result and the second original processing duration to obtain the reference processing duration.
[0126] For example Figure 7 in, the reference processing duration is the total batch quantity 4 The third original processing duration 2T + the second original processing duration T = 9T.
[0127] After obtaining the reference processing duration, calculate the ratio of the third original processing duration to the processing rate ratio corresponding to each adjustment frequency to obtain the second post-frequency-reduction processing duration corresponding to each adjustment frequency (for example Figure 7 in, the ratio of the third original processing duration 2T to the processing rate ratio x corresponding to each adjustment frequency to obtain the second post-frequency-reduction processing duration 2T / x for each adjustment frequency); determine the product of the second post-frequency-reduction processing duration corresponding to each adjustment frequency and the target batch quantity to obtain the second calculation result corresponding to each adjustment frequency (for example Figure 7 in, the product of the second post-frequency-reduction processing duration 2T / x corresponding to each adjustment frequency and the target batch quantity 3 to obtain the second calculation result 6T / x corresponding to each adjustment frequency); obtain the difference between the total batch quantity and the target batch quantity to obtain the non-frequency-reduced batch quantity (for example Figure 7 in, the difference between the total batch quantity 4 and the target batch quantity 3 to obtain the non-frequency-reduced batch quantity as 1); obtain the product of the non-frequency-reduced batch quantity and the third original processing duration to obtain the third calculation result (for example Figure 7 in, the product of the non-frequency-reduced batch quantity 1 and the third original processing duration 2T to obtain the third calculation result 2T); determine the sum value of the second calculation result corresponding to each adjustment frequency and the third calculation result to obtain the total processing duration corresponding to each adjustment frequency in the backpropagation stage (for example Figure 7 in, the sum value of the second calculation result 6T / x corresponding to each adjustment frequency and the third calculation result 2T to obtain the total processing duration 6T / x + 2T corresponding to each adjustment frequency in the backpropagation stage).
[0128] Determine the adjustment frequencies with the total processing duration not exceeding the reference processing duration as candidate adjustment frequencies (for example Figure 7 in 6T / x + 2T 9T, that is, 6T / x 7T, x 6 / 7, and the adjustment frequency with a processing rate ratio x not less than 6 / 7 is determined as the candidate adjustment frequency).
[0129] In some embodiments, determining the second power consumption value after frequency reduction corresponding to each candidate adjustment frequency based on the second batch quantity, the third batch quantity, each candidate adjustment frequency and the corresponding processor voltage, and the processing rate ratio corresponding to each candidate adjustment frequency includes:
[0130] (1.1) Obtain the capacitance load constant;
[0131] (1.2) Calculate the ratio of the second original processing duration to the processing rate ratio corresponding to each candidate adjustment frequency to obtain the third processing duration after frequency reduction corresponding to each candidate adjustment frequency;
[0132] (1.2) Based on the third processing duration after frequency reduction corresponding to each candidate adjustment frequency, the corresponding processor voltage, the capacitance load constant, and the second batch quantity, determine the first sub-power consumption value after frequency reduction of each candidate adjustment frequency in the forward propagation stage;
[0133] (1.4) Based on the second processing duration after frequency reduction corresponding to each candidate adjustment frequency, the corresponding processor voltage, the capacitance load constant, and the third batch quantity, determine the second sub-power consumption value after frequency reduction of each candidate adjustment frequency in the backward propagation stage;
[0134] (1.5) Add the first sub-power consumption value after frequency reduction and the second sub-power consumption value after frequency reduction of the same candidate adjustment frequency to obtain the second power consumption value after frequency reduction corresponding to each candidate adjustment frequency.
[0135] Among them, in order to ensure that on the basis of not exceeding the total processing duration before frequency reduction, it is necessary to screen out the second adjustment frequency with the lowest energy consumption value from the candidate adjustment frequencies, then it is necessary to obtain the second power consumption value after frequency reduction corresponding to each candidate adjustment frequency. First, obtain the capacitance load constant , calculate the ratio of the second original processing duration to the processing rate ratio corresponding to each candidate adjustment frequency to obtain the third processing duration after frequency reduction corresponding to each candidate adjustment frequency (for example Figure 7 the ratio of the second original processing duration T to the processing rate ratio x corresponding to each candidate adjustment frequency to obtain the third processing duration after frequency reduction T / x corresponding to each candidate adjustment frequency); according to the power consumption value formula combined with the third processing duration after frequency reduction corresponding to each candidate adjustment frequency, the corresponding processor voltage, the capacitance load constant, and the second batch quantity, determine the first sub-power consumption value after frequency reduction of each candidate adjustment frequency in the forward propagation stage (for example Figure 7 the first sub-power consumption value after frequency reduction in is 1 / ); According to the power consumption value formula Combined with the second post-downscaling processing duration, the corresponding processor voltage, the capacitance load constant, and the number of the third batch corresponding to each candidate adjustment frequency, determine the second sub-downscaling post-power consumption value of each candidate adjustment frequency in the backward propagation stage (for example Figure 7 the second sub-downscaling post-power consumption value in [example] is 9 / ); Add the first sub-downscaling post-power consumption value and the second sub-downscaling post-power consumption value of the same candidate adjustment frequency to obtain the second post-downscaling power consumption value corresponding to each candidate adjustment frequency (for example Figure 7 the second post-downscaling power consumption value corresponding to each candidate adjustment frequency in [example] is the first sub-downscaling post-power consumption value 1 / + the second sub-downscaling post-power consumption value is 9 / = 19 / ).
[0136] In some embodiments, the second adjustment frequency can also be determined according to the total processing duration of the entire training process, for example Figure 7 the total processing power consumption value of the entire training process corresponding to each candidate adjustment frequency in [example] 19 / + 29 . Substitute the processing rate ratio x of each candidate adjustment frequency into this formula to obtain the total processing power consumption value of the entire training process corresponding to each candidate adjustment frequency, and select the candidate adjustment frequency with the smallest total processing power consumption value as the second adjustment frequency.
[0137] As can be seen from the above, in the embodiment of the present application, the regulation delay of the frequency regulation of any neural network processor required for model training is obtained, and multiple neural network processors sequentially process multiple batches of tasks; the regulation delay is compared with a regulation delay threshold to obtain a comparison result; when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, each neural network processor is controlled to sequentially process multiple batches of tasks, and at least one neural network processor is downscaled when processing at least one batch of tasks in the forward propagation stage; when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, each neural network processor is controlled to sequentially process multiple batches of tasks in the forward propagation stage of the model training process, and each neural network processor is controlled to process multiple batches of tasks at intervals in the backward propagation stage, and at least one neural network processor is downscaled when processing at least one batch of tasks in the forward propagation stage and the backward propagation stage. Compared with the related art where a single NPU frequency and voltage are set for one computing task, in the embodiment of the present application, adaptive downscaling is performed in different stages of model training according to the regulation delay of the frequency regulation of the neural network processor, thereby improving energy efficiency.
[0138] For the specific implementation of each of the above steps, reference may be made to the previous embodiments, and details will not be elaborated herein.
[0139] To facilitate better implementation of the model training method provided by the embodiment of the present application, the embodiment of the present application also provides a device based on the above model training method. The meanings of the terms are the same as those in the above model training method, and the specific implementation details can be referred to the description in the method embodiment.
[0140] Please refer to Figure 8 , Figure 8 FIG. is a schematic structural diagram of the model training device provided by the embodiment of the present application. The model training device is applied to a computer device. The model training device may include a first acquisition unit 601, a comparison unit 602, a first downscaling unit 603, a second downscaling unit 604, etc.
[0141] The first acquisition unit 601 is configured to acquire the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple neural network processors sequentially process multiple batches of tasks;
[0142] The comparison unit 602 is configured to compare the regulation delay with a regulation delay threshold to obtain a comparison result;
[0143] The first downscaling unit 603 is configured to, when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each neural network processor to sequentially process multiple batches of tasks, and downscale at least one neural network processor when processing at least one batch of tasks in the forward propagation stage;
[0144] A second frequency reduction unit 604, configured to, when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, control each of the neural network processors to sequentially process multiple batches of tasks in the forward propagation stage of the model training process, control each of the neural network processors to process multiple batches of tasks at intervals in the backward propagation stage, and reduce the frequency when at least one neural network processor processes at least one batch of tasks in the forward propagation stage and the backward propagation stage.
[0145] In some embodiments, the first frequency reduction unit 603 includes:
[0146] A first determination subunit, configured to determine a first batch number of batch tasks that need to be frequency-reduced in the forward propagation stage;
[0147] A first acquisition subunit, configured to acquire a plurality of adjustment frequencies and corresponding processor voltages, and acquire a processing rate ratio corresponding to each of the adjustment frequencies;
[0148] A second determination subunit, configured to determine a first power consumption value after frequency reduction corresponding to each of the adjustment frequencies in the forward propagation stage based on the first batch number, each of the adjustment frequencies and the corresponding processor voltages, and the processing rate ratio corresponding to each of the adjustment frequencies;
[0149] A first screening subunit, configured to screen out a first adjustment frequency corresponding to the smallest first power consumption value after frequency reduction from the plurality of adjustment frequencies;
[0150] A first adjustment subunit, configured to adjust the frequency to the first adjustment frequency before each of the neural network processors processes the batch tasks that need to be frequency-reduced.
[0151] In some embodiments, the second determination subunit is configured to:
[0152] Acquire a first original processing duration for processing one batch of tasks in the forward propagation stage and a capacitance load constant;
[0153] Calculate a ratio of the first original processing duration to the processing rate ratio corresponding to each of the adjustment frequencies to obtain a first processing duration after frequency reduction corresponding to each of the adjustment frequencies;
[0154] Determine a first power consumption value after frequency reduction corresponding to each of the adjustment frequencies in the forward propagation stage based on the first batch number, the capacitance load constant, the first processing duration after frequency reduction, each of the adjustment frequencies, the first processing duration after frequency reduction corresponding to each of the adjustment frequencies, and the processor voltage corresponding to each of the adjustment frequencies.
[0155] In some embodiments, the first acquisition subunit is configured to:
[0156] Obtain the processing rate corresponding to each adjustment frequency and the processing rate corresponding to the highest frequency for any one of the neural network processors;
[0157] Determine the ratio of the processing rate corresponding to different adjustment frequencies to the processing rate corresponding to the highest frequency to obtain the processing rate proportion corresponding to each adjustment frequency.
[0158] In some embodiments, the second frequency reduction unit 604 includes:
[0159] A third determination subunit, configured to determine the second batch quantity of the batch tasks that need to be frequency-reduced in the forward propagation stage and the third batch quantity of the batch tasks that need to be frequency-reduced in the backward propagation stage;
[0160] A second acquisition subunit, configured to acquire the total batch quantity of the multiple batch tasks, the second original processing duration for processing one batch task in the forward propagation stage, and the third original processing duration for processing one batch task in the backward propagation stage;
[0161] A second screening subunit, configured to screen out the target batch quantity with the largest batch quantity from the batch quantities of the batch tasks that need to be frequency-reduced by each neural network processor in the backward propagation stage;
[0162] A third screening subunit, configured to screen out candidate adjustment frequencies from the multiple adjustment frequencies based on the total batch quantity, the target batch quantity, the second original processing duration, the third original processing duration, and the processing rate proportion corresponding to each adjustment frequency;
[0163] A fourth determination subunit, configured to determine the second power consumption value after frequency reduction corresponding to each candidate adjustment frequency based on the second batch quantity, the third batch quantity, each candidate adjustment frequency and the corresponding processor voltage, and the processing rate proportion corresponding to each candidate adjustment frequency;
[0164] A fourth screening subunit, configured to screen out the second adjustment frequency with the smallest second total power consumption value after frequency reduction from the multiple candidate adjustment frequencies;
[0165] A second adjustment subunit, configured to adjust the frequency to the second adjustment frequency before each neural network processor processes the batch tasks that need to be frequency-reduced.
[0166] In some embodiments, the third screening subunit is configured to:
[0167] Obtain the product of the total batch quantity and the third original processing duration to obtain a first calculation result;
[0168] Calculate the sum of the first calculation result and the second original processing duration to obtain a reference processing duration;
[0169] Calculate the ratio of the third original processing duration to the processing rate ratio corresponding to each adjustment frequency to obtain the second post-downclocking processing duration corresponding to each adjustment frequency;
[0170] Determine the product of the second post-downclocking processing duration corresponding to each adjustment frequency and the target batch quantity to obtain the second calculation result corresponding to each adjustment frequency;
[0171] Obtain the difference between the total batch quantity and the target batch quantity to obtain the non-downclocked batch quantity;
[0172] Obtain the product of the non-downclocked batch quantity and the third original processing duration to obtain a third calculation result;
[0173] Determine the sum of the second calculation result corresponding to each adjustment frequency and the third calculation result to obtain the total processing duration corresponding to each adjustment frequency in the backpropagation stage;
[0174] Determine the adjustment frequencies with the total processing duration not exceeding the reference processing duration as candidate adjustment frequencies.
[0175] In some embodiments, a fourth determination subunit is configured to:
[0176] Obtain a capacitance load constant;
[0177] Calculate the ratio of the second original processing duration to the processing rate ratio corresponding to each candidate adjustment frequency to obtain the third post-downclocking processing duration corresponding to each candidate adjustment frequency;
[0178] Based on the third post-downclocking processing duration corresponding to each candidate adjustment frequency, the corresponding processor voltage, the capacitance load constant, and the second batch quantity, determine the first sub-post-downclocking power consumption value of each candidate adjustment frequency in the forward propagation stage;
[0179] Based on the second post-downclocking processing duration corresponding to each candidate adjustment frequency, the corresponding processor voltage, the capacitance load constant, and the third batch quantity, determine the second sub-post-downclocking power consumption value of each candidate adjustment frequency in the backpropagation stage;
[0180] Add the first sub-post-downclocking power consumption value and the second sub-post-downclocking power consumption value of the same candidate adjustment frequency to obtain the second post-downclocking power consumption value corresponding to each candidate adjustment frequency.
[0181] For the specific implementation of each of the above units, reference may be made to the previous embodiments and will not be elaborated herein.
[0182] As can be seen from the above, in the embodiment of the present application, the first acquisition unit 601 acquires the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple neural network processors sequentially process multiple batches of tasks; the comparison unit 602 compares the regulation delay with the regulation delay threshold to obtain a comparison result; when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, the first frequency reduction unit 603 controls each neural network processor to sequentially process multiple batches of tasks, and reduces the frequency of at least one neural network processor processing at least one batch of tasks in the forward propagation stage; when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, the second frequency reduction unit 604 controls each neural network processor to sequentially process multiple batches of tasks in the forward propagation stage of the model training process, controls each neural network processor to process multiple batches of tasks at intervals in the backward propagation stage, and reduces the frequency when at least one neural network processor processes at least one batch of tasks in the forward propagation stage and the backward propagation stage. Compared with the related art where a single NPU frequency and voltage are set for one computing task, the embodiment of the present application adaptively reduces the frequency at different stages of model training according to the regulation delay of the frequency regulation of the neural network processor, thereby improving energy efficiency.
[0183] For the specific implementation of each of the above units, reference may be made to the previous embodiments, which will not be elaborated here.
[0184] Refer to Figure 9 , Figure 9 FIG. is a block diagram of a part of the computer device 1000 for implementing the embodiment of the present disclosure. The computer device 1000 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 622 (for example, one or more processors) and a memory 632, and one or more storage media 630 (for example, one or more mass storage devices) for storing application programs 642 or data 644. Among them, the memory 632 and the storage media 630 may be transient storage or persistent storage. The program stored in the storage media 630 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device 1000. Further, the central processor 622 may be configured to communicate with the storage media 630 and execute a series of instruction operations in the storage media 630 on the computer device 1000.
[0185] The computer device 1000 may further include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0186] The central processing unit 622 in the computer device 1000 may be used to execute the model training method of the embodiments of the present disclosure. For example:
[0187] Obtain the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple said neural network processors sequentially process multiple batches of tasks;
[0188] Compare the regulation delay with a regulation delay threshold to obtain a comparison result;
[0189] When the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each of the neural network processors to sequentially process multiple batches of tasks, and reduce the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage;
[0190] When the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, control each of the neural network processors to sequentially process multiple batches of tasks during the forward propagation stage of the model training process, control each of the neural network processors to process multiple batches of tasks at intervals during the backward propagation stage, and reduce the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage and the backward propagation stage.
[0191] The embodiments of the present disclosure also provide a computer-readable storage medium for storing program codes for executing the model training methods of the foregoing various embodiments.
[0192] The embodiments of the present disclosure also provide a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes to implement the above-mentioned model training method. For example:
[0193] Obtain the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple said neural network processors sequentially process multiple batches of tasks;
[0194] Compare the regulation delay with a regulation delay threshold to obtain a comparison result;
[0195] When the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each of the neural network processors to sequentially process multiple batches of tasks, and downscale at least one neural network processor when processing at least one batch of tasks during the forward propagation phase;
[0196] When the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, control each of the neural network processors to sequentially process multiple batches of tasks during the forward propagation phase of the model training process, control each of the neural network processors to process multiple batches of tasks at intervals during the backward propagation phase, and downscale at least one neural network processor when processing at least one batch of tasks during the forward propagation phase and the backward propagation phase.
[0197] In addition, the terms "including" and "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0198] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one)" or similar expressions thereof refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0199] It should be understood that in the description of the embodiments of this application, the meaning of "multiple (or multiple items)" is more than two, and understandings such as "greater than", "less than", "exceeding", etc. do not include the present number, and understandings such as "above", "below", "within", etc. include the present number.
[0200] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, indirect couplings or communication connections of devices or units, and can be in electrical, mechanical, or other forms.
[0201] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0202] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0203] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0204] It should also be understood that the various embodiments provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0205] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of that module or unit.
[0206] The above is a specific description of the implementation manner of the present application, but the present application is not limited to the above implementation manner. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
Claims
1. A model training method, characterized in that, Including: Obtaining the regulation delay of the frequency regulation of any neural network processor required for model training. Multiple neural network processors sequentially process multiple batches of tasks. The regulation delay is the difference between the time point when any neural network processor reaches the set high frequency when the frequency is adjusted from the low frequency to the set high frequency through a script while in a stable low-frequency state and the time when the frequency adjustment instruction is issued by the script. Comparing the regulation delay with a regulation delay threshold to obtain a comparison result. When the comparison result indicates that the regulation delay is greater than the regulation delay threshold, controlling each neural network processor to sequentially process multiple batches of tasks, and reducing the frequency when at least one neural network processor processes at least one batch of tasks in the forward propagation stage. When the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, controlling each neural network processor to sequentially process multiple batches of tasks in the forward propagation stage of the model training process, controlling each neural network processor to process multiple batches of tasks at intervals in the backward propagation stage, and reducing the frequency when at least one neural network processor processes at least one batch of tasks in the forward propagation stage and the backward propagation stage.
2. The model training method according to claim 1, wherein The reducing the frequency when at least one neural network processor processes at least one batch of tasks in the forward propagation stage includes: Determining the first batch quantity of the batch of tasks that need to have their frequency reduced in the forward propagation stage. Obtaining multiple adjusted frequencies and the corresponding processor voltages, and obtaining the processing rate ratio corresponding to each adjusted frequency. Based on the first batch quantity, each adjusted frequency and the corresponding processor voltage, and the processing rate ratio corresponding to each adjusted frequency, determining the first power consumption value after frequency reduction corresponding to each adjusted frequency in the forward propagation stage. Selecting, from multiple adjusted frequencies, the first adjusted frequency corresponding to the smallest first power consumption value after frequency reduction. Before each neural network processor processes the batch of tasks that need to have their frequency reduced, adjusting the frequency to the first adjusted frequency.
3. The model training method according to claim 2, wherein The determining the first power consumption value after frequency reduction corresponding to each adjusted frequency in the forward propagation stage based on the first batch quantity, each adjusted frequency and the corresponding processor voltage, and the processing rate ratio corresponding to each adjusted frequency includes: Obtaining the first original processing duration for processing one batch of tasks in the forward propagation stage and the capacitance load constant. Calculating the ratio of the first original processing duration to the processing rate ratio corresponding to each adjusted frequency to obtain the first processing duration after frequency reduction corresponding to each adjusted frequency. Based on the first batch quantity, the capacitance load constant, the first processing duration after frequency reduction, each adjusted frequency, the first processing duration after frequency reduction corresponding to each adjusted frequency, and the processor voltage corresponding to each adjusted frequency, determining the first power consumption value after frequency reduction corresponding to each adjusted frequency in the forward propagation stage.
4. The model training method according to claim 2, wherein The obtaining the processing rate ratio corresponding to each adjusted frequency includes: Obtaining the processing rate of any neural network processor corresponding to each adjusted frequency and the processing rate corresponding to the highest frequency. Determine the ratio of the processing rate corresponding to different adjustment frequencies to the processing rate corresponding to the highest frequency, and obtain the processing rate proportion corresponding to each adjustment frequency.
5. The model training method according to claim 1, wherein, The downscaling during the forward propagation stage and the backward propagation stage when at least one neural network processor processes at least one batch of tasks includes: Determine the second batch quantity of the batch of tasks that need to be downscaled during the forward propagation stage, and determine the third batch quantity of the batch of tasks that need to be downscaled during the backward propagation stage; Obtain the total batch quantity of multiple batches of tasks, the second original processing duration for processing one batch of tasks during the forward propagation stage, and the third original processing duration for processing one batch of tasks during the backward propagation stage; Screen out the target batch quantity with the largest batch quantity from the batch quantities of the batches of tasks that need to be downscaled by each neural network processor during the backward propagation stage; Based on the total batch quantity, the target batch quantity, the second original processing duration, the third original processing duration, and the processing rate proportion corresponding to each adjustment frequency, screen out candidate adjustment frequencies from multiple adjustment frequencies; Based on the second batch quantity, the third batch quantity, each candidate adjustment frequency and the corresponding processor voltage, and the processing rate proportion corresponding to each candidate adjustment frequency, determine the second post-downscaling power consumption value corresponding to each candidate adjustment frequency; From multiple candidate adjustment frequencies, screen out the second adjustment frequency corresponding to the smallest second total post-downscaling power consumption value; Before each neural network processor processes the batch of tasks that need to be downscaled, adjust the frequency to the second adjustment frequency.
6. The model training method according to claim 5, wherein The screening out candidate adjustment frequencies from multiple adjustment frequencies based on the total batch quantity, the target batch quantity, the second original processing duration, the third original processing duration, and the processing rate proportion corresponding to each adjustment frequency includes: Obtain the product of the total batch quantity and the third original processing duration to obtain a first calculation result; Calculate the sum value of the first calculation result and the second original processing duration to obtain a reference processing duration; Calculate the ratio of the third original processing duration to the processing rate proportion corresponding to each adjustment frequency to obtain the second post-downscaling processing duration corresponding to each adjustment frequency; Determine the product of the second post-downscaling processing duration corresponding to each adjustment frequency and the target batch quantity to obtain a second calculation result corresponding to each adjustment frequency; Obtain the difference between the total batch quantity and the target batch quantity to obtain the non-downscaled batch quantity; Obtain the product of the non-downscaled batch quantity and the third original processing duration to obtain a third calculation result; Determine the sum value of the second calculation result corresponding to each adjustment frequency and the third calculation result to obtain the total processing duration corresponding to each adjustment frequency during the backward propagation stage; Determine the adjustment frequencies with a total processing duration not exceeding the reference processing duration as candidate adjustment frequencies.
7. The model training method according to claim 6, wherein Determining the second power consumption value after downscaling corresponding to each of the candidate adjustment frequencies based on the second batch quantity, the third batch quantity, each of the candidate adjustment frequencies, and the corresponding processor voltage, and the proportion of the processing rate corresponding to each of the candidate adjustment frequencies includes: Obtaining a capacitance load constant; Calculating the ratio of the second original processing duration to the proportion of the processing rate corresponding to each of the candidate adjustment frequencies to obtain the third processing duration after downscaling corresponding to each of the candidate adjustment frequencies; Determining the first sub-power consumption value after downscaling of each of the candidate adjustment frequencies in the forward propagation stage based on the third processing duration after downscaling corresponding to each of the candidate adjustment frequencies, the corresponding processor voltage, the capacitance load constant, and the second batch quantity; Determining the second sub-power consumption value after downscaling of each of the candidate adjustment frequencies in the backward propagation stage based on the second processing duration after downscaling corresponding to each of the candidate adjustment frequencies, the corresponding processor voltage, the capacitance load constant, and the third batch quantity; Adding the first sub-power consumption value after downscaling and the second sub-power consumption value after downscaling of the same candidate adjustment frequency to obtain the second power consumption value after downscaling corresponding to each of the candidate adjustment frequencies.
8. A model training device, characterized in that, Including: A first acquisition unit for acquiring the regulation delay of the frequency regulation of any neural network processor required for model training. Multiple neural network processors sequentially process multiple batches of tasks. The regulation delay is the difference between the time point when any neural network processor reaches the set high frequency when the frequency is adjusted from the low frequency to the set high frequency through a script when it is in a stable low-frequency state and the time when the script issues a frequency adjustment instruction; A comparison unit for comparing the regulation delay with a regulation delay threshold to obtain a comparison result; A first downscaling unit for, when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, controlling each neural network processor to sequentially process multiple batches of tasks and downscaling at least one neural network processor to process at least one batch of tasks in the forward propagation stage; A second downscaling unit for, when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, controlling each neural network processor to sequentially process multiple batches of tasks in the forward propagation stage of the model training process, controlling each neural network processor to process multiple batches of tasks at intervals in the backward propagation stage, and downscaling at least one neural network processor to process at least one batch of tasks in the forward propagation stage and the backward propagation stage.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the model training method according to any one of claims 1 to 7.
10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the model training method according to any one of claims 1 to 7.
Citation Information
Patent Citations
NPU power consumption optimization system and method based on neural network structure
CN114217688A
Frequency modulation method, device and equipment and readable storage medium
CN118819860A