Model training method and device, storage medium and computer equipment

By obtaining the frequency regulation delay of the neural network processor and adaptively adjusting the frequency to reduce energy consumption, the problem of high energy consumption of smart computing clusters in AI large model training is solved and energy efficiency is improved.

CN120086025AActive Publication Date: 2025-06-03PENG CHENG LAB
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510572168.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-03
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

When smart computing clusters are training AI big models, due to the huge model scale, the computing resources of a single computing node are difficult to meet the requirements, resulting in high energy consumption and low energy efficiency.

Method used

By obtaining the frequency regulation delay of the neural network processor, the frequency of each neural network processor is adaptively adjusted to reduce the frequency at different stages of model training and improve energy efficiency.

Benefits of technology

Adaptive frequency reduction is achieved according to dynamic changes in the calculation load, which improves the energy efficiency of model training and reduces energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086025A_ABST
    Figure CN120086025A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a model training method and device, a storage medium and computer equipment, and the method comprises the steps: obtaining the regulation delay of frequency regulation of any one NPU in model training, and enabling a plurality of NPUs to process a plurality of batches of tasks in sequence. Comparing the regulation and control delay with a regulation and control delay threshold value to obtain a comparison result; and if the regulation delay is greater than a threshold value, enabling each NPU to process multiple batches of tasks in sequence, and performing frequency reduction on at least one batch of tasks processed by at least one NPU in a forward propagation stage. If the regulation delay is smaller than or equal to a threshold value, each NPU is still enabled to process multiple batches of tasks in sequence in a forward propagation stage of model training; and in the back propagation stage, each NPU is controlled to process multiple batches of tasks at intervals. Meanwhile, in the forward propagation stage and the back propagation stage, frequency reduction is carried out on at least one batch of tasks processed by at least one NPU. The mode of self-adaptive frequency reduction in different stages of model training according to NPU frequency regulation delay can effectively improve energy efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a model training method, device, storage medium, and computer device. Background Art

[0002] Today, with the rapid development of artificial intelligence (AI) technology, intelligent computing clusters (hereinafter referred to as intelligent computing clusters) have become the key pillars for promoting AI scientific research and industrial applications. Intelligent computing clusters are built by adopting advanced AI processors, including graphics processing units (GPUs), neural network processing units (NPUs), etc., and can efficiently process complex AI tasks and the computing requirements of large-scale data. However, the popularization and wide application of intelligent computing clusters have also brought severe energy consumption challenges, which pose a huge pressure on global energy use and environmental protection.

[0003] When training large AI models on intelligent computing clusters, due to the huge model scale, the computing resources of a single computing node are difficult to meet the requirements. Therefore, parallel training methods need to be adopted to accelerate the training process and make full use of the computing power resources of intelligent computing clusters. Parallel training mainly includes several common methods such as data parallelism, pipeline parallelism, and tensor parallelism. These methods usually set a single NPU frequency and voltage for a computing task according to the severity of the computing task. Although this method reduces the implementation complexity, it ignores the dynamic changes of the computing load, resulting in low energy efficiency. Therefore, related technologies urgently need to propose a model training method to solve the above technical problems. Summary of the Invention

[0004] The main purpose of this application is to provide a model training method, device, storage medium, and computer device, which can adaptively reduce the frequency in different stages of model training according to the regulation delay of the frequency regulation of the neural network processor, and improve energy efficiency.

[0005] In a first aspect, an embodiment of this application provides a model training method, including: Obtain the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple said neural network processors sequentially process multiple batches of tasks; Compare the regulation delay with a regulation delay threshold to obtain a comparison result; When the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each said neural network processor to sequentially process multiple batches of tasks, and reduce the frequency of at least one neural network processor processing at least one batch of tasks in the forward propagation stage; When the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, during the forward propagation stage of the model training process, control each of the neural network processors to sequentially process multiple batches of tasks, during the backward propagation stage, control each of the neural network processors to process multiple batches of tasks at intervals, and during the forward propagation stage and the backward propagation stage, downscale the frequency when at least one neural network processor processes at least one batch of tasks.

[0006] In a second aspect, an embodiment of the present application provides a model training device, including: A first acquisition unit, configured to acquire the regulation delay of frequency regulation of any neural network processor required for model training, and multiple batches of tasks are sequentially processed among the multiple neural network processors; A comparison unit, configured to compare the regulation delay with the regulation delay threshold to obtain a comparison result; A first downscaling unit, configured to, when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each of the neural network processors to sequentially process multiple batches of tasks, and downscale the frequency when at least one neural network processor processes at least one batch of tasks during the forward propagation stage; A second downscaling unit, configured to, when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, during the forward propagation stage of the model training process, control each of the neural network processors to sequentially process multiple batches of tasks, during the backward propagation stage, control each of the neural network processors to process multiple batches of tasks at intervals, and during the forward propagation stage and the backward propagation stage, downscale the frequency when at least one neural network processor processes at least one batch of tasks.

[0007] In a third aspect, an embodiment of the present application provides a storage medium. The computer-readable storage medium stores multiple instructions, and these instructions are suitable for being loaded by a processor to execute the model training method as described in any one of the above.

[0008] In a fourth aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the model training method as described in any one of the above is implemented.

[0009] In an embodiment of the present application, by obtaining the regulation delay of the frequency regulation of any neural network processor required for model training, multiple neural network processors sequentially process multiple batches of tasks; comparing the regulation delay with a regulation delay threshold to obtain a comparison result; when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, controlling each neural network processor to sequentially process multiple batches of tasks, and reducing the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage; when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, controlling each neural network processor to sequentially process multiple batches of tasks during the forward propagation stage of the model training process, controlling each neural network processor to process multiple batches of tasks at intervals during the backward propagation stage, and reducing the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage and the backward propagation stage. Compared with the related art where a single NPU frequency and voltage are set for one computing task, the embodiment of the present application adaptively reduces the frequency during different stages of model training according to the regulation delay of the frequency regulation of the neural network processor, thereby improving energy efficiency.

[0010] Other features and advantages of the present disclosure will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present disclosure. The objectives and other advantages of the present disclosure can be realized and obtained by the structures specifically pointed out in the specification, the claims, and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0012] Figure 1 It is a schematic diagram of the scenario of the model training system provided by the embodiment of the present application.

[0013] Figure 2 It is a schematic flowchart of the model training method provided by the embodiment of the present application.

[0014] Figure 3 It is a schematic diagram of the pipeline parallel training of the AI large model provided by the embodiment of the present application.

[0015] Figure 4 It is a schematic diagram of the bubble merging scheduling under high regulation delay and the training before frequency regulation provided by the embodiment of the present application.

[0016] Figure 5The bubble merging scheduling under high regulation delay provided by the embodiments of this application, and the training schematic diagram after frequency regulation.

[0017] Figure 6 The bubble merging scheduling under low regulation delay provided by the embodiments of this application, and the training schematic diagram before frequency regulation.

[0018] Figure 7 The bubble merging scheduling under low regulation delay provided by the embodiments of this application, and the training schematic diagram after frequency regulation.

[0019] Figure 8 The structural schematic diagram of the model training device provided by the embodiments of this application.

[0020] Figure 9 The structural schematic diagram of the computer device provided by the embodiments of this application. Detailed implementation manners

[0021] In order to enable those skilled in the art of this technology to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope protected by this application.

[0022] It should be noted that in some processes described in the specification, claims, and the above-mentioned drawings, there are multiple steps that appear in a specific order. However, it should be clearly understood that these steps may not be executed in the order in which they appear in this article or may be executed in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, descriptions such as "first", "second", or "target" in this article are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence.

[0023] Before further elaborating on the embodiments of this disclosure, the nouns and terms involved in the embodiments of this disclosure are described. The nouns and terms involved in the embodiments of this disclosure are applicable to the following explanations: Neural - Processing Unit (NPU): A processor designed specifically for neural network operations, especially for a large number of matrix and vector operations in deep learning. With the rapid development of deep learning technology in the field of artificial intelligence, traditional CPUs and GPUs gradually expose problems of low efficiency and high energy consumption when processing deep learning tasks. The NPU emerged as the times require, and it has the following characteristics: Highly Parallel Computing: A large number of computing units are integrated internally, capable of processing numerous data simultaneously, significantly accelerating the training and inference speeds of deep learning models. For example, in image recognition tasks, it can quickly complete the extraction and analysis of image features.

[0024] Specialized Architecture: Designed based on the characteristics of neural network algorithms, it has specialized optimizations for common operations such as convolution and activation, and can efficiently execute various operations of neural networks.

[0025] Low Power Consumption: Compared with general-purpose processors, it can operate with lower power consumption when performing deep learning tasks, making it suitable for mobile devices, smart home devices, etc. with strict power consumption requirements.

[0026] Today, NPUs are widely used in fields such as intelligent security, autonomous driving, and intelligent voice assistants, promoting the popularization and development of artificial intelligence applications.

[0027] Forward Propagation: It refers to passing the input data into the input layer of the neural network. The data is sequentially calculated through each layer according to the network hierarchy. For each layer, first calculate the weighted sum of the input data and add the bias, then apply the activation function to obtain the output of that layer. Finally, the predicted value of the model is obtained from the output layer. Then, use a loss function (such as mean squared error, cross-entropy, etc.) to calculate the gap between the predicted value and the actual label, that is, the loss value, and the gradient of the loss with respect to the output is also calculated. Simply put, it is the process from input to generating a prediction result and calculating the loss.

[0028] Backward Propagation: It is the core algorithm for training deep neural networks, aiming to optimize model parameters by calculating and propagating gradients. Its core is the chain rule. Using the chain rule, the gradient of the loss function with respect to the model output is propagated backward layer by layer to each parameter in the network, and the gradients of the loss function with respect to the weights and biases of each layer are calculated. Then, based on the calculated gradients, use the gradient descent algorithm or other optimization algorithms to update the weights and biases of each layer, thereby reducing the value of the loss function and making the predicted result of the model closer to the actual label.

[0029] Forward propagation provides the intermediate results required for calculating gradients in backward propagation, and backward propagation optimizes the model by adjusting parameters based on the loss obtained from forward propagation. The two complement each other and jointly complete the model training process.

[0030] When training large AI models on an intelligent computing cluster, due to the huge scale of the model, the computing resources of a single computing node are difficult to meet the requirements. Therefore, parallel training methods need to be adopted to accelerate the training process and make full use of the computing power resources of the intelligent computing cluster. Parallel training mainly includes several common methods such as data parallelism, pipeline parallelism, and tensor parallelism. Each method has its applicable scenarios and advantages, which will be introduced one by one below.

[0031] (1)Data Parallelism Data parallelism is one of the most commonly used parallel training methods. Its core idea is to keep the model parameters consistent across all computing devices, and then divide the training data into multiple batches, which are separately assigned to different computing nodes for independent calculation. Each node performs the same forward and backward propagation, but uses different subsets of data. After each batch of training is completed, the nodes will merge the gradients they have calculated respectively through a gradient synchronization mechanism and update the global parameters of the model. Its advantage is that data parallelism can be extended to multiple nodes without changing the model structure, is suitable for training large-scale datasets, and is easy to implement. Frameworks such as PyTorch and TensorFlow provide support for data parallelism.

[0032] (2)Tensor Parallelism Tensor parallelism is a fine-grained implementation method of model parallelism. In large models, many computational operations (such as matrix multiplication) involve operations on large tensors. Tensor parallelism divides these large tensors into smaller parts and distributes them to different computing nodes for parallel calculation. For example, a large matrix can be split by rows or columns among multiple NPUs, and each NPU is responsible for processing a part of the matrix calculation and finally merging the results. This method is suitable for extremely large models, especially those cases where a single computing node cannot accommodate the complete model parameters, and can make full use of hardware resources by disassembling the computational tasks into fine-grained parts and executing them on multiple devices.

[0033] (3)Pipeline Parallelism Pipeline parallelism is another model parallelism method. Its core idea is to split the model hierarchically and assign different layers to different NPUs. During training, the input data passes through each part of the model in a pipeline manner. The first batch of data is immediately passed to the next NPU after passing through the first NPU, and this NPU can start processing the second batch of data. This is suitable for cases where the model has an obvious hierarchical structure and can be split by layer, such as deep neural networks, and reduces the memory footprint because each node only needs to store a part of the model parameters.

[0034] (4)Hybrid Parallelism The hybrid parallel method combines data parallelism, tensor parallelism, and pipeline parallelism to perform parallel computing at different levels. For very large-scale models, a single parallel method may not be able to fully utilize the efficiency of computing resources, so the hybrid parallel method is usually adopted. For example, tensor parallelism can be adopted within a computing node first, and the tensors of each layer are sliced and distributed to multiple devices for processing; at the same time, different layers of the entire model can be distributed to multiple nodes through pipeline parallelism; finally, data parallelism is used to slice the large dataset and distribute it to different computing nodes for processing.

[0035] The above method usually sets a single NPU frequency and voltage for a computing task according to the severity of the computing task. Although this method reduces the implementation complexity, it ignores the dynamic changes of the computing load, resulting in low energy efficiency.

[0036] To solve the above problems, the embodiments of the present application obtain the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple neural network processors sequentially process multiple batches of tasks; compare the regulation delay with the regulation delay threshold to obtain a comparison result; when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each neural network processor to sequentially process multiple batches of tasks, and reduce the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage; when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, control each neural network processor to sequentially process multiple batches of tasks during the forward propagation stage of the model training process, control each neural network processor to process multiple batches of tasks at intervals during the backward propagation stage, and reduce the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage and the backward propagation stage. Compared with the related art of setting a single NPU frequency and voltage for a computing task, the embodiments of the present application adaptively reduce the frequency according to the regulation delay of the frequency regulation of the neural network processor during different stages of model training, thereby improving energy efficiency. For details, please continue to refer to the following specific embodiments.

[0037] Please refer to Figure 1 , Figure 1 which is a scenario schematic diagram of the model training system provided by the embodiments of the present application. It includes a terminal 140, the Internet 130, a gateway 120, a server 110, etc.

[0038] The terminal 140 includes, but is not limited to, pre-configured laptop computers, tablet computers, desktop computers, and other electronic devices with data reporting capabilities. In addition, it can be a single device or a set of multiple devices. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.

[0039] The terminal 140 refers to a computer system that can report data to the server 110. Compared with ordinary terminals, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, etc.

[0040] The gateway 120 is also called an internetwork connector and a protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. Between two systems using different communication protocols, data formats, or languages, and even with completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. The message sent by the terminal 140 to the server 110 needs to be sent to the corresponding server 110 through the gateway 120. The message sent by the server 110 to the terminal 140 also needs to be sent to the corresponding terminal 140 through the gateway 120.

[0041] The model training method of the embodiments of the present disclosure can be implemented on the server 110.

[0042] It should be noted that Figure 1 The scenario schematic diagram of the model training system shown is only an example. The model training system and scenario described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of image processing technology and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0043] In this embodiment, the description will be made from the perspective of the model training device, which can be specifically integrated in a computer device with a storage unit and installed with a microprocessor and having computing capabilities.

[0044] Please refer to Figure 2 , Figure 2 This is a schematic flowchart of the model training method provided by the embodiments of the present application. The model training method includes: In step 201, obtain the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple neural network processors sequentially process multiple batches of tasks.

[0045] Among them, in the training process of the embodiments of the present application, multiple neural network processors (NPUs) with the same regulation delay are used for training. Therefore, only by obtaining the regulation delay of frequency regulation of any one neural network processor required for model training can the regulation delay of each NPU used for model training be known. In view of the differences in different NPU hardware conditions, in order to obtain the accurate frequency-voltage regulation delay values required for energy efficiency optimization of the AI large model pipeline parallel training, the following specific measurement methods for frequency-voltage regulation delay are proposed: Turn on the NPU frequency reading module. First, ensure that the frequency reading module of the NPU is in the on state to monitor the NPU frequency change in real time.

[0046] Maintain the low-frequency state. Use a script to set the frequency of the NPU to a low frequency and maintain it for 5 seconds in the low-frequency state. This step ensures that the NPU is in a stable low-frequency state and provides a benchmark for subsequent operations.

[0047] Frequency switching operation. Immediately adjust the frequency of the NPU from the low frequency to the high frequency through a script: record the time when the frequency adjustment instruction is issued by the script, denoted as t_1. At this time, the frequency-voltage regulation mechanism inside the NPU starts to operate to adjust to the set high frequency.

[0048] Monitor the high-frequency arrival time. Through the NPU frequency reading module, monitor the time point when the NPU reaches the set high frequency, denoted as t_2.

[0049] Calculate the frequency-voltage regulation delay. Calculate the NPU frequency-voltage regulation delay according to the following formula: (frequency regulation delay) delay = t_2 - t_1.

[0050] Delay range and analysis. Through multiple experiments, verify and record the regulation delay data in different scenarios. Generally, the NPU frequency-voltage regulation delay is between 10 milliseconds and 1000 milliseconds. The specific value of the delay depends on the model, design, and operating environment of the NPU hardware.

[0051] Through the above method, the accurate NPU frequency-voltage regulation delay can be obtained, providing key data support for the subsequent energy efficiency optimization of AI large model training. Based on obtaining the regulation delay of each NPU before model training, a group of NPUs with the same or regulation delays differing by no more than a preset difference can be used as the NPUs required for training the model.

[0052] Specifically, please refer to Figure 3 , Figure 3 which is the schematic diagram of the AI large model pipeline parallel training provided by the embodiments of the present application. As shown in Figure 3As shown, the figure shows the process of training a micro-batch in parallel through a pipeline. The micro-batch needs to run on NPU1 - NPU4 in sequence. The forward propagation is represented in light color, and the backward propagation is represented in dark color. The tasks of a micro-batch are carried out in the order of NPU1 - NPU2 - NPU3 - NPU4 during the forward propagation stage (for example Figure 3 the forward propagation of micro-batch 1 in Figure 3 is in the order of NPU1 - NPU2 - NPU3 - NPU4), and the tasks of a micro-batch are carried out in the order of NPU4 - NPU3 - NPU2 - NPU1 during the backward propagation stage (for example

[0053] the backward propagation 1B of micro-batch 1 in

[0054] is in the order of NPU4 - NPU3 - NPU2 - NPU1). Since the amount of calculation in the backward propagation is about twice that of the forward propagation, the calculation duration is also about twice the relationship. Due to the logical dependencies between these 8 calculations, the calculation arrangement cannot be further compressed and must be carried out sequentially.

[0055] Specifically, if it belongs to low regulation delay, the comparison result is that the regulation delay is less than or equal to the regulation delay threshold; if it belongs to high regulation delay, the comparison result is that the regulation delay is greater than the regulation delay threshold.

[0056] In step 203, when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each of the neural network processors to process multiple batches of tasks in sequence, and reduce the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage.

[0057] Among them, please refer to Figure 4 , Figure 4 which is the training schematic diagram before frequency regulation of the bubble merging scheduling under high regulation delay provided by the embodiments of the present application. In order to further improve the training efficiency of large models, the embodiments of the present application generally use not only 1 micro-batch, but a group composed of multiple micro-batches (for example Figure 4 the micro-batch tasks 1, 2, 3, and 4 inFigure 4 The bubble merging scheduling method in Figure 4 performs frequency-voltage regulation by arranging and executing. This scheduling method can merge bubbles (e.g.,

[0058] the uncalculated part in the central white area) together, reducing the number of times of frequency regulation execution, thereby avoiding the impact of high regulation latency. Figure 5 Specifically, as Figure 5 shown, Figure 5 this is a training schematic diagram of bubble merging scheduling under high regulation latency provided by an embodiment of the present application after frequency regulation. In the scenario of high regulation latency, the calculation part in the backpropagation stage is not downscaled because the calculation time in the backpropagation stage is long, and downscaling will significantly increase the total calculation duration, affecting the overall calculation duration. Therefore, downscaling is only performed for at least one neural network processor to process at least one batch of tasks. For example

[0059] In some embodiments, downscaling at least one neural network processor to process at least one batch of tasks in the forward propagation stage includes: (1) Determining the first batch quantity of the batch of tasks that need to be downscaled in the forward propagation stage; (2) Obtaining multiple adjusted frequencies and corresponding processor voltages, and obtaining the processing rate ratio corresponding to each of the adjusted frequencies; (3) Based on the first batch quantity, each of the adjusted frequencies and the corresponding processor voltages, and the processing rate ratio corresponding to each of the adjusted frequencies, determining the first post-downscaling power consumption value corresponding to each of the adjusted frequencies in the forward propagation stage; (4) Screening out the first adjusted frequency with the smallest corresponding first post-downscaling power consumption value from the multiple adjusted frequencies; (5) Before each neural network processor processes the batch of tasks that need to be downscaled, adjusting the frequency to the first adjusted frequency.

[0060] Among them, although it is known that in the scenario of high regulation latency, downscaling is performed on the batch of tasks in the forward propagation stage, but specifically what frequency to reduce to in order to minimize the power consumption value during the training process needs to be determined through specific calculations. The determination method is: Determine the first batch quantity of the batch tasks that need to be downclocked during the forward propagation stage of model training. Here, the first batch quantity is the sum of the task quantities of the batch tasks that need to be downclocked in the batch tasks corresponding to each NPU. For example, for the forward propagation calculation of batch task 4 on NPU2, the forward propagation calculations of batch tasks 3 and 4 on NPU3, and the forward propagation calculations of batch tasks 2, 3, and 4 on NPU4 are downclocked. Then the first batch quantity is the sum of the task quantity 1 corresponding to NPU2, the task quantity 2 corresponding to NPU3, and the task quantity 3 corresponding to NPU4, that is, 1 + 2 + 3 = 6.

[0061] Specifically, for an NPU chip, it is necessary to measure its corresponding voltage, processing rate, and power under different frequency settings. For this purpose, a frequency-voltage-processing rate-power relationship mapping table needs to be measured.

[0062] To perform the measurement, first find the forward propagation operators in the training load of the AI large model. For each adjustable frequency f_0, f_1, f_2, f_3,... of the NPU chip, measure its corresponding processor voltage V_0, V_1, V_2, V_3,... and the corresponding operator processing rate s_0, s_1, s_2, s_3,... and the power P_0, P_1, P_2, P_3,... of the NPU. Suppose f_0 is the highest frequency of the NPU chip. Then normalize the operator processing rate to obtain the processing rate ratios corresponding to f_0, f_1, f_2, f_3,... as 1, s_1 / s_0, s_2 / s_0, s_3 / s_0,....

[0063] Using this rate table, the frequency-voltage setting for a specific processing rate can be found. In the following text, assume that for a processing rate ratio x, the required frequency setting is f(x) and the voltage is V(x). For example, for a certain NPU chip, f_3 = 1500 MHz, V_3 = 810 mV, s_3 / s_0 = 0.75, and P_3 = 280 W. Then when it is necessary to reduce the NPU processing rate to 0.75, according to this table, it can be found that the frequency needs to be set to f(0.75) = 1500 MHz and the voltage to V(0.75) = 810 mV.

[0064] Based on this, different adjustment frequencies f(x) and the corresponding processor voltages V(x) can be obtained, and the processing rate ratio x = s(x) / s_0 corresponding to each adjustment frequency f. Based on this information, the first power consumption value E1 after frequency reduction corresponding to each adjustment frequency in the forward propagation stage can be calculated. The first adjustment frequency P1 with the smallest first power consumption value E1min after frequency reduction is selected from multiple adjustment frequencies. It shows that after reducing the frequency according to the first adjustment frequency P1, the processing power consumption value is the lowest. Therefore, before each neural network processor processes a batch of tasks that require frequency reduction, the frequency is adjusted to the first adjustment frequency P1.

[0065] For example, there are adjustment frequencies P1, P2, P3, and P4. The first power consumption value after frequency reduction corresponding to P1 is 150W, the first power consumption value after frequency reduction corresponding to P2 is 120W, the first power consumption value after frequency reduction corresponding to P3 is 130W, and the first power consumption value after frequency reduction corresponding to P4 is 140W. Then, the adjustment frequency P2 corresponding to the smallest first power consumption value after frequency reduction of 120W is determined as the first adjustment frequency. Before each neural network processor processes a batch of tasks that require frequency reduction, the frequency is adjusted to P2.

[0066] In some embodiments, the determining the first power consumption value after frequency reduction corresponding to each adjustment frequency in the forward propagation stage based on the first batch quantity, each adjustment frequency and the corresponding processor voltage, and the processing rate ratio corresponding to each adjustment frequency includes: (1.1) Obtain the first original processing duration for processing a batch of tasks in the forward propagation stage and the capacitance load constant; (1.2) Calculate the ratio of the first original processing duration to the processing rate ratio corresponding to each adjustment frequency to obtain the first processing duration after frequency reduction corresponding to each adjustment frequency; (1.3) Based on the first batch quantity, the capacitance load constant, the first processing duration after frequency reduction, each adjustment frequency, the first processing duration after frequency reduction corresponding to each adjustment frequency, and the processor voltage corresponding to each adjustment frequency, determine the first power consumption value after frequency reduction corresponding to each adjustment frequency in the forward propagation stage.

[0067] Among them, the specific method for determining the first power consumption value after frequency reduction corresponding to each adjustment frequency in the forward propagation stage is: obtain the first original processing duration T for processing a batch of tasks in the forward propagation stage and the capacitance load constant C, calculate the ratio of the first original processing duration T to the processing rate corresponding to each adjustment frequency x to obtain the first processing duration after frequency reduction T / x corresponding to each adjustment frequency. According to the power consumption value calculation formula It can be known that the power consumption value after frequency reduction corresponding to each adjustment frequency f(x) for processing each batch of tasks is , where t is the processing duration, and combined with the quantity of the first batch, the first power consumption value after frequency reduction corresponding to each adjustment frequency f(x) can be calculated, that is, the quantity of the first batch P(x).

[0068] Taking Figure 5 as an example, Figure 5 in it, the quantity of the first batch is 6, then the first power consumption value after frequency reduction corresponding to each adjustment frequency f(x) is 6 .

[0069] In some embodiments, the first total power consumption value after frequency reduction corresponding to each adjustment frequency can also be calculated to determine the first adjustment frequency. For example in it, the calculation duration of a single forward propagation is T, and the reverse calculation duration is generally 2 times that of the forward propagation, so the reverse calculation duration is 2T. The processing rate before frequency reduction is 1, the corresponding frequency is f(1), the regulation frequencies of 3 NPUs regulated by voltage V(1) are f(x), the voltage is V(x), and the calculation duration of each block becomes T / x, then the total first power consumption value after frequency reduction Figure 5 is , 6 is the quantity of the first batch, the remaining task quantity of the forward propagation stage without frequency reduction is 10, and the remaining task quantity of the backward propagation stage without frequency reduction is 16, where the calculation duration corresponding to 16 is 2T, and the calculation duration corresponding to 10 is T, then 42 in the formula is obtained by 16 ×2 + 10. By calculating the first total power consumption value after frequency reduction corresponding to each adjustment frequency to determine the first adjustment frequency with the smallest first total power consumption value after frequency reduction . In some embodiments, obtaining the ratio of the processing rate corresponding to each adjustment frequency includes:

[0070] (1.1) Obtaining the processing rate corresponding to each adjustment frequency of any one of the neural network processors and the processing rate corresponding to the highest frequency; (1.2) Determining the ratio of the processing rate corresponding to different adjustment frequencies to the processing rate corresponding to the highest frequency to obtain the ratio of the processing rate corresponding to each adjustment frequency. Among them, the method for determining the ratio of the processing rate corresponding to each adjustment frequency is to obtain the processing rate s(x) corresponding to each adjustment frequency of any one of the neural network processors and the processing rate s(0) corresponding to the highest frequency, and determine the ratio of the processing rate corresponding to different adjustment frequencies to the processing rate corresponding to the highest frequency to obtain the processing rate s(x) / s(0) corresponding to each adjustment frequency.

[0071]

[0072] In step 204, when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, during the forward propagation stage of the model training process, control each of the neural network processors to sequentially process multiple batches of tasks, during the backward propagation stage, control each of the neural network processors to process multiple batches of tasks at intervals, and downscale the frequency when at least one neural network processor processes at least one batch of tasks during the forward propagation stage and the backward propagation stage.

[0073] Among them, as Figure 6 shown, Figure 6 This is the schematic diagram of the bubble merging scheduling under low regulation delay provided by the embodiment of the present application before frequency regulation. If the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, it means that the current is in a low regulation delay scenario. During the forward propagation stage of the model training process, control each neural network processor to sequentially process multiple batches of tasks. For example, control NPU1 to sequentially process batch tasks 1, 2, 3, and 4; during the backward propagation stage, control each of the neural network processors to process multiple batches of tasks at intervals. For example, control NPU3 to process batch task 2 at an interval after processing batch task 1. This interval is because NPU4 immediately processes the backward propagation calculation of batch task 1 after processing batch task 1, and waits until the backward propagation calculation of batch task 1 is completed before processing batch task 2, resulting in an interval between the backward propagation calculations of batch task 1 and batch task 2 of NPU3.

[0074] The bubble reduction scheduling in the low regulation delay scenario is different from the Figure 4 bubble merging scheduling in Figure 4 in terms of the scheduling method. Its purpose is to reduce the size of each bubble, so that more computing units can perform frequency regulation. Figure 6 In

[0075] Please refer to Figure 7 Figure 7 This is the schematic diagram of the bubble merging scheduling under low regulation delay provided by the embodiment of the present application after frequency regulation. During the forward propagation stage and the backward propagation stage, when at least one neural network processor processes at least one batch of tasks, downscale its NPU frequency voltage, so that the calculation duration is appropriately extended, and the remaining calculation duration remains unchanged. In this way, the power consumption of the underlined calculation can be reduced, but at the same time, the overall calculation progress is not affected.

[0076] For example, Figure 7For the 3B calculations on NPU1, NPU2, and NPU3, the operation speed ratio is reduced to 0.86 of the original by setting the frequency to f(0.86) and the voltage to V(0.86); the frequencies of the remaining calculations are also set in a similar manner. Through this strategy, the energy consumption can be effectively reduced without affecting the overall training time, achieving an improvement in energy efficiency.

[0077] In some embodiments, downscaling the frequency when at least one neural network processor processes at least one batch of tasks in the forward propagation stage and the backward propagation stage includes: (1) Determining the second batch quantity of the batch tasks that need to be downscaled in the forward propagation stage, and determining the third batch quantity of the batch tasks that need to be downscaled in the backward propagation stage; (2) Obtaining the total batch quantity of multiple said batch tasks, the second original processing duration for processing one batch of tasks in the forward propagation stage, and the third original processing duration for processing one batch of tasks in the backward propagation stage; (3) Screening out the target batch quantity with the largest batch quantity from the batch quantities of the batch tasks that need to be downscaled in the backward propagation stage for each said neural network processor; (4) Based on the total batch quantity, the target batch quantity, the second original processing duration, the third original processing duration, and the processing rate ratio corresponding to each adjustment frequency, screening out candidate adjustment frequencies from multiple said adjustment frequencies; (5) Based on the second batch quantity, the third batch quantity, each said candidate adjustment frequency and the corresponding processor voltage, and the processing rate ratio corresponding to each said candidate adjustment frequency, determining the second power consumption value after downscaling corresponding to each said candidate adjustment frequency; (6) Screening out the second adjustment frequency with the smallest second total power consumption value after downscaling corresponding to it from multiple said candidate adjustment frequencies; (7) Before each said neural network processor processes the batch tasks that need to be downscaled, adjusting the frequency to the second adjustment frequency.

[0078] Among them, to avoid the total processing duration of the entire processing process after downscaling exceeding the original total processing duration before downscaling when the processor is frequency-adjusted according to certain adjustment frequencies, it is necessary to screen out candidate adjustment frequencies that do not affect the total processing duration from the adjustment frequencies, and according to the second power consumption value after downscaling corresponding to each candidate adjustment frequency, screen out the second adjustment frequency with the smallest power consumption value, so as to adjust the frequency to the second adjustment frequency before each neural network processor processes the batch tasks that need to be downscaled. Thus, the lowest power consumption is achieved without affecting the total processing duration.

[0079] Specifically, the screening process of candidate adjustment frequencies needs to determine the second batch quantity of batch tasks that need to be downclocked during the forward propagation stage, and the third batch quantity of batch tasks that need to be downclocked during the backward propagation stage; obtain the total batch quantity of multiple batch tasks, the second original processing duration for processing one batch task during the forward propagation stage, and the third original processing duration for processing one batch task during the backward propagation stage; screen out the target batch quantity with the largest batch quantity from the batch quantities of batch tasks that need to be downclocked by each neural network processor during the backward propagation stage; and screen out candidate adjustment frequencies from multiple said adjustment frequencies based on the total batch quantity, the target batch quantity, the second original processing duration, the third original processing duration, and the processing rate ratio corresponding to each adjustment frequency.

[0080] For example Figure 7 Among them, the multiple batch tasks are batch task 1, batch task 2, batch task 3, and batch task 4 respectively, so the total batch quantity is 4. During the forward propagation stage, it is determined that the batch task that needs to be downclocked is the batch task in NPU3 4 , so the second batch quantity is 1; during the backward propagation stage, the batch tasks that need to be downclocked are respectively in NPU1, NPU2, and NPU3 1B , 2B , 3B , so the third batch quantity is 9; the second original processing duration for processing one batch task during the forward propagation stage is T, and the third original processing duration for processing one batch task during the backward propagation stage is 2T. The batch quantities that need to be downclocked in NPU1, NPU2, and NPU3 are the largest, all being 3, so the target batch quantity is 3; finally, based on the total batch quantity 4, the target batch quantity 3, the second original processing duration T, the third original processing duration 2T, and the processing rate ratio x corresponding to each adjustment frequency, candidate adjustment frequencies are screened out from multiple said adjustment frequencies.

[0081] The calculation method for the second power consumption value after downclocking corresponding to each candidate adjustment frequency is to determine the second power consumption value after downclocking corresponding to each candidate adjustment frequency based on the second batch quantity, the third batch quantity, each candidate adjustment frequency and the corresponding processor voltage, and the processing rate ratio corresponding to each candidate adjustment frequency.

[0082] In some embodiments, the screening out of candidate adjustment frequencies from multiple said adjustment frequencies based on the total batch quantity, the target batch quantity, the second original processing duration, the third original processing duration, and the processing rate ratio corresponding to each adjustment frequency includes: (1.1) Obtain the product of the total batch quantity and the third original processing duration to get the first calculation result; (1.2) Calculate the sum of the first calculation result and the second original processing duration to obtain a reference processing duration; (1.3) Calculate the ratio of the third original processing duration to the processing rate ratio corresponding to each adjustment frequency to obtain the second post-downscaling processing duration corresponding to each adjustment frequency; (1.4) Determine the product of the second post-downscaling processing duration corresponding to each adjustment frequency and the target batch quantity to obtain the second calculation result corresponding to each adjustment frequency; (1.5) Obtain the difference between the total batch quantity and the target batch quantity to obtain the non-downscaled batch quantity; (1.6) Obtain the product of the non-downscaled batch quantity and the third original processing duration to obtain a third calculation result; (1.7) Determine the sum of the second calculation result corresponding to each adjustment frequency and the third calculation result to obtain the total processing duration corresponding to each adjustment frequency in the backward propagation stage; (1.8) Determine the adjustment frequencies whose total processing duration does not exceed the reference processing duration as candidate adjustment frequencies.

[0083] Among them, the specific method of screening candidate adjustment frequencies from multiple adjustment frequencies according to the total batch quantity, target batch quantity, second original processing duration, third original processing duration, and the processing rate ratio corresponding to each adjustment frequency is as follows: First, determine a reference processing duration. Since this duration is mainly the reference duration in the backward propagation stage, obtain the processing duration in the backward propagation stage before downscaling, that is, the product of the total batch quantity and the third original processing duration, to obtain a first calculation result. Calculate the sum of the first calculation result and the second original processing duration to obtain the reference processing duration.

[0084] For example Figure 7 in, the reference processing duration is the total batch quantity 4 The third original processing duration 2T + the second original processing duration T = 9T.

[0085] After obtaining the reference processing duration, calculate the ratio of the third original processing duration to the processing rate ratio corresponding to each adjustment frequency to obtain the second post-downscaling processing duration corresponding to each adjustment frequency (for example Figure 7 in, the ratio of the third original processing duration 2T to the processing rate ratio x corresponding to each adjustment frequency to obtain the second post-downscaling processing duration 2T / x for each adjustment frequency); determine the product of the second post-downscaling processing duration corresponding to each adjustment frequency and the target batch quantity to obtain the second calculation result corresponding to each adjustment frequency (for example Figure 7Multiply the second post-downclock processing duration 2T / x corresponding to each adjustment frequency by the target batch quantity 3 to obtain the second calculation result 6T / x corresponding to each adjustment frequency; obtain the difference between the total batch quantity and the target batch quantity to obtain the non-downclocked batch quantity (for example Figure 7 In the difference between the total batch quantity 4 and the target batch quantity 3, the non-downclocked batch quantity is obtained as 1); obtain the product of the non-downclocked batch quantity and the third original processing duration to obtain the third calculation result (for example Figure 7 In the product of the non-downclocked batch quantity 1 and the third original processing duration 2T, the third calculation result 2T is obtained); determine the sum value of the second calculation result and the third calculation result corresponding to each adjustment frequency to obtain the total processing duration corresponding to each adjustment frequency in the backpropagation stage (for example Figure 7 In the sum value of the second calculation result 6T / x and the third calculation result 2T corresponding to each adjustment frequency, the total processing duration 6T / x + 2T corresponding to each adjustment frequency in the backpropagation stage is obtained).

[0086] Determine the adjustment frequencies with the total processing duration not exceeding the reference processing duration as candidate adjustment frequencies (for example Figure 7 In 6T / x + 2T 9T, that is, 6T / x 7T, x 6 / 7, and thus determine the adjustment frequencies with the processing rate ratio x not less than 6 / 7 as candidate adjustment frequencies).

[0087] In some embodiments, the determining the second post-downclock power consumption value corresponding to each candidate adjustment frequency based on the second batch quantity, the third batch quantity, each candidate adjustment frequency and the corresponding processor voltage, and the processing rate ratio corresponding to each candidate adjustment frequency includes: (1.1) Obtain the capacitance load constant; (1.2) Calculate the ratio of the second original processing duration to the processing rate ratio corresponding to each candidate adjustment frequency to obtain the third post-downclock processing duration corresponding to each candidate adjustment frequency; (1.2) Based on the third post-downclock processing duration corresponding to each candidate adjustment frequency, the corresponding processor voltage, the capacitance load constant, and the second batch quantity, determine the first sub-post-downclock power consumption value of each candidate adjustment frequency in the forward propagation stage; (1.4) Based on the second post-downclock processing duration corresponding to each candidate adjustment frequency, the corresponding processor voltage, the capacitance load constant, and the third batch quantity, determine the second sub-post-downclock power consumption value of each candidate adjustment frequency in the backpropagation stage; (1.5) Add the first sub-frequency-down power consumption value and the second sub-frequency-down power consumption value of the same candidate adjustment frequency to obtain the second frequency-down power consumption value corresponding to each candidate adjustment frequency.

[0088] Among them, in order to ensure that the total processing duration before frequency down is not exceeded, it is necessary to screen out the second adjustment frequency with the lowest energy consumption value from the candidate adjustment frequencies. Then, it is necessary to obtain the second frequency-down power consumption value corresponding to each candidate adjustment frequency. First, obtain the capacitance load constant , calculate the ratio of the second original processing duration to the processing rate ratio corresponding to each candidate adjustment frequency, and obtain the third frequency-down processing duration corresponding to each candidate adjustment frequency (for example Figure 7 in the ratio of the second original processing duration T to the processing rate ratio x corresponding to each candidate adjustment frequency, to obtain the third frequency-down processing duration T / x corresponding to each candidate adjustment frequency); according to the power consumption value formula Combine the third frequency-down processing duration, the corresponding processor voltage, the capacitance load constant, and the second batch quantity corresponding to each candidate adjustment frequency to determine the first sub-frequency-down power consumption value of each candidate adjustment frequency in the forward propagation stage (for example Figure 7 in the first sub-frequency-down power consumption value is 1 / ); according to the power consumption value formula Combine the second frequency-down processing duration, the corresponding processor voltage, the capacitance load constant, and the third batch quantity corresponding to each candidate adjustment frequency to determine the second sub-frequency-down power consumption value of each candidate adjustment frequency in the backward propagation stage (for example Figure 7 in the second sub-frequency-down power consumption value is 9 / ); add the first sub-frequency-down power consumption value and the second sub-frequency-down power consumption value of the same candidate adjustment frequency to obtain the second frequency-down power consumption value corresponding to each candidate adjustment frequency (for example Figure 7 in the second frequency-down power consumption value corresponding to each candidate adjustment frequency is the first sub-frequency-down power consumption value 1 / + the second sub-frequency-down power consumption value is 9 / = 19 / ).

[0089] In some embodiments, the second adjustment frequency can also be determined according to the total processing duration of the entire training process. For example Figure 7 in the total processing power consumption value of the entire training process corresponding to each candidate adjustment frequency 19 / + 29 For each candidate adjustment frequency, substitute the processing rate ratio x into this formula to obtain the total processing power consumption value corresponding to the entire training process for each candidate adjustment frequency. Then, select the candidate adjustment frequency with the smallest total processing power consumption value as the second adjustment frequency.

[0090] As can be seen from the above, in the embodiment of the present application, by obtaining the regulation delay of the frequency regulation of any neural network processor required for model training, multiple neural network processors sequentially process multiple batches of tasks; compare the regulation delay with the regulation delay threshold to obtain a comparison result; when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each neural network processor to sequentially process multiple batches of tasks, and reduce the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage; when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, control each neural network processor to sequentially process multiple batches of tasks during the forward propagation stage of the model training process, control each neural network processor to process multiple batches of tasks at intervals during the backward propagation stage, and reduce the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage and the backward propagation stage. Compared with the related art where a single NPU frequency and voltage are set for a computing task, in the embodiment of the present application, by adaptively reducing the frequency according to the regulation delay of the frequency regulation of the neural network processor at different stages of model training, the energy efficiency is improved.

[0091] For the specific implementation of each of the above steps, reference can be made to the previous embodiments, which will not be elaborated here.

[0092] To facilitate better implementation of the model training method provided by the embodiment of the present application, the embodiment of the present application also provides a device based on the above model training method. The meanings of the nouns are the same as those in the above model training method, and the specific implementation details can refer to the description in the method embodiment.

[0093] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of the model training device provided by the embodiment of the present application. The model training device is applied to a computer device. The model training device may include a first acquisition unit 601, a comparison unit 602, a first frequency reduction unit 603, a second frequency reduction unit 604, etc.

[0094] The first acquisition unit 601 is used to obtain the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple neural network processors sequentially process multiple batches of tasks; The comparison unit 602 is used to compare the regulation delay with the regulation delay threshold to obtain a comparison result; The first frequency reduction unit 603 is configured to, when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each of the neural network processors to process multiple batches of tasks in sequence, and reduce the frequency of at least one neural network processor when processing at least one batch of tasks in the forward propagation stage; The second frequency reduction unit 604 is configured to, when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, control each of the neural network processors to process multiple batches of tasks in sequence in the forward propagation stage of the model training process, control each of the neural network processors to process multiple batches of tasks at intervals in the backward propagation stage, and reduce the frequency of at least one neural network processor when processing at least one batch of tasks in the forward propagation stage and the backward propagation stage.

[0095] In some embodiments, the first frequency reduction unit 603 includes: The first determination subunit is configured to determine the first batch quantity of the batch of tasks that need to be frequency-reduced in the forward propagation stage; The first acquisition subunit is configured to acquire multiple adjustment frequencies and the corresponding processor voltages, and acquire the processing rate ratio corresponding to each of the adjustment frequencies; The second determination subunit is configured to determine the first power consumption value after frequency reduction corresponding to each of the adjustment frequencies in the forward propagation stage based on the first batch quantity, each of the adjustment frequencies and the corresponding processor voltages, and the processing rate ratio corresponding to each of the adjustment frequencies; The first screening subunit is configured to screen out the first adjustment frequency corresponding to the smallest first power consumption value after frequency reduction from the multiple adjustment frequencies; The first adjustment subunit is configured to adjust the frequency to the first adjustment frequency before each of the neural network processors processes the batch of tasks that need to be frequency-reduced.

[0096] In some embodiments, the second determination subunit is configured to: Acquire the first original processing duration for processing one batch of tasks in the forward propagation stage and the capacitance load constant; Calculate the ratio of the first original processing duration to the processing rate ratio corresponding to each of the adjustment frequencies to obtain the first processing duration after frequency reduction corresponding to each of the adjustment frequencies; Determine the first power consumption value after frequency reduction corresponding to each of the adjustment frequencies in the forward propagation stage based on the first batch quantity, the capacitance load constant, the first processing duration after frequency reduction, each of the adjustment frequencies, the first processing duration after frequency reduction corresponding to each of the adjustment frequencies, and the processor voltage corresponding to each of the adjustment frequencies.

[0097] In some embodiments, the first acquisition subunit is configured to: Obtain the processing rate corresponding to each adjustment frequency and the processing rate corresponding to the highest frequency for any one of the neural network processors; Determine the ratio of the processing rate corresponding to different adjustment frequencies to the processing rate corresponding to the highest frequency, and obtain the processing rate proportion corresponding to each adjustment frequency.

[0098] In some embodiments, the second frequency reduction unit 604 includes: A third determination subunit, configured to determine the second batch quantity of the batch tasks that need to be frequency-reduced in the forward propagation stage, and determine the third batch quantity of the batch tasks that need to be frequency-reduced in the backward propagation stage; A second acquisition subunit, configured to acquire the total batch quantity of the multiple batch tasks, the second original processing duration for processing one batch task in the forward propagation stage, and the third original processing duration for processing one batch task in the backward propagation stage; A second screening subunit, configured to screen out the target batch quantity with the largest batch quantity from the batch quantities of the batch tasks that need to be frequency-reduced by each neural network processor in the backward propagation stage; A third screening subunit, configured to screen out candidate adjustment frequencies from the multiple adjustment frequencies based on the total batch quantity, the target batch quantity, the second original processing duration, the third original processing duration, and the processing rate proportion corresponding to each adjustment frequency; A fourth determination subunit, configured to determine the second power consumption value after frequency reduction corresponding to each candidate adjustment frequency based on the second batch quantity, the third batch quantity, each candidate adjustment frequency and the corresponding processor voltage, and the processing rate proportion corresponding to each candidate adjustment frequency; A fourth screening subunit, configured to screen out the second adjustment frequency with the smallest second total power consumption value after frequency reduction from the multiple candidate adjustment frequencies; A second adjustment subunit, configured to adjust the frequency to the second adjustment frequency before each neural network processor processes the batch tasks that need to be frequency-reduced.

[0099] In some embodiments, the third screening subunit is configured to: Obtain the product of the total batch quantity and the third original processing duration to obtain a first calculation result; Calculate the sum value of the first calculation result and the second original processing duration to obtain a reference processing duration; Calculate the ratio of the third original processing duration to the processing rate proportion corresponding to each adjustment frequency to obtain the second processing duration after frequency reduction corresponding to each adjustment frequency; Determine the product of the second post-downscaling processing duration corresponding to each of the adjustment frequencies and the target batch quantity to obtain the second calculation result corresponding to each of the adjustment frequencies; Obtain the difference between the total batch quantity and the target batch quantity to obtain the non-downscaled batch quantity; Obtain the product of the non-downscaled batch quantity and the third original processing duration to obtain the third calculation result; Determine the sum value of the second calculation result corresponding to each of the adjustment frequencies and the third calculation result to obtain the total processing duration corresponding to each adjustment frequency in the backpropagation stage; Determine the adjustment frequencies with the total processing duration not exceeding the reference processing duration as the candidate adjustment frequencies.

[0100] In some embodiments, a fourth determination subunit is configured to: Obtain the capacitance load constant; Calculate the ratio of the second original processing duration to the processing rate ratio corresponding to each of the candidate adjustment frequencies to obtain the third post-downscaling processing duration corresponding to each of the candidate adjustment frequencies; Based on the third post-downscaling processing duration corresponding to each of the candidate adjustment frequencies, the corresponding processor voltage, the capacitance load constant, and the second batch quantity, determine the first sub-post-downscaling power consumption value of each of the candidate adjustment frequencies in the forward propagation stage; Based on the second post-downscaling processing duration corresponding to each of the candidate adjustment frequencies, the corresponding processor voltage, the capacitance load constant, and the third batch quantity, determine the second sub-post-downscaling power consumption value of each of the candidate adjustment frequencies in the backpropagation stage; Add the first sub-post-downscaling power consumption value and the second sub-post-downscaling power consumption value of the same candidate adjustment frequency to obtain the second post-downscaling power consumption value corresponding to each of the candidate adjustment frequencies.

[0101] For the specific implementation of each of the above units, reference may be made to the previous embodiments and will not be elaborated here.

[0102] As can be seen from the above, in the embodiment of the present application, the first acquisition unit 601 acquires the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple neural network processors sequentially process multiple batches of tasks; the comparison unit 602 compares the regulation delay with the regulation delay threshold to obtain a comparison result; when the comparison result indicates that the regulation delay is greater than the regulation delay threshold, the first frequency reduction unit 603 controls each neural network processor to sequentially process multiple batches of tasks, and reduces the frequency of at least one neural network processor when processing at least one batch of tasks in the forward propagation stage; when the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, the second frequency reduction unit 604 controls each neural network processor to sequentially process multiple batches of tasks in the forward propagation stage of the model training process, controls each neural network processor to process multiple batches of tasks at intervals in the backward propagation stage, and reduces the frequency when at least one neural network processor processes at least one batch of tasks in the forward propagation stage and the backward propagation stage. Compared with the related art where a single NPU frequency and voltage are set for one computing task, the embodiment of the present application adaptively reduces the frequency at different stages of model training according to the regulation delay of the frequency regulation of the neural network processor, thereby improving energy efficiency.

[0103] For the specific implementation of each of the above units, reference may be made to the previous embodiments, which will not be elaborated here.

[0104] Refer to Figure 9 , Figure 9 FIG. is a block diagram of a part of a computer device 1000 for implementing the embodiment of the present disclosure. The computer device 1000 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 622 (for example, one or more processors) and a memory 632, and one or more storage media 630 (for example, one or more mass storage devices) for storing application programs 642 or data 644. Among them, the memory 632 and the storage media 630 may be transient storage or persistent storage. The program stored in the storage media 630 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device 1000. Further, the central processing unit 622 may be configured to communicate with the storage media 630 and execute a series of instruction operations in the storage media 630 on the computer device 1000.

[0105] The computer device 1000 may further include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0106] The central processing unit 622 in the computer device 1000 can be used to execute the model training method of the embodiments of the present disclosure. For example: Obtain the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple said neural network processors sequentially process multiple batches of tasks; Compare the regulation delay with a regulation delay threshold to obtain a comparison result; When the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each of the neural network processors to sequentially process multiple batches of tasks, and reduce the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage; When the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, control each of the neural network processors to sequentially process multiple batches of tasks during the forward propagation stage of the model training process, control each of the neural network processors to process multiple batches of tasks at intervals during the backward propagation stage, and reduce the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage and the backward propagation stage.

[0107] The embodiments of the present disclosure further provide a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the model training methods of the foregoing various embodiments.

[0108] The embodiments of the present disclosure further provide a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes to implement the above-mentioned model training method. For example: Obtain the regulation delay of the frequency regulation of any neural network processor required for model training, and multiple said neural network processors sequentially process multiple batches of tasks; Compare the regulation delay with a regulation delay threshold to obtain a comparison result; When the comparison result indicates that the regulation delay is greater than the regulation delay threshold, control each of the neural network processors to sequentially process multiple batches of tasks, and reduce the frequency of at least one neural network processor processing at least one batch of tasks during the forward propagation stage; When the comparison result indicates that the regulation delay is less than or equal to the regulation delay threshold, during the forward propagation stage of the model training process, control each of the neural network processors to sequentially process multiple batches of tasks, during the backward propagation stage, control each of the neural network processors to process multiple batches of tasks at intervals, and downscale at least one neural network processor when processing at least one batch of tasks during the forward propagation stage and the backward propagation stage.

[0109] In addition, the terms "including" and "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0110] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0111] It should be understood that in the description of the embodiments of this application, the meaning of "multiple (or multiple items)" is more than two. Understandings such as "greater than", "less than", and "exceeding" do not include the present number, and understandings such as "above", "below", and "within" include the present number.

[0112] In several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, indirect coupling or communication connection of devices or units, and can be in electrical, mechanical or other forms.

[0113] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0114] In addition, in various embodiments of the present application, each functional unit may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0115] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs, etc., which can store program codes.

[0116] It should also be understood that the various embodiments provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.

[0117] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of this module or unit.

[0118] The above is a specific description of the embodiments of the present application, but the present application is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.

Claims

1. A model training method, characterized in that: include: Obtaining a control delay of frequency control of any neural network processor required for model training, wherein multiple neural network processors sequentially process multiple batches of tasks; Comparing the control delay with a control delay threshold to obtain a comparison result; When the comparison result indicates that the control delay is greater than the control delay threshold, controlling each of the neural network processors to process multiple batches of tasks in sequence, and reducing the frequency of at least one neural network processor processing at least one batch of tasks in the forward propagation stage; When the comparison result indicates that the control delay is less than or equal to the control delay threshold, each of the neural network processors is controlled to process multiple batches of tasks in sequence during the forward propagation stage of the model training process, and each of the neural network processors is controlled to process multiple batches of tasks at intervals during the backward propagation stage, and at least one neural network processor is frequency-reduced when processing at least one batch of tasks during the forward propagation stage and the backward propagation stage.

2. The model training method according to claim 1, characterized in that: The step of downclocking at least one batch of tasks processed by at least one neural network processor during the forward propagation phase includes: Determine the first batch number of batch tasks that need to be down-sampled in the forward propagation phase; Obtaining multiple adjustment frequencies and corresponding processor voltages, and obtaining a processing rate ratio corresponding to each of the adjustment frequencies; Determine a first post-frequency reduction power consumption value corresponding to each of the adjustment frequencies in the forward propagation phase based on the first batch quantity, each of the adjustment frequencies and the corresponding processor voltage, and the processing rate ratio corresponding to each of the adjustment frequencies; Filtering out, from the plurality of adjustment frequencies, a first adjustment frequency with the smallest power consumption value after the first frequency reduction; Before each neural network processor processes a batch of tasks that require frequency reduction, the frequency is adjusted to the first adjustment frequency.

3. The model training method according to claim 2, characterized in that: The determining, based on the first batch quantity, each of the adjustment frequencies and the corresponding processor voltage, and the processing rate ratio corresponding to each of the adjustment frequencies, a first power consumption value after frequency reduction corresponding to each of the adjustment frequencies in the forward propagation phase includes: Obtain the first original processing time of a batch of tasks in the forward propagation phase, and the capacitive load constant; Calculating a ratio of the first original processing duration to a processing rate ratio corresponding to each of the adjusted frequencies to obtain a first post-frequency reduction processing duration corresponding to each of the adjusted frequencies; Based on the first batch quantity, the capacitive load constant, the first post-frequency reduction processing time, each of the adjustment frequencies, the first post-frequency reduction processing time corresponding to each of the adjustment frequencies, and the processor voltage corresponding to each of the adjustment frequencies, determine the first post-frequency reduction power consumption value corresponding to each of the adjustment frequencies in the forward propagation stage.

4. The model training method according to claim 2, characterized in that: The obtaining of the processing rate proportion corresponding to each of the adjustment frequencies includes: Obtaining a processing rate corresponding to each adjustment frequency and a processing rate corresponding to the highest frequency of any of the neural network processors; The ratio of the processing rates corresponding to the different adjustment frequencies to the processing rate corresponding to the highest frequency is determined to obtain the processing rate ratio corresponding to each adjustment frequency.

5. The model training method according to claim 1, characterized in that: The step of reducing the frequency of at least one neural network processor when processing at least one batch of tasks in the forward propagation stage and the back propagation stage includes: Determine the second batch number of batch tasks that need to be frequency-reduced in the forward propagation phase, and determine the third batch number of batch tasks that need to be frequency-reduced in the backward propagation phase; Obtaining a total batch number of the plurality of batch tasks, a second original processing time for processing a batch task in the forward propagation phase, and a third original processing time for processing a batch task in the backward propagation phase; Filter out the target batch number with the largest batch number from the batch numbers of the batch tasks that need to be down-converted in the back-propagation phase of each neural network processor; Based on the total batch quantity, the target batch quantity, the second original processing time, the third original processing time, and the processing rate ratio corresponding to each adjustment frequency, a candidate adjustment frequency is screened out from the plurality of adjustment frequencies; Determine a second post-frequency reduction power consumption value corresponding to each candidate adjustment frequency based on the second batch quantity, the third batch quantity, each candidate adjustment frequency and the corresponding processor voltage, and the processing rate ratio corresponding to each candidate adjustment frequency; Filtering out, from the plurality of candidate adjustment frequencies, a second adjustment frequency with the smallest second total power consumption value after frequency reduction; Before each neural network processor processes a batch of tasks that require frequency reduction, the frequency is adjusted to the second adjustment frequency.

6. The model training method according to claim 5, characterized in that: The selecting a candidate adjustment frequency from the plurality of adjustment frequencies based on the total batch quantity, the target batch quantity, the second original processing time, the third original processing time, and a processing rate ratio corresponding to each adjustment frequency includes: Obtaining the product of the total batch quantity and the third original processing time to obtain a first calculation result; Calculating the sum of the first calculation result and the second original processing duration to obtain a reference processing duration; Calculate the ratio of the third original processing duration to the processing rate proportion corresponding to each of the adjusted frequencies to obtain the second post-frequency reduction processing duration corresponding to each of the adjusted frequencies; Determine the product of the second post-frequency reduction processing time corresponding to each of the adjustment frequencies and the target batch quantity to obtain a second calculation result corresponding to each of the adjustment frequencies; Obtaining the difference between the total batch quantity and the target batch quantity to obtain the number of batches that are not down-converted; Obtaining the product of the number of non-downclocked batches and the third original processing time to obtain a third calculation result; Determine the sum of the second calculation result and the third calculation result corresponding to each of the adjustment frequencies to obtain a total processing time corresponding to each adjustment frequency in the backward propagation stage; The adjustment frequencies whose total processing duration does not exceed the reference processing duration are determined as candidate adjustment frequencies.

7. The model training method according to claim 6, characterized in that: The determining, based on the second batch number, the third batch number, each of the candidate adjustment frequencies and the corresponding processor voltage, and the processing rate ratio corresponding to each of the candidate adjustment frequencies, a second post-frequency reduction power consumption value corresponding to each of the candidate adjustment frequencies includes: Get the capacitive load constant; Calculate the ratio of the second original processing duration to the processing rate proportion corresponding to each of the candidate adjustment frequencies to obtain a third post-frequency reduction processing duration corresponding to each of the candidate adjustment frequencies; Determine a first sub-frequency-reduced power consumption value of each candidate adjustment frequency in a forward propagation phase based on a third frequency-reduced processing duration corresponding to each candidate adjustment frequency, a corresponding processor voltage, the capacitive load constant, and the second batch quantity; Determine a second sub-frequency-reduced power consumption value of each candidate adjustment frequency in a backward propagation phase based on a second post-frequency-reduction processing duration corresponding to each candidate adjustment frequency, a corresponding processor voltage, the capacitive load constant, and the third batch quantity; The first sub-frequency-reduced power consumption value and the second sub-frequency-reduced power consumption value of the same candidate adjustment frequency are added together to obtain the second frequency-reduced power consumption value corresponding to each candidate adjustment frequency.

8. A model training device, characterized in that: include: A first acquisition unit is used to acquire the control delay of the frequency control of any neural network processor required for model training, wherein multiple batches of tasks are processed sequentially between multiple neural network processors; A comparison unit, used for comparing the control delay with a control delay threshold to obtain a comparison result; A first frequency reduction unit, configured to control each of the neural network processors to process multiple batches of tasks in sequence when the comparison result indicates that the control delay is greater than the control delay threshold, and to reduce the frequency of at least one neural network processor processing at least one batch of tasks in a forward propagation phase; The second frequency reduction unit is used to control each of the neural network processors to process multiple batches of tasks in sequence during the forward propagation stage of the model training process when the comparison result indicates that the control delay is less than or equal to the control delay threshold, and to control each of the neural network processors to process multiple batches of tasks at intervals during the backward propagation stage, and to reduce the frequency of at least one neural network processor when processing at least one batch of tasks during the forward propagation stage and the backward propagation stage.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, which are suitable for loading by a processor to execute the model training method described in any one of claims 1 to 7.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the model training method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • NPU power consumption optimization system and method based on neural network structure

    CN114217688A

  • Frequency modulation method and related equipment

    CN118202714A

  • Frequency modulation method, device and equipment and readable storage medium

    CN118819860A

  • Method and apparatus with neural network device performance and power efficiency prediction

    US20250068938A1