Frequency Adjustment Method, Device, Storage Medium and Computer Equipment
By adjusting the processor frequency in the artificial intelligence computing cluster, the load imbalance problem caused by chip heterogeneity is solved, and energy saving and energy efficiency are achieved without affecting the computing progress.
Patent Information
- Application Number
- CN202510601521.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-12
AI Technical Summary
In artificial intelligence computing clusters, load imbalance between computing nodes caused by chip manufacturing heterogeneity slows down the overall computing progress and causes energy waste.
By obtaining the historical calculation time of each artificial intelligence processor, filter out the target calculation time with the largest duration, and adjust the frequency of other processors according to performance indicators to align their calculation time with the target duration, reducing overall energy consumption.
Without affecting the overall calculation progress, energy saving, calculation efficiency and energy consumption are reduced.
Smart Images

Figure CN120122800B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a frequency adjustment method, device, storage medium, and computer device. Background Art
[0002] Today, with the rapid development of artificial intelligence (AI) technology, intelligent computing clusters (hereinafter referred to as intelligent computing clusters) have become the key pillars for promoting AI scientific research and industrial applications. Intelligent computing clusters are built by adopting advanced AI processors, including graphics processing units (GPUs), neural network processing units (NPUs), etc., and can efficiently process complex AI tasks and the computing requirements of large-scale data. However, the popularization and wide application of intelligent computing clusters have also brought severe energy consumption challenges, which pose a huge pressure on global energy use and environmental protection.
[0003] In related technologies, even in chips with the same model, architecture, and completely identical microcircuit design, due to factors such as material errors, dimensional errors, and process fluctuations during the manufacturing process, there will be differences in the performance, power consumption, etc. of the chips. Although this heterogeneity difference can be ignored in personal electronic devices with a single or a small number of chips, it will have a significant impact in large-scale computing clusters.
[0004] In parallel computing tasks, if these performance differences are ignored, it will lead to uneven load among computing nodes, slow down the overall computing progress, and thus cause energy waste. Therefore, related technologies urgently need to propose a frequency adjustment method to solve the above technical problems. Summary of the Invention
[0005] The main purpose of the present application is to provide a frequency adjustment method, device, storage medium, and computer device, which can save energy and improve energy efficiency without affecting the overall computing progress.
[0006] In a first aspect, an embodiment of the present application provides a frequency adjustment method, including:
[0007] Obtain the first historical computing duration of each artificial intelligence processor during the computing stage of model training;
[0008] Screen out the target computing duration with the largest duration from the multiple first historical computing durations;
[0009] Determine the artificial intelligence processors other than the artificial intelligence processor corresponding to the target computing duration among the multiple artificial intelligence processors as other artificial intelligence processors;
[0010] Determine the ratio of the first historical computing duration of each of the other artificial intelligence processors to the target computing duration to obtain the first performance index of each of the other artificial intelligence processors;
[0011] Substitute the first performance metric of each of the other artificial intelligence processors into the corresponding frequency function to obtain the first adjusted frequency corresponding to each of the other artificial intelligence processors;
[0012] Perform frequency adjustment according to the first adjusted frequency corresponding to each of the other artificial intelligence processors.
[0013] In a second aspect, an embodiment of the present application provides a frequency adjustment device, including:
[0014] A first acquisition unit, configured to acquire the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training;
[0015] A screening unit, configured to screen out the target calculation duration with the maximum duration from the multiple first historical calculation durations;
[0016] A first determination unit, configured to determine, as other artificial intelligence processors, the artificial intelligence processors among the multiple artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration;
[0017] A second determination unit, configured to determine the ratio of the first historical calculation duration of each of the other artificial intelligence processors to the target calculation duration to obtain the first performance metric of each of the other artificial intelligence processors;
[0018] A substitution unit, configured to substitute the first performance metric of each of the other artificial intelligence processors into the corresponding frequency function to obtain the first adjusted frequency corresponding to each of the other artificial intelligence processors;
[0019] An adjustment unit, configured to perform frequency adjustment according to the first adjusted frequency corresponding to each of the other artificial intelligence processors.
[0020] In a third aspect, an embodiment of the present application provides a storage medium. The computer-readable storage medium stores multiple instructions, and these instructions are suitable for being loaded by a processor to execute the frequency adjustment method as described in any one of the above.
[0021] In a fourth aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the frequency adjustment method as described in any one of the above is implemented.
[0022] In the embodiments of the present application, by obtaining the first historical computing duration of each artificial intelligence processor during the computing stage of model training; screening out the target computing duration with the maximum duration from the multiple first historical computing durations; determining, as other artificial intelligence processors, the artificial intelligence processors among the multiple artificial intelligence processors except the artificial intelligence processor corresponding to the target computing duration; determining the ratio of the first historical computing duration of each other artificial intelligence processor to the target computing duration to obtain the first performance index of each other artificial intelligence processor; substituting the first performance index of each other artificial intelligence processor into the corresponding frequency function to obtain the first adjustment frequency corresponding to each other artificial intelligence processor; and performing frequency adjustment according to the first adjustment frequency corresponding to each other artificial intelligence processor, compared with the related art, where the load imbalance between computing nodes caused by the heterogeneity of artificial intelligence processors slows down the overall computing progress and thus causes energy waste, it is possible to save energy and improve energy efficiency without affecting the overall computing progress.
[0023] Other features and advantages of the present disclosure will be described in the following specification, and will, in part, be obvious from the specification, or can be learned by implementing the present disclosure. The objectives and other advantages of the present disclosure can be achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0025] Figure 1 It is a schematic diagram of the model training process in the ideal state provided by the embodiments of the present application.
[0026] Figure 2 It is a schematic diagram of the model training process in the actual state provided by the embodiments of the present application.
[0027] Figure 3 It is a schematic diagram of the scenario of the frequency adjustment system provided by the embodiments of the present application.
[0028] Figure 4 It is a schematic flowchart of the frequency adjustment method provided by the embodiments of the present application.
[0029] Figure 5 It is a schematic diagram of using frequency adjustment to reduce the power consumption of model training provided by the embodiments of the present application.
[0030] Figure 6 This is a schematic structural diagram of the frequency adjustment device provided by the embodiment of the present application.
[0031] Figure 7 This is a schematic structural diagram of the computer device provided by the embodiment of the present application. Detailed implementation manners
[0032] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0033] It should be noted that in some processes described in the specification, claims and the above-mentioned drawings, there are multiple steps that appear in a specific order. However, it should be clearly understood that these steps may not be executed in the order in which they appear in this document or may be executed in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, descriptions such as "first", "second" or "target" in this document are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0034] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations:
[0035] Artificial Intelligence Processor: Also known as an artificial intelligence accelerator, it is high-performance hardware designed specifically for artificial intelligence (AI) applications, especially deep learning, machine learning, and neural networks.
[0036] Unique design: Adopts a multi-core architecture with numerous processing cores, which can divide matrix operations in deep learning into multiple small tasks and execute them in parallel on different cores; in addition to traditional CPU or GPU cores, it may include dedicated hardware accelerators, such as tensor processing units (TPUs) or neural processing units (NPUs), to optimize common operations in deep learning such as matrix multiplication and convolution; is equipped with a high-speed memory and cache system, such as high-bandwidth memory (HBM) or on-chip SRAM cache, to reduce data transfer bottlenecks; adopts advanced interconnection technologies, such as mesh interconnection or crossbar switches, to ensure efficient communication and collaboration among different parts; to cope with high power consumption, it includes a complex power management system that uses technologies such as dynamic voltage and frequency adjustment, load balancing, and thermal throttling to maximize energy efficiency.
[0037] Operating principle: First, preprocess and partition the input data to make it adapt to the parallel architecture of the processor; then allocate the data and operations to multiple processing cores or dedicated accelerators; then each core or accelerator accesses the data using high-speed memory while executing the allocated tasks; then collect and integrate the processed data results to form the final output; finally, perform iterative adjustment according to the model feedback to optimize the model performance.
[0038] Application scenarios: Can quickly train large-scale neural network models; accelerate the model inference process in actual applications and provide prediction results in real time; used for big data analysis, such as playing a role in fields such as image recognition and natural language processing.
[0039] In the data parallel training of large AI models, each training cycle usually consists of two stages: computing and communication. The computing part is carried out independently on each AI processor (such as NPU), while the communication part cannot start until all AI processors have completed the computing. Therefore, the reasonable arrangement of computing and communication is crucial for the overall training efficiency and energy efficiency.
[0040] The ideal parallel training arrangement is as Figure 1 shown, Figure 1 which is a schematic diagram of the model training process in the ideal state provided by the embodiment of this application. The computing tasks are carried out in parallel on 4 NPU processors, and after all processors have completed the computing, they enter the communication stage and repeat periodically.
[0041] However, due to the existence of chip manufacturing heterogeneity, the actual situation is not like this. There are differences in the computing performance of different chips, resulting in inconsistent computing completion times, as specifically shown in Figure 2 shown, Figure 2 which is a schematic diagram of the model training process in the actual state provided by the embodiment of this application. In this figure, NPU1 is the chip with the slowest computing speed, and its delay slows down the start time of the overall communication stage. Other processors (such as NPU2, NPU3, and NPU4) must wait for the result of NPU1 after they have completed the computing. Although there is no effective computing task during this waiting time, a large amount of energy is still consumed, resulting in energy consumption waste.
[0042] In order to solve the above problems, embodiments of the present application obtain the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training; screen out the target calculation duration with the longest duration from the multiple first historical calculation durations; determine the artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration among the multiple artificial intelligence processors as other artificial intelligence processors; determine the ratio of the first historical calculation duration of each other artificial intelligence processor to the target calculation duration to obtain the first performance index of each other artificial intelligence processor; substitute the first performance index of each other artificial intelligence processor into the corresponding frequency function to obtain the first adjustment frequency corresponding to each other artificial intelligence processor; perform frequency adjustment according to the first adjustment frequency corresponding to each other artificial intelligence processor. Compared with the related art, due to the heterogeneity of artificial intelligence processors, the load imbalance between computing nodes slows down the overall computing progress, resulting in energy waste. It is possible to save energy and improve energy efficiency without affecting the overall computing progress. For specific details, please continue to refer to the following specific embodiments.
[0043] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the scenario of the frequency adjustment system provided by the embodiments of the present application. It includes a terminal 140, the Internet 130, a gateway 120, a server 110, etc.
[0044] The terminal 140 includes, but is not limited to, a pre-configured laptop computer, a tablet computer, a desktop computer, or other electronic devices with data reporting capabilities. In addition, it can be a single device or a collection of multiple devices. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.
[0045] The terminal 140 refers to a computer system that can report data to the server 110. Compared with ordinary terminals, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, or a combination of parts (such as virtual machines) allocated from multiple high-performance computers.
[0046] The gateway 120 is also known as an internetwork connector or protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. Between two systems using different communication protocols, data formats, or languages, or even with completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. Messages sent from the terminal 140 to the server 110 need to be sent to the corresponding server 110 through the gateway 120. Messages sent from the server 110 to the terminal 140 also need to be sent to the corresponding terminal 140 through the gateway 120.
[0047] The frequency adjustment method of the embodiments of the present disclosure can be implemented in the server 110.
[0048] It should be noted that Figure 3 The scenario schematic diagram of the frequency adjustment system shown is only an example. The frequency adjustment system and scenario described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of image processing technology and the emergence of new service scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0049] In this embodiment, the description will be made from the perspective of the frequency adjustment device, which can be specifically integrated in a computer device with a storage unit and installed with a microprocessor and having computing capabilities.
[0050] Please refer to Figure 4 , Figure 4 which is the flowchart of the frequency adjustment method provided by the embodiments of the present application. The frequency adjustment method includes:
[0051] In step 201, obtain the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training.
[0052] Among them, in the training of the artificial intelligence model, multiple stages will be experienced. The calculation stage of the model training process mainly refers to the process of performing mathematical operations such as matrix multiplication, convolution operations, and activation function calculations, which are used to update the parameters of the model. The first historical calculation duration is the time spent by the artificial intelligence processor in the past during the calculation stage of model training.
[0053] Specifically, through the log recording system or monitoring tool, collect the time data of each artificial intelligence processor during the calculation stage of multiple previous model trainings. For example, in a deep learning training platform, after each training task is completed by each processor, the time spent in the calculation stage is recorded in a log file. The system regularly extracts the required data from these log files as the first historical calculation duration.
[0054] The purpose of this step is to provide basic data for subsequent frequency adjustment, so as to understand the performance of each processor during the calculation stage.
[0055] In some embodiments, obtaining the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training includes:
[0056] (1) Obtaining the actual calculation duration of each artificial intelligence processor during each calculation stage in the historical model training process;
[0057] (2) Determining the sum value of the multiple actual calculation durations of each artificial intelligence processor to obtain the first total calculation duration;
[0058] (3) Determining the ratio of the first total calculation duration of each artificial intelligence processor to the number of stages of the calculation stage, to obtain the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training.
[0059] Among them, a dedicated hardware monitoring tool or system-level monitoring software is used to collect the calculation time of the processor. These tools can monitor the running state of the processor in real time, and automatically record the corresponding time information when detecting the start and end events of the calculation stage. For example, some server management systems can provide monitoring of processor performance metrics, including the duration information of the calculation stage. Obtaining the actual calculation duration of each calculation stage is the basis for subsequent calculations. By recording the actual calculation duration of each stage in detail, the performance of the processor under different calculation tasks can be accurately reflected, providing accurate data for subsequent calculation of the average duration. Different calculation stages may have different complexities and calculation amounts, and recording the duration of each stage helps to analyze the performance of the processor more meticulously.
[0060] Specifically, read all the actual calculation duration data of each processor from the stored log files or monitoring data. Calculate the first total calculation duration The overall calculation time consumption of the processor in multiple calculation stages can be comprehensively considered. Different calculation stages may have different durations due to factors such as task complexity and data volume. By summing up to obtain the total duration, the calculation ability and time overhead of the processor during the entire historical model training process can be understood macroscopically. Given the first total calculation duration of each processor and the number of calculation stages n, the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training can be determined Calculating the average value of the first historical calculation duration can eliminate the influence of duration fluctuations in different calculation stages and obtain a more representative processor calculation duration indicator. This indicator can more accurately reflect the average performance of the processor during the model training calculation stage, providing a reliable reference basis for subsequent optimization operations such as frequency adjustment. Through the average duration, the calculation performance between different processors can be compared more fairly, and then more reasonable resource scheduling and optimization can be carried out.
[0061] In some embodiments, obtaining the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training includes:
[0062] (1) Obtaining the actual calculation duration of each artificial intelligence processor during each calculation stage in the historical model training process;
[0063] (2) For each artificial intelligence processor, obtaining the calculation busy degree of each calculation stage;
[0064] (3) Calculating the product of each calculation busy degree and the actual calculation duration of the corresponding calculation stage to obtain the calculation result of each calculation stage;
[0065] (4) Determining the sum value of the calculation results of multiple calculation stages to obtain the second total calculation duration;
[0066] (5) Determining the ratio of the second total calculation duration to the number of stages of the calculation stage to obtain the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training.
[0067] Among them, since different calculation stages may have different task complexities, the calculation duration is affected by the task complexity. Considering the calculation busy degree can more accurately evaluate the actual workload of the processor in each calculation stage. Therefore, it can be quantified according to influencing factors such as the task type or complexity executed in the calculation stage. The calculation busy degree can be expressed in the form of a percentage, such as 30%.
[0068] Specifically, by multiplying the calculation busy degree by the actual calculation duration, the duration and resource occupancy situation of the calculation stage are comprehensively considered. The calculation result obtained in this way can more accurately reflect the actual impact of this calculation stage on the overall performance of the processor, providing a more reasonable basis for subsequent calculation of the total duration. Calculating the average value can eliminate the differences between different calculation stages and obtain a more representative processor calculation duration indicator. This indicator considers the calculation busy degree and can more accurately reflect the average performance of the processor during the model training calculation stage, providing a reliable reference for subsequent optimization operations such as frequency adjustment and resource scheduling.
[0069] For example, there is an AI processor A. The actual computing duration of computing stage 1 is 10 minutes, the actual computing duration of computing stage 2 is 15 minutes, and the actual computing duration of computing stage 3 is 8 minutes. The busy degree of computing stage 1 is 80%, the busy degree of computing stage 2 is 60%, and the busy degree of computing stage 3 is 70%. Then the computing result of computing stage 1 is 10 * 80% = 8 minutes, the computing result of computing stage 2 is 15 * 60% = 9 minutes, and the computing result of computing stage 3 is 8 * 70% = 5.6 minutes. The second total computing duration is 8 + 9 + 5.6 = 22.6 minutes. The first historical computing duration of the computing stage in the model training process of AI processor A is 22.6 / 3 = 7.53 minutes.
[0070] In step 202, the target computing duration with the longest duration is screened out from the multiple first historical computing durations.
[0071] Among them, the target computing duration is the time value with the longest duration among all the first historical computing durations. The target computing duration is used as the benchmark duration, and subsequent adjustments are made to the frequencies of other processors based on it to align the computing times of each processor as much as possible and reduce the waiting time.
[0072] For example Figure 2 As shown, since the first historical computing duration of computing stage NPU1 is the longest among NPU1, NPU2, NPU3, and NPU4, the computing duration of NPU1 is also the longest in the actual computing process, that is Figure 2 As shown. To avoid increasing the original computing duration during the process of adjusting the frequency, the first historical computing duration of the largest NPU1 is used as the benchmark duration, and the frequencies of NPU2, NPU3, and NPU4 are adjusted to avoid the computing duration after adjustment exceeding the target computing duration.
[0073] In step 203, the AI processors other than the AI processor corresponding to the target computing duration among the multiple AI processors are determined as other AI processors.
[0074] Among them, other AI processors are the remaining processors among all the processors performing model training except the one with the longest computing duration (i.e., the one corresponding to the target computing duration). By traversing the dataset recording the processors and their corresponding first historical computing durations, the processor corresponding to the target computing duration is found and excluded, and the remaining processors are the other AI processors.
[0075] The purpose of this step is to clarify the range of processors that need to have their frequencies adjusted. Since the processor corresponding to the target computing duration is already in the state of having the longest computing time and does not need to be adjusted, while other processors need to be adjusted to optimize the overall performance.
[0076] In step 204, determine the ratio of the first historical computing duration of each of the other artificial intelligence processors to the target computing duration, and obtain the first performance metric of each of the other artificial intelligence processors.
[0077] The first performance metric is a quantitative value used to represent the relative level of the computing performance of the other artificial intelligence processors relative to the processor corresponding to the target computing duration, and is obtained by the ratio of the computing durations of the two.
[0078] Specifically, for the other artificial intelligence processor i, use its first historical computing duration divided by the target computing duration to obtain the first performance metric as .
[0079] In this way, by quantifying the performance difference between the other artificial intelligence processors and the artificial intelligence processor corresponding to the target computing duration, it provides a basis for subsequent calculation of the adjustment frequency, facilitating targeted frequency adjustment according to the performance difference.
[0080] In step 205, substitute the first performance metric of each of the other artificial intelligence processors into the corresponding frequency function to obtain the first adjustment frequency corresponding to each of the other artificial intelligence processors.
[0081] The frequency function is a pre-established mathematical relationship model that describes the corresponding relationship between the performance metric and the processor adjustment frequency. The first adjustment frequency is calculated through the frequency function based on the first performance metric and is a value used to adjust the working frequency of the other artificial intelligence processors.
[0082] In this way, a suitable adjustment frequency is calculated according to the processor performance difference, and by adjusting the frequency, the computing time of the other processors is made close to that of the processor corresponding to the target computing duration, improving the overall computing efficiency.
[0083] In step 206, perform frequency adjustment according to the first adjustment frequency corresponding to each of the other artificial intelligence processors.
[0084] Frequency adjustment refers to changing the working frequency of the artificial intelligence processor through hardware control instructions or software system settings, thereby adjusting its computing duration. Make the computing speed of the other artificial intelligence processors the same as the target computing duration of the target processor, reduce the computing waiting time caused by the processor performance difference, improve the efficiency of the entire computing system in the model training calculation stage, and will not increase the computing duration of the entire computing process, and reduce the energy consumption.
[0085] Please refer to Figure 5 , Figure 5Schematic diagram of reducing model training power consumption through frequency adjustment provided by embodiments of this application. For the i-th AI processor, the computing duration thereon is . Therefore, when using M AI processors for parallel training, the computing duration of each cycle is determined by the processor with the longest duration, which is , where is the maximum value among all , = max( , , …, ).
[0086] To achieve time alignment, for each processor i, it is necessary to reduce its frequency so that its computing duration is adjusted from to . At this time, the normalized performance of this processor changes from 1 to . Thus, the computing durations of different processors are aligned, as shown in Figure 5 . Since the power also decreases during the process of reducing the frequency of the processor, therefore, the computing duration after optimization remains unchanged, but the total energy consumption will decrease.
[0087] Through the frequency function, the corresponding power after frequency reduction can be calculated as:
[0088] ;
[0089] Among them, , , and are fitting coefficients. Multiply the power by the duration T_max to obtain the power consumption of the i-th processor in the entire segment of computing:
[0090] ;
[0091] Then sum up all M processors to obtain the total energy consumption of all processors in the entire segment of computing:
[0092] ;
[0093] Through the above method, this patent effectively solves the problem of uneven computing caused by chip manufacturing heterogeneity, realizes the precise alignment of the computing time of AI processors, and thus significantly reduces the overall energy consumption on the premise of ensuring the computing efficiency of the system. The optimized energy consumption value can be calculated using the provided formula.
[0094] In some embodiments, before obtaining the first historical computing duration of each artificial intelligence processor in the computing stage of model training, it further includes:
[0095] (1)Obtain the second historical computing duration of multiple candidate artificial intelligence processors, and obtain the number of processors required for model training;
[0096] (2)Sort the multiple candidate artificial intelligence processors in ascending order of the second historical computing duration to obtain a processor sequence;
[0097] (3)Based on the number of processors and the processor sequence, determine the processor combinations and the corresponding total computing power consumption values corresponding to each first artificial intelligence processor whose sorting serial number is not less than the number of processors from the processor sequence;
[0098] (4)From the multiple first artificial intelligence processors, screen out the target first artificial intelligence processor with the smallest corresponding total computing power consumption value;
[0099] (5)Determine each artificial intelligence processor in the processor combination corresponding to the target first artificial intelligence processor as the artificial intelligence processor required in the model training process.
[0100] Among them, before model training is performed by multiple artificial intelligence processors, a group of processor combinations need to be selected from multiple candidate artificial intelligence processors. The purpose of the selection is to ensure the lowest power consumption value in the entire model training process. Based on this purpose, first obtain the second historical computing duration of multiple candidate artificial intelligence processors, and obtain the number of processors required for this model training. The second historical computing duration is the computing duration data of the candidate processors in the computing stage of past model training, and is used to evaluate their computing speed. The processor sequence is an ordered list formed by sorting the candidate artificial intelligence processors in ascending order of the second historical computing duration. Each element is a candidate processor, and its order reflects the relative magnitude of the computing duration. Based on the number of processors and the processor sequence, determine the processor combinations and the corresponding total computing power consumption values corresponding to each first artificial intelligence processor whose sorting serial number is not less than the number of processors from the processor sequence. After obtaining the total computing power consumption value of each first artificial intelligence processor, screen out the target first artificial intelligence processor with the smallest corresponding total computing power consumption value, and determine each artificial intelligence processor in the processor combination corresponding to the target first artificial intelligence processor as the artificial intelligence processor required in the model training process. It shows that each artificial intelligence processor in the processor combination corresponding to the target first artificial intelligence processor can achieve the smallest total computing power consumption value as the processor required for this model training.
[0101] Specifically, the first artificial intelligence processor refers to a processor selected starting from a position in the processor sequence where the sorting serial number is not less than the number of processors. The processor combinations are different combination forms composed of these candidate artificial intelligence processors. Each combination may be used as a processor configuration scheme for model training, and the number of processors in each processor combination is the number of artificial intelligence processors required for this model training. The total computing power consumption value is the total power value expected to be consumed by each processor combination when performing the model training task, and is used to measure the energy consumption of different combinations.
[0102] In some embodiments, determining, based on the number of processors and the processor sequence, each processor combination corresponding to the first artificial intelligence processor whose sorting serial number is not less than the number of processors and the corresponding total computing power consumption value from the processor sequence includes:
[0103] (1.1) Screening out the first artificial intelligence processors from the processor sequence whose sorting serial numbers are the same as the number of processors;
[0104] (1.2) Obtaining the target second historical computing duration of the first artificial intelligence processor;
[0105] (1.3) Determining the ratio of the second historical computing duration of each second artificial intelligence processor before the first artificial intelligence processor to the target second historical computing duration to obtain a second performance index;
[0106] (1.4) Substituting the second performance index of each second artificial intelligence processor into the corresponding frequency function to obtain the second adjusted frequency corresponding to each second artificial intelligence processor;
[0107] (1.5) Calculating the product of the second adjusted frequency of each second artificial intelligence processor and the original frequency of the first artificial intelligence processor and the target second historical computing duration to obtain the computing power consumption value of each second artificial intelligence processor and the first artificial intelligence processor;
[0108] (1.6) Based on the computing power consumption values of each second artificial intelligence processor and the first artificial intelligence processor, determining the processor combination corresponding to the first artificial intelligence processor and the corresponding total computing power consumption value;
[0109] (1.7) Determine the next artificial intelligence processor after the first artificial intelligence processor in the processor sequence as the first artificial intelligence processor, return to execute the step of obtaining the target second historical computing time of the first artificial intelligence processor, until there is no next artificial intelligence processor after the first artificial intelligence processor in the processor sequence, and obtain the processor combination corresponding to each of the first artificial intelligence processors and the corresponding total computing power consumption value.
[0110] The specific method for determining the processor combination corresponding to each first artificial intelligence processor and the corresponding total computing power consumption value is as follows: select the first artificial intelligence processor with the same sorting number and the same processor number from the processor sequence, and determine the corresponding target second historical computing time; calculate the ratio of the second historical computing time of each second artificial intelligence processor before the first artificial intelligence processor to the target second historical computing time to obtain the second performance index; the second performance index of each second artificial intelligence processor is known, substitute it into the corresponding frequency function for calculation, and obtain the second adjustment frequency of each second artificial intelligence processor; calculate the second adjustment frequency of each second artificial intelligence processor and the product of the original frequency of the first artificial intelligence processor and the target second historical computing time to obtain the computing power consumption value of each second artificial intelligence processor and the first artificial intelligence processor; determine the processor combination corresponding to the first artificial intelligence processor and the corresponding total computing power consumption value based on the computing power consumption value of each second artificial intelligence processor and the first artificial intelligence processor; determine the next artificial intelligence processor after the first artificial intelligence processor in the processor sequence as the first artificial intelligence processor, and return to execute the step of obtaining the target second historical computing time of the first artificial intelligence processor until there is no next artificial intelligence processor after the first artificial intelligence processor in the processor sequence, and obtain the processor combination corresponding to each first artificial intelligence processor and the corresponding total computing power consumption value.
[0111] Specifically, to select M chips from all N chips to perform tasks, let the M chips be numbered from small to large as , ,…, , so the calculation time is , , …, At this time, the maximum duration = At this time, fix The chip remains unchanged and the energy consumption E is optimized under this condition. = No change, just select A minimum of M-1 chips will suffice.
[0112] For this purpose, for all chips numbered from i = 1, 2, … calculate their corresponding , and select the smallest M - 1 out of all values. When making this selection, calculate the total power consumption of all M chips: ;
[0113] ;
[0114] Note that in this step, the chips are fixed, so the calculated total energy consumption E is related to the specific values. Therefore, the total energy consumption is denoted as E( ).
[0115] Next, instead of fixing the chips as unchanged, set the chips numbered M, M + 1, …, N to in sequence, repeat the steps, and calculate the corresponding E(M), E(M + 1), …, E(N) for each value in turn. Note that the is different for each calculation, so it is necessary to calculate and select the smallest M - 1 chips from the values in each case and sum them up.
[0116] Finally, compare the magnitudes of E(M), E(M + 1), …, E(N). Select the smallest power consumption , and the corresponding chip selection is the optimal energy - efficiency selection. The calculated total energy consumption is the value of the minimum energy consumption. In this way, the energy - efficiency optimization based on chip - manufacturing heterogeneity is completed. The method proposed in the embodiments of the present application avoids brute - force enumeration of all combinations, greatly reducing the computational complexity. In addition, this method can dynamically adjust the optimal processor combination according to the heterogeneity of the chips, achieve precise optimization of the energy consumption of the computing cluster, and effectively improve the overall energy efficiency of the system.
[0117] In some embodiments, determining the processor combination corresponding to the first artificial intelligence processor and the corresponding total computing power consumption value based on the computing power consumption values of each of the second artificial intelligence processors and the first artificial intelligence processor includes:
[0118] (1.1) Sort the computing power consumption values of each of the second artificial intelligence processors in ascending order of the computing power consumption values to obtain a sequence of processors to be combined;
[0119] (1.2) Screen out each target second artificial intelligence processor whose sorting serial number is before the number of processors from the to-be-combined processor sequence, and determine each said target second artificial intelligence processor and the first artificial intelligence processor as the processor combination corresponding to the first neural network processing;
[0120] (1.3) Determine the target sum value of the computing power consumption value of the target second artificial intelligence processor and the computing power consumption value of the first artificial intelligence processor, and obtain the total computing power consumption value corresponding to the first artificial intelligence processor.
[0121] Among them, sort the computing power consumption values of each second artificial intelligence processor in ascending order of the computing power consumption value to obtain the to-be-combined processor sequence. The artificial intelligence processors included in this to-be-combined processor sequence are those that can be combined with the first artificial intelligence processor and may be used for the current model training. Since the number of processors has been specified for the model training, excluding the first artificial intelligence processor, M - 1 processors are still needed. That is, screen out each target second artificial intelligence processor whose sorting serial number is before the number of processors from the to-be-combined processor sequence, so as to determine each target second artificial intelligence processor and the first artificial intelligence processor as the processor combination corresponding to the first neural network processing.
[0122] For example, the number of processors is 4, the first artificial intelligence processor is , and the second artificial intelligence processors are respectively , , , , ; The corresponding computing power consumption value is 30, The corresponding computing power consumption value is 20, The corresponding computing power consumption value is 40, The corresponding computing power consumption value is 35, The corresponding computing power consumption value is 15. Then the to-be-combined processor sequence is , then the with sorting serial numbers before the number of processors 4 (i.e., 1, 2, and 3) is determined as the target second artificial intelligence processor. Then the processor combination corresponding to the first artificial intelligence processor is , and the total computing power consumption value corresponding to the first artificial intelligence processor is , the sum of the computing power consumption values.
[0123] In some embodiments, before substituting the second performance metrics of each of the second AI processors into the corresponding frequency function to obtain the second adjusted frequency corresponding to each of the second AI processors, the following steps are further included:
[0124] (1.1) Obtain the power consumption values of each of the second AI processors under multiple preset performance metrics;
[0125] (1.2) For each of the second AI processors, substitute the power consumption values corresponding to each of the preset performance metrics into a cubic function for fitting, determine the fitting coefficients of each term in the cubic function for each of the second AI processors, and obtain the frequency function corresponding to each of the second AI processors.
[0126] Among them, before substituting the second performance metrics of each of the second AI processors into the corresponding frequency function, it is first necessary to clarify the frequency function of each second AI processor. The specific determination method is as follows: for each second AI processor i, measure the power consumption data point values of the function at multiple (for example, at least 10) preset performance metrics x , and then further perform cubic function fitting, that is, fit the curve through the measured data points = . Among them, , , and are fitting coefficients, which are determined by calculation through measured data. Through this fitting function, the power consumption at any performance x can be quickly estimated , thereby providing accurate power consumption estimation for subsequent energy efficiency optimization and laying an optimization foundation. Since chip heterogeneity is generated during the manufacturing process and remains unchanged during use, the performance-power curve only needs to be measured once for each chip and then reused.
[0127] As can be seen from the above, in the embodiment of the present application, the first historical calculation duration of each artificial intelligence processor in the calculation stage of model training is obtained; the target calculation duration with the maximum duration is screened out from the multiple first historical calculation durations; the artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration among the multiple artificial intelligence processors are determined as other artificial intelligence processors; the ratio of the first historical calculation duration of each other artificial intelligence processor to the target calculation duration is determined to obtain the first performance index of each other artificial intelligence processor; the first performance index of each other artificial intelligence processor is substituted into the corresponding frequency function to obtain the first adjusted frequency corresponding to each other artificial intelligence processor; frequency adjustment is performed according to the first adjusted frequency corresponding to each other artificial intelligence processor. Compared with the related art, due to the heterogeneity of artificial intelligence processors, the load between computing nodes is unbalanced, which slows down the overall computing progress and thus causes energy waste. It is possible to save energy and improve energy efficiency without affecting the overall computing progress.
[0128] For the specific implementation of each of the above steps, reference may be made to the previous embodiments, which will not be elaborated here.
[0129] To facilitate better implementation of the frequency adjustment method provided by the embodiment of the present application, the embodiment of the present application also provides a device based on the above frequency adjustment method. The meanings of the nouns are the same as those in the above frequency adjustment method, and the specific implementation details can refer to the description in the method embodiment.
[0130] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of the frequency adjustment device provided by the embodiment of the present application. The frequency adjustment device is applied to a computer device. The frequency adjustment device may include a first acquisition unit 601, a screening unit 602, a first determination unit 603, a second determination unit 604, a substitution unit 605, an adjustment unit 606, etc.
[0131] The first acquisition unit 601 is configured to acquire the first historical calculation duration of each artificial intelligence processor in the calculation stage of model training;
[0132] The screening unit 602 is configured to screen out the target calculation duration with the maximum duration from the multiple first historical calculation durations;
[0133] The first determination unit 603 is configured to determine the artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration among the multiple artificial intelligence processors as other artificial intelligence processors;
[0134] A second determination unit 604, configured to determine a ratio of a first historical calculation duration of each of the other artificial intelligence processors to the target calculation duration, so as to obtain a first performance metric of each of the other artificial intelligence processors;
[0135] A substitution unit 605, configured to substitute the first performance metric of each of the other artificial intelligence processors into a corresponding frequency function, so as to obtain a first adjusted frequency corresponding to each of the other artificial intelligence processors;
[0136] An adjustment unit 606, configured to perform frequency adjustment according to the first adjusted frequency corresponding to each of the other artificial intelligence processors.
[0137] In some embodiments, the apparatus further includes:
[0138] A second acquisition unit, configured to acquire a second historical calculation duration of a plurality of candidate artificial intelligence processors, and acquire the number of processors of the artificial intelligence processors required for model training;
[0139] A sorting unit, configured to sort the plurality of candidate artificial intelligence processors in ascending order of the second historical calculation duration, so as to obtain a processor sequence;
[0140] A third determination unit, configured to determine, based on the number of processors and the processor sequence, a processor combination and a corresponding total calculation power consumption value corresponding to each first artificial intelligence processor whose sorting serial number is not less than the number of processors from the processor sequence;
[0141] A second screening unit, configured to screen out a target first artificial intelligence processor with the smallest corresponding total calculation power consumption value from the plurality of first artificial intelligence processors;
[0142] A fourth determination unit, configured to determine each artificial intelligence processor in the processor combination corresponding to the target first artificial intelligence processor as the artificial intelligence processor required in the model training process.
[0143] In some embodiments, the third determination unit includes:
[0144] A screening subunit, configured to screen out a first artificial intelligence processor whose sorting serial number is the same as the number of processors from the processor sequence;
[0145] A first acquisition subunit, configured to acquire a target second historical calculation duration of the first artificial intelligence processor;
[0146] A first determination subunit, configured to determine a ratio of a second historical calculation duration of each second artificial intelligence processor before the first artificial intelligence processor to the target second historical calculation duration, so as to obtain a second performance metric;
[0147] A substitution subunit, configured to substitute the second performance metrics of each of the second artificial intelligence processors into corresponding frequency functions to obtain second adjusted frequencies corresponding to each of the second artificial intelligence processors;
[0148] A first calculation subunit, configured to calculate the product of the second adjusted frequencies of each of the second artificial intelligence processors and the original frequency of the first artificial intelligence processor and the target second historical calculation duration, to obtain calculation power consumption values of each of the second artificial intelligence processors and the first artificial intelligence processor;
[0149] A second determination subunit, configured to determine a processor combination corresponding to the first artificial intelligence processor and a corresponding total calculation power consumption value based on the calculation power consumption values of each of the second artificial intelligence processors and the first artificial intelligence processor;
[0150] An execution subunit, configured to determine the next artificial intelligence processor after the first artificial intelligence processor in the processor sequence as the first artificial intelligence processor, and return to execute the step of obtaining the target second historical calculation duration of the first artificial intelligence processor until there is no next artificial intelligence processor after the first artificial intelligence processor in the processor sequence, to obtain a processor combination corresponding to each of the first artificial intelligence processors and a corresponding total calculation power consumption value.
[0151] In some embodiments, the second determination subunit is configured to:
[0152] Sort the calculation power consumption values of each of the second artificial intelligence processors in ascending order of the calculation power consumption values to obtain a to-be-combined processor sequence;
[0153] Screen out each target second artificial intelligence processor whose sorting serial number is before the number of processors from the to-be-combined processor sequence, and determine each of the target second artificial intelligence processors and the first artificial intelligence processor as the processor combination corresponding to the first neural network processing;
[0154] Determine the sum value of the calculation power consumption value of the target second artificial intelligence processor and the calculation power consumption value of the first artificial intelligence processor to obtain the total calculation power consumption value corresponding to the first artificial intelligence processor.
[0155] In some embodiments, the third determination unit further includes:
[0156] A second acquisition subunit, configured to acquire the power consumption values of each of the second artificial intelligence processors under multiple preset performance metrics;
[0157] A third determination subunit, configured to, for each of the second artificial intelligence processors, substitute the power consumption values corresponding to each of the preset performance metrics into a cubic function for fitting, determine the fitting coefficients of each term in the cubic function for each of the second artificial intelligence processors, and obtain a frequency function corresponding to each of the second artificial intelligence processors.
[0158] In some embodiments, the first acquisition unit 601 includes:
[0159] A third acquisition subunit, configured to acquire the actual calculation duration of each artificial intelligence processor in each calculation stage during the historical model training;
[0160] A fourth determination subunit, configured to determine the sum value of the multiple actual calculation durations of each artificial intelligence processor to obtain a first total calculation duration;
[0161] A fifth determination subunit, configured to determine the ratio of the first total calculation duration of each artificial intelligence processor to the number of stages of the calculation stage, and obtain the first historical calculation duration of each artificial intelligence processor in the calculation stage during the model training.
[0162] In some embodiments, the first acquisition unit 601 includes:
[0163] A fourth acquisition subunit, configured to acquire the actual calculation duration of each artificial intelligence processor in each calculation stage during the historical model training;
[0164] A fifth acquisition subunit, configured to, for each artificial intelligence processor, acquire the calculation busy degree of each calculation stage;
[0165] A second calculation subunit, configured to calculate the product of each calculation busy degree and the actual calculation duration of the corresponding calculation stage to obtain a calculation result of each calculation stage;
[0166] A sixth determination subunit, configured to determine the sum value of the calculation results of the multiple calculation stages to obtain a second total calculation duration;
[0167] A seventh determination subunit, configured to determine the ratio of the second total calculation duration to the number of stages of the calculation stage, and obtain the first historical calculation duration of each artificial intelligence processor in the calculation stage during the model training.
[0168] For the specific implementation of each of the above units, reference may be made to the previous embodiments, which will not be elaborated herein.
[0169] As can be seen from the above, in the embodiment of the present application, the first acquisition unit 601 acquires the first historical calculation duration of each artificial intelligence processor in the calculation stage of the model training process; the screening unit 602 screens out the target calculation duration with the maximum duration from the multiple first historical calculation durations; the first determination unit 603 determines, as other artificial intelligence processors, the artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration among the multiple artificial intelligence processors; the second determination unit 604 determines the ratio of the first historical calculation duration of each of the other artificial intelligence processors to the target calculation duration, to obtain the first performance index of each of the other artificial intelligence processors; the substitution unit 605 substitutes the first performance index of each of the other artificial intelligence processors into the corresponding frequency function, to obtain the first adjusted frequency corresponding to each of the other artificial intelligence processors; and the adjustment unit 606 performs frequency adjustment according to the first adjusted frequency corresponding to each of the other artificial intelligence processors. Compared with the related art, in which the load imbalance between computing nodes caused by the heterogeneity of artificial intelligence processors slows down the overall computing progress and further causes energy waste, it is possible to save energy and thus improve energy efficiency without affecting the overall computing progress.
[0170] For the specific implementation of each of the above units, reference may be made to the foregoing embodiments, which will not be elaborated herein.
[0171] Refer to Figure 7 , Figure 7 FIG. is a block diagram of a part of a computer device 1000 for implementing the embodiments of the present disclosure. The computer device 1000 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 622 (for example, one or more processors) and a memory 632, and one or more storage media 630 (for example, one or more mass storage devices) for storing application programs 642 or data 644. Among them, the memory 632 and the storage media 630 may be transient storage or persistent storage. The program stored in the storage media 630 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 600. Further, the central processing unit 622 may be configured to communicate with the storage media 630 and execute a series of instruction operations in the storage media 630 on the server 600.
[0172] The computer device 1000 may further include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0173] The central processing unit 622 in the computer device 1000 may be used to execute the frequency adjustment method of the embodiments of the present disclosure. For example:
[0174] Obtain the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training;
[0175] Screen out the target calculation duration with the maximum duration from the multiple first historical calculation durations;
[0176] Determine the artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration among the multiple artificial intelligence processors as other artificial intelligence processors;
[0177] Determine the ratio of the first historical calculation duration of each of the other artificial intelligence processors to the target calculation duration to obtain the first performance index of each of the other artificial intelligence processors;
[0178] Substitute the first performance index of each of the other artificial intelligence processors into the corresponding frequency function to obtain the first adjustment frequency corresponding to each of the other artificial intelligence processors;
[0179] Perform frequency adjustment according to the first adjustment frequency corresponding to each of the other artificial intelligence processors.
[0180] The embodiments of the present disclosure further provide a computer-readable storage medium for storing program codes for executing the frequency adjustment method of the foregoing various embodiments.
[0181] The embodiments of the present disclosure further provide a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes to implement the above frequency adjustment method. For example:
[0182] Obtain the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training;
[0183] Screen out the target calculation duration with the maximum duration from the multiple first historical calculation durations;
[0184] Determine the artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration among the multiple artificial intelligence processors as other artificial intelligence processors;
[0185] Determine the ratio of the first historical calculation duration of each of the other artificial intelligence processors to the target calculation duration to obtain the first performance index of each of the other artificial intelligence processors;
[0186] Substitute the first performance index of each of the other artificial intelligence processors into the corresponding frequency function to obtain the first adjusted frequency corresponding to each of the other artificial intelligence processors;
[0187] Perform frequency adjustment according to the first adjusted frequency corresponding to each of the other artificial intelligence processors.
[0188] In addition, the terms "include" and "comprise" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product, or device.
[0189] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0190] It should be understood that in the description of the embodiments of this application, the meaning of multiple (or multiple items) is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.
[0191] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0192] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0193] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0194] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0195] It should also be understood that the various embodiments provided in the embodiments of this application can be combined arbitrarily to achieve different technical effects.
[0196] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.
[0197] The above is a specific description of the implementation manner of the present application. However, the present application is not limited to the above implementation manner. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
Claims
1. A frequency adjustment method, characterized in that, Including: Obtaining the first historical computing duration of each artificial intelligence processor during the computing stage of model training; Screening out the target computing duration with the largest duration from the multiple first historical computing durations; Determining, among the multiple artificial intelligence processors, the artificial intelligence processors other than the artificial intelligence processor corresponding to the target computing duration as other artificial intelligence processors; Determine the ratio of the first historical computing duration of each of the other AI processors to the target computing duration to obtain the first performance metric of each of the other AI processors. For the other AI processor i, divide its first historical computing duration by the target computing duration to obtain the first performance metric as ; Substituting the first performance index of each of the other artificial intelligence processors into the corresponding frequency function to obtain the first adjusted frequency corresponding to each of the other artificial intelligence processors; Performing frequency adjustment according to the first adjusted frequency corresponding to each of the other artificial intelligence processors; Through the frequency function, the corresponding power after frequency reduction can be calculated as: ; Among them, , , and are fitting coefficients, and the power multiplied by the duration , that is, the power consumption of the i-th processor in the whole calculation is obtained: ; Summing up all M processors to obtain the total energy consumption of all processors during the entire segment of computing; 。 2. The frequency adjustment method according to claim 1, wherein Before the step of obtaining the first historical computing duration of each artificial intelligence processor during the computing stage of model training, it further includes: Obtaining the second historical computing duration of multiple candidate artificial intelligence processors, and obtaining the number of processors of the artificial intelligence processors required for model training; Sorting the multiple candidate artificial intelligence processors in ascending order of the second historical computing duration to obtain a processor sequence; Based on the number of processors and the processor sequence, determining the processor combinations and the corresponding total computing power consumption values corresponding to each first artificial intelligence processor whose sorting serial number is not less than the number of processors from the processor sequence; Screening out the target first artificial intelligence processor with the smallest corresponding total computing power consumption value from the multiple first artificial intelligence processors; Determining each artificial intelligence processor in the processor combination corresponding to the target first artificial intelligence processor as the artificial intelligence processor required for model training.
3. The frequency adjustment method according to claim 2, wherein The step of determining, based on the number of processors and the processor sequence, the processor combinations and the corresponding total computing power consumption values corresponding to each first artificial intelligence processor whose sorting serial number is not less than the number of processors from the processor sequence includes: Screening out the first artificial intelligence processor whose sorting serial number is the same as the number of processors from the processor sequence; Obtaining the target second historical computing duration of the first artificial intelligence processor; Determining the ratio of the second historical computing duration of each second artificial intelligence processor before the first artificial intelligence processor to the target second historical computing duration to obtain the second performance index; Substituting the second performance index of each of the second artificial intelligence processors into the corresponding frequency function to obtain the second adjusted frequency corresponding to each of the second artificial intelligence processors; Calculating the product of the second adjusted frequency of each of the second artificial intelligence processors and the original frequency of the first artificial intelligence processor and the target second historical computing duration to obtain the computing power consumption values of each of the second artificial intelligence processors and the first artificial intelligence processor; Based on the computing power consumption values of each of the second artificial intelligence processors and the first artificial intelligence processor, determining the processor combination and the corresponding total computing power consumption value corresponding to the first artificial intelligence processor; Determine the next artificial intelligence processor after the first artificial intelligence processor in the processor sequence as the first artificial intelligence processor, and return to execute the step of obtaining the target second historical computing duration of the first artificial intelligence processor until there is no next artificial intelligence processor after the first artificial intelligence processor in the processor sequence, so as to obtain the processor combination corresponding to each first artificial intelligence processor and the corresponding total computing power consumption value.
4. The frequency adjustment method according to claim 3, characterized in that, The determining the processor combination corresponding to the first artificial intelligence processor and the corresponding total computing power consumption value based on the computing power consumption values of each second artificial intelligence processor and the first artificial intelligence processor includes: Sort the computing power consumption values of each second artificial intelligence processor in ascending order of the computing power consumption value to obtain a sequence of processors to be combined; Select each target second artificial intelligence processor whose sorting serial number is before the number of processors from the sequence of processors to be combined, and determine each target second artificial intelligence processor and the first artificial intelligence processor as the processor combination corresponding to the first neural network processing; Determine the sum value of the computing power consumption value of the target second artificial intelligence processor and the computing power consumption value of the first artificial intelligence processor to obtain the total computing power consumption value corresponding to the first artificial intelligence processor.
5. The frequency adjustment method according to claim 3, wherein Before substituting the second performance index of each second artificial intelligence processor into the corresponding frequency function to obtain the second adjusted frequency corresponding to each second artificial intelligence processor, it further includes: Obtain the power consumption values of each second artificial intelligence processor under multiple preset performance indexes; For each second artificial intelligence processor, substitute the power consumption value corresponding to each preset performance index into a cubic function for fitting, determine the fitting coefficient of each item in the cubic function for each second artificial intelligence processor, and obtain the frequency function corresponding to each second artificial intelligence processor.
6. The frequency adjustment method according to claim 1, wherein The obtaining the first historical computing duration of each artificial intelligence processor in the computing stage of the model training process includes: Obtain the actual computing duration of each artificial intelligence processor in each computing stage of the historical model training process; Determine the sum value of the multiple actual computing durations of each artificial intelligence processor to obtain the first total computing duration; Determine the ratio of the first total computing duration of each artificial intelligence processor to the number of stages of the computing stage, and obtain the first historical computing duration of each artificial intelligence processor in the computing stage of the model training process.
7. The frequency adjustment method according to claim 1, characterized in that The obtaining the first historical computing duration of each artificial intelligence processor in the computing stage of the model training process includes: Obtain the actual computing duration of each artificial intelligence processor in each computing stage of the historical model training process; For each artificial intelligence processor, obtain the computing busy degree of each computing stage; Calculate the product of each computing busy degree and the actual computing duration of the corresponding computing stage to obtain the computing result of each computing stage. Determine the sum of the calculation results of multiple said calculation stages to obtain a second total calculation duration; Determine the ratio of the second total calculation duration to the number of stages of the calculation stage to obtain a first historical calculation duration of the calculation stage of each artificial intelligence processor during model training.
8. A frequency adjustment device, characterized in that, It includes: A first acquisition unit, configured to acquire the first historical calculation duration of the calculation stage of each artificial intelligence processor during model training; A screening unit, configured to screen out a target calculation duration with the longest duration from multiple said first historical calculation durations; A first determination unit, configured to determine, as other artificial intelligence processors, the artificial intelligence processors among multiple said artificial intelligence processors except the artificial intelligence processor corresponding to the target calculation duration; A second determination unit, configured to determine a ratio of a first historical calculation duration of each of the other artificial intelligence processors to the target calculation duration, so as to obtain a first performance metric of each of the other artificial intelligence processors. For the other artificial intelligence processor i, use its first historical calculation duration divided by the target calculation duration to obtain a first performance metric as ; A substitution unit, configured to substitute the first performance index of each said other artificial intelligence processor into the corresponding frequency function to obtain a first adjusted frequency corresponding to each said other artificial intelligence processor; An adjustment unit, configured to perform frequency adjustment according to the first adjusted frequency corresponding to each said other artificial intelligence processor; Through the frequency function, the corresponding power after frequency reduction can be calculated as: ; Among them, , , and are fitting coefficients, and the power multiplied by the duration , that is, the power consumption of the i-th processor in the whole calculation is obtained: ; Sum up all M processors to obtain the total energy consumption of all processors in the entire calculation; 。 9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the frequency adjustment method according to any one of claims 1 to 7.
10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the frequency adjustment method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Working frequency adjustment method and device, electronic equipment and storage medium
CN115617500A
Frequency adjustment method and device, storage medium and computer equipment
CN117666754A