Frequency adjustment method and device, storage medium and computer equipment
By screening the historical computing time and performance indicators of artificial intelligence processors and adjusting the processor frequency, the calculation imbalance caused by chip heterogeneity in the artificial intelligence computing cluster is solved, and energy conservation and energy efficiency improvement are achieved.
Patent Information
- Application Number
- CN202510601521.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-12
AI Technical Summary
Performance differences in the artificial intelligence computing cluster due to chip manufacturing heterogeneity lead to unbalanced load between computing nodes, slowing down the overall computing progress and causing energy waste.
By obtaining the historical calculation time of each artificial intelligence processor, filtering out the target calculation time with the largest duration, and calculating its performance indicators based on the ratio of the historical calculation time of other processors to the target duration, and then adjusting the frequency of each processor through the frequency function to achieve alignment of the calculation time.
Without affecting the overall computing progress, energy saving, energy efficiency improves, and reduce computing waiting time and energy waste caused by processor performance differences.
Smart Images

Figure CN120122800A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a frequency adjustment method, device, storage medium, and computer device. Background Art
[0002] Today, with the rapid development of artificial intelligence (AI) technology, intelligent computing clusters (hereinafter referred to as intelligent computing clusters) have become the key pillars for promoting AI scientific research and industrial applications. Intelligent computing clusters are built by adopting advanced AI processors, including graphics processing units (GPUs), neural network processing units (NPUs), etc., and can efficiently process complex AI tasks and the computing requirements of large-scale data. However, the popularization and wide application of intelligent computing clusters have also brought severe energy consumption challenges, which pose a huge pressure on global energy use and environmental protection.
[0003] In related technologies, even in chips with the same model, architecture, and completely consistent microcircuit design, due to factors such as material errors, dimensional errors, and process fluctuations in the manufacturing process, there will be differences in chip performance, power consumption, etc. Although this heterogeneity difference can be ignored in personal electronic devices with a single or small number of chips, it will have a significant impact in large-scale computing clusters.
[0004] In parallel computing tasks, if these performance differences are ignored, it will lead to uneven load among computing nodes, slow down the overall computing progress, and thus cause energy waste. Therefore, related technologies urgently need to propose a frequency adjustment method to solve the above technical problems. Summary of the Invention
[0005] The main purpose of this application is to provide a frequency adjustment method, device, storage medium, and computer device, which can save energy and thus improve energy efficiency without affecting the overall computing progress.
[0006] In a first aspect, an embodiment of this application provides a frequency adjustment method, including: Obtain the first historical computing duration of each artificial intelligence processor during the computing stage of model training; Screen out the target computing duration with the longest duration from the multiple first historical computing durations; Determine the artificial intelligence processors other than the artificial intelligence processor corresponding to the target computing duration among the multiple artificial intelligence processors as other artificial intelligence processors; Determine the ratio of the first historical computing duration of each of the other artificial intelligence processors to the target computing duration to obtain the first performance index of each of the other artificial intelligence processors; Substitute the first performance metric of each of the other artificial intelligence processors into the corresponding frequency function to obtain the first adjusted frequency corresponding to each of the other artificial intelligence processors; Perform frequency adjustment according to the first adjusted frequency corresponding to each of the other artificial intelligence processors.
[0007] In a second aspect, an embodiment of the present application provides a frequency adjustment device, including: A first acquisition unit configured to acquire the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training; A screening unit configured to screen out the target calculation duration with the maximum duration from the multiple first historical calculation durations; A first determination unit configured to determine, among the multiple artificial intelligence processors, the artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration as other artificial intelligence processors; A second determination unit configured to determine the ratio of the first historical calculation duration of each of the other artificial intelligence processors to the target calculation duration to obtain the first performance metric of each of the other artificial intelligence processors; A substitution unit configured to substitute the first performance metric of each of the other artificial intelligence processors into the corresponding frequency function to obtain the first adjusted frequency corresponding to each of the other artificial intelligence processors; An adjustment unit configured to perform frequency adjustment according to the first adjusted frequency corresponding to each of the other artificial intelligence processors.
[0008] In a third aspect, an embodiment of the present application provides a storage medium. The computer-readable storage medium stores multiple instructions, and these instructions are suitable for being loaded by a processor to execute the frequency adjustment method as described in any one of the above.
[0009] In a fourth aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the frequency adjustment method as described in any one of the above is implemented.
[0010] In an embodiment of the present application, by obtaining the first historical computing duration of each artificial intelligence processor during the computing phase of model training; screening out the target computing duration with the maximum duration from the multiple first historical computing durations; determining, as other artificial intelligence processors, the artificial intelligence processors among the multiple artificial intelligence processors other than the artificial intelligence processor corresponding to the target computing duration; determining the ratio of the first historical computing duration of each of the other artificial intelligence processors to the target computing duration to obtain the first performance metric of each of the other artificial intelligence processors; substituting the first performance metric of each of the other artificial intelligence processors into the corresponding frequency function to obtain the first adjustment frequency corresponding to each of the other artificial intelligence processors; and performing frequency adjustment according to the first adjustment frequency corresponding to each of the other artificial intelligence processors, compared with the related art where the load imbalance between computing nodes caused by the heterogeneity of artificial intelligence processors slows down the overall computing progress and thus causes energy waste, it is possible to save energy and improve energy efficiency without affecting the overall computing progress.
[0011] Other features and advantages of the present disclosure will be described in the following specification, and, in part, will be apparent from the specification or learned by implementing the present disclosure. The objectives and other advantages of the present disclosure can be achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 It is a schematic diagram of the model training process in an ideal state provided by an embodiment of the present application.
[0014] Figure 2 It is a schematic diagram of the model training process in an actual state provided by an embodiment of the present application.
[0015] Figure 3 It is a schematic diagram of the scenario of the frequency adjustment system provided by an embodiment of the present application.
[0016] Figure 4 It is a schematic flowchart of the frequency adjustment method provided by an embodiment of the present application.
[0017] Figure 5 It is a schematic diagram of using frequency adjustment to reduce the power consumption of model training provided by an embodiment of the present application.
[0018] Figure 6 The structural schematic diagram of the frequency adjustment device provided by the embodiment of the present application.
[0019] Figure 7 The structural schematic diagram of the computer device provided by the embodiment of the present application. Detailed implementation manners
[0020] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0021] It should be noted that in some processes described in the specification, claims and the above-mentioned drawings, there are multiple steps that appear in a specific order. However, it should be clearly understood that these steps may not be executed in the order in which they appear in this document or may be executed in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, descriptions such as "first", "second" or "target" in this document are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.
[0022] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are explained. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations: Artificial Intelligence Processor: Also known as an artificial intelligence accelerator, it is high-performance hardware designed specifically for artificial intelligence (AI) applications, especially for deep learning, machine learning, and neural networks.
[0023] Unique design: Adopting a multi-core architecture with numerous processing cores, it can divide matrix operations in deep learning into multiple small tasks and execute them in parallel on different cores; in addition to traditional CPU or GPU cores, it may include dedicated hardware accelerators such as tensor processing units (TPUs) or neural processing units (NPUs) to optimize common operations in deep learning such as matrix multiplication and convolution; equipped with a high-speed memory and cache system such as high-bandwidth memory (HBM) or on-chip SRAM cache to reduce data transfer bottlenecks; adopting advanced interconnection technologies such as mesh interconnection or crossbar switches to ensure efficient communication and cooperation among different parts; to cope with high power consumption, it includes a complex power management system that uses technologies such as dynamic voltage and frequency adjustment, load balancing, and thermal throttling to maximize energy efficiency.
[0024] Operating principle: First, preprocess and partition the input data to adapt it to the parallel architecture of the processor; then allocate the data and operations to multiple processing cores or dedicated accelerators; then each core or accelerator accesses the data using high-speed memory while executing the allocated tasks; then collect and integrate the processed data results to form the final output; finally, perform iterative adjustment according to the model feedback to optimize the model performance.
[0025] Application scenarios: It can quickly train large-scale neural network models; accelerate the model inference process in actual applications and provide prediction results in real time; be used for big data analysis, such as playing a role in fields such as image recognition and natural language processing.
[0026] In the data parallel training of AI large models, each training cycle usually consists of two stages: computing and communication. The computing part is carried out independently on each AI processor (such as NPU), while the communication part cannot start until all AI processors have completed the computing. Therefore, the reasonable arrangement of computing and communication is crucial for the overall training efficiency and energy efficiency.
[0027] The ideal parallel training arrangement is as Figure 1 shown, Figure 1 which is a schematic diagram of the model training process in the ideal state provided by the embodiments of this application. The computing tasks are carried out in parallel on 4 NPU processors, and all processors enter the communication stage after completing the computing, repeating periodically.
[0028] However, due to the existence of chip manufacturing heterogeneity, the actual situation is not like this. There are differences in the computing performance of different chips, resulting in inconsistent computing completion times, as specifically shown in Figure 2 shown, Figure 2 which is a schematic diagram of the model training process in the actual state provided by the embodiments of this application. In this figure, NPU1 is the chip with the slowest computing speed, and its delay slows down the start time of the overall communication stage. Other processors (such as NPU2, NPU3, and NPU4) must wait for the result of NPU1 after completing the computing. Although there is no effective computing task during this waiting time, a large amount of energy is still consumed, resulting in energy consumption waste.
[0029] In order to solve the above problems, embodiments of the present application obtain the first historical computing duration of each artificial intelligence processor during the computing stage of model training; screen out the target computing duration with the maximum duration from the multiple first historical computing durations; determine, as other artificial intelligence processors, the artificial intelligence processors other than the artificial intelligence processor corresponding to the target computing duration among the multiple artificial intelligence processors; determine the ratio of the first historical computing duration of each of the other artificial intelligence processors to the target computing duration to obtain the first performance index of each of the other artificial intelligence processors; substitute the first performance index of each of the other artificial intelligence processors into the corresponding frequency function to obtain the first adjusted frequency corresponding to each of the other artificial intelligence processors; perform frequency adjustment according to the first adjusted frequency corresponding to each of the other artificial intelligence processors. Compared with the related art, due to the heterogeneity of artificial intelligence processors, the load imbalance between computing nodes slows down the overall computing progress, thereby causing energy waste. It is possible to save energy and improve energy efficiency without affecting the overall computing progress. For details, please continue to refer to the following specific embodiments.
[0030] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the scenario of the frequency adjustment system provided by the embodiments of the present application. It includes a terminal 140, the Internet 130, a gateway 120, a server 110, etc.
[0031] The terminal 140 includes, but is not limited to, pre-configured laptop computers, tablet computers, desktop computers, and other electronic devices with data reporting capabilities. In addition, it can be a single device or a collection of multiple devices. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.
[0032] The terminal 140 refers to a computer system that can report data to the server 110. Compared with ordinary terminals, the server 110 has higher requirements in terms of stability, security, performance, etc. The server 110 can be a high-performance computer in a network platform, a cluster of multiple high-performance computers, a part (such as a virtual machine) allocated from a high-performance computer, a combination of parts (such as virtual machines) allocated from multiple high-performance computers, etc.
[0033] The gateway 120 is also known as an internetwork connector or protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. Between two systems using different communication protocols, data formats, or languages, or even with completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. Messages sent from the terminal 140 to the server 110 need to be sent to the corresponding server 110 through the gateway 120. Messages sent from the server 110 to the terminal 140 also need to be sent to the corresponding terminal 140 through the gateway 120.
[0034] The frequency adjustment method of the embodiments of the present disclosure can be implemented in the server 110.
[0035] It should be noted that Figure 3 The schematic diagram of the scenario of the shown frequency adjustment system is only an example. The frequency adjustment system and scenario described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of image processing technology and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0036] In this embodiment, it will be described from the perspective of the frequency adjustment device, which can be specifically integrated in a computer device with a storage unit and installed with a microprocessor and having computing capabilities.
[0037] Please refer to Figure 4 , Figure 4 , which is the flowchart of the frequency adjustment method provided by the embodiments of the present application. The frequency adjustment method includes: In step 201, obtain the first historical calculation duration of each artificial intelligence processor in the calculation stage of the model training process.
[0038] Among them, in the artificial intelligence model training, multiple stages will be experienced. The calculation stage of the model training process mainly refers to the process of performing mathematical operations such as matrix multiplication, convolution operation, and activation function calculation, which are used to update the parameters of the model. The first historical calculation duration is the time spent by the artificial intelligence processor in the past in the calculation stage of model training.
[0039] Specifically, through the log recording system or monitoring tool, collect the time data of each artificial intelligence processor in the calculation stage of multiple previous model trainings. For example, in a deep learning training platform, after each training task is completed by each processor, the time spent in the calculation stage will be recorded in the log file. The system regularly extracts the required data from these log files as the first historical calculation duration.
[0040] The purpose of this step is to provide basic data for subsequent frequency adjustment in order to understand the performance of each processor during the calculation phase.
[0041] In some embodiments, obtaining the first historical calculation duration of each artificial intelligence processor during the calculation phase of model training includes: (1) Obtaining the actual calculation duration of each artificial intelligence processor during each calculation phase in the historical model training process; (2) Determining the sum of the multiple actual calculation durations of each artificial intelligence processor to obtain the first total calculation duration; (3) Determining the ratio of the first total calculation duration of each artificial intelligence processor to the number of stages of the calculation phase to obtain the first historical calculation duration of each artificial intelligence processor during the calculation phase of model training.
[0042] Among them, specialized hardware monitoring tools or system-level monitoring software are used to collect the calculation time of the processor. These tools can monitor the running state of the processor in real time and automatically record the corresponding time information when detecting the start and end events of the calculation phase. For example, some server management systems can provide monitoring of processor performance metrics, including the duration information of the calculation phase. Obtaining the actual calculation duration of each calculation phase is the basis for subsequent calculations. By detailed recording of the actual calculation duration of each phase, the performance of the processor under different calculation tasks can be accurately reflected, providing accurate data for subsequent calculation of the average duration. Different calculation phases may have different complexities and amounts of calculation, and recording the duration of each phase helps to analyze the performance of the processor more meticulously.
[0043] Specifically, all the actual calculation duration data of each processor is read from the stored log files or monitoring data. Calculate the first total calculation duration The overall calculation time consumption of the processor during multiple calculation phases can be comprehensively considered. Different calculation phases may have different durations due to factors such as task complexity and data volume. By summing up to obtain the total duration, the calculation ability and time overhead of the processor during the entire historical model training process can be understood macroscopically. Given the first total calculation duration of each processor and the number of stages n of the calculation phase, the first historical calculation duration of each artificial intelligence processor during the calculation phase of model training can be determined Calculating the average value of the first historical calculation duration can eliminate the influence of duration fluctuations in different calculation stages and obtain a more representative processor calculation duration metric. This metric can more accurately reflect the average performance of the processor during the model training calculation stage and provide a reliable reference basis for subsequent optimization operations such as frequency adjustment. Through the average duration, the calculation performance between different processors can be compared more fairly, and then more reasonable resource scheduling and optimization can be carried out.
[0044] In some embodiments, obtaining the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training includes: (1) Obtaining the actual calculation duration of each artificial intelligence processor during each calculation stage in the historical model training process; (2) For each artificial intelligence processor, obtaining the calculation busy degree of each calculation stage; (3) Calculating the product of each calculation busy degree and the actual calculation duration of the corresponding calculation stage to obtain the calculation result of each calculation stage; (4) Determining the sum value of the calculation results of multiple calculation stages to obtain the second total calculation duration; (5) Determining the ratio of the second total calculation duration to the number of stages of the calculation stage to obtain the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training.
[0045] Among them, since different calculation stages may have different task complexities, the calculation duration is affected by the task complexity. Considering the calculation busy degree can more accurately evaluate the actual workload of the processor in each calculation stage. Therefore, it can be quantified according to influencing factors such as the task type or complexity executed in the calculation stage. The calculation busy degree can be expressed in the form of a percentage, such as 30%.
[0046] Specifically, by multiplying the calculation busy degree by the actual calculation duration, the duration and resource occupancy situation of the calculation stage are comprehensively considered. The calculation result obtained in this way can more accurately reflect the actual impact of this calculation stage on the overall performance of the processor and provide a more reasonable basis for subsequent calculation of the total duration. Calculating the average value can eliminate the differences between different calculation stages and obtain a more representative processor calculation duration metric. This metric considers the calculation busy degree and can more accurately reflect the average performance of the processor during the model training calculation stage, providing a reliable reference for subsequent optimization operations such as frequency adjustment and resource scheduling.
[0047] For example, there is an artificial intelligence processor A. The actual calculation duration in calculation stage 1 is 10 minutes, the actual calculation duration in calculation stage 2 is 15 minutes, and the actual calculation duration in calculation stage 3 is 8 minutes; the busy degree in calculation stage 1 is 80%, the busy degree in calculation stage 2 is 60%, and the busy degree in calculation stage 3 is 70%. Then the calculation result of calculation stage 1 is 10 * 80% = 8 minutes, the calculation result of calculation stage 2 is 15 * 60% = 9 minutes, and the calculation result of calculation stage 3 is 8 * 70% = 5.6 minutes; the second total calculation duration is 8 + 9 + 5.6 = 22.6 minutes; the first historical calculation duration of the artificial intelligence processor A in the calculation stage during the model training process is 22.6 / 3 = 7.53 minutes.
[0048] In step 202, the target calculation duration with the longest duration is selected from the multiple first historical calculation durations.
[0049] Among them, the target calculation duration is the time value with the longest duration among all the first historical calculation durations. The target calculation duration is used as the benchmark duration, and subsequently, the frequencies of other processors are adjusted based on it to make the calculation times of each processor as aligned as possible and reduce the waiting time.
[0050] For example Figure 2 As shown, since the first historical calculation duration of calculation stage NPU1 is the largest among NPU1, NPU2, NPU3, and NPU4, the calculation duration of NPU1 is also the largest during the actual calculation process, that is Figure 2 As shown. To avoid increasing the original calculation duration during the process of adjusting the frequency, the first historical calculation duration of the largest NPU1 is used as the benchmark duration to adjust the frequencies of NPU2, NPU3, and NPU4, so as to avoid the calculation duration after adjustment exceeding the target calculation duration.
[0051] In step 203, the artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration among the multiple artificial intelligence processors are determined as other artificial intelligence processors.
[0052] Among them, other artificial intelligence processors are the remaining processors among all the processors performing model training, excluding the processor with the longest calculation duration (i.e., the one corresponding to the target calculation duration). By traversing the data set recording the processors and their corresponding first historical calculation durations, the processor corresponding to the target calculation duration is found and excluded, and the remaining processors are the other artificial intelligence processors.
[0053] The purpose of this step is to clarify the range of processors that need to have their frequencies adjusted. Since the processor corresponding to the target calculation duration is already in the state of having the longest calculation time and does not need to be adjusted, while other processors need to be adjusted to optimize the overall performance.
[0054] In step 204, determine the ratio of the first historical computing duration of each of the other artificial intelligence processors to the target computing duration, and obtain the first performance metric of each of the other artificial intelligence processors.
[0055] The first performance metric is a quantitative value used to represent the relative level of the computing performance of the other artificial intelligence processors compared to the processor corresponding to the target computing duration, and is obtained through the ratio of their computing durations.
[0056] Specifically, for the other artificial intelligence processor i, divide its first historical computing duration by the target computing duration to obtain the first performance metric as .
[0057] In this way, by quantifying the performance difference between the other artificial intelligence processors and the processor corresponding to the target computing duration, it provides a basis for calculating the adjustment frequency in the subsequent calculation, facilitating targeted frequency adjustment according to the performance difference.
[0058] In step 205, substitute the first performance metric of each of the other artificial intelligence processors into the corresponding frequency function to obtain the first adjustment frequency corresponding to each of the other artificial intelligence processors.
[0059] The frequency function is a pre-established mathematical relationship model that describes the corresponding relationship between the performance metric and the processor adjustment frequency. The first adjustment frequency is calculated through the frequency function based on the first performance metric and is a value used to adjust the working frequency of the other artificial intelligence processors.
[0060] In this way, a suitable adjustment frequency is calculated based on the processor performance difference, and by adjusting the frequency, the computing time of the other processors is made close to that of the processor corresponding to the target computing duration, improving the overall computing efficiency.
[0061] In step 206, perform frequency adjustment according to the first adjustment frequency corresponding to each of the other artificial intelligence processors.
[0062] Frequency adjustment refers to changing the working frequency of the artificial intelligence processor through hardware control instructions or software system settings, thereby adjusting its computing duration. This makes the computing speed of the other artificial intelligence processors the same as the target computing duration of the target processor, reduces the computing waiting time caused by processor performance differences, improves the efficiency of the entire computing system during the model training calculation stage, does not increase the computing duration of the entire computing process, and reduces energy consumption.
[0063] Please refer to Figure 5 , Figure 5Schematic diagram for reducing the power consumption of model training by frequency adjustment provided in the embodiments of the present application. For the i-th AI processor, the computing duration thereon is . Therefore, when using M AI processors for parallel training, the computing duration of each cycle is determined by the processor with the longest duration, which is , where is the maximum value among all , = max( , , …, ).
[0064] To achieve time alignment, for each processor i, it is necessary to reduce its frequency so that its computing duration is adjusted from to . At this time, the normalized performance of this processor changes from 1 to . Thus, the computing durations of different processors are aligned, as shown in Figure 5 . Since the power will also decrease during the process of the processor reducing its frequency, therefore, the computing duration remains unchanged after optimization, but the total energy consumption will decrease.
[0065] Through the frequency function, the corresponding power after the frequency reduction can be calculated as: ; where , , and are fitting coefficients. Multiplying the power by the duration T_max, the power consumption of the i-th processor in the entire segment of computing can be obtained: ; Then summing up all M processors, the total energy consumption of all processors in the entire segment of computing can be obtained: ; Through the above method, this patent effectively solves the problem of uneven computing caused by chip manufacturing heterogeneity, realizes the precise alignment of the computing time of AI processors, and thus significantly reduces the overall energy consumption on the premise of ensuring the computing efficiency of the system. The optimized energy consumption value can be calculated using the provided formula.
[0066] In some embodiments, before obtaining the first historical computing duration of each artificial intelligence processor in the computing stage of model training, it further includes: (1) Obtaining the second historical computing durations of multiple candidate artificial intelligence processors, and obtaining the number of processors of artificial intelligence processors required for model training; (2) Sort the multiple candidate artificial intelligence processors in ascending order of the second historical calculation duration to obtain a processor sequence; (3) Based on the number of processors and the processor sequence, determine the processor combinations and the corresponding total computing power consumption values corresponding to each first artificial intelligence processor whose sorting serial number is not less than the number of processors from the processor sequence; (4) From the multiple first artificial intelligence processors, screen out the target first artificial intelligence processor with the smallest corresponding total computing power consumption value; (5) Determine each artificial intelligence processor in the processor combination corresponding to the target first artificial intelligence processor as the artificial intelligence processor required for the model training process.
[0067] Among them, before training a model with multiple artificial intelligence processors, a set of processor combinations needs to be selected from multiple candidate artificial intelligence processors. The purpose of the selection is to ensure the lowest power consumption value during the entire model training process. For this purpose, first obtain the second historical calculation duration of multiple candidate artificial intelligence processors, and obtain the number of processors required for this model training. The second historical calculation duration is the calculation duration data of the candidate processor during the calculation stage of the past model training, which is used to evaluate its calculation speed. The processor sequence is an ordered list formed by sorting the candidate artificial intelligence processors in ascending order of the second historical calculation duration. Each element is a candidate processor, and its order reflects the relative size of the calculation duration. Based on the number of processors and the processor sequence, determine the processor combinations and the corresponding total computing power consumption values corresponding to each first artificial intelligence processor whose sorting serial number is not less than the number of processors from the processor sequence. After obtaining the total computing power consumption value of each first artificial intelligence processor, screen out the target first artificial intelligence processor with the smallest corresponding total computing power consumption value, and determine each artificial intelligence processor in the processor combination corresponding to the target first artificial intelligence processor as the artificial intelligence processor required for the model training process. It shows that each artificial intelligence processor in the processor combination corresponding to the target first artificial intelligence processor can achieve the smallest total computing power consumption value as the processor required for this model training.
[0068] Specifically, the first artificial intelligence processor refers to the processor selected starting from the position where the sorting serial number is not less than the number of processors in the processor sequence. The processor combination is different combination forms composed of these candidate artificial intelligence processors. Each combination may be used as the processor configuration scheme for model training, and the number of processors in each processor combination is the number of artificial intelligence processors required for this model training. The total computing power consumption value is the total power value expected to be consumed by each processor combination when executing the model training task, which is used to measure the energy consumption of different combinations.
[0069] In some embodiments, determining, based on the number of processors and the processor sequence, the processor combination and the corresponding total computing power consumption value corresponding to each first artificial intelligence processor whose sorting serial number in the processor sequence is not less than the number of processors includes: (1.1) Screening out the first artificial intelligence processors in the processor sequence whose sorting serial numbers are the same as the number of processors; (1.2) Obtaining the target second historical computing duration of the first artificial intelligence processor; (1.3) Determining the ratio of the second historical computing duration of each second artificial intelligence processor before the first artificial intelligence processor to the target second historical computing duration to obtain a second performance metric; (1.4) Substituting the second performance metric of each second artificial intelligence processor into the corresponding frequency function to obtain the second adjusted frequency corresponding to each second artificial intelligence processor; (1.5) Calculating the product of the second adjusted frequency of each second artificial intelligence processor and the original frequency of the first artificial intelligence processor and the target second historical computing duration to obtain the computing power consumption values of each second artificial intelligence processor and the first artificial intelligence processor; (1.6) Based on the computing power consumption values of each second artificial intelligence processor and the first artificial intelligence processor, determining the processor combination and the corresponding total computing power consumption value corresponding to the first artificial intelligence processor; (1.7) Determining the next artificial intelligence processor after the first artificial intelligence processor in the processor sequence as the first artificial intelligence processor, and returning to execute the step of obtaining the target second historical computing duration of the first artificial intelligence processor until there is no next artificial intelligence processor after the first artificial intelligence processor in the processor sequence, so as to obtain the processor combination and the corresponding total computing power consumption value corresponding to each first artificial intelligence processor.
[0070] Among them, the specific method for determining the processor combination corresponding to each first artificial intelligence processor and the corresponding total computing power consumption value is as follows: Screen out the first artificial intelligence processors whose sorting serial numbers are the same as the number of processors from the processor sequence, and determine the corresponding target second historical computing duration; Calculate the ratio of the second historical computing duration of each second artificial intelligence processor before the first artificial intelligence processor to the target second historical computing duration to obtain the second performance index; Given the second performance index of each second artificial intelligence processor, substitute it into the corresponding frequency function for calculation to obtain the second adjusted frequency of each second artificial intelligence processor; Calculate the product of the second adjusted frequency of each second artificial intelligence processor and the original frequency of the first artificial intelligence processor and the target second historical computing duration to obtain the computing power consumption value of each second artificial intelligence processor and the first artificial intelligence processor; Based on the computing power consumption values of each second artificial intelligence processor and the first artificial intelligence processor, determine the processor combination corresponding to the first artificial intelligence processor and the corresponding total computing power consumption value; Determine the next artificial intelligence processor after the first artificial intelligence processor in the processor sequence as the first artificial intelligence processor, and return to execute the step of obtaining the target second historical computing duration of the first artificial intelligence processor until there is no next artificial intelligence processor after the first artificial intelligence processor in the processor sequence, so as to obtain the processor combination corresponding to each first artificial intelligence processor and the corresponding total computing power consumption value.
[0071] Specifically, to select M chips from all N chips to execute tasks, let the numbers of the M chips be arranged in ascending order as , ,…, , so their computing durations are respectively , , …, . At this time, the maximum duration = . At this time, fix the chip unchanged, and optimize the energy consumption E under this condition. At this time = remains unchanged, and only need to select the smallest M - 1 chips.
[0072] For this reason, calculate the corresponding for all chips numbered from i = 1, 2, …, , select the smallest M - 1 among all values. When making this selection, calculate the total power consumption of all M chips: ; ; Note that this step fixes the The chip remains unchanged. Therefore, the total energy consumption E calculated is related to the specific value. Thus, the total energy consumption is denoted as E( ).
[0073] Next, instead of fixing the chip unchanged, the chips numbered M, M + 1, …, N are sequentially set to , and the steps are repeated. For each value, the corresponding E(M), E(M + 1), …, E(N) are calculated in turn. Note that each during the calculation is different. Therefore, it is necessary to calculate separately and select the smallest M - 1 chips from the numerical values in each case and sum them up.
[0074] Finally, compare the magnitudes of E(M), E(M + 1), …, E(N). Select the smallest power consumption , and the corresponding chip selection is the optimal energy - efficiency selection. The total energy consumption calculated is the value of the minimum energy consumption. In this way, the energy - efficiency optimization based on chip - manufacturing heterogeneity is completed. The method proposed in the embodiments of the present application avoids brute - force enumeration of all combinations, greatly reducing the computational complexity. In addition, this method can dynamically adjust the optimal processor combination according to the heterogeneity of the chips, achieve precise optimization of the energy consumption of the computing cluster, and effectively improve the overall energy efficiency of the system.
[0075] In some embodiments, determining the processor combination corresponding to the first artificial intelligence processor and the corresponding total computing power consumption value based on the computing power consumption values of each of the second artificial intelligence processors and the first artificial intelligence processor includes: (1.1) Sort the computing power consumption values of each of the second artificial intelligence processors in ascending order of the computing power consumption values to obtain a sequence of processors to be combined; (1.2) Screen out each target second artificial intelligence processor whose sorting serial number is before the number of processors from the sequence of processors to be combined, and determine each target second artificial intelligence processor and the first artificial intelligence processor as the processor combination corresponding to the first neural network processing; (1.3) Determine the target sum value of the computing power consumption value of the target second artificial intelligence processor and the computing power consumption value of the first artificial intelligence processor to obtain the total computing power consumption value corresponding to the first artificial intelligence processor.
[0076] Among them, the computing power consumption values of each second artificial intelligence processor are sorted in ascending order to obtain a sequence of processors to be combined. The artificial intelligence processors included in this sequence of processors to be combined are those that can be combined with the first artificial intelligence processor to possibly be used for the current model training. Since the number of processors has been specified for model training, excluding the first artificial intelligence processor, M - 1 processors are still needed. That is, each target second artificial intelligence processor with a sorting serial number before the number of processors is selected from the sequence of processors to be combined, so as to determine each target second artificial intelligence processor and the first artificial intelligence processor as the processor combination corresponding to the first neural network processing.
[0077] For example, the number of processors is 4, and the first artificial intelligence processor is , and the second artificial intelligence processors are respectively 、 、 、 、 ; The corresponding computing power consumption value is 30, The corresponding computing power consumption value is 20, The corresponding computing power consumption value is 40, The corresponding computing power consumption value is 35, The corresponding computing power consumption value is 15. Then the sequence of processors to be combined is , and then the with sorting serial numbers before the number of processors 4 (i.e., 1, 2, and 3) is determined as the target second artificial intelligence processor. Then the processor combination corresponding to the first artificial intelligence processor is , and the total computing power consumption value corresponding to the first artificial intelligence processor is the sum of the computing power consumption values of 、 .
[0078] In some embodiments, before substituting the second performance indicators of each second artificial intelligence processor into the corresponding frequency function to obtain the second adjusted frequency corresponding to each second artificial intelligence processor, it further includes: (1.1) Obtain the power consumption values of each second artificial intelligence processor under multiple preset performance indicators; (1.2) For each second artificial intelligence processor, substitute the power consumption values corresponding to each preset performance indicator into a cubic function for fitting, determine the fitting coefficients of each term in the cubic function for each second artificial intelligence processor, and obtain the frequency function corresponding to each second artificial intelligence processor.
[0079] Among them, before substituting the second performance metrics of each of the second AI processors into the corresponding frequency function, it is first necessary to clarify the frequency function of each second AI processor. The specific determination method is as follows: For each second AI processor i, power consumption data point values of the function at multiple (e.g., at least 10) preset performance metrics x are measured and obtained. , and then further through cubic function fitting, that is, fitting a curve through the measured data points = . Among them, , , and are fitting coefficients, which are determined by calculating through the measured data. Through this fitting function, the power consumption at any performance x can be quickly estimated , thereby providing accurate power consumption estimation for subsequent energy efficiency optimization and laying an optimization foundation. Since chip heterogeneity is generated during the manufacturing process and remains unchanged during use, the performance-power curve only needs to be measured once for each chip and then reused.
[0080] As can be seen from the above, in the embodiment of the present application, the first historical calculation duration of each AI processor in the calculation stage of model training is obtained; the target calculation duration with the maximum duration is selected from the multiple first historical calculation durations; the AI processors other than the AI processor corresponding to the target calculation duration among the multiple AI processors are determined as other AI processors; the ratio of the first historical calculation duration of each of the other AI processors to the target calculation duration is determined to obtain the first performance metric of each of the other AI processors; the first performance metric of each of the other AI processors is substituted into the corresponding frequency function to obtain the first adjusted frequency corresponding to each of the other AI processors; frequency adjustment is performed according to the first adjusted frequency corresponding to each of the other AI processors. Compared with the related art, due to the heterogeneity of AI processors, the load imbalance between computing nodes slows down the overall computing progress, thereby causing energy waste. It is possible to save energy without affecting the overall computing progress, and thus improve energy efficiency.
[0081] For the specific implementation of each of the above steps, reference can be made to the previous embodiments, and details will not be elaborated here.
[0082] To facilitate better implementation of the frequency adjustment method provided in the embodiment of the present application, the embodiment of the present application also provides a device based on the above frequency adjustment method. The meanings of the terms therein are the same as those in the above frequency adjustment method, and the specific implementation details can refer to the description in the method embodiment.
[0083] Please refer to Figure 6 , Figure 6The figure is a schematic structural diagram of a frequency adjustment device provided by an embodiment of the present application. The frequency adjustment device is applied to a computer device. The frequency adjustment device may include a first acquisition unit 601, a screening unit 602, a first determination unit 603, a second determination unit 604, a substitution unit 605, an adjustment unit 606, etc.
[0084] The first acquisition unit 601 is configured to acquire the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training. The screening unit 602 is configured to screen out the target calculation duration with the maximum duration from the multiple first historical calculation durations. The first determination unit 603 is configured to determine, as other artificial intelligence processors, the artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration among the multiple artificial intelligence processors. The second determination unit 604 is configured to determine the ratio of the first historical calculation duration of each of the other artificial intelligence processors to the target calculation duration, and obtain the first performance index of each of the other artificial intelligence processors. The substitution unit 605 is configured to substitute the first performance index of each of the other artificial intelligence processors into the corresponding frequency function, and obtain the first adjustment frequency corresponding to each of the other artificial intelligence processors. The adjustment unit 606 is configured to perform frequency adjustment according to the first adjustment frequency corresponding to each of the other artificial intelligence processors.
[0085] In some embodiments, the device further includes: The second acquisition unit is configured to acquire the second historical calculation duration of multiple candidate artificial intelligence processors, and acquire the number of processors of the artificial intelligence processors required for model training. The sorting unit is configured to sort the multiple candidate artificial intelligence processors in ascending order of the second historical calculation duration, and obtain a processor sequence. The third determination unit is configured to determine, based on the number of processors and the processor sequence, the processor combination and the corresponding total calculation power consumption value corresponding to each first artificial intelligence processor whose sorting serial number is not less than the number of processors from the processor sequence. The second screening unit is configured to screen out the target first artificial intelligence processor with the minimum corresponding total calculation power consumption value from the multiple first artificial intelligence processors. The fourth determination unit is configured to determine, as the artificial intelligence processors required for the model training process, each artificial intelligence processor in the processor combination corresponding to the target first artificial intelligence processor.
[0086] In some embodiments, the third determination unit includes: A screening subunit, configured to screen out a first artificial intelligence processor from the processor sequence whose sorting serial number is the same as the number of processors; A first obtaining subunit, configured to obtain a target second historical computing duration of the first artificial intelligence processor; A first determining subunit, configured to determine a ratio of the second historical computing duration of each second artificial intelligence processor before the first artificial intelligence processor to the target second historical computing duration, to obtain a second performance metric; A substituting subunit, configured to substitute the second performance metric of each second artificial intelligence processor into a corresponding frequency function, to obtain a second adjusted frequency corresponding to each second artificial intelligence processor; A first calculating subunit, configured to calculate a product of the second adjusted frequency of each second artificial intelligence processor and the original frequency of the first artificial intelligence processor and the target second historical computing duration, to obtain a computing power consumption value of each second artificial intelligence processor and the first artificial intelligence processor; A second determining subunit, configured to determine a processor combination corresponding to the first artificial intelligence processor and a corresponding total computing power consumption value based on the computing power consumption values of each second artificial intelligence processor and the first artificial intelligence processor; An executing subunit, configured to determine the next artificial intelligence processor after the first artificial intelligence processor in the processor sequence as the first artificial intelligence processor, and return to execute the step of obtaining the target second historical computing duration of the first artificial intelligence processor until there is no next artificial intelligence processor after the first artificial intelligence processor in the processor sequence, to obtain a processor combination corresponding to each first artificial intelligence processor and a corresponding total computing power consumption value.
[0087] In some embodiments, the second determining subunit is configured to: Sort the computing power consumption values of each second artificial intelligence processor in ascending order of the computing power consumption value, to obtain a to-be-combined processor sequence; Screen out each target second artificial intelligence processor whose sorting serial number is before the number of processors from the to-be-combined processor sequence, and determine each target second artificial intelligence processor and the first artificial intelligence processor as the processor combination corresponding to the first neural network processing; Determine a sum value of the computing power consumption value of the target second artificial intelligence processor and the computing power consumption value of the first artificial intelligence processor, to obtain the total computing power consumption value corresponding to the first artificial intelligence processor.
[0088] In some embodiments, the third determining unit further includes: A second acquisition subunit, configured to acquire the power consumption values of each of the second artificial intelligence processors under multiple preset performance metrics; A third determination subunit, configured to, for each of the second artificial intelligence processors, substitute the power consumption values corresponding to each of the preset performance metrics into a cubic function for fitting, determine the fitting coefficients of each term in the cubic function for each of the second artificial intelligence processors, and obtain a frequency function corresponding to each of the second artificial intelligence processors.
[0089] In some embodiments, the first acquisition unit 601 includes: A third acquisition subunit, configured to acquire the actual calculation duration of each of the artificial intelligence processors in each calculation stage during the historical model training process; A fourth determination subunit, configured to determine the sum value of the multiple actual calculation durations of each of the artificial intelligence processors to obtain a first total calculation duration; A fifth determination subunit, configured to determine the ratio of the first total calculation duration of each of the artificial intelligence processors to the number of stages of the calculation stage, and obtain the first historical calculation duration of each artificial intelligence processor in the calculation stage during the model training process.
[0090] In some embodiments, the first acquisition unit 601 includes: A fourth acquisition subunit, configured to acquire the actual calculation duration of each of the artificial intelligence processors in each calculation stage during the historical model training process; A fifth acquisition subunit, configured to, for each of the artificial intelligence processors, acquire the calculation busy degree of each calculation stage; A second calculation subunit, configured to calculate the product of each of the calculation busy degrees and the actual calculation duration of the corresponding calculation stage to obtain a calculation result of each calculation stage; A sixth determination subunit, configured to determine the sum value of the calculation results of the multiple calculation stages to obtain a second total calculation duration; A seventh determination subunit, configured to determine the ratio of the second total calculation duration to the number of stages of the calculation stage, and obtain the first historical calculation duration of each artificial intelligence processor in the calculation stage during the model training process.
[0091] For the specific implementation of each of the above units, reference may be made to the previous embodiments, which will not be elaborated here.
[0092] As can be seen from the above, in the embodiment of the present application, the first acquisition unit 601 acquires the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training; the screening unit 602 screens out the target calculation duration with the maximum duration from the multiple first historical calculation durations; the first determination unit 603 determines, as other artificial intelligence processors, the artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration among the multiple artificial intelligence processors; the second determination unit 604 determines the ratio of the first historical calculation duration of each of the other artificial intelligence processors to the target calculation duration, to obtain the first performance index of each of the other artificial intelligence processors; the substitution unit 605 substitutes the first performance index of each of the other artificial intelligence processors into the corresponding frequency function, to obtain the first adjustment frequency corresponding to each of the other artificial intelligence processors; and the adjustment unit 606 performs frequency adjustment according to the first adjustment frequency corresponding to each of the other artificial intelligence processors. Compared with the related art, in which the load imbalance between computing nodes caused by the heterogeneity of artificial intelligence processors slows down the overall computing progress and thus causes energy waste, energy can be saved without affecting the overall computing progress, thereby improving energy efficiency.
[0093] For the specific implementation of each of the above units, reference may be made to the foregoing embodiments, which will not be elaborated herein.
[0094] Refer to Figure 7 , Figure 7 FIG. is a partial structural block diagram of a computer device 1000 for implementing the embodiment of the present disclosure. The computer device 1000 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 622 (for example, one or more processors) and a memory 632, and one or more storage media 630 (for example, one or more mass storage devices) for storing application programs 642 or data 644. Among them, the memory 632 and the storage media 630 may be transient storage or persistent storage. The program stored in the storage media 630 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 600. Further, the central processing unit 622 may be configured to communicate with the storage media 630 and execute a series of instruction operations in the storage media 630 on the server 600.
[0095] The computer device 1000 may further include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0096] The central processing unit 622 in the computer device 1000 may be used to execute the frequency adjustment method of the embodiments of the present disclosure, for example: Obtain the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training; Screen out the target calculation duration with the maximum duration from the multiple first historical calculation durations; Determine, as other artificial intelligence processors, the artificial intelligence processors among the multiple artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration; Determine the ratio of the first historical calculation duration of each of the other artificial intelligence processors to the target calculation duration to obtain the first performance index of each of the other artificial intelligence processors; Substitute the first performance index of each of the other artificial intelligence processors into the corresponding frequency function to obtain the first adjustment frequency corresponding to each of the other artificial intelligence processors; Perform frequency adjustment according to the first adjustment frequency corresponding to each of the other artificial intelligence processors.
[0097] The embodiments of the present disclosure further provide a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the frequency adjustment methods of the foregoing various embodiments.
[0098] The embodiments of the present disclosure further provide a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes to implement the above frequency adjustment method. For example: Obtain the first historical calculation duration of each artificial intelligence processor during the calculation stage of model training; Screen out the target calculation duration with the maximum duration from the multiple first historical calculation durations; Determine, as other artificial intelligence processors, the artificial intelligence processors among the multiple artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration; Determine the ratio of the first historical calculation duration of each of the other artificial intelligence processors to the target calculation duration to obtain the first performance index of each of the other artificial intelligence processors; Substitute the first performance metric of each of the other artificial intelligence processors into the corresponding frequency function to obtain the first adjusted frequency corresponding to each of the other artificial intelligence processors; Perform frequency adjustment according to the first adjusted frequency corresponding to each of the other artificial intelligence processors.
[0099] In addition, the terms "including" and "comprising" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0100] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0101] It should be understood that in the description of the embodiments of this application, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.
[0102] In several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0103] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of these units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0104] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0105] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0106] It should also be understood that the various embodiments provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0107] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of this module or unit.
[0108] The above is a specific description of the embodiments of the present application, but the present application is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
Claims
1. A frequency adjustment method, characterized in that: include: Obtain the first historical computing time of each AI processor in the computing phase of the model training process; Filter out the target calculation duration with the largest duration from the plurality of first historical calculation durations; Determine the artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration among the plurality of artificial intelligence processors as other artificial intelligence processors; Determine a ratio of a first historical computing time of each of the other artificial intelligence processors to the target computing time, and obtain a first performance indicator of each of the other artificial intelligence processors; Substituting the first performance indicator of each of the other artificial intelligence processors into the corresponding frequency function to obtain a first adjustment frequency corresponding to each of the other artificial intelligence processors; Frequency adjustment is performed according to the first adjustment frequency corresponding to each of the other artificial intelligence processors.
2. The frequency adjustment method according to claim 1, characterized in that: Before obtaining the first historical computing time of each artificial intelligence processor in the computing phase of the model training process, the method further includes: Obtaining second historical computing durations of multiple candidate artificial intelligence processors, and obtaining the number of artificial intelligence processors required for model training; Sort the plurality of candidate artificial intelligence processors in ascending order according to the second historical calculation duration to obtain a processor sequence; Based on the number of processors and the processor sequence, determine from the processor sequence a processor combination corresponding to each first artificial intelligence processor whose sorting number is not less than the number of processors and a corresponding total computing power consumption value; Filtering out a target first artificial intelligence processor with the smallest total computing power consumption value from the plurality of first artificial intelligence processors; Each artificial intelligence processor in the processor combination corresponding to the target first artificial intelligence processor is determined as the artificial intelligence processor required for the model training process.
3. The frequency adjustment method according to claim 2, characterized in that: The step of determining, based on the number of processors and the processor sequence, from the processor sequence, a processor combination corresponding to each first artificial intelligence processor having a sorting number not less than the number of processors and a corresponding total computing power consumption value, comprises: Filter out a first artificial intelligence processor whose sorting number is the same as the number of processors from the processor sequence; Obtain a target second historical computing duration of the first artificial intelligence processor; Determine a ratio of a second historical computing time of each second artificial intelligence processor before the first artificial intelligence processor to the target second historical computing time to obtain a second performance indicator; Substituting the second performance indicator of each of the second artificial intelligence processors into the corresponding frequency function to obtain a second adjustment frequency corresponding to each of the second artificial intelligence processors; Calculate the second adjusted frequency of each of the second artificial intelligence processors and the product of the original frequency of the first artificial intelligence processor and the target second historical computing time to obtain computing power consumption values of each of the second artificial intelligence processors and the first artificial intelligence processor; Determine, based on the computing power consumption values of each of the second artificial intelligence processors and the first artificial intelligence processor, a processor combination corresponding to the first artificial intelligence processor and a corresponding total computing power consumption value; The next artificial intelligence processor after the first artificial intelligence processor in the processor sequence is determined as the first artificial intelligence processor, and the step of obtaining the target second historical computing time of the first artificial intelligence processor is returned to execute until there is no next artificial intelligence processor after the first artificial intelligence processor in the processor sequence, and the processor combination corresponding to each of the first artificial intelligence processors and the corresponding total computing power consumption value are obtained.
4. The frequency adjustment method according to claim 3, characterized in that: The determining, based on the computing power consumption values of each of the second artificial intelligence processors and the first artificial intelligence processor, a processor combination corresponding to the first artificial intelligence processor and a corresponding total computing power consumption value, comprises: Sort the computing power consumption values of each of the second artificial intelligence processors in ascending order of computing power consumption values to obtain a sequence of processors to be combined; Filter out each target second artificial intelligence processor whose sorting sequence number is before the number of processors from the sequence of processors to be combined, and determine each target second artificial intelligence processor and the first artificial intelligence processor as a processor combination corresponding to the first neural network processing; Determine the sum of the computing power consumption value of the target second artificial intelligence processor and the computing power consumption value of the first artificial intelligence processor to obtain a total computing power consumption value corresponding to the first artificial intelligence processor.
5. The frequency adjustment method according to claim 3, characterized in that: Before substituting the second performance indicator of each of the second artificial intelligence processors into the corresponding frequency function to obtain the second adjustment frequency corresponding to each of the second artificial intelligence processors, the method further includes: Obtaining power consumption values of each of the second artificial intelligence processors under a plurality of preset performance indicators; For each of the second artificial intelligence processors, the power consumption value corresponding to each of the preset performance indicators is substituted into the cubic function for fitting, the fitting coefficient of each item in the cubic function of each of the second artificial intelligence processors is determined, and the frequency function corresponding to each of the second artificial intelligence processors is obtained.
6. The frequency adjustment method according to claim 1, characterized in that: The obtaining of the first historical computing duration of each artificial intelligence processor in the computing phase of the model training process includes: Obtaining the actual computing time of each computing stage of each artificial intelligence processor during the historical model training process; Determine a sum of multiple actual computing durations of each of the artificial intelligence processors to obtain a first total computing duration; Determine the ratio of the first total computing time of each artificial intelligence processor to the number of stages in the computing stage, and obtain the first historical computing time of each artificial intelligence processor in the computing stage of the model training process.
7. The frequency adjustment method according to claim 1, characterized in that: The obtaining of the first historical computing duration of each artificial intelligence processor in the computing phase of the model training process includes: Obtaining the actual computing time of each computing stage of each artificial intelligence processor during the historical model training process; For each of the artificial intelligence processors, obtaining the computational busyness of each computational stage; Calculate the product of each of the calculation busyness and the actual calculation time of the corresponding calculation stage to obtain the calculation result of each calculation stage; Determine the sum of the calculation results of the plurality of calculation stages to obtain a second total calculation duration; Determine the ratio of the second total computing time to the number of stages in the computing stage, and obtain the first historical computing time of each artificial intelligence processor in the computing stage of the model training process.
8. A frequency adjustment device, characterized in that: include: A first acquisition unit, used to acquire a first historical calculation duration of each artificial intelligence processor in a calculation phase of a model training process; A screening unit, configured to screen out a target calculation duration with a maximum duration from the plurality of first historical calculation durations; A first determining unit is used to determine the artificial intelligence processors other than the artificial intelligence processor corresponding to the target calculation duration among the plurality of artificial intelligence processors as other artificial intelligence processors; A second determination unit is used to determine the ratio of the first historical computing time of each of the other artificial intelligence processors to the target computing time, and obtain a first performance indicator of each of the other artificial intelligence processors; a substitution unit, used to substitute the first performance indicator of each of the other artificial intelligence processors into the corresponding frequency function to obtain a first adjustment frequency corresponding to each of the other artificial intelligence processors; An adjustment unit is used to perform frequency adjustment according to a first adjustment frequency corresponding to each of the other artificial intelligence processors.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the frequency adjustment method according to any one of claims 1 to 7.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the frequency adjustment method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Working frequency adjustment method and device, electronic equipment and storage medium
CN115617500A
Frequency adjustment method and device, storage medium and computer equipment
CN117666754A
Frequency adjustment method and device, electronic equipment and readable storage medium
CN119472971A