Evaluation Method and Device for Operator Optimization

By setting multiple target frequencies in the artificial intelligence chip and determining the total execution time, combining bandwidth and computing power utilization, the problem of bottleneck determination in operator optimization is solved, and efficient optimization evaluation and tuning efficiency improvement is achieved.

CN118312394BActive Publication Date: 2025-07-22SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410489536.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-22
Publication Date
2025-07-22
Estimated Expiration
2044-04-22

AI Technical Summary

Technical Problem

In artificial intelligence chips, how to simply, quickly and accurately determine the limit bottlenecks of operator execution to guide the direction of optimization is a technical problem that needs to be solved urgently.

Method used

By setting a plurality of target frequencies of the first processor, the total execution time of the target operator at each frequency is determined, and the evaluation is performed based on the total execution time, bandwidth utilization rate and computing power utilization rate, indicating the optimization direction.

Benefits of technology

The operator optimization evaluation is simple, fast and efficient, which improves the operator development and tuning efficiency and improves the overall process speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118312394B_ABST
    Figure CN118312394B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of artificial intelligence technologies, and particularly to a method and apparatus for evaluating operator optimization. The method includes: setting a plurality of target frequencies required for a first processor to execute a target operator; determining the total execution duration consumed by the first processor to execute the target operator at each target frequency; evaluating the target operator based on the determined plurality of total execution durations to obtain an evaluation result, where the evaluation result is used to indicate the optimization direction of the target operator; wherein the target frequency includes at least one of a first frequency related to data reading and writing of the first processor and a second frequency related to data calculation of the first processor. It can achieve simple, fast, efficient, and accurate evaluation of operator optimization, improve the development, tuning efficiency and speed of operators, and achieve speed improvement of the entire process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to an evaluation method and device for operator optimization. Background Art

[0002] In the process of optimizing the operator performance of artificial intelligence chips, it is challenging to evaluate whether the operator (or algorithm) is optimized to the theoretical performance limit for the following reasons: on the one hand, the processors in artificial intelligence chips are becoming more and more complex, and the hardware involves data interaction between multiple computing modules and memory modules; on the other hand, when the operator is more complex, the performance bottleneck of the operator may be reflected in different modules in different time windows. How to simply, quickly and accurately determine the limiting bottleneck of the operator execution in the processor and guide the optimization direction of subsequent operators is a technical problem that needs to be solved urgently. Summary of the invention

[0003] In view of this, the present disclosure proposes an evaluation method and device for operator optimization.

[0004] According to one aspect of the present disclosure, there is provided an evaluation method for operator optimization, the method comprising:

[0005] Setting a plurality of target frequencies required for the first processor to execute a target operator;

[0006] Determining a total execution time consumed by the first processor to execute the target operator at each of the target frequencies;

[0007] According to the determined multiple total execution times, the target operator is evaluated to obtain an evaluation result, wherein the evaluation result is used to indicate an optimization direction of the target operator;

[0008] The target frequency includes at least one of a first frequency associated with data reading and writing by the first processor and a second frequency associated with data calculation by the first processor.

[0009] In a possible implementation, evaluating the target operator according to the determined multiple total execution times to obtain an evaluation result includes:

[0010] The target operator is evaluated according to the determined multiple total execution times and the bandwidth utilization and computing power utilization of the target operator under each of the total execution times to obtain an evaluation result.

[0011] In a possible implementation, the method further includes:

[0012] When it is determined that a target frequency needs to be selected from the first frequency and the second frequency, the first frequency or the second frequency is determined as the target frequency according to the data reading / writing speed and data calculation speed of the first processor.

[0013] In a possible implementation manner, determining the first frequency or the second frequency as the target frequency according to the data reading / writing speed and data calculation speed of the first processor includes:

[0014] When it is determined according to the data reading / writing speed and the data calculation speed that the transmission duration for the first processor to perform data reading / writing is less than the calculation duration for the first processor to perform data calculation, the first frequency is determined as the target frequency;

[0015] When it is determined according to the data reading / writing speed and the data calculation speed that the transmission duration for the first processor to perform data reading / writing is greater than the calculation duration for the first processor to perform data calculation, the second frequency is determined as the target frequency.

[0016] In a possible implementation manner, the first frequency includes the reading / writing frequency for the load / store unit in the first processor to perform data reading / writing;

[0017] Alternatively, the first frequency includes the reading / writing frequency and the memory frequency of the memory accessed by the load / store unit. When the target frequency is the first frequency and the first frequency includes the reading / writing frequency and the memory frequency, at least one of the reading / writing frequencies and memory frequencies of each target frequency is different.

[0018] In a possible implementation manner, when the target frequency includes the first frequency, the multiple total execution durations include multiple second total execution durations; or,

[0019] When the target frequency includes the second frequency, the multiple total execution durations include multiple first total execution durations; or,

[0020] When the target frequency includes the first frequency and the second frequency, the multiple total execution durations include multiple first total execution durations and multiple second total execution durations;

[0021] Wherein, the multiple first total execution durations are the durations consumed by the first processor to respectively execute the target operator when the first frequency remains unchanged and the second frequency is different; the multiple second total execution durations are the second total execution durations consumed by the first processor to respectively execute the target operator when the second frequency remains unchanged and the first frequency is different.

[0022] In a possible implementation, based on the determined multiple total execution durations, the bandwidth utilization rate and computing power utilization rate of the target operator at each of the total execution durations, the target operator is evaluated to obtain an evaluation result, including:

[0023] Based on the determined multiple total execution durations, determine the importance of the influence of the transmission duration and the computing duration on the total execution duration;

[0024] Based on the importance of the influence of the transmission duration and the computing duration, and the bandwidth utilization rate and computing power utilization rate of the target operator at each of the total execution durations, determine the current performance limiting bottleneck of the target operator when executed in the first processor, where the current performance limiting bottleneck includes a memory bottleneck or a computing bottleneck;

[0025] Based on the current performance limiting bottleneck, determine an optimization direction to form an evaluation result.

[0026] According to another aspect of the present disclosure, there is provided an evaluation device for operator optimization, the device including:

[0027] A frequency setting module, configured to set multiple target frequencies required for the first processor to execute a target operator;

[0028] A duration determination module, configured to determine the total execution duration consumed by the first processor to execute the target operator at each of the target frequencies;

[0029] An optimization evaluation module, configured to evaluate the target operator according to the determined multiple total execution durations to obtain an evaluation result, where the evaluation result is used to indicate the optimization direction of the target operator;

[0030] Wherein, the target frequency includes at least one of a first frequency related to data reading and writing of the first processor and a second frequency related to data calculation of the first processor.

[0031] In a possible implementation, the optimization evaluation module includes:

[0032] An evaluation sub-module, configured to evaluate the target operator according to the determined multiple total execution durations, the bandwidth utilization rate and computing power utilization rate of the target operator at each of the total execution durations, to obtain an evaluation result.

[0033] In a possible implementation, the device further includes:

[0034] A frequency selection module, configured to, when it is determined that a target frequency needs to be selected from the first frequency and the second frequency, determine the first frequency or the second frequency as the target frequency according to the data reading and writing speed and data calculation speed of the first processor.

[0035] In a possible implementation, determining the first frequency or the second frequency as the target frequency according to the data read / write speed and the data calculation speed of the first processor includes:

[0036] When it is determined according to the data read / write speed and the data calculation speed that the transmission duration for the first processor to perform data read / write is less than the calculation duration for the first processor to perform data calculation, determining the first frequency as the target frequency;

[0037] When it is determined according to the data read / write speed and the data calculation speed that the transmission duration for the first processor to perform data read / write is greater than the calculation duration for the first processor to perform data calculation, determining the second frequency as the target frequency.

[0038] In a possible implementation, the first frequency includes the read / write frequency for the load / store unit in the first processor to perform data read / write;

[0039] Alternatively, the first frequency includes the read / write frequency and the memory frequency of the memory accessed by the load / store unit. When the target frequency is the first frequency and the first frequency includes the read / write frequency and the memory frequency, at least one of the read / write frequency and the memory frequency of each target frequency is different.

[0040] In a possible implementation, when the target frequency includes the first frequency, the multiple total execution durations include multiple second total execution durations; or,

[0041] when the target frequency includes the second frequency, the multiple total execution durations include multiple first total execution durations; or,

[0042] when the target frequency includes the first frequency and the second frequency, the multiple total execution durations include multiple first total execution durations and multiple second total execution durations;

[0043] Wherein, the multiple first total execution durations are the durations consumed by the first processor to execute the target operator respectively when the first frequency remains unchanged and the second frequency is different; the multiple second total execution durations are the second total execution durations consumed by the first processor to execute the target operator respectively when the second frequency remains unchanged and the first frequency is different.

[0044] In a possible implementation, evaluating the target operator according to the determined multiple total execution durations, the bandwidth utilization rate and the computing power utilization rate of the target operator under each total execution duration to obtain an evaluation result, includes:

[0045] Based on the determined multiple total execution durations, determine the importance of the influence of the transmission duration and the calculation duration on the total execution duration;

[0046] Based on the importance of the influence of the transmission duration and the calculation duration, and the bandwidth utilization rate and computing power utilization rate of the target operator under each of the total execution durations, determine the current performance limitation bottleneck for the target operator to execute in the first processor, where the current performance limitation bottleneck includes a memory bottleneck or a computing bottleneck;

[0047] Based on the current performance limitation bottleneck, determine an optimization direction to form an evaluation result.

[0048] According to another aspect of the present disclosure, there is provided an electronic device, including: a second processor; a memory for storing instructions executable by the second processor; wherein, the second processor is configured to implement the above method when executing the instructions stored in the memory.

[0049] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a third processor.

[0050] According to another aspect of the present disclosure, there is provided a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, when the computer-readable code runs in a fourth processor of an electronic device, the fourth processor in the electronic device executes the above method.

[0051] Through the embodiments of the present disclosure, there is provided an evaluation method and device for operator optimization. A plurality of target frequencies required for a first processor to execute a target operator are preset; the total execution durations consumed by the first processor to execute the target operator at each target frequency are determined; based on the determined multiple total execution durations, an evaluation result is obtained for the target operator, and the evaluation result is used to indicate the optimization direction of the target operator; wherein, the target frequency includes at least one of a first frequency for the first processor to perform data reading and writing and a second frequency for the first processor to perform data calculation. It can realize simple, fast, efficient, and accurate evaluation of operator optimization, improve the development, tuning efficiency and speed of operators, and achieve the speedup of the entire process.

[0052] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. Description of the Drawings

[0053] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure together with the specification.

[0054] Figure 1 A flowchart showing an evaluation method for operator optimization according to an embodiment of the present disclosure.

[0055] Figure 2 A block diagram showing a first processor according to an embodiment of the present disclosure.

[0056] Figure 3 A timing diagram showing a target operator under memory bottleneck according to an embodiment of the present disclosure.

[0057] Figure 4 A frequency - performance curve diagram showing a target operator under memory bottleneck according to an embodiment of the present disclosure.

[0058] Figure 5 A timing diagram showing a target operator under computing bottleneck according to an embodiment of the present disclosure.

[0059] Figure 6 A frequency - performance curve diagram showing a target operator under computing bottleneck according to an embodiment of the present disclosure.

[0060] Figure 7 A block diagram showing an evaluation device for operator optimization according to an embodiment of the present disclosure.

[0061] Figure 8 A block diagram of a device 1900 for an electronic device or a server shown according to an exemplary embodiment. Detailed implementation manners

[0062] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0063] The special word "exemplary" here means "serving as an example, an embodiment, or illustrative". Any embodiment described as "exemplary" here does not necessarily have to be construed as superior to or better than other embodiments.

[0064] In addition, for a better description of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well - known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0065] In the process of analyzing and evaluating the performance of an operator, the performance bottleneck of the operator can usually be attributed to two situations: memory bound (memory bottleneck, also known as memory - limited, memory - access - limited, etc.) and compute bound (computing bottleneck, also known as compute - limited, etc.). The memory bottleneck means that the performance bottleneck of the operator is mainly reflected in the memory - access limitation, which is caused by insufficient memory - access bandwidth leading to performance problems of the operator. The compute bottleneck means that the performance bottleneck of the operator is mainly reflected in data computing, which is caused by insufficient hardware computing performance leading to performance problems of the operator. Only by determining the current limiting bottleneck (memory bottleneck or compute bottleneck) of the operator can we continue to optimize the operator, improve its performance, and increase the bandwidth utilization rate and computing - power utilization rate when the operator is executed in the processor. In related technologies, in the process of optimizing the operator, there are the following problems in the way of determining the current limiting bottleneck of the operator:

[0066] For the first - type solution, directly determine the current limiting bottleneck of the operator according to the bandwidth utilization rate and computing - power utilization rate of the operator for the processor in the artificial - intelligence chip. However, the determination method of this type of solution is too simple. Although the process of determining the result is very efficient, there are problems such as low accuracy.

[0067] For the second - type solution, through the method of performance simulation, realize the simulation test of the operator running on the processor in the artificial - intelligence chip. Although the determination accuracy of this type of solution is high, it is too dependent on specific hardware and has problems such as long time - consumption and complex process.

[0068] To solve the above - mentioned technical problems, the embodiments of the present disclosure provide an evaluation method and device for operator optimization. In this method, multiple target frequencies required for a first processor to execute a target operator are preset in advance; determine the total execution duration consumed by the first processor to execute the target operator at each of the target frequencies; evaluate the target operator according to the determined multiple total execution durations to obtain an evaluation result, and the evaluation result is used to indicate the optimization direction of the target operator; wherein, the target frequencies include at least one of a first frequency at which the first processor reads and writes data and a second frequency at which the first processor performs data computing. It can realize the simple, fast, efficient, and accurate evaluation of operator optimization.

[0069] As Figure 1 shown, the evaluation method for operator optimization provided by the embodiments of the present disclosure includes steps S101 - S103. This method can be applied to a server or an electronic device.

[0070] In step S101, multiple target frequencies required for a first processor to execute a target operator are set. Among them, the target frequency includes at least one of a first frequency related to data reading and writing of the first processor and a second frequency related to data calculation of the first processor. Among them, the target operator can be an operator for processing various types of user input data such as audio, video, image, text, etc., and the target operator can be applied to fields such as scientific computing, machine learning, data analysis, artificial intelligence, and financial modeling.

[0071] In this embodiment, the first processor that executes the target operator may be an artificial intelligence chip, such as a Graphics Processing Unit (GPU), a General-Purpose computing on Graphics Processing Unit (GPGPU), etc., and the present disclosure does not limit this. As Figure 2 shown, the first processor may include at least one load / store unit (LSU) 11 and multiple compute units 12. The load / store unit 11 is used to access the memory 13 of the first processor for data reading and writing. That is, the load / store unit 11 reads data from the memory 13 and writes the result obtained after the compute units 12 calculate into the memory 13. Among them, for simplicity Figure 2 only one load / store unit 11 and one compute unit 12 of the first processor are schematically shown, and the remaining load / store units 11 and compute units 12 are not shown.

[0072] Among them, the first frequency may be a related frequency that affects the speed of data reading and writing of the first processor. For example, the first frequency may include the read / write frequency of the load / store unit 11 for data reading and writing and the data reading and writing frequency of the memory 13 of the first processor (which can be called the memory frequency, the memory frequency). During the process of the first processor executing the target operator, the load / store unit 11 and the multiple compute units 12 execute the related steps of the operator in parallel. Since the frequencies of data reading and storage in the first processor are the same, the read / write frequency may refer to the frequency of data reading of each load / store unit 11 in the first processor or the frequency of data storage of the load / store unit 11. The second frequency may be a related frequency that affects the speed of data calculation of the first processor. For example, the second frequency may include the frequency of data calculation of each compute unit 12.

[0073] In a possible implementation, the method may further include: before step S101, when it is determined that a target frequency needs to be selected from the first frequency and the second frequency, determining the first frequency or the second frequency as the target frequency according to the data reading / writing speed and data calculation speed of the first processor.

[0074] In this implementation, the target frequency can be selected in the following manner: when it is determined according to the data reading / writing speed and the data calculation speed that the transmission duration of the first processor for data reading / writing is less than the calculation duration of the first processor for data calculation, determining the first frequency as the target frequency. When it is determined according to the data reading / writing speed and the data calculation speed that the transmission duration of the first processor for data reading / writing is greater than the calculation duration of the first processor for data calculation, determining the second frequency as the target frequency.

[0075] In this embodiment, multiple target frequencies can be configured based on the computing power and bandwidth of the first processor. After each optimization of the target operator is completed, multiple target frequency settings can be performed based on the current reading / writing frequency f10, second frequency f20, and memory frequency f30 of the first processor that executes the target operator. For example, if the target frequency is the first frequency and the first frequency only includes the reading / writing frequency or the memory frequency (wherein, one of the reading / writing frequency and the memory frequency that mainly affects the transmission duration can be used as the target frequency, that is, the one with a relatively greater impact on the transmission duration among the reading / writing frequency and the memory frequency), then multiple reading / writing frequencies different from f10 can be set based on the current reading / writing frequency f10, such as 0.9×f10, 0.8×f10, 0.7×f10..., and the second frequency and memory frequency of the first processor remain unchanged, still being the current f20 and f30. If the target frequency is the second frequency, then multiple second frequencies different from f20 can be set based on the current second frequency f20, such as 0.9×f20, 0.8×f20, 0.7×f20..., and the reading / writing frequency and memory frequency of the first processor remain unchanged, still being the current f10 and f30. If the target frequency is the first frequency and the first frequency includes the reading / writing frequency and the memory frequency, then multiple groups of reading / writing frequencies and memory frequencies different from at least one of f10 and f30 can be set based on the current f10 and f30, such as 0.9×f10 and 0.9×f13, 0.9×f10 and 0.8×f13, 0.8×f10 and 0.9×f13, 0.7×f10 and 0.8×f13..., and the second frequency of the first processor remains unchanged, still being the current f20.

[0076] In step S102, determine the total execution duration consumed by the first processor for executing the target operator at each of the target frequencies.

[0077] In this embodiment, when the target frequency includes the second frequency, the multiple total execution durations include multiple first total execution durations. Alternatively, when the target frequency includes the first frequency, the multiple total execution durations include multiple second total execution durations. Alternatively, when the target frequency includes the first frequency and the second frequency, the multiple total execution durations include multiple first total execution durations and multiple second total execution durations. The multiple first total execution durations are the durations consumed by the first processor to execute the target operator respectively when the first frequency remains unchanged and the second frequency is different; the multiple second total execution durations are the second total execution durations consumed by the first processor to execute the target operator respectively when the second frequency remains unchanged and the first frequency is different. That is, the first frequencies corresponding to different first total execution durations are the same but the second frequencies are different, and the first frequencies corresponding to different second total execution durations are different (the different first frequencies may mean that at least one of the read / write frequency and the memory frequency is different) but the second frequencies are the same.

[0078] In step S103, according to the determined multiple total execution durations, the target operator is evaluated to obtain an evaluation result, and the evaluation result is used to indicate the optimization direction of the target operator. According to the determined multiple total execution durations, etc., the current performance limitation bottleneck of the target operator when executed in the first processor can be determined first, and then the optimization direction can be determined based on the current performance limitation bottleneck, and then the evaluation result can be generated. Among them, the optimization direction can correspond to the current performance limitation bottleneck of the target operator when executed in the current processor. If the current performance limitation bottleneck is a memory bottleneck, the optimization direction can be the operator optimization direction to solve the memory bottleneck; if the current performance limitation bottleneck is a calculation bottleneck, the optimization direction can be the operator optimization direction to solve the calculation bottleneck.

[0079] In this embodiment, there are several different implementation manners for step S103. To further illustrate the rationality of different implementation manners of step S103, the following first combines Figures 2-6 to schematically illustrate the implementable principle and basis of the operator optimization evaluation method provided by the embodiments of the present disclosure.

[0080] Assume that for a certain target operator, the first processor that executes the target operator has a structure as Figure 2 shown.

[0081] In the first case, if it is assumed that the current performance limitation bottleneck of the target operator when executed in the first processor is a memory bottleneck, it is assumed that the input data is divided into n data blocks for processing respectively. Then, theoretically analyzed, in the memory bottleneck scenario, the transmission duration of each data block should be greater than the calculation duration, and then the timing diagram as Figure 3 shown can be obtained. Figure 3Let \(L_i\) denote the duration for the first processor to load the \(i\)-th data block, \(C_i\) denote the computing duration for the first processor to perform computations on the \(i\)-th data block to obtain the \(i\)-th result, and \(S_i\) denote the storage duration for the first processor to store the \(i\)-th result corresponding to the \(i\)-th data block. The value of \(i\) ranges from 1, 2, …, \(n\). Then the total execution duration of the target operator in the first processor is “\(L_1 + L_2 + \cdots + L_n + C_n + S_n\)”. It can be seen that in the memory bottleneck scenario, theoretically, the main factor affecting the total execution duration is the total data loading duration “\(L_1 + L_2 + \cdots + L_n\)”. Then the following conclusion can be inferred: If the second frequency of the computing unit 12 is reduced, the total execution duration of the target operator in the first processor will necessarily change little.

[0082] To verify the above conclusion, we assume that the input data is divided into 16 blocks (i.e., \(n = 16\)), and set the duration for a specific first processor to load data, perform data computations, and store data for executing a specific target operator as \(t_1\), \(t_2\), and \(t_3\) respectively, that is, \(L_i = t_1\), \(C_i = t_2\), \(S_i = t_3\). Further, set \(t_1 = t_3\) and \(t_2 = k_1\times t_1\). Then the total execution duration of the specific target operator in the specific first processor \(L_1 + L_2 + \cdots + L_n + C_n + S_n=n\times t_1 + t_2 + t_3\). Assume the frequency reduction ratio is \(q\), that is, the second frequency will be \(q\) times the original. Then the total execution duration of the target operator after frequency reduction is \(n\times t_1+\frac{t_2}{q}+t_3\). The difference between the two total execution durations is \((\frac{1}{q}-1)t_2\). At this time, the performance is reduced to \(1-\frac{(\frac{1}{q}-1)t_2}{n\times t_1 + t_2 + t_3}\). According to this formula, we can obtain Figure 4 the curve of the performance reduction ratio versus the frequency reduction ratio as shown. Then, combined with Figure 4 it can be known that: when \(k_1\) (i.e., \(\frac{t_2}{t_1}\)) is 0.1, 0.2, and 0.3, even if the computing frequency is reduced to 30% of the original, its performance can still remain above 95% of the original, which is indeed as the above inferred conclusion.

[0083] In other words, in the above step S103, if the target frequency is the second frequency, after setting multiple second frequencies, taking the obtained multiple first total execution durations as the total execution duration, if the difference between the multiple first total execution durations is within the first difference range, it proves that the change in the second frequency has little impact on the total execution duration, and it can be determined that the current performance limitation bottleneck for the target operator to execute in the first processor is the memory bottleneck. Among them, the first difference range can be set based on the data read / write speed and data computing speed of the first processor for executing the target operator.

[0084] In the second case, if it is assumed that the current performance limitation bottleneck for the target operator executed in the first processor is the computing bottleneck, and it is assumed that the input data is divided into n data blocks for separate processing. Then, theoretically analyzed, in the computing bottleneck scenario, the transmission duration of each data block should be less than the computing duration. Thus, it can be obtained that Figure 5 the timing diagram shown in Figure 5 where Li represents the duration for the first processor to load the i-th data block, Ci represents the computing duration for the first processor to perform computations on the i-th data block to obtain the i-th result, Si represents the duration for the first processor to store the i-th result corresponding to the i-th data block, and the value of i is 1, 2, …, n. Then the total execution duration of the target operator in the first processor is “L1 + C1 + C2 … + Cn + Sn”. It can be seen that in the computing bottleneck scenario, theoretically speaking, the main factor affecting the total execution duration is the total computing duration of the data “C1 + C2 … + Cn”. Then the following conclusion can be inferred: If the first frequency of the load / store unit 11 is reduced, the total execution duration of the target operator in the first processor will surely change little.

[0085] To verify the above conclusion, we assume that the input data is divided into 16 blocks (i.e., n = 16), and set the duration for a specific first processor to load data for executing a specific target operator as t1, the duration for data computation as t2, and the duration for data storage as t3, that is, Li = t1, Ci = t2, Si = t3. Further set t1 = t3, t1 = k2 × t2. Then the total execution duration of the specific target operator in the specific first processor is L1 + C1 + C2 … + Cn + Sn = t1 + n × t2 + t3. Assume the frequency reduction ratio is p, that is, the first frequency will be p times the original. Then the total execution duration of the target operator after frequency reduction is t1 / p + n × t2 + t3 / p, and the time difference between the two times is (1 / p - 1)(t1 + t3). At this time, the performance is reduced to 1 - ((1 / p - 1)(t1 + t3)) / (t1 + n × t2 + t3). According to this formula, it can be obtained that Figure 6 the curve of the performance reduction ratio versus the frequency reduction ratio shown in Figure 6 It can be known that: in the cases where k2 (i.e., t1 / t2) is 0.1, 0.2, and 0.3, even if the computing frequency is reduced to 30% of the original, its performance can still remain above 90% of the original, which is indeed as the above inferred conclusion.

[0086] In other words, in the above step S103, if the target frequency is the first frequency, after setting multiple first frequencies, the obtained multiple second total execution durations are used as the total execution duration. If the difference between the multiple second total execution durations is within the second difference range, it proves that the change in the first frequency has little impact on the total execution duration, and it can be determined that the current performance limitation bottleneck of the target operator when executed in the first processor is the computing bottleneck. Among them, the second difference range can be set based on the data read / write speed and data computing speed of the first processor for executing the target operator.

[0087] In this embodiment, the implementation logic of the entire method can be as follows:

[0088] If the proportional relationship among t1, t2, and t3 when the first processor executes the target operator is clear, that is, when the data read / write speed and data computing speed of the first processor can be known, one of the "first frequency" and "second frequency" can be first determined as the target frequency. At this time, it is actually speculated that the current performance limitation bottleneck of the target operator is the bottleneck corresponding to the target frequency (that is, if the target frequency is the first frequency, the current performance limitation bottleneck is the computing bottleneck; if the target frequency is the second frequency, the current performance limitation bottleneck is the memory bottleneck).

[0089] If the difference between the obtained multiple total execution durations happens to be within the corresponding difference range (the first difference range or the second difference range), it can be determined that among the importance of the memory bottleneck and the computing bottleneck affecting the performance of the target operator currently, the bottleneck corresponding to the target frequency is more important, and the current performance limitation bottleneck can be determined as the bottleneck corresponding to the target frequency. Alternatively, based on the importance of the memory bottleneck and the computing bottleneck, the bandwidth utilization rate and computing power utilization rate of the target operator at each of the total execution durations, it can be comprehensively evaluated which of the memory bottleneck and the computing bottleneck can be used as the current performance limitation bottleneck. The strategy of this comprehensive evaluation can be adaptively set according to the situation of the first processor and the target operator, and the present disclosure does not limit this.

[0090] If the difference between the obtained multiple total execution durations does not fall within the corresponding difference range (the first difference range or the second difference range), or in other words, it is impossible to determine which of the memory bottleneck and the computing bottleneck is more important, then the remaining one of the frequencies that has not been determined as the target frequency can be further determined as the new target frequency. Then, based on the finally obtained multiple first total execution durations and multiple second total execution durations, the importance of the memory bottleneck and the computing bottleneck in the performance impact on the target operator is determined. Then, the more important one among the importance of the memory bottleneck and the computing bottleneck is determined as the current performance limiting bottleneck. Alternatively, based on the importance of the memory bottleneck and the computing bottleneck, the bandwidth utilization rate and the computing power utilization rate of the target operator under each of the total execution durations, it can be comprehensively evaluated which of the memory bottleneck and the computing bottleneck can be used as the current performance limiting bottleneck. The strategy of this comprehensive evaluation can be adaptively set according to the situation of the first processor and the target operator, and the present disclosure does not limit this.

[0091] That is to say, there are several different implementation manners for step S103:

[0092] Manner 1: When the multiple total execution durations are multiple first total execution durations or multiple second total execution durations, the importance of the transmission duration and the computing duration to the total execution duration can be directly determined according to the multiple first total execution durations or the multiple second total execution durations. Then, according to the importance of the transmission duration and the computing duration, the current performance limiting bottleneck of the target operator when executed in the first processor is determined, that is, the bottleneck with greater importance among the memory bottleneck and the computing bottleneck is used as the current performance limiting bottleneck. Finally, based on the current performance limiting bottleneck, an optimization direction is determined to form an evaluation result.

[0093] Among them, Manner 1 actually applies to the scenario where the multiple first total execution durations or the multiple second total execution durations can determine which of the memory bottleneck and the computing bottleneck is more important.

[0094] Manner 2: When the total execution duration is multiple first total execution durations or multiple second total execution durations, the importance of the transmission duration and the computing duration to the total execution duration can be directly determined according to the multiple first total execution durations or the multiple second total execution durations. Then, based on the importance of the transmission duration and the computing duration, the bandwidth utilization rate and the computing power utilization rate of the target operator under each of the total execution durations, it is comprehensively evaluated which of the memory bottleneck and the computing bottleneck can be used as the current performance limiting bottleneck of the target operator when executed in the first processor. Finally, based on the current performance limiting bottleneck, an optimization direction is determined to form an evaluation result.

[0095] Among them, Method 2 is actually applicable to scenarios where multiple first total execution durations or multiple second total execution durations can determine which of the memory bottleneck and the computing bottleneck is more important. Compared with Method 1, the evaluation results determined are more accurate.

[0096] Method 3: When the total execution duration is multiple first total execution durations (or multiple second total execution durations), it is determined that the importance gap of the transmission duration and the computing duration on the total execution duration is not significant according to the multiple first total execution durations (or multiple second total execution durations); then multiple target frequencies need to be added to obtain multiple second total execution durations (or multiple first total execution durations). Then, based on the obtained multiple first total execution durations and multiple second total execution durations, the importance of the memory bottleneck and the computing bottleneck in the performance impact on the target operator is determined. Then, the more important one among the importance of the memory bottleneck and the computing bottleneck is determined as the current performance limiting bottleneck for the target operator to execute in the first processor. Finally, based on the current performance limiting bottleneck, an optimization direction is determined to form an evaluation result.

[0097] Among them, Method 3 is actually applicable to scenarios where a single multiple first total execution durations or multiple second total execution durations cannot determine which of the memory bottleneck and the computing bottleneck is more important, and it is possible to comprehensively evaluate and determine which of the memory bottleneck and the computing bottleneck is more important by further combining multiple first total execution durations and multiple second total execution durations.

[0098] Method 4: When the total execution duration is multiple first total execution durations (or multiple second total execution durations), it is determined that the importance gap of the transmission duration and the computing duration on the total execution duration is not significant according to the multiple first total execution durations (or multiple second total execution durations); then multiple target frequencies need to be added to obtain multiple second total execution durations (or multiple first total execution durations). Then, based on the obtained multiple first total execution durations and multiple second total execution durations, the importance of the memory bottleneck and the computing bottleneck in the performance impact on the target operator is determined. Then, based on the importance of the transmission duration and the computing duration, the bandwidth utilization rate and the computing power utilization rate of the target operator under each total execution duration, it is comprehensively evaluated which of the memory bottleneck and the computing bottleneck can be used as the current performance limiting bottleneck for the target operator to execute in the first processor. Finally, based on the current performance limiting bottleneck, an optimization direction is determined to form an evaluation result.

[0099] Among them, Method 4 is actually applicable to scenarios where a single multiple first total execution durations or multiple second total execution durations cannot determine which of the memory bottleneck and the computing bottleneck is more important, and it is possible to comprehensively evaluate and determine which of the memory bottleneck and the computing bottleneck is more important by further combining multiple first total execution durations and multiple second total execution durations. Compared with Method 3, the evaluation results determined are more accurate.

[0100] In some embodiments, the evaluation result may further include a performance evaluation result for the target operator. The performance evaluation result may be excellent, medium, poor, etc. Different performance evaluation results corresponding to different total execution durations, bandwidth utilization rates, and computing power utilization rates may be preset, and the present disclosure does not limit this.

[0101] In this way, each time after the optimization of the target operator is completed, the evaluation method for operator optimization in the embodiments of the present disclosure can be used to evaluate the optimization of the target operator. Then, according to the evaluation result, it can be further determined whether the optimization of the target operator can be stopped. And when it is determined that further optimization is required, the target operator can be correspondingly optimized based on the optimization direction in the evaluation result, which can improve the optimization speed and efficiency of the operator and the entire process speed of operator development and tuning.

[0102] As Figure 7 shown, the embodiments of the present disclosure further provide an evaluation device for operator optimization. This device is used to execute the above-mentioned evaluation method for operator optimization. The device includes a frequency setting module 71, a duration determination module 72, and an optimization evaluation module 73.

[0103] The frequency setting module 71 is used to set a plurality of target frequencies required for the first processor to execute the target operator.

[0104] The duration determination module 72 is used to determine the total execution duration consumed by the first processor to execute the target operator at each of the target frequencies.

[0105] The optimization evaluation module 73 is used to evaluate the target operator based on the determined multiple total execution durations to obtain an evaluation result, and the evaluation result is used to indicate the optimization direction of the target operator.

[0106] Wherein, the target frequency includes at least one of a first frequency related to data reading and writing of the first processor and a second frequency related to data calculation of the first processor.

[0107] In a possible implementation manner, the optimization evaluation module 73 includes:

[0108] An evaluation sub-module is used to evaluate the target operator based on the determined multiple total execution durations, the bandwidth utilization rate and the computing power utilization rate of the target operator at each of the total execution durations to obtain an evaluation result.

[0109] In a possible implementation manner, the device further includes:

[0110] A frequency selection module, configured to determine a target frequency from the first frequency and the second frequency according to the data reading / writing speed and data calculation speed of the first processor when it is determined that a target frequency needs to be selected from the first frequency and the second frequency.

[0111] In a possible implementation manner, determining the first frequency or the second frequency as the target frequency according to the data reading / writing speed and data calculation speed of the first processor includes:

[0112] When it is determined according to the data reading / writing speed and the data calculation speed that the transmission duration for the first processor to perform data reading / writing is less than the calculation duration for the processor to perform data calculation, determining the first frequency as the target frequency;

[0113] When it is determined according to the data reading / writing speed and the data calculation speed that the transmission duration for the first processor to perform data reading / writing is greater than the calculation duration for the first processor to perform data calculation, determining the second frequency as the target frequency.

[0114] In a possible implementation manner, the first frequency includes the reading / writing frequency for the load / store unit in the first processor to perform data reading / writing;

[0115] Alternatively, the first frequency includes the reading / writing frequency and the memory frequency of the memory accessed by the load / store unit. When the target frequency is the first frequency and the first frequency includes the reading / writing frequency and the memory frequency, at least one of the reading / writing frequency and the memory frequency of each target frequency is different.

[0116] In a possible implementation manner, when the target frequency includes the first frequency, the multiple total execution durations include multiple second total execution durations; or,

[0117] When the target frequency includes the second frequency, the multiple total execution durations include multiple first total execution durations; or,

[0118] When the target frequency includes the first frequency and the second frequency, the multiple total execution durations include multiple first total execution durations and multiple second total execution durations;

[0119] Wherein, the multiple first total execution durations are the durations consumed by the first processor to respectively execute the target operator when the first frequency remains unchanged and the second frequency is different; the multiple second total execution durations are the second total execution durations consumed by the first processor to respectively execute the target operator when the second frequency remains unchanged and the first frequency is different.

[0120] In a possible implementation, the target operator is evaluated based on the determined multiple total execution durations, the bandwidth utilization rate and the computing power utilization rate of the target operator at each of the total execution durations, and an evaluation result is obtained, including:

[0121] Based on the determined multiple total execution durations, determine the importance of the influence of the transmission duration and the computing duration on the total execution duration;

[0122] Based on the importance of the influence of the transmission duration and the computing duration, and the bandwidth utilization rate and the computing power utilization rate of the target operator at each of the total execution durations, determine the current performance limitation bottleneck of the target operator when executing in the first processor, where the current performance limitation bottleneck includes a memory bottleneck or a computing bottleneck;

[0123] Based on the current performance limitation bottleneck, determine an optimization direction to form an evaluation result.

[0124] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0125] It should be noted that although the above embodiments are used as examples to introduce the evaluation method and device for operator optimization as above, those skilled in the art can understand that the present disclosure should not be limited thereto. In fact, users can flexibly set each step and module according to personal preferences and / or actual application scenarios as long as it conforms to the technical solution of the present disclosure.

[0126] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a third processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0127] The embodiments of the present disclosure also propose an electronic device, including: a second processor; a memory for storing executable instructions of the second processor; wherein, the second processor is configured to implement the above methods when executing the instructions stored in the memory.

[0128] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a fourth processor of an electronic device, the processor in the electronic device executes the above methods.

[0129] It should be noted that the first processor, the second processor, the third processor, the fourth processor, and the fifth processor below involved in the embodiments of the present disclosure may be the same or different processors, and the first processor, the second processor, the third processor, the fourth processor, and the fifth processor can be set according to the requirements for the processor in different situations. The present disclosure does not limit this.

[0130] Figure 8 FIG. 4 is a block diagram of an apparatus 1900 for an electronic device or a server according to an exemplary embodiment. That is, the apparatus 1900 may be provided as a server or a terminal device. Referring to Figure 8 , the apparatus 1900 includes a processing component 1922, which further includes one or more fifth processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0131] The apparatus 1900 may further include a power component 1926 configured to perform power management of the apparatus 1900, a wired or wireless network interface 1950 configured to connect the apparatus 1900 to a network, and an input / output interface 1958 (I / O interface). The apparatus 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , MacOS X TM , Unix TM , Linux TM , FreeBSD TM or the like.

[0132] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the apparatus 1900 to complete the above method.

[0133] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0134] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example - but not limited to - an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being an instantaneous signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0135] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0136] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0137] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.

[0138] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0139] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0140] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0141] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. An evaluation method for operator optimization, characterized in that The method includes: Setting a plurality of target frequencies required for a first processor to execute a target operator, where the target frequencies are configured based on the computing power and bandwidth of the first processor; Determining the total execution duration consumed by the first processor to execute the target operator at each of the target frequencies; Evaluating the target operator based on the determined plurality of total execution durations to obtain an evaluation result, where the evaluation result is used to indicate the optimization direction of the target operator, and the optimization direction corresponds to the current performance limiting bottleneck of the target operator when executed in the first processor; Wherein, the target frequencies include at least one of a first frequency related to data reading and writing of the first processor and a second frequency related to data calculation of the first processor; Wherein, evaluating the target operator based on the determined plurality of total execution durations to obtain an evaluation result includes: Evaluating the target operator according to the determined plurality of total execution durations, the bandwidth utilization rate and computing power utilization rate of the target operator at each total execution duration, and the importance of the transmission duration and calculation duration to the total execution duration to obtain an evaluation result.

2. The method according to claim 1, wherein The method further includes: When it is determined that a target frequency needs to be selected from the first frequency and the second frequency, determining the first frequency or the second frequency as the target frequency according to the data reading and writing speed and data calculation speed of the first processor.

3. The method according to claim 2, wherein Determining the first frequency or the second frequency as the target frequency according to the data reading and writing speed and data calculation speed of the first processor includes: When it is determined according to the data reading and writing speed and the data calculation speed that the transmission duration of the first processor for data reading and writing is less than the calculation duration of the first processor for data calculation, determining the first frequency as the target frequency; When it is determined according to the data reading and writing speed and the data calculation speed that the transmission duration of the first processor for data reading and writing is greater than the calculation duration of the first processor for data calculation, determining the second frequency as the target frequency.

4. The method according to claim 3, wherein The first frequency includes the reading and writing frequency of the load / store unit in the first processor for data reading and writing; Alternatively, the first frequency includes the reading and writing frequency and the memory frequency of the memory accessed by the load / store unit. When the target frequency is the first frequency and the first frequency includes the reading and writing frequency and the memory frequency, at least one of the reading and writing frequency and the memory frequency of each target frequency is different.

5. The method according to claim 1, wherein When the target frequency includes the first frequency, the plurality of total execution durations include a plurality of second total execution durations; or When the target frequency includes the second frequency, the plurality of total execution durations include a plurality of first total execution durations; or When the target frequency includes the first frequency and the second frequency, the plurality of total execution durations include a plurality of first total execution durations and a plurality of second total execution durations; Among them, the multiple first total execution durations are the durations consumed by the first processor to execute the target operator respectively under the condition that the first frequency remains unchanged and the second frequency is different; the multiple second total execution durations are the second total execution durations consumed by the first processor to execute the target operator respectively under the condition that the second frequency remains unchanged and the first frequency is different.

6. The method according to claim 4, characterized in that Based on the determined multiple total execution durations, the bandwidth utilization rate and computing power utilization rate of the target operator under each total execution duration, and the importance of the transmission duration and computing duration to the total execution duration, evaluate the target operator to obtain an evaluation result, including: Based on the determined multiple total execution durations, determine the importance of the transmission duration and the computing duration to the total execution duration; Based on the importance of the transmission duration and the computing duration, and the bandwidth utilization rate and computing power utilization rate of the target operator under each total execution duration, determine the current performance limitation bottleneck when the target operator is executed in the first processor, and the current performance limitation bottleneck includes a memory bottleneck or a computing bottleneck; Based on the current performance limitation bottleneck, determine an optimization direction to form an evaluation result.

7. An evaluation device for operator optimization, characterized in that The device includes: A frequency setting module, configured to set multiple target frequencies required for the first processor to execute the target operator, and the target frequencies are configured based on the computing power and bandwidth of the first processor; A duration determination module, configured to determine the total execution duration consumed by the first processor to execute the target operator under each target frequency; An optimization evaluation module, configured to evaluate the target operator according to the determined multiple total execution durations to obtain an evaluation result, and the evaluation result is used to indicate the optimization direction of the target operator, and the optimization direction corresponds to the current performance limitation bottleneck when the target operator is executed in the first processor; Among them, the target frequencies include at least one of a first frequency related to data reading and writing of the first processor and a second frequency related to data calculation of the first processor; Among them, the optimization evaluation module includes: An evaluation sub-module, configured to evaluate the target operator according to the determined multiple total execution durations, the bandwidth utilization rate and computing power utilization rate of the target operator under each total execution duration, and the importance of the transmission duration and computing duration to the total execution duration, to obtain an evaluation result.

8. An electronic device, characterized in that, Includes: A second processor; A memory for storing instructions executable by the second processor; Among them, the second processor is configured to implement the method according to any one of claims 1 to 6 when executing the instructions stored in the memory.

9. A non - volatile computer - readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions implement the method according to any one of claims 1 to 6 when executed by a third processor.

10. A computer program product, comprising computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, characterized in that, When the computer-readable code runs in the fourth processor of the electronic device, the fourth processor in the electronic device executes the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Voltage frequency adjustment method and device, neural network accelerator and storage medium

    CN116611484A

  • Model-free GPU online energy efficiency optimization method and system

    CN117891677A