Artificial intelligence accelerator, chip, data processing method and electronic device
By using a computation scheduling module in the artificial intelligence accelerator to adjust the computation time of the operator, the problems of system instability caused by unstable current and high cost of power management module are solved, thus achieving current stability and system stability.
Patent Information
- Application Number
- CN202210565284.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2042-05-23
AI Technical Summary
Artificial intelligence accelerators are prone to unstable current during data processing, which can lead to unstable system operating frequency or even system crash. At the same time, the power management module has high design costs.
The operation scheduling module obtains the attribute information of the operators and adjusts the operation time of each operator to ensure that the current output by the power management module is consistent and to avoid current instability.
This improved the system's operational stability and reduced the design cost of the power management module.
Smart Images

Figure CN117172295B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular, to an artificial intelligence accelerator, an artificial intelligence acceleration chip, a data processing method and an electronic device. BACKGROUND
[0002] An artificial intelligence accelerator is a special hardware accelerator or computer system designed to accelerate the application of artificial intelligence and process massive data. The artificial intelligence accelerator includes an operation engine and a power management module. The operation engine can process data, and the power management module supplies power to the operation engine.
[0003] At present, in the data processing process, the artificial intelligence accelerator uses multiple operation engines for operation at different operation times, thereby realizing efficient processing of massive data. However, when the above artificial intelligence accelerator processes data, the current is not stable, which causes the system frequency to be unstable, thereby affecting the system stability. SUMMARY
[0004] The purpose of the present disclosure is to provide an artificial intelligence accelerator, an artificial intelligence acceleration chip, a data processing method and an electronic device, thereby at least to some extent overcoming the problem of system instability caused by the instability of current due to the limitations and defects of the related art.
[0005] According to a first aspect of the present disclosure, an artificial intelligence accelerator is provided, comprising: an operation engine set, a power management module and an operation scheduling module. The operation engine set includes two or more operation engines for operating each operator in a data processing network; the power management module is used to supply power to the operation engines in the operation engine set; the operation scheduling module is used to obtain attribute information of each operator in the data processing network, and adjust the operation time of each operator by using the attribute information of each operator, so that the current output by the power management module is consistent in the data processing process.
[0006] According to a second aspect of the present disclosure, a data processing method is provided, which can be executed by the artificial intelligence accelerator described above.
[0007] According to a third aspect of the present disclosure, an artificial intelligence acceleration chip is provided, which encapsulates the artificial intelligence accelerator described above.
[0008] According to a fourth aspect of the present disclosure, an electronic device is provided, which includes the artificial intelligence accelerator described above.
[0009] In the technical solution provided by some embodiments of the present disclosure, the artificial intelligence accelerator can include a set of operation engines, a power management module, and an operation scheduling module. The set of operation engines includes two or more operation engines configured to perform operations on operators in a data processing network. The power management module is configured to supply power to the operation engines in the set of operation engines. The operation scheduling module is configured to obtain attribute information of the operators in the data processing network, and adjust operation time of the operators based on the attribute information of the operators, so that the current output by the power management module is consistent during the data processing. The operation scheduling module in the artificial intelligence accelerator can obtain the attribute information of the operators and adjust the operation time of the operators based on the attribute information, so as to maintain the stability of the current of the artificial intelligence accelerator during the data processing, thereby avoiding the problem that the power management module provides unstable current to the operation engines at different operation times due to the operation of multiple operation engines at the same operation time in some technologies. The artificial intelligence accelerator provided by the present disclosure can solve the problem of system instability or even crash caused by unstable current, and avoid the problem of high design cost of the power management module caused by unstable current.
[0010] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0011] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure. It is clear that the drawings in the following description are only some embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor. In the drawings:
[0012] Figure 1 A schematic diagram of an artificial intelligence accelerator hardware structure according to an exemplary embodiment of the present disclosure is shown schematically;
[0013] Figure 2 A schematic diagram of a data processing network of an artificial intelligence accelerator according to an exemplary embodiment of the present disclosure is shown schematically;
[0014] Figure 3 A utilization rate diagram of operation engines used by different computing layers according to an exemplary embodiment of the present disclosure is shown schematically;
[0015] Figure 4 A current change diagram according to an exemplary embodiment of the present disclosure is shown schematically;
[0016] Figure 5A system architecture diagram of an artificial intelligence accelerator is illustratively shown according to an example embodiment of the present disclosure;
[0017] Figure 6 A data processing network diagram with adjusted operator computation time is illustratively shown according to an example embodiment of the present disclosure;
[0018] Figure 7 A diagram of adjusting operator computation time within a maximum computation time is illustratively shown according to an example embodiment of the present disclosure;
[0019] Figure 8 A process diagram of adjusting operator computation time by dividing time periods is illustratively shown according to an example embodiment of the present disclosure;
[0020] Figure 9 Another data processing network diagram with adjusted operator computation time is illustratively shown according to an example embodiment of the present disclosure;
[0021] Figure 10 Another utilization diagram of different computation layers using operation engines is illustratively shown according to an example embodiment of the present disclosure;
[0022] Figure 11 A flowchart of a data processing method is illustratively shown according to an example embodiment of the present disclosure;
[0023] Figure 12 A block diagram of an electronic device is illustratively shown according to an example embodiment of the present disclosure. DETAILED DESCRIPTION
[0024] Example embodiments now will be described more fully hereinafter with reference to the accompanying drawings. Example embodiments, however, can be implemented in many different forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the example embodiments to those skilled in the art. The features, structures, or characteristics described in connection with the embodiments can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of embodiments of the present disclosure. One skilled in the relevant art will recognize, however, that the techniques described herein can be practiced without one or more of the specific details, or with other methods, components, materials, and so forth. Other instances of known teachings, procedures and components can be used as well. The disclosure is not limited to the implementation described herein, but can be implemented in any number of ways, including using the alternative method described herein.
[0025] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0026] The flowchart shown in the attached diagram is merely an illustrative example and does not necessarily include all steps. For example, some steps may be broken down, while others may be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0027] The artificial intelligence accelerator provided in this disclosure can be applied to application scenarios that process massive amounts of data. For example, when training a neural network model, it is necessary to acquire a large amount of sample data and learn the neural network model through the sample data. The sample data can be image, audio, video data, etc. Taking an embedded neural network processor (NPU) as an example, the NPU adopts a data-driven parallel architecture, which enables the internal computing engine of the NPU to execute in parallel, thereby achieving efficient processing of massive amounts of video and image data.
[0028] like Figure 1 As shown, Figure 1 The illustration schematically depicts a hardware structure diagram of an artificial intelligence accelerator according to an exemplary embodiment of the present disclosure. Reference Figure 1 The hardware structure of an AI accelerator mainly consists of a set of computing engines and a power management module. The computing engines can include tensor engines, vector engines, and partial summation units. The AI accelerator can process large amounts of data using these computing engines, while the power management module primarily supplies power to the computing engines. The computing engines consume a significant portion of the current during computation.
[0029] exist Figure 1 Based on the AI accelerator shown, it will combine Figure 2 For use Figure 1 The data processing network of the AI accelerator shown is described in detail.
[0030] Figure 2A data processing network diagram of an artificial intelligence accelerator is shown schematically according to an example embodiment of the present disclosure. In Figure 2 The data processing network shown includes a plurality of operators, each of which can perform operations by invoking a set of operation engines, and operators at the same operation time can invoke different types of operation engines in the set of operation engines. Among them, the operator is a computing unit, and the conversion of input data can be realized through the operator. The computing layer includes operators, and the operators in the same computing layer can perform operation processes in parallel. For example, the computing layer can be a convolution computing layer, a pooling computing layer, etc., and the convolution algorithm in the convolution computing layer can be an operator.
[0031] As shown in Figure 2 , it is assumed that operators 2, 4, and 8 need to invoke tensor processing engines to complete operations on the T time axis, operators 3, 5, 7, and 9 need to invoke vector processing engines to complete operations, and operator 6 needs to invoke a partial sum accumulation unit to complete operations. At the same time, in the data processing network, operator 1 is at the first operation time, and it does not invoke the set of operation engines; operators 2 and 3 invoke the tensor processing engine and the vector processing engine respectively and operate in parallel at the second operation time; similarly, operators 4 and 5 invoke the tensor processing engine and the vector processing engine respectively and operate in parallel at the third operation time; operators 6 and 7 invoke the partial sum accumulation unit and the vector processing engine respectively and operate in parallel at the fourth operation time; operators 8 and 9 invoke the tensor processing engine and the vector processing engine respectively and operate in parallel at the fifth operation time; and operator 10 invokes the vector processing engine and operates at the sixth operation time. The artificial intelligence processor can improve the overall operation efficiency by parallel operation of a plurality of operators.
[0032] Based on the hardware structure of the artificial intelligence accelerator shown in Figure 1 , the utilization rate of different operation engines and the current intensity provided by the corresponding power module of each computing layer at different operation times will be described in detail below. Figure 3 , Figure 4
[0033] Figure 3 A utilization rate diagram of operation engines used by different computing layers is shown schematically according to an example embodiment of the present disclosure. The utilization rate of operation engines used by different computing layers causes Figure 3 the current change shown in Figure 4 , Figure 4 A current change diagram is shown schematically according to an example embodiment of the present disclosure.
[0034] Reference is made to Figure 3 In the first operation time, the operation engine set is not called to operate the operator, and the operation engine utilization rate of the artificial intelligence accelerator is 0 at this moment. In the second operation time, part of the tensor processing engines in the artificial intelligence accelerator hardware structure are called to operate the operator in the second calculation layer. Similarly, in the third operation time, all the tensor processing engines in the artificial intelligence accelerator hardware structure are called to operate the operator in the third calculation layer. In the fourth operation time, all the vector processing engines in the artificial intelligence accelerator hardware structure are called to operate the operator in the fourth calculation layer. In the fifth operation time, all the tensor processing engines, vector processing engines, part and accumulation units in the artificial intelligence accelerator hardware structure are called to operate the operator in the fifth calculation layer. In the sixth operation time, all the tensor processing engines and vector processing engines in the artificial intelligence accelerator hardware structure are called to operate the operator in the sixth calculation layer.
[0035] From the above Figure 3 It can be seen that, for a data processing network, the types of operation engines operating in parallel at the same operation time and the utilization rates of the operation engines are different, which can cause Figure 4 the current change process provided by the power management module.
[0036] Figure 4 A current change schematic diagram according to an example embodiment of the present disclosure is schematically shown. Referring to Figure 4 , the operation engine utilization rate of the first operation time is 0, and the intensity of the current is 0. The utilization rates of all types of operation engines of the fifth operation time are 100%, and the intensity of the current is the largest. At the same time, the tensor processing engines are called in the second operation time and the third operation time, and the utilization rate of the tensor processing engine of the second operation time is less than that of the third operation time. At this time, the current intensity of the second operation time is less than that of the third operation time. The types of operation engines called in the second operation time and the fourth operation time are different, which causes different intensities of the current.
[0037] From the above Figure 3 , Figure 4 It can be seen that the intensity of the current is related to the type of the operation engine and the utilization rate of the operation engine, and for the same type of operation engine, the higher the utilization rate of the operation engine, the greater the intensity of the current. Therefore, when the above artificial intelligence accelerator operates the operator in the data processing network, the phenomenon of unstable current is easy to appear. On the one hand, it is easy to cause the instability of the working frequency of the system, and even can cause the system to be paralyzed. On the other hand, the design of the power management module needs to consider that the maximum current intensity is occupied when the artificial intelligence accelerator performs data processing, thereby increasing the design cost of the power management module.
[0038] This exemplary embodiment addresses the aforementioned problems and proposes an artificial intelligence accelerator. First, the accelerator can obtain attribute information of each operator in the data processing network through a computation scheduling module. This operator attribute information may include operator parameters and the computation engine to be invoked for computation. For a given hardware architecture, the utilization rate of the computation engine for that operator can be directly determined through the operator parameters. Then, the computation scheduling module adjusts the computation time of each operator using its attribute information to ensure consistent current output from the power management module during data processing. This process avoids the current instability problem that occurs in the aforementioned artificial intelligence accelerator during data processing, thereby ensuring the stability of the power management module's output current. This not only improves the overall system stability but also reduces the design cost of the power management module.
[0039] Figure 5 A system architecture diagram of an artificial intelligence accelerator provided in this disclosure embodiment is shown below. Figure 5 As shown, the system includes a power management module 50, a set of computing engines 52, and a computing scheduling module 54. The power management module 50 provides power to the set of computing engines 52 and the computing scheduling module 54; the set of computing engines 52 includes at least two types of computing engines, and each computing engine can perform calculations on various operators in the data processing network.
[0040] In a data processing network, the operation scheduling module 54 can acquire attribute information of each operator in the network. This attribute information may include the operator's parameters and the computation engine invoked for that operation. The operator's parameters determine the utilization rate of the computation engine, which represents the utilization rate of hardware resources during the operation. After acquiring the attribute information of each operator, the operation scheduling module 54 can adjust the execution time of each operator to ensure consistent current output from the power management module during data processing. This consistency can be achieved by ensuring that the current variation range between two adjacent operation times does not exceed a threshold, or by ensuring that the current variation range across the entire time axis does not exceed a threshold.
[0041] In the technical solution provided by some embodiments of the present disclosure, the artificial intelligence accelerator can include a set of operation engines, a power management module, and an operation scheduling module. The set of operation engines includes two or more operation engines for performing operations on operators in a data processing network. The power management module is configured to supply power to the operation engines in the set of operation engines. The operation scheduling module is configured to obtain attribute information of the operators in the data processing network, and adjust operation times of the operators based on the attribute information of the operators, so that the current output by the power management module is consistent during the data processing. The operation scheduling module in the artificial intelligence accelerator can adjust the operation times of the operators based on the attribute information of the operators, thereby maintaining the stability of the current during the operation of the artificial intelligence accelerator, avoiding the problem of unstable current provided by the power management module at the same operation time due to the operation of multiple operation engines in some technologies, thereby on the one hand causing system instability or even collapse, and on the other hand increasing the cost of power management module design.
[0042] In an exemplary embodiment of the present disclosure, the set of operation engines can include at least two of a tensor processing engine, a vector processing engine, and a partial and accumulation unit. For example, in the example shown in FIG. 1, the set of operation engines includes the tensor processing engine 101, the vector processing engine 102, and the partial and accumulation unit 103. Figure 2 For example, the operators 2 and 3 at the same operation time are taken as an example. It is assumed that the operator 2 can call the tensor calculation engine for operation, and the operator 3 can call the vector calculation engine for operation. It should be understood that the types of operation engines called by the multiple operators at the same operation time are different, and therefore the types of operation engines included in the set of operation engines can be determined according to the operators operating in parallel at the same operation time in the data processing network. For example, if three operators are executed at the same operation time, the set of operation engines includes the tensor processing engine, the vector processing engine, and the partial and accumulation unit used by the three operators, respectively.
[0043] By calling different types of operation engines to perform parallel operation on the operators at the same operation time, the efficiency of data processing by the artificial intelligence accelerator can be improved. Meanwhile, in order to consider the problem of unstable current caused by using the above parallel operation mode, the operation scheduling module can be used to adjust the operation times of the operators, so that the current output by the power management module is consistent during the data processing.
[0044] The process of adjusting the operation times of the operators by using the operation scheduling module will be described in detail below.
[0045] In an exemplary embodiment of the present disclosure, the data processing network includes a first operator and a second operator. If it is determined that the first operator and the second operator are executed in parallel based on the attribute information of the first operator and the attribute information of the second operator, the operation scheduling module controls the operation time of the first operator to be staggered with the operation time of the second operator.
[0046] Specifically, the operation scheduling module can obtain attribute information of the first operator and the second operator in the data processing network, and determine whether the first operator and the second operator belong to parallel execution through the attribute information. If the first operator and the second operator do not belong to the parallel execution case, the operation scheduling module does not make any operation; if they belong to the parallel execution case, the operation scheduling module can control the operation time of the first operator to be staggered with the operation time of the second operator.
[0047] In this example embodiment, the operation scheduling module can stagger the operation time of the operators in parallel execution through the attribute information of the operators. This method can avoid the problem of current instability caused by the large difference in utilization rate of the operation engine when multiple operators are in parallel execution, on the one hand, ensuring the stability of the system operating frequency, thereby maintaining the system stability, on the other hand, when designing the power management module, there is no need to consider the maximum current that may occur when the artificial intelligence accelerator processes data, thereby reducing the cost of the power management module.
[0048] In the example embodiment of the present disclosure, the operation scheduling module can control the first operation engine to be closed after the first operation engine operates the first operator, and start the second operation engine to operate the second operator after the first operation engine is closed.
[0049] Specifically, the operation scheduling module can determine the first operation engine called when the first operator operates and the second operation engine called when the second operator operates by obtaining the attribute information of the first operator and the second operator. When the operation time of the first operator is staggered with the operation time of the second operator, the operation scheduling module can close the first operation engine after the first operation engine operates the first operator, and then start the second operation engine to operate the second operator. Alternatively, the second operation engine can be called first to operate the second operator, and then the first operation engine can be started to operate the first operator after the second operation engine is closed.
[0050] It should be understood that the above-mentioned way of closing the first operation engine or the second operation engine can release the hardware resources occupied by the operation operator in the hardware structure, or temporarily lock the hardware resources occupied by the operation operator. The present disclosure does not make any limitation on the way of closing the operation engine, which belongs to the scope of protection of the present disclosure.
[0051] By calling the first computing engine to perform operations on the first operator, then shutting down the first computing engine before triggering the second computing engine to perform operations on the second operator, the operation times of the first and second operators are staggered. This ensures that the current occupied by the computing engine during the same operation time will not be much greater than the current during other operation times, thus ensuring the stability of the output current of the power management module, improving system stability, and reducing the design cost of the power management module.
[0052] exist Figure 2 Based on the data processing network shown, and in conjunction with the above exemplary embodiments and Figure 6 An example is provided for a data processing network after the operation time of each operator is adjusted using an operation scheduling module. Figure 6 The illustration schematically shows a data processing network diagram with adjusted operation times for each operator according to an exemplary embodiment of the present disclosure.
[0053] by Figure 2 Taking operators 2 and 3 as an example, assuming the first operator is operator 2 and the second operator is operator 3, the operation scheduling module can determine that operators 2 and 3 are executed in parallel by obtaining their attribute information. Specifically, operator 2 calls the tensor processing engine as its first operation engine, while operator 3 calls the vector processing engine as its second operation engine. Then, the operation scheduling module uses the obtained attribute information to stagger the computation time of operator 2 and operator 3.
[0054] Specifically, such as Figure 6 As shown, firstly, the operation scheduling module can send the output data of operator 1 to the tensor processing engine. Then, after the tensor processing engine completes the operation on operator 2, it sends the output data to the vector processing engine of operator 3, while simultaneously controlling the tensor processing engine to shut down. Finally, it starts the vector processing engine to perform the operation on operator 3. Similarly, Figure 2 The operation times of operators 4 and 5, operators 6 and 7, and operators 8 and 9, which are executed in parallel as shown, can also be staggered by using the above-mentioned operation scheduling module.
[0055] It should be understood that, taking operators 2 and 3 as examples, the operation scheduling module can also prioritize calling the vector processing engine to perform the operation on operator 3, then shut down the vector processing engine and restart the tensor processing engine to perform the operation on operator 2.
[0056] Furthermore, in order to maintain current stability while ensuring high efficiency in the computation of each operator in the data processing network, it is also necessary to ensure that the entire computation process is completed within the maximum computation time required for this data processing procedure. The following will provide a detailed explanation of the process of completing the entire data processing procedure within the maximum computation time and maintaining current stability.
[0057] In an example embodiment of the present disclosure, the operation scheduling module is further configured to obtain a maximum operation time required by the data processing process, and if the operation scheduling module calculates in advance that the operation time required by the data processing process after adjusting the operation time of the first operator and the operation time of the second operator is greater than the maximum operation time, the operation scheduling module is configured to start the second operation engine to operate the second operator during the process in which the first operation engine operates the first operator.
[0058] The maximum operation time is the operation time required by the entire data processing process. Comparing the total operation time required by the data processing process after adjusting the operators with the maximum operation time required by the data processing process, if it does not exceed the maximum operation time, the first operation engine is called to operate the first operator, and then the first operation engine is closed and the second operation engine is started to operate the second operator. If it exceeds the maximum operation time, the second operation engine is started to operate the second operator during the process in which the first operation engine operates the first operator. This process can ensure the stability of the output current of the power management module without affecting the efficiency of the operation, and can also reduce the design cost of the power management module.
[0059] In an example embodiment of the present disclosure, after the operation scheduling module obtains the maximum operation time required by the data processing process, if the operation scheduling module calculates in advance that the operation time required by the data processing process after adjusting the operation time of the first operator and the operation time of the second operator is greater than a threshold, the operation scheduling module is configured to start the second operation engine to operate the second operator during the process in which the first operation engine operates the first operator; if the operation scheduling module calculates in advance that the operation time required by the data processing process after adjusting the operation time of the first operator and the operation time of the second operator is not greater than a threshold, the operation scheduling module controls the first operation engine to be closed after the first operation engine operates the first operator, and then starts the second operation engine to operate the second operator after the first operation engine is closed.
[0060] The following will be described in conjunction with Figure 7 The process of controlling the operation time of the first operator and the operation time of the second operator to be staggered in the case that the operation time required by the data processing process after adjusting the operation time of the first operator and the operation time of the second operator is greater than the maximum operation time is described.
[0061] Figure 7A schematic diagram of adjusting operation time of each operator within maximum operation time is shown according to an example embodiment of the present disclosure. It is assumed that in the third operation time t1, the first operator and the second operator contained in the data processing network are No. 1 operator and No. 2 operator respectively, the first processing engine called is a tensor processing engine to operate No. 1 operator, the second processing engine called is a vector processing engine to operate No. 2 operator, and the running time of No. 1 operator and No. 2 operator is t1-t2 time period. In the fifth operation time t3, the first operator and the second operator contained in the data processing network are No. 3 operator and No. 4 operator respectively, the first processing engine called is a vector processing engine to operate No. 3 operator, the second processing engine called is a partial sum accumulation unit to operate No. 4 operator, and the running time of No. 3 operator and No. 4 operator is t3-t6 time period. Compared with other operation times, the third operation time t1 and the fifth operation time t3 have the problem of unstable current. In order to ensure the stable growth of current and not affect the operation time required by the whole data processing process, the operation scheduling module can adjust the operation time of the first operator and the operation time of the second operator.
[0062] As shown in Figure 7 It is assumed that in the third operation time t1, the operation scheduling module pre-calculates that the operation time required by the data processing process after adjusting the operation time of No. 1 operator and the operation time of No. 2 operator is less than a threshold value from the operation time of No. 1 operator and No. 2 operator executed in parallel, the operation scheduling module first calls the tensor processing engine of No. 1 operator to operate, then closes the tensor processing engine and starts the vector processing engine of No. 2 operator to operate. In this process, the threshold value can be set according to demand, but in order to ensure that the operation efficiency of the artificial intelligence accelerator is not affected, it should be set to be small.
[0063] In the fifth operation time t3, the operation scheduling module pre-calculates that the operation time required by the data processing process after adjusting the operation time of No. 3 operator and the operation time of No. 4 operator is greater than a threshold value from the operation time of No. 3 operator and No. 4 operator executed in parallel, which leads to the operation time required by the whole data processing process being greater than the maximum operation time. In the process of calling the vector processing engine to operate No. 4 operator, the operation scheduling module starts the partial sum accumulation unit to operate No. 5 operator, and controls the corresponding operation engine to be closed after No. 4 operator and No. 5 operator are respectively operated. Figure 7 t3-t4 is the time occupied by only calling the vector processing engine to operate No. 4 operator, t5-t6 is the time occupied by only calling the partial sum accumulation unit to operate No. 5 operator, and t4-t5 is the time occupied by simultaneously calling the vector processing engine and the partial sum accumulation unit to respectively operate No. 4 operator and No. 5 operator.
[0064] In the embodiment, the operation time required by the data processing process after adjustment of each operator is compared with the operation time required by the data processing process before adjustment of each operator, and a scheme for adjusting the operation time of each operator is determined according to the maximum operation time. The scheme can not only ensure the efficiency of the operation of each operator by the artificial intelligence accelerator, but also make the output current of the power management module tend to be stable, thereby improving the stability of the system and the design cost of the power management module.
[0065] In another embodiment of the present disclosure, the operation scheduling module can also determine the operation engine corresponding to each operator and the time occupied by the operation engine for the operation of each operator by using the attribute information of each operator, and adjust the operation time of each operator according to the operation engine corresponding to each operator and the time occupied by the operation engine for the operation of each operator.
[0066] Specifically, the operation time of each operator is the time point at which the operation engine starts to operate the operator, and the time occupied by the operation engine for the operation of the operator is the time period from the start of the operation of the operation engine for the operation of the operator to the stop of the operation of the operation engine for the operation of the operator. By adjusting the operation time of each operator according to the time occupied by each operator for the operation, the stable transformation of the current can be realized, and the stability of the system can be improved.
[0067] In some example embodiments, in the data processing process, for each divided time period, the operation scheduling module adjusts the operation time of each operator according to the operation engine corresponding to each operator and the time occupied by the operation engine for the operation of each operator, so that the current output by the power management module in each time period is consistent.
[0068] Specifically, the time period can be divided according to the operation time of the operator, and the time period can be divided according to actual needs. When dividing the time period, the time period can be divided at equal time intervals, or the time period can be divided according to the utilization rate of the called operation engine. For example, if there are few operators operating in parallel or the utilization rate of the operation engine is low in the entire data processing network, the divided time period is increased; if there are many operators operating in parallel or the utilization rate of the operation engine is high, the divided time period is reduced.
[0069] In the exemplary embodiment of the present disclosure, after the time period is divided, the operation time of each operator can be adjusted according to the utilization rate of the operation engine. For example, in the time period t1-t2, the operator requiring the utilization rate of the tensor processing engine of 50% and the operator requiring the utilization rate of the vector processing engine of 70% are adjusted to the same parallel operation time. In the time period t2-t3, the operator requiring the utilization rate of the tensor processing engine of 40% and the operator requiring the utilization rate of the vector processing engine of 90% are adjusted to the same parallel operation time. At this time, in the time period t1-t3, the utilization rate of the operation engine called by each operator operation time tends to be stable, and the current output by the power management module also tends to be stable, thereby ensuring the stability of the system.
[0070] Next, a process of adjusting the operation time of each operator by dividing the time period will be described with reference to the accompanying drawings. Figure 8 The process of adjusting the operation time of each operator by dividing the time period will be described. Figure 8 A process diagram of adjusting the operation time of each operator by dividing the time period according to an exemplary embodiment of the present disclosure is schematically shown.
[0071] In the above Figure 2 In the data processing network shown in the above Figure 8 As an example of the method of dividing the time period at equal time intervals, it is assumed that the divided time periods are t1-t2 and t2-t3, and t1-t2 is the operation time occupied by the operators 2, 3, 4, and 5, and t2-t3 is the operation time occupied by the operators 6, 7, 8, and 9. Next, the time period t1-t2 will be described as an example.
[0072] In an implementable manner, the operation scheduling module can schedule the operation time of the operators 2, 3, 4, and 5 in the time period t1-t2. For example, the operation scheduling module determines that the total utilization rate of the operation engine called by the parallel operation of the operators 2 and 3 is 80%, and the total utilization rate of the operation engine called by the parallel operation of the operators 4 and 5 is 150%, wherein the utilization rate of the operation engine called by the operator 4 is 70%, and the utilization rate of the operation engine called by the operator 4 is 80%. Then, the operation scheduling module can make the operators 2 and 3 operate in parallel, and then immediately trigger the operation engine of the operator 4 to operate after the operation engine of the operators 2 and 3 is turned off, and then turn off the operation engine after the operation of the operator 4 is completed, and then trigger the operation time of the operator 5, thereby ensuring the stability of the current in the time period t1-t2 and improving the stability of the system.
[0073] In an implementable manner, in the Figure 2In the shown data processing network, it is assumed that the utilization rate of the operation engine called by the operator No. 2 is 30%, the utilization rate of the operation engine called by the operator No. 3 is 90%, the utilization rate of the operation engine called by the operator No. 4 is 110%, and the utilization rate of the operation engine called by the operator No. 5 is 100%. Since the utilization rates of the operation engines of the operators No. 4 and No. 5 are relatively high, the operation of the operator No. 3 can be started in the process of the operation of the operator No. 2 in the time period t1-t2, the operation engine of the operator No. 2 is closed after the operation of the operator No. 2 is completed, the operation engine of the operator No. 4 is started, and the operation engine of the operator No. 5 is started after the operation of the operator No. 3 is completed. This process ensures the current stability in the time period t1-t2, and improves the system stability.
[0074] In another implementable manner, the current stability in the time period t3-t4 can also be achieved by the above scheme, and the operation scheduling module also needs to ensure the current stability between the time period t1-t2 and the time period t2-t3. For example, in the time period t1-t2, the tensor operation engine is executed at a utilization rate of 50% and the vector operation engine is executed at a utilization rate of 70%, and in the time period t2-t3, the tensor operation engine is executed at a utilization rate of 40% and the vector operation engine is executed at a utilization rate of 90%, so as to ensure that the resources of the operation engines are balanced in each time period. On the one hand, the current can be kept stable in the entire data processing process, avoiding the problem that the current intensity exceeds the current intensity that can be provided by the power management module, and improving the stability and safety of the system operating frequency; on the other hand, the power management module design does not need to consider the occasional large current intensity, thereby reducing the cost of the power management module design.
[0075] In the above Figure 8 After the operation time of each operator is adjusted in the exemplary embodiment, it can be obtained that Figure 9 the data processing network shown in Figure 10 and the schematic diagram of another operation engine set working at different operation times shown in. It should be understood that Figure 9 the data processing network shown in Figure 10 and the schematic diagram of another operation engine set working at different operation times shown in. It should be understood that
[0076] Figure 9 Another data processing network schematic diagram after the operation time of each operator is adjusted according to the exemplary embodiment of the present disclosure is schematically shown. As Figure 9As shown, in the t1-t2 time period, the operation scheduling module pre-calculates the operation time of the No. 2 operator and the operation time of the No. 3 operator, and the operation time of the data processing process after adjusting the operation time of the No. 3 operator is greater than the maximum operation time. The operation engine is called to operate the No. 2 and No. 3 operators in parallel, and then the operation engine of the No. 2 and No. 3 operators is closed, and the operation engine of the No. 4 and No. 5 operators is started, wherein the operation time of the No. 4 and No. 5 operators is staggered, that is, after the operation engine operates the No. 4 operator, the operation scheduling module controls the operation engine to be closed, and after the operation engine is closed, the operation engine of the No. 5 operator is started to operate. Similarly, in the t3-t4 time period, the No. 6, No. 7, No. 8 and No. 9 operators are operated in parallel by calling the operation engine, and then the operation engine of the No. 6 and No. 7 operators is closed, and the operation engine of the No. 8 and No. 9 operators is started. The operation time adjustment process of each operator can adjust the structure of the data processing network, so that the current output by the power management module in the data processing process is consistent.
[0077] Figure 10 Another different utilization of the utilization rate of the operation engine of another different calculation layer according to the example embodiment of the present disclosure is schematically shown. As shown in FIG. 6, the operation engine of the No. 1 operator is started to operate the No. 1 operator, and then the operation engine of the No. 1 operator is closed, and the operation engine of the No. 2 operator is started to operate the No. 2 operator. Similarly, the operation engine of the No. 3 operator is started to operate the No. 3 operator, and then the operation engine of the No. 3 operator is closed, and the operation engine of the No. 4 operator is started to operate the No. 4 operator. Similarly, the operation engine of the No. 5 operator is started to operate the No. 5 operator, and then the operation engine of the No. 5 operator is closed, and the operation engine of the No. 6 operator is started to operate the No. 6 operator. Similarly, the operation engine of the No. 7 operator is started to operate the No. 7 operator, and then the operation engine of the No. 7 operator is closed, and the operation engine of the No. 8 operator is started to operate the No. 8 operator. Similarly, the operation engine of the No. 9 operator is started to operate the No. 9 operator, and then the operation engine of the No. 9 operator is closed, and the operation engine of the No. 10 operator is started to operate the No. 10 operator. Figure 10 Figure 8 Figure 9 As shown in the above
[0078] Further, the example embodiment also provides an artificial intelligence acceleration chip.
[0079] The artificial intelligence acceleration chip provided by the example embodiment of the present disclosure can encapsulate the artificial intelligence accelerator in any of the above embodiments. The implementation principle and beneficial effects of the artificial intelligence acceleration chip are similar to those of the artificial intelligence accelerator. For details, refer to the implementation principle and beneficial effects of the artificial intelligence accelerator, which will not be described here.
[0080] Further, the example embodiment also provides a data processing method.
[0081] Figure 11 A flowchart of a data processing method according to the example embodiment of the present disclosure is schematically shown. As shown in FIG. 8, the data processing method includes the following steps. Figure 11 The data processing method can be executed by an electronic device, and the electronic device comprises the artificial intelligence accelerator provided in any of the example embodiments. Figure 11 As shown in the figure, the data processing method provided by the embodiments of the present disclosure can comprise the following steps:
[0082] S110: obtaining to-be-processed data.
[0083] The to-be-processed data can include but is not limited to video, audio, image, text and the like.
[0084] S112: sending the to-be-processed data to an artificial intelligence acceleration chip, so as to process the to-be-processed data by an artificial intelligence accelerator packaged in the artificial intelligence acceleration chip, and obtain a data processing result.
[0085] In the example embodiments of the present disclosure, the artificial intelligence acceleration chip can be integrated in a terminal device, which can be a mobile phone, a tablet computer, a server and the like. The artificial intelligence acceleration chip can execute an artificial intelligence task, which can be a data classification task, a prediction task, a detection task and the like.
[0086] In the technical solution provided by some embodiments of the present disclosure, the data processing method comprises obtaining to-be-processed data, sending the to-be-processed data to an artificial intelligence acceleration chip, processing the to-be-processed data by an artificial intelligence accelerator packaged in the artificial intelligence chip, and obtaining a data processing result. By processing the to-be-processed data by the artificial intelligence accelerator provided by the present disclosure, the stability of the system in the data processing process can be improved.
[0087] In the example embodiments of the present disclosure, a computer readable storage medium is also provided, and the computer readable storage medium stores a program product capable of implementing the method described in the present specification. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which comprises program code for causing a terminal device to execute the steps described in the “example method” section of the present specification according to various example embodiments of the present disclosure when the program product is executed on the terminal device.
[0088] The program product for implementing the above method according to the embodiments of the present disclosure can adopt a portable compact disc read-only memory (CD-ROM) and comprises program code, and can be executed on a terminal device such as a personal computer. However, the program product of the present disclosure is not limited to this, and in this document, the readable storage medium can be any tangible medium containing or storing a program, which can be used by or in conjunction with an instruction execution system, device or apparatus.
[0089] The program product can employ any combination of one or more computer readable media or storage media. The computer readable media or storage media can be a computer readable signal medium or a computer readable storage medium. A computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0090] The computer readable signal medium can include a computer readable data signal embodied in a carrier wave, or a propagated signal, that is based, at least in part, on the computer readable program code. The computer readable signal medium can further be a computer readable storage medium, or any combination of the foregoing. The computer readable storage medium can be a tangible storage medium that is not a signal. The computer readable storage medium can be, for example, but not limited to, a hard disk, a floppy disk, a magnetic disk, an optical disk, a tape, a cassette, a card, a portable memory chip, or any suitable combination of the foregoing.
[0091] The program code embodied on the computer readable media can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0092] The program code can be executed by one or more programmable processors, which can be implemented in one or more computer devices, which can be discrete or integrated. The program code can be stored in a computer readable storage medium, which can be implemented using any appropriate medium for storage of data and / or computer code including non-transitory computer readable media (e.g., magnetic, optical, quantum, or semiconductor storages). The computer program product can also be downloaded to the computer device from, for example, the Internet or a
[0093] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above-described artificial intelligence accelerator is also provided.
[0094] Those skilled in the art can understand that each aspect of the present application can be implemented as a system, a method or a program product. Therefore, each aspect of the present application can be embodied in a form of entirely hardware, entirely software (including firmware, microcode, etc.), or a combination of hardware and software, which can be generically referred to as "circuitry", "module" or "system".
[0095] The electronic device 1200 according to this embodiment of the present application will be described below with reference to Figure 12 Figure 12 The electronic device 1200 shown is merely an example and should not be taken as limiting the functionality or use of embodiments of the present application.
[0096] As shown in Figure 12 The electronic device 1200 is in the form of a general computing device. Components of the electronic device 1200 can include, but are not limited to, the at least one processing unit 1210 described above, the at least one storage unit 1220 described above, a bus 1230 that connects different system components, including the storage unit 1220 and the processing unit 1210, and a display unit 1240.
[0097] The storage unit stores program code that can be executed by the processing unit 1210, so that the processing unit 1210 performs the steps according to various exemplary embodiments of the present application described in the "Exemplary Method" section of the present specification. For example, the processing unit 1210 can perform the data processing method shown in steps S110 and S112.
[0098] The storage unit 1220 can include a readable medium in the form of a volatile storage unit, such as a random access memory (RAM) 12201 and / or a cache memory 12202, and can further include a read-only memory (ROM) 12203.
[0099] The storage unit 1220 can also include program / utility 12204 having a set of at least one program modules 12205, such as an operating system, one or more application programs, other program modules, and program data, and each of these examples, or some combination thereof, can include implementation of a network environment.
[0100] The bus 1230 can represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of a variety of bus structures.
[0101] The electronic device 1200 can also communicate with one or more external devices 1300 such as a keyboard, a pointing device, a Bluetooth device, etc.; and can communicate with one or more devices that enable a user to interact with the electronic device 1200 and / or one or more devices (e.g. routers, modems, etc.) that enable the electronic device 1200 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interface 1250. Still yet, the electronic device 1200 can communicate with one or more networks such as a local area network (LAN), a wide area network (WAN), and / or the Internet through network adapter 1260. As depicted, network adapter 1260 communicates with the other components of the electronic device 1200 through bus 1230. It should be appreciated that although not shown, other hardware and / or software modules could be used in connection with the electronic device 1200. Such modules include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.
[0102] From the above description of embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software, or by software in combination with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash disk, a mobile hard disk, or the like) or a network, and includes a number of instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to perform the methods according to the embodiments of the present disclosure.
[0103] In addition, the above-described diagrams are only schematic illustrations of the processes included in the method according to the example embodiments of the present application, and are not intended to be limiting. It is easy to understand that the processes shown in the above-described diagrams do not indicate or limit the time sequence of the processes. In addition, it is also easy to understand that the processes can be executed synchronously or asynchronously, for example, in a plurality of modules.
[0104] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. Indeed, according to embodiments of the present disclosure, the features and functionalities of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functionalities of one module or unit described above can be further divided into embodied by a plurality of modules or units.
[0105] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the features disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the disclosure being indicated by the following claims.
[0106] It should be understood that the present disclosure is not limited to the precise structures herein described and illustrated in the drawings, and that various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the claims that follow.
Claims
1. An artificial intelligence accelerator, characterized in that, include: A collection of computing engines, including two or more computing engines, used to perform operations on various operators in the data processing network; The power management module is used to supply power to the computing engines in the computing engine set; The operation scheduling module is used to obtain the attribute information of each operator in the data processing network, and use the attribute information of each operator to adjust the operation time of each operator so that the current output by the power management module is consistent during the data processing.
2. The artificial intelligence accelerator according to claim 1, characterized in that, The data processing network includes a first operator and a second operator; wherein, the process by which the operation scheduling module adjusts the operation time of each operator using the attribute information of each operator includes: If the attribute information of the first operator and the attribute information of the second operator determine that the first operator and the second operator are executed in parallel, the operation scheduling module controls the operation time of the first operator to be staggered from the operation time of the second operator.
3. The artificial intelligence accelerator according to claim 2, characterized in that, The first operator calls the first computing engine in the computing engine set to perform the operation, and the second operator calls the second computing engine in the computing engine set to perform the operation; wherein, the process by which the computing scheduling module controls the computing time of the first operator and the computing time of the second operator to be staggered includes: After the first computing engine performs calculations on the first operator, the computing scheduling module controls the first computing engine to shut down, and after the first computing engine shuts down, it starts the second computing engine to perform calculations on the second operator.
4. The artificial intelligence accelerator according to claim 2, characterized in that, The first operator calls the first computing engine in the computing engine set to perform the operation, and the second operator calls the second computing engine in the computing engine set to perform the operation; wherein, the computing scheduling module is also used to obtain the maximum computing time required by the data processing process; If the computation time required for the data processing after adjusting the computation time of the first operator and the computation time of the second operator is calculated in advance to be greater than the maximum computation time, then the computation scheduling module is used to start the second computation engine to perform computation on the second operator while the first computation engine is performing computation on the first operator.
5. The artificial intelligence accelerator according to claim 1, characterized in that, The process by which the operation scheduling module adjusts the operation time of each operator using the attribute information of each operator includes: The operation scheduling module uses the attribute information of each operator to determine the operation engine corresponding to each operator and the time occupied by the operation engine in performing operations on each operator, and adjusts the operation time of each operator according to the operation engine corresponding to each operator and the time occupied by the operation engine in performing operations on each operator.
6. The artificial intelligence accelerator according to claim 5, characterized in that, The process by which the operation scheduling module adjusts the operation time of each operator based on the operation engine corresponding to each operator and the time occupied by the operation engine in performing operations on each operator includes: During the data processing, for each divided time period, the operation scheduling module adjusts the operation time of each operator according to the operation engine corresponding to each operator and the time occupied by the operation engine to perform operations on each operator, so as to make the current output by the power management module consistent in each time period.
7. The artificial intelligence accelerator according to any one of claims 1 to 6, characterized in that, The set of computation engines includes at least two of the following: a tensor processing engine, a vector processing engine, and partial and accumulator units.
8. A data processing method, characterized in that, The data processing method is performed by the artificial intelligence accelerator as described in any one of claims 1-7.
9. An artificial intelligence acceleration chip, characterized in that, The artificial intelligence acceleration chip contains an artificial intelligence accelerator as described in any one of claims 1-7.
10. An electronic device, characterized in that, Including the artificial intelligence accelerator as described in any one of claims 1-7.
Citation Information
Patent Citations
Database System with Methodology for Parallel Schedule Generation in a Query Optimizer
US20060080285A1
Host CPU-assisted audio processing method and computing system performing the same
US20170147282A1