Artificial intelligence chip based on data flow and driving method and device thereof
By introducing multiplexing circuits and second multiplexing circuits into the artificial intelligence chip, direct data transmission between the storage module and the computing circuit is realized, solving the problems of computing speed and power consumption, and improving the overall computing efficiency and storage space utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN CORERAIN TECH CO LTD
- Filing Date
- 2021-10-25
- Publication Date
- 2026-05-19
AI Technical Summary
The computation speed of existing dataflow-based artificial intelligence chips is relatively low, mainly because different computing circuits need to move data from one storage module to another when performing calculations, which leads to increased data transmission latency and power consumption.
A multiplexing circuit is used to connect the storage module and the computing circuit. Data is directly transmitted between the storage module and the computing circuit through drive signals, avoiding the need to move data between storage modules. A second multiplexing circuit optimizes the utilization of storage space and allows for flexible data storage.
It improves the computing speed of artificial intelligence chips, reduces computing power consumption, increases the utilization rate of storage modules, and reduces the frequency of access to external storage structures.
Smart Images

Figure CN116029386B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence, and in particular to a data flow-based artificial intelligence chip and its driving method and apparatus. Background Technology
[0002] As a crucial component of artificial intelligence technology, machine learning is widely applied across various industries to improve productivity. However, the complexity of machine learning algorithms and the large amount of computation involved limit the running speed of machine learning algorithm models.
[0003] In related technologies, customized dataflow-based artificial intelligence (AI) chips can be used to execute machine learning algorithms, thereby improving the running speed of machine learning algorithm models. Such AI chips include multiple computing circuits for performing various types of computations within the machine learning algorithms. These computing circuits are connected to storage units within the AI chip, allowing data needed for computation to be read directly from the storage units when computation is required. This reduces data transmission latency, thereby increasing the computational speed of the AI chip. Summary of the Invention
[0004] The inventors noted that, under the methods employed in related technologies, the computational speed of artificial intelligence chips remains relatively low.
[0005] The inventors discovered through analysis that the storage unit in this artificial intelligence chip is designed as multiple storage modules, each dedicated to storing a specific type of data. To facilitate data retrieval by the computing circuit, each storage module is connected to a computing circuit that requires the corresponding type of data during computation.
[0006] However, in real-world applications, it's common for different computing circuits to require the same data when performing calculations. In such cases, time is spent transferring data from one storage module to another, thus reducing the computing speed of the AI chip.
[0007] To address the aforementioned problems, the present disclosure proposes the following solutions.
[0008] According to one aspect of the present disclosure, a dataflow-based artificial intelligence chip is provided, comprising: a plurality of storage modules; a plurality of computing circuits, the different computing circuits being configured to perform different types of computations in a machine learning algorithm model; and a first multiplexing circuit configured to, in response to a first driving signal corresponding to a task, read first data from a first group of storage modules corresponding to the first driving signal, and transmit the first data to the first group of computing circuits corresponding to the first driving signal, the first multiplexing circuit comprising: a plurality of first input terminals connected one-to-one with the plurality of storage modules, and a plurality of first output terminals, including a first group of first output terminals connected one-to-one with the plurality of computing circuits.
[0009] In some embodiments, the first multiplexing circuit is further configured to: read the first data in the first set of storage modules in response to another first drive signal corresponding to another task, and transmit the first data to a second set of computing circuits corresponding to the other first drive signal.
[0010] In some embodiments, the artificial intelligence chip further includes: a second multiplexing circuit configured to, in response to a second driving signal corresponding to the task, read second data from a storage structure outside the artificial intelligence chip and store the second data in a second set of storage modules corresponding to the second driving signal, wherein the second data is included in the first data and the second set of storage modules is included in the first set of storage modules, wherein the second multiplexing circuit includes: a plurality of second input terminals, including a first set of second input terminals connected to the storage structure, and a plurality of second output terminals, which are connected one-to-one with the plurality of storage modules.
[0011] In some embodiments, the second multiplexing circuit is further configured to: in response to the second drive signal, read data of the same type from the second data and store the data of the same type in at least two storage modules in the second group of storage modules.
[0012] In some embodiments, the plurality of first output terminals further include a second group of first output terminals connected to the storage structure; the first multiplexing circuit is further configured to: in response to a third driving signal, read third data in a third group of storage modules corresponding to the third driving signal, and store the third data in the storage structure; wherein at least one of the third group of storage modules is included in the second group of storage modules.
[0013] In some embodiments, the plurality of second input terminals further include a second group of second input terminals connected one-to-one with the output terminals of the plurality of computing circuits; the second multiplexing circuit is further configured to: in response to a fourth driving signal, store the output of a computing circuit into a storage module corresponding to the fourth driving signal.
[0014] In some embodiments, the machine learning algorithm model includes a neural network algorithm model, and the task is the computation of one or at least two consecutive computational layers of the neural network algorithm model.
[0015] According to another aspect of the embodiments of this disclosure, a driving method for a dataflow-based artificial intelligence chip according to any of the above embodiments is provided, comprising: determining a first set of computing circuits required to perform a task, the task including at least one type of computation corresponding to the first set of computing circuits; determining a first set of storage modules corresponding to first data required to perform the task; and sending a first driving signal corresponding to the first set of computing circuits and the first set of storage modules to a first multiplexing circuit, so that the first multiplexing circuit reads the first data and transmits the first data to the first set of computing circuits.
[0016] In some embodiments, the method further includes: determining a second set of computing circuits required to perform another task, the other task including at least one type of computation corresponding to the second set of computing circuits; sending another first drive signal corresponding to the second set of computing circuits and the first set of storage modules to the first multiplexing circuit, so that the first multiplexing circuit reads the first data and transmits the first data to the second set of computing circuits.
[0017] In some embodiments, the artificial intelligence chip further includes a second multiplexing circuit, the second multiplexing circuit including a first set of second input terminals connected to a storage structure outside the artificial intelligence chip, and a plurality of second output terminals connected one-to-one with the plurality of storage modules; the method further includes: determining a first capacity required to store each type of data in the second data in the storage structure, the second data being included in the first data; determining a fourth set of idle storage modules among the plurality of storage modules; determining a second capacity of each storage module in the fourth set of storage modules; determining a second set of storage modules corresponding to the second data from the fourth set of storage modules based on the first capacity and the second capacity, the second set of storage modules being included in the first set of storage modules; sending a second driving signal corresponding to the second set of storage modules to the second multiplexing circuit, so that the second multiplexing circuit reads the second data and stores the second data in the second set of storage modules.
[0018] In some embodiments, determining the second group of storage modules corresponding to the second data from the fourth group of storage modules based on the first capacity and the second capacity includes: determining at least two storage modules corresponding to a certain type of data in the second data when the first capacity of a certain type of data is greater than the second capacity of each storage module in the fourth group of storage modules; wherein the second group of storage modules includes the at least two storage modules.
[0019] In some embodiments, the plurality of first output terminals further include a second group of first output terminals connected to the storage structure; the method further includes: before determining the fourth group of storage modules, determining a fifth group of idle storage modules among the plurality of storage modules, wherein the total capacity of the fifth group of storage modules is less than the total capacity required to store the second data, or the number of storage modules in the fifth group of storage modules is less than the number of data types in the second data; determining a third group of storage modules corresponding to the third data required by other tasks after the task; sending a third driving signal corresponding to the third group of storage modules to the first multiplexing circuit, so that the first multiplexing circuit reads the third data and stores the third data in the storage structure.
[0020] In some embodiments, the plurality of second input terminals further include a second group of second input terminals connected one-to-one with the output terminals of the plurality of computing circuits; the method further includes: determining a third capacity required to store the output of the computing circuit performing the calculation; determining a storage module corresponding to the output based on the second capacity and the third capacity; sending a fourth drive signal corresponding to the computing circuit and the storage module to the second multiplexing circuit, so that the second multiplexing circuit stores the output in the storage module.
[0021] In some embodiments, the task is the computation of one or at least two consecutive computational layers of a neural network algorithm model.
[0022] According to another aspect of the present disclosure, a driving device for a dataflow-based artificial intelligence chip according to any of the above embodiments is provided, comprising: a determining module configured to determine a first set of computing circuits required to perform a task, the task including at least one type of computation corresponding to the first set of computing circuits; determining a first set of storage modules corresponding to first data required to perform the task; and a sending module configured to send a first driving signal corresponding to the first set of computing circuits and the first set of storage modules to a first multiplexing circuit, so that the first multiplexing circuit reads the first data and transmits the first data to the first set of computing circuits.
[0023] According to another aspect of the present disclosure, a driving device for a dataflow-based artificial intelligence chip according to any of the above embodiments is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute the driving method described in any of the above embodiments based on instructions stored in the memory.
[0024] According to another aspect of the present disclosure, an artificial intelligence accelerator is provided, comprising: a data flow-based artificial intelligence chip as described in any of the foregoing embodiments; and a driving device as described in any of the foregoing embodiments.
[0025] According to another aspect of the present disclosure, a server is provided, including: the artificial intelligence accelerator described in any of the foregoing embodiments.
[0026] According to another aspect of the present disclosure, a computer-readable storage medium is provided, including computer program instructions, wherein the computer program instructions, when executed by a processor, implement the driving method described in any of the above embodiments.
[0027] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program, when executed by a processor, implements the driving method described in any of the above embodiments.
[0028] In this embodiment, multiple storage modules are connected to multiple computing circuits via a first multiplexing circuit, rather than being directly connected to the computing circuits. The first multiplexing circuit can read the required data from any storage module according to a drive signal and transmit the read data to any computing circuit, eliminating the need for time-consuming data transfer between storage modules. This improves the computing speed of the artificial intelligence chip.
[0029] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0030] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1A This is a schematic diagram of the structure of a dataflow-based artificial intelligence chip according to some embodiments of the present disclosure;
[0032] Figure 1BThis is a flowchart illustrating a driving method for a dataflow-based artificial intelligence chip 100 according to some embodiments of the present disclosure;
[0033] Figure 2A This is a schematic diagram of the structure of a dataflow-based artificial intelligence chip according to other embodiments of this disclosure;
[0034] Figure 2B This is a flowchart illustrating a driving method for a dataflow-based artificial intelligence chip 200 according to some embodiments of the present disclosure;
[0035] Figure 3 This is a schematic diagram of the structure of a driving device for a data stream-based artificial intelligence chip according to some embodiments of the present disclosure;
[0036] Figure 4 This is a schematic diagram of the structure of a driving device for a data stream-based artificial intelligence chip according to other embodiments of this disclosure;
[0037] Figure 5 This is a schematic diagram of the structure of an artificial intelligence accelerator according to some embodiments of the present disclosure. Detailed Implementation
[0038] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0039] Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0040] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0041] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0042] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0043] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0044] Figure 1A This is a schematic diagram of the structure of a dataflow-based artificial intelligence chip according to some embodiments of the present disclosure.
[0045] like Figure 1A As shown, the data flow-based artificial intelligence chip 100 includes multiple storage modules ( Figure 1A Four are schematically shown: storage module 11a, storage module 11b, storage module 11c, and storage module 11d, and multiple computing circuits. Figure 1A Four circuits are schematically shown: computing circuit 12a, computing circuit 12b, computing circuit 12c, and computing circuit 12d, and a first multiplexing circuit 13.
[0046] Multiple storage modules can be, for example, random access memory (RAM).
[0047] Different computing circuits are configured to perform different types of computations in machine learning algorithm models. For example, different computing circuits can be configured to perform different types of computations in neural network algorithm models, including, for example, convolution computations, shortcut computations, etc.
[0048] The first multiplexing circuit 13 includes multiple first input terminals 131 and multiple first output terminals.
[0049] Multiple first input terminals 131 are connected to multiple storage modules in a one-to-one correspondence. Multiple first output terminals include a first group of first output terminals 132a connected to multiple computing circuits in a one-to-one correspondence.
[0050] The first multiplexing circuit 13 is configured to read first data from the first set of storage modules corresponding to the first drive signal in response to the first drive signal, and transmit the first data to the first set of computing circuits corresponding to the first drive signal.
[0051] In some embodiments, the task is to compute one of the multiple computational layers of a neural network algorithm model or at least two consecutive computational layers.
[0052] It should be understood that the first set of storage modules may include one or more of a plurality of storage modules, and the first set of computing circuits may include one or more of a plurality of computing circuits.
[0053] In the above embodiments, multiple storage modules are connected to multiple computing circuits via a first multiplexing circuit, rather than being directly connected to the multiple computing circuits. The first multiplexing circuit can read the required data from any storage module according to a drive signal and transmit the read data to any computing circuit, eliminating the need for time-consuming data transfer between storage modules. This improves the computing speed of the artificial intelligence chip.
[0054] Furthermore, in related technologies, data cannot be directly moved from one storage module to another. Instead, it needs to be moved first from the original storage module to a storage structure outside the AI chip, and then from that storage structure to another storage module. This increases the frequency with which the AI chip accesses external storage structures, thereby increasing the power consumption of the AI chip during computation. In some embodiments of this disclosure, different computing circuits can use the same data when performing computations without complex data transfer, thereby reducing the frequency with which the AI chip accesses external storage structures and reducing the power consumption of the AI chip during computation.
[0055] Can be executed Figure 1B The driving method shown is used to drive the artificial intelligence chip 100. Figure 1B This is a flowchart illustrating a driving method for a dataflow-based artificial intelligence chip 100 according to some embodiments of the present disclosure.
[0056] like Figure 1B As shown, the driving method of the data flow-based artificial intelligence chip 100 includes steps 110 to 130.
[0057] In step 110, the first set of computing circuits required to perform the task is determined.
[0058] For example, the specific architecture of the machine learning algorithm model to be run can be obtained from a host computer (e.g., a computer, server, etc.) via a bus. Taking a neural network algorithm model as an example, the type of computation of the computation layer corresponding to the current task can be determined based on the specific architecture of the neural network algorithm model. Then, the first set of computation circuits to perform the corresponding type of computation can be determined.
[0059] In step 120, the first set of storage modules corresponding to the first data required to perform the task is determined.
[0060] For example, multiple storage modules have multiple preset addresses, each address belonging to a specific storage module. The address of the first piece of data within these multiple storage modules can be determined first, and the corresponding first group of storage modules can be identified accordingly.
[0061] In step 130, a first drive signal corresponding to the first set of computing circuits and the first set of storage modules is sent to the first multiplexing circuit, so that the first multiplexing circuit reads the first data and transmits the first data to the first set of computing circuits.
[0062] For example, it can be via the signal input terminal 133 of the first multiplexing circuit 13 (see...) Figure 1A A first drive signal is sent to the first multiplexing circuit 13. In response to the first drive signal corresponding to the task, the first multiplexing circuit 13 can first adjust the connection relationship between multiple first input terminals 131 and the first group of first output terminals 132a, so that the first group of storage modules is connected to the first group of computing circuits via the first multiplexing circuit 13. Then, the first multiplexing circuit 13 can read first data from the first group of storage modules and transmit the first data to the first group of computing circuits.
[0063] It should be understood that, in response to different first drive signals, the first multiplexing circuit 13 can adjust the different connection relationships between multiple first input terminals 131 and the first group of first output terminals 132a.
[0064] For example, in response to a first drive signal carrying an instruction of "1100", the first multiplexing circuit 13 can connect the storage module 11a to the computing circuit 12a and the storage module 11b to the computing circuit 12b; in response to a first drive signal carrying an instruction of "1101", the first multiplexing circuit 13 can connect the storage module 11b to the computing circuit 12a and the storage module 11a to the computing circuit 12b; in response to a first drive signal carrying an instruction of "1110", the first multiplexing circuit 13 can connect the computing circuit 12a to the storage modules 11a and 11b respectively.
[0065] These are just some examples. In practical applications, the first drive signal can be written into the AI chip 100 according to the register configuration of the AI chip 100.
[0066] In some embodiments, the first drive signal carries address information representing the address of the first data in multiple storage modules. The first multiplexing circuit 13 can decode the address information to obtain the address of the first data in the multiple storage modules. Then, the first multiplexing circuit 13 can read the first data from this address.
[0067] After the first data is transmitted to the first set of computing circuits, the first set of computing circuits can process the first data to perform tasks.
[0068] In the above embodiments, after determining the storage module corresponding to the data required for the task and the computing circuit for the task, a corresponding drive signal is sent to the first multiplexing circuit. This allows the first multiplexing circuit to accurately read the data required for the task and accurately transmit the data to the computing circuit, thereby enabling the artificial intelligence chip to accurately perform calculations.
[0069] The artificial intelligence chip 100 and its driving method are further illustrated below with reference to some embodiments.
[0070] In some embodiments, the first multiplexing circuit 13 may also be configured to read first data from the first set of storage modules in response to another first drive signal corresponding to another task, and transmit the first data to the second set of computing circuits corresponding to the other first drive signal.
[0071] It should be understood that the second set of computing circuits is different from the first set of computing circuits. For example, the first set of computing circuits includes computing circuit 12a, while the second set of computing circuits includes computing circuit 12d.
[0072] For example, the task in step 110 is to compute a certain computational layer of the neural network algorithm model, while another task could be to compute other computational layers before or after that computational layer in the same neural network algorithm model.
[0073] It can be done in a manner similar to Figure 1B The first multiplexing circuit 13 sends another first drive signal corresponding to another task in a certain manner. That is, it can determine the second set of computing circuits required to perform the other task. Similarly, the other task includes at least one type of computation corresponding to the second set of computing circuits. Then, another first drive signal corresponding to the second set of computing circuits and the first set of storage modules can be sent to the first multiplexing circuit 13 so that the first multiplexing circuit 13 reads the first data and transmits the first data to the second set of computing circuits.
[0074] In the above embodiment, two first drive signals corresponding to each task are sent to the first multiplexing circuit 13. Thus, different first and second sets of computing circuits can use the same first data without data transfer, thereby improving the computing speed of the AI chip and reducing the power consumption of the AI chip during computation.
[0075] In some embodiments, the first set of computing circuits includes at least two computing circuits that perform calculations sequentially, and the first data includes at least two sets of first data corresponding one-to-one with the at least two computing circuits.
[0076] In these embodiments, the first multiplexing circuit 13 may also be configured to read at least two sets of first data corresponding one-to-one with at least two computing circuits in this order.
[0077] The order in which at least two computing circuits perform calculations can be determined, and a first drive signal carrying information indicating this order is sent to the first multiplexing circuit 13, so that the first multiplexing circuit 13 reads at least two sets of first data in this order.
[0078] For example, the first set of computing circuits includes computing circuits 12a and 12b that perform calculations sequentially. Computing circuit 12b needs to process the calculation results of computing circuit 12a. The information indicating the order carried by the first drive signal can be, for example, the time A at which to start reading a set of data required by computing circuit 12a and the time B at which to start reading another set of data required by computing circuit 12b. The time interval between time A and time B can be the duration required for computing circuit 12a to complete at least part of the calculation. The duration required for computing circuit 12a to complete at least part of the calculation can be estimated based on the amount of data to be processed by computing circuit 12a.
[0079] In the above embodiment, the first multiplexing circuit 13 reads the first data in the order in which the first group of computing circuits perform calculations. This allows the computing circuits to perform calculations in a data stream-driven manner, thereby improving the computing speed of the artificial intelligence chip.
[0080] The inventors also noted that in related technologies, the storage modules for storing a certain type of data are fixed; that is, the available capacity for storing that type of data is also fixed. This means that when the capacity of a certain storage module is insufficient to store the corresponding data, the data needs to be stored and processed in batches. This further limits the computing speed of artificial intelligence chips.
[0081] However, the inventors discovered that performing the computation corresponding to a task only requires a few types of data. In other words, when performing the computation corresponding to the task, only the storage space of a few storage modules is utilized, while the storage space of the remaining storage modules is wasted. In view of this, this disclosure also provides the following solution.
[0082] Figure 2A This is a schematic diagram of the structure of a dataflow-based artificial intelligence chip according to other embodiments of this disclosure.
[0083] like Figure 2A As shown, in addition to multiple storage modules, multiple computing circuits and the first multiplexing circuit 13, the artificial intelligence chip 200 also includes a second multiplexing circuit 21.
[0084] The second multiplexing circuit 21 includes multiple second input terminals and multiple second output terminals 212.
[0085] Multiple second input terminals include a first set of second input terminals 211a connected to a storage structure SS (e.g., computer memory) outside the artificial intelligence chip 200. Figure 2A (Two are shown schematically). Multiple second output terminals 212 are connected to multiple storage modules one by one.
[0086] The second multiplexing circuit 21 is configured to read the second data in the storage structure SS in response to the second drive signal corresponding to the task, and store the second data into the second set of storage modules corresponding to the second drive signal.
[0087] Here, the second data is contained within the first data, and the second set of storage modules is contained within the first set of storage modules.
[0088] For example, before the task begins, all the data required for computation by computing circuits 12a and 12b is stored in storage structure SS. This data can be stored as second data in the second set of storage modules (e.g., storage modules 11a and 11b) by sending a second drive signal to the second multiplexing circuit 21. After this data begins to be stored in storage modules 11a and 11b, sending a first drive signal corresponding to the first set of storage modules (i.e., storage modules 11a and 11b) and the first set of computing circuits (i.e., computing circuits 12a and 12b) to the first multiplexing circuit 13 transmits this data as first data to computing circuits 12a and 12b for computation.
[0089] In the above embodiments, the second multiplexing circuit, connected to multiple storage modules, can store one type of data into any storage module according to the drive signal, eliminating the need to store one type of data in a fixed corresponding storage module. This allows for more efficient use of the storage space of multiple storage modules, improving the problem of data needing to be stored and processed in batches, thereby further increasing the computing speed of the artificial intelligence chip.
[0090] Can be executed Figure 2B The driving method shown is used to drive the artificial intelligence chip 200. Figure 2B This is a flowchart illustrating a driving method for a dataflow-based artificial intelligence chip 200 according to some embodiments of the present disclosure.
[0091] like Figure 2B As shown, the driving method of the artificial intelligence chip 200 includes steps 210 to 250.
[0092] In step 210, the first capacity required for each type of data in the second data in the storage structure is determined.
[0093] For example, the second data in the storage structure SS is the data required for the current task in the neural network algorithm model. The initial capacity required to store each type of data in the second data can be determined based on the specific architecture of the neural network algorithm model and the amount of data input to the neural network algorithm model.
[0094] Taking convolution computation as an example, the types of data required for convolution computation include, for example, input feature map data and bias data. The lengths of the input feature map data and bias data corresponding to the current task can be determined separately, thus obtaining the initial capacity required to store each type of data.
[0095] In step 220, a fourth group of storage modules that is free among the multiple storage modules is identified.
[0096] For example, it can be determined whether subsequent tasks will still need the data currently stored in a particular storage module. If it is no longer needed, the storage module is determined to be idle; if it is needed, the storage module is determined to be non-idle.
[0097] In step 230, the second capacity of each storage module in the fourth group of storage modules is determined.
[0098] In step 240, the second group of storage modules corresponding to the second data is determined from the fourth group of storage modules based on the first capacity and the second capacity.
[0099] The storage module corresponding to each type of data in the second data can be determined to obtain the second set of storage modules corresponding to the second data.
[0100] For example, the second set of data includes input feature map data and bias data. The required capacity for storing the input feature map data is 180, and the required capacity for storing the bias data is 120. The fourth set of storage modules includes storage modules 11a to d, with capacities of 50, 80, 100, and 200, respectively. It can be determined that the input feature map data corresponds to storage module 11d with a capacity of 200, and the bias data corresponds to two storage modules 11a and 11b with capacities of 50 and 80, respectively.
[0101] In step 250, a second drive signal corresponding to the second set of storage modules is sent to the second multiplexing circuit, so that the second multiplexing circuit reads the second data and stores the second data in the second set of storage modules.
[0102] For example, it can be via the signal input terminal 213 of the second multiplexing circuit 21 (see...) Figure 2AA second drive signal is sent to the second multiplexing circuit 21. In response to the second drive signal, the second multiplexing circuit 21 can first adjust the connection relationship between the first group of second input terminals 211a and the multiple second output terminals 212, so that the storage structure SS is connected to the second group of storage modules via the second multiplexing circuit 21. Then, the second multiplexing circuit 21 can read the second data and transmit the second data to the second group of storage modules.
[0103] In some embodiments, the second drive signal carries address information representing the address of the second data in the storage structure SS. The second multiplexing circuit 21 can decode the address information to obtain the address of the second data in the storage structure SS. Then, the second multiplexing circuit 21 can read the second data from this address.
[0104] In some embodiments, the second drive signal also carries information indicating the address corresponding to each type of data in the second data. The second multiplexing circuit 21 can store the second data into the corresponding address based on this information.
[0105] In the above embodiments, the storage module corresponding to the data is determined based on the capacity required to store each type of data and the capacity of the available storage modules, and a corresponding second drive signal is sent to the second multiplexing circuit. In this way, the second multiplexing circuit can be driven to flexibly store data in any suitable available storage module, thereby improving the utilization rate of the storage space of multiple storage modules and ensuring that the problem of data needing to be stored and processed in batches can be alleviated.
[0106] The following examples further illustrate the artificial intelligence chip 200 and its driving method.
[0107] In some embodiments, if the first capacity of a certain type of data in the second data is greater than the second capacity of each storage module in the fourth group of storage modules, at least two storage modules corresponding to that type of data can be determined. Here, the determined at least two storage modules corresponding to that type of data are included in the second group of storage modules corresponding to the second data. For example, a corresponding second drive signal can be sent to the second multiplexing circuit 21 to cause the second multiplexing circuit 21 to read the data of that type in the second data and store the data of that type in the determined at least two storage modules.
[0108] In the above embodiments, when a certain type of data exceeds the available capacity of each free storage module, multiple storage modules corresponding to that type of data can be identified, and corresponding drive signals can be sent to cause the second multiplexing circuit to store this data in the identified multiple storage modules. Subsequently, when this type of data is used, the first multiplexing circuit reads this data from the multiple storage modules for the computing circuit to complete the calculation in one go. This further improves the problem of data needing to be stored and calculated in batches, thereby further increasing the computing speed of the artificial intelligence chip.
[0109] In some embodiments, see Figure 2A The first multiplexer circuit 13 of the artificial intelligence chip 200 also includes a second set of first output terminals 132b connected to the storage structure SS. Figure 2A (One is shown schematically).
[0110] In these embodiments, the first multiplexing circuit 13 may also be configured to, in response to a third drive signal, read third data from a third set of storage modules corresponding to the third drive signal and store the third data in the storage structure SS. Here, at least one of the third set of storage modules is included in the second set of storage modules.
[0111] For example, the first multiplexing circuit 13 can receive the third driving signal before the second multiplexing circuit 21 receives the second driving signal. Alternatively, the first multiplexing circuit 13 can receive the third driving signal simultaneously with the second multiplexing circuit 21 receiving the second driving signal. In this case, the first multiplexing circuit 13 can begin performing operations in response to the third driving signal first, and the second multiplexing circuit 21 can begin performing operations in response to the second driving signal later.
[0112] As previously described, in response to the second drive signal, the second multiplexing circuit 21 can read a certain type of data and store that type of data in, for example, two storage modules 11c and 11d. In this case, the first multiplexing circuit 13 can read this type of data as third data from both storage modules 11c and 11d and store it in the storage structure SS. In other words, the storage space of the two storage modules 11c and 11d (i.e., the third set of storage modules) is released simultaneously. Therefore, in this case, the second set of storage modules can include only one storage module 11c or one storage module 11d, or the second set of storage modules can include both storage modules 11c and 11d simultaneously.
[0113] The driving method for the AI chip 200 also includes the following steps.
[0114] Before determining the fourth group of storage modules, first determine the available fifth group of storage modules among the multiple storage modules. Here, the total capacity of the fifth group of storage modules is less than the total capacity required to store the second data, or the number of storage modules in the fifth group is less than the number of data types in the second data.
[0115] For example, the second data includes two types of data: input feature map data and bias data, while the fifth group of storage modules has only one storage module 11a. This means that the number of storage modules in the fifth group of storage modules is less than the number of data types in the second data.
[0116] Then, the third set of storage modules corresponding to the third data required by other tasks after the task is determined (i.e., the third data is currently stored in the storage module of the artificial intelligence chip 200). After that, a third drive signal corresponding to the third set of storage modules is sent to the first multiplexing circuit 13, so that the first multiplexing circuit 13 reads the third data and stores the third data in the storage structure SS.
[0117] After the third data is stored in the storage structure SS, the third group of storage modules is adjusted from non-idle to idle. In other words, the third group of storage modules is included in the subsequently determined idle fourth group of storage modules. For example, if the determined fifth group of storage modules includes storage module 11a and the third group of storage modules includes storage module 11b, then the fourth group of storage modules can be determined to include storage modules 11a and 11b. In this way, the problem of data needing to be stored and computed in batches can be further improved, thereby further increasing the computing speed of artificial intelligence chips.
[0118] In some embodiments, see Figure 2A The second multiplexing circuit 21 of the artificial intelligence chip 200 also includes a second group of second input terminals 211b that are connected one-to-one with the output terminals of multiple computing circuits.
[0119] In these embodiments, the second multiplexing circuit 21 may also be configured to store the output of a computing circuit into a storage module corresponding to the fourth drive signal in response to the fourth drive signal.
[0120] The driving method for the AI chip 200 also includes the following steps.
[0121] First, the third capacity required to store the output of the computing circuit performing the calculation can be determined. Then, the corresponding storage module for the output can be determined based on the second and third capacities of each available storage module. Afterward, a fourth drive signal corresponding to the computing circuit and the storage module can be sent to the second multiplexing circuit 21, causing the second multiplexing circuit 21 to store the output in the determined storage module.
[0122] In the above embodiments, the output of the computing circuit is stored in the storage module of the artificial intelligence chip, and can then be directly used by other tasks. This can further improve the computing speed of the artificial intelligence chip and further reduce the frequency of the artificial intelligence chip accessing the storage structure, thereby further reducing the power consumption of the artificial intelligence chip in performing calculations.
[0123] It should be understood that the various embodiments of this disclosure do not limit the number and capacity of the multiple storage modules in the artificial intelligence chip 100 / 200, but the number and capacity of the multiple storage modules can be arbitrarily set.
[0124] As one implementation approach, the architecture of multiple machine learning algorithm models to be used can be analyzed to determine the various types of data required by each model and the storage capacity needed for each type. The number and capacity of multiple storage modules can then be configured based on the frequency of use and corresponding capacity of different data types within the multiple machine learning algorithm models. In this way, the AI chip 100 / 200 can better meet practical application needs.
[0125] As another implementation method, since the storage capacity of various types of data required for a certain task is usually different, multiple storage modules with different capacities can be set so that each type of data can be stored in a storage module that is more compatible with the length of that type of data.
[0126] For example, the fourth set of idle storage modules includes storage modules 11a to 11c, with capacities of 30, 50, and 100 respectively. If the required storage capacity is 70, then this data can be stored in storage modules 11a and 11b with capacities of 30 and 50, instead of storing it in storage module 11c with a capacity of 100.
[0127] This allows for further improvement in the utilization rate of the storage module. This further addresses the issue of data needing to be stored and processed in batches, thereby enabling an even greater increase in the computing speed of AI chips.
[0128] Figure 3 This is a schematic diagram of the structure of a driving device for a data stream-based artificial intelligence chip according to some embodiments of the present disclosure.
[0129] like Figure 3 As shown, the drive device 300 includes a determining module 301 and a sending module 302.
[0130] The determination module 301 is configured to determine the first set of computing circuits required to perform the task, the task including at least one type of computation corresponding to the first set of computing circuits; and to determine the first set of storage modules corresponding to the first data required to perform the task.
[0131] The sending module 302 is configured to send a first driving signal corresponding to the first set of computing circuits and the first set of storage modules to the first multiplexing circuit 13, so that the first multiplexing circuit 13 reads the first data and transmits the first data to the first set of computing circuits.
[0132] It should be understood that the determining module 301 and the sending module 302 can also be configured to perform other operations so that the driving device 300 can execute the driving method of the artificial intelligence chip 100 / 200 of any of the above embodiments.
[0133] Figure 4 This is a schematic diagram of the structure of a driving device for a data stream-based artificial intelligence chip according to other embodiments of this disclosure.
[0134] like Figure 4 As shown, the driving device 400 includes a memory 401 and a processor 402 coupled to the memory 401. The processor 402 is configured to execute the driving method of any of the above embodiments based on instructions stored in the memory 401.
[0135] The memory 401 may include, for example, system memory, fixed non-volatile storage media, etc. The system memory may store, for example, an operating system, application programs, a boot loader, and other programs.
[0136] The drive device 400 may also include an input / output interface 403, a network interface 404, and a storage interface 405. These interfaces 403, 404, and 405, as well as the memory 401 and processor 402, can be connected, for example, via a bus 406. The input / output interface 403 provides a connection interface for input / output devices such as monitors, mice, keyboards, and touchscreens. The network interface 404 provides a connection interface for various networked devices. The storage interface 405 provides a connection interface for external storage devices such as cards and USB flash drives.
[0137] Figure 5 This is a schematic diagram of the structure of an artificial intelligence accelerator according to some embodiments of the present disclosure.
[0138] like Figure 5 As shown, the artificial intelligence accelerator includes a data flow-based artificial intelligence chip 100 / 200 of any of the above embodiments and a driving device 300 / 400 of any of the above embodiments.
[0139] This disclosure also provides a server that includes the artificial intelligence accelerator of any of the above embodiments.
[0140] This disclosure also provides a computer-readable storage medium including computer program instructions that, when executed by a processor, implement the driving method of any of the above embodiments.
[0141] This disclosure also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the driving method of any of the above embodiments.
[0142] The embodiments of this disclosure have now been described in detail. To avoid obscuring the concept of this disclosure, some details known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.
[0143] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the embodiments of the driving device, artificial intelligence accelerator, and server, since they largely correspond to the embodiments of the data stream-based artificial intelligence chip and its driving method, the descriptions are relatively simple; relevant parts can be referred to in the descriptions of the embodiments of the artificial intelligence chip and its driving method.
[0144] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable non-transitory storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0145] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that the functions specified in one or more flowchart illustrations and / or one or more block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means configured to implement the functions specified in one or more flowchart illustrations and / or one or more block diagrams.
[0146] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0147] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps configured to implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0148] While specific embodiments of this disclosure have been described in detail by way of examples, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments or equivalent substitutions can be made to some technical features without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.
Claims
1. A dataflow-based artificial intelligence chip, comprising: Multiple storage modules; Multiple computing circuits, each configured to perform different types of computations in a machine learning algorithm model; A first multiplexing circuit is configured to, in response to a first drive signal corresponding to a task, read first data from a first set of storage modules corresponding to the first drive signal, and transmit the first data to a first set of computing circuits corresponding to the first drive signal. The first multiplexing circuit includes: Multiple first input terminals are connected one-to-one with the multiple storage modules, and Multiple first output terminals, including a first group of first output terminals connected one-to-one with the multiple computing circuits; A second multiplexing circuit is configured to, in response to a second drive signal corresponding to the task, read second data from a storage structure outside the AI chip and store the second data in a second set of storage modules corresponding to the second drive signal. The second data is included in the first data, and the second set of storage modules is included in the first set of storage modules. The second multiplexing circuit includes: Multiple second input terminals, including a first set of second input terminals connected to the storage structure, and Multiple second output terminals are connected to the multiple storage modules one by one.
2. The artificial intelligence chip according to claim 1, wherein, The first multiplexing circuit is also configured to: In response to another first drive signal corresponding to another task, the first data in the first set of storage modules is read and the first data is transmitted to the second set of computing circuits corresponding to the other first drive signal.
3. The artificial intelligence chip according to claim 1, wherein, The second multiplexing circuit is further configured to: in response to the second drive signal, read data of the same type from the second data and store the data of the same type in at least two storage modules in the second group of storage modules.
4. The artificial intelligence chip according to claim 1, wherein, The plurality of first output terminals also includes a second set of first output terminals connected to the storage structure; The first multiplexing circuit is further configured to: in response to a third driving signal, read third data from a third set of storage modules corresponding to the third driving signal, and store the third data in the storage structure; At least one of the third group of storage modules is included in the second group of storage modules.
5. The artificial intelligence chip according to claim 1, wherein, The plurality of second input terminals also include a second group of second input terminals that are connected one-to-one with the output terminals of the plurality of computing circuits; The second multiplexing circuit is further configured to: in response to a fourth drive signal, store the output of a computing circuit into a storage module corresponding to the fourth drive signal.
6. The artificial intelligence chip according to claim 1, wherein, The machine learning algorithm model includes a neural network algorithm model, and the task is the computation of one or at least two consecutive computation layers of the neural network algorithm model.
7. A driving method for a dataflow-based artificial intelligence chip according to any one of claims 1-6, comprising: Determine a first set of computing circuits required to perform a task, wherein the task includes at least one type of computation corresponding to the first set of computing circuits; Determine the first group of storage modules corresponding to the first data required to perform the task; A first drive signal corresponding to the first set of computing circuits and the first set of storage modules is sent to the first multiplexing circuit, so that the first multiplexing circuit reads the first data and transmits the first data to the first set of computing circuits.
8. The method according to claim 7, further comprising: Determine a second set of computing circuits required to perform another task, the other task including at least one type of computation corresponding to the second set of computing circuits; Send another first drive signal corresponding to the second set of computing circuits and the first set of storage modules to the first multiplexing circuit, so that the first multiplexing circuit reads the first data and transmits the first data to the second set of computing circuits.
9. The method according to claim 7 or 8, wherein, The artificial intelligence chip also includes a second multiplexing circuit, which includes a first group of second input terminals connected to a storage structure outside the artificial intelligence chip, and a plurality of second output terminals connected to the plurality of storage modules one by one. The method further includes: Determine the first capacity required to store each type of data in the second data contained within the first data in the storage structure; Identify the fourth group of idle storage modules among the plurality of storage modules; Determine the second capacity of each storage module in the fourth group of storage modules; Based on the first capacity and the second capacity, the second group of storage modules corresponding to the second data is determined from the fourth group of storage modules, and the second group of storage modules is included in the first group of storage modules; A second drive signal corresponding to the second set of storage modules is sent to the second multiplexing circuit so that the second multiplexing circuit reads the second data and stores the second data in the second set of storage modules.
10. The method according to claim 9, wherein, The second group of storage modules corresponding to the second data, determined from the fourth group of storage modules based on the first capacity and the second capacity, includes: If the first capacity of a certain type of data in the second data is greater than the second capacity of each storage module in the fourth group of storage modules, at least two storage modules corresponding to that type of data are determined. The second group of storage modules includes the at least two storage modules.
11. The method according to claim 9, wherein, The plurality of first output terminals also includes a second set of first output terminals connected to the storage structure; The method further includes: Before determining the fourth group of storage modules, a fifth group of idle storage modules is determined among the plurality of storage modules, wherein the total capacity of the fifth group of storage modules is less than the total capacity required to store the second data, or the number of storage modules in the fifth group of storage modules is less than the number of data types in the second data; The third group of storage modules corresponding to the third data required by other tasks after the aforementioned task is determined; A third driving signal corresponding to the third group of storage modules is sent to the first multiplexing circuit so that the first multiplexing circuit reads the third data and stores the third data in the storage structure.
12. The method according to claim 9, wherein, The plurality of second input terminals also include a second group of second input terminals that are connected one-to-one with the output terminals of the plurality of computing circuits; The method further includes: Determine the third capacity required to store the output of the computational circuit that is performing the calculation; The storage module corresponding to the output is determined based on the second capacity and the third capacity; A fourth drive signal corresponding to the computing circuit and the storage module is sent to the second multiplexing circuit so that the second multiplexing circuit stores the output into the storage module.
13. The method according to claim 7, wherein, The task is the computation of one or at least two consecutive computational layers in a neural network algorithm model.
14. A driving device for a dataflow-based artificial intelligence chip according to any one of claims 1-6, comprising: A determination module is configured to determine a first set of computing circuits required to perform a task, the task including at least one type of computation corresponding to the first set of computing circuits. Determine the first group of storage modules corresponding to the first data required to perform the task; The transmitting module is configured to send a first driving signal corresponding to the first set of computing circuits and the first set of storage modules to the first multiplexing circuit, so that the first multiplexing circuit reads the first data and transmits the first data to the first set of computing circuits.
15. A driving device for a dataflow-based artificial intelligence chip according to any one of claims 1-6, comprising: Memory; as well as A processor coupled to the memory is configured to execute the driving method of any one of claims 7-13 based on instructions stored in the memory.
16. An artificial intelligence accelerator, comprising: The data flow-based artificial intelligence chip according to any one of claims 1-6; and The driving device according to claim 14 or 15.
17. A server, comprising: The artificial intelligence accelerator as described in claim 16.
18. A computer-readable storage medium comprising computer program instructions, wherein, When the computer program instructions are executed by the processor, they implement the driving method according to any one of claims 7-13.
19. A computer program product comprising a computer program, wherein, When the computer program is executed by the processor, it implements the driving method according to any one of claims 7-13.