General data preprocessing circuit, computing array, AI chip and equipment
By adopting a general data preprocessing circuit in high parallelism and large computing power scenarios, and using the combined design of combined selection modules and selection submodules, the problem of large resource overhead in sparse mode is solved, compatible processing of sparse and non-sparse data is achieved, and hardware resource overhead is reduced.
Patent Information
- Application Number
- CN202510080556.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
In high parallelism and large computing power scenarios, the resource overhead brought by supporting sparse mode is relatively large, and it is difficult for the existing technology to effectively reduce hardware resource overhead.
A general data preprocessing circuit is employed, including a plurality of combined selection modules, each of which comprises a first selection submodule and a second selection submodule. The first selection submodule selects the first index or the second index according to the input data mode and provides the target index to the second selection submodule. The second selection submodule selects the target number of data from the input data based on the received target index.
Through this general data preprocessing circuit, hardware resource overhead can be significantly reduced when compatible with sparse mode and non-sparse mode, and is suitable for high parallelism and large computing power scenarios.
Smart Images

Figure CN120012852A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence (AI) technology, in particular to the field of chips, and specifically to a universal data preprocessing circuit, a computing array, an AI chip and an electronic device. Background Art
[0002] High-performance computing chips require high computing power, and therefore, the computing parallelism, number of computing units, and bit width of data paths of such chips are relatively large. In many designs, such as systolic arrays, data reuse is exploited to reduce the number of transfers, thereby reducing chip power consumption.
[0003] In sparse mode, although data can be reused, the index corresponding to each piece of data is different and the index cannot be reused. Therefore, each computing unit needs to be configured with a separate data preprocessing module to implement the data preprocessing logic. The main function of the data preprocessing module is to select all non-zero data in the sparse data of a specific sparse structure (for example: 2:4 or 4:8, etc.) to perform calculations. At the same time, the data preprocessing module also needs to be compatible with the original data path in non-sparse mode, and choose one of the two modes according to the mode.
[0004] The resource overhead of the data preprocessing module is linearly related to the computational parallelism of the computing unit. The greater the computational parallelism, the greater the resource overhead of a single data preprocessing module. At the same time, the number of computing units also determines the number of data preprocessing modules. Therefore, in high-parallelism and high-computing-power scenarios, the resource overhead brought by supporting the sparse mode will be very large. Summary of the invention
[0005] The present disclosure provides a general data preprocessing circuit, a computing array, an AI chip and an electronic device.
[0006] According to one aspect of the present disclosure, a universal data preprocessing circuit is provided, comprising: a plurality of combination selection modules, each combination selection module comprising a first selection submodule and a second selection submodule;
[0007] The first selection submodule is used to select in the first index or the second index according to the input data mode, and provide the selected target index to the second selection submodule in the same module;
[0008] The second selection submodule is used to select a target number of data from a set of input data according to the received target index.
[0009] According to another aspect of the present disclosure, there is provided a computing array, comprising a plurality of computing units, and a plurality of universal data preprocessing circuits as described in any one of the embodiments of the present disclosure, respectively pre-connected to each of the computing units;
[0010] The calculation unit is used to receive the data selected from multiple groups of data by the connected general data preprocessing circuit, and perform calculation after combining the data.
[0011] According to another aspect of the present disclosure, an AI chip is also provided, comprising a computing array as described in any one of the embodiments of the present disclosure.
[0012] According to another aspect of the present disclosure, an electronic device is also provided, comprising the AI chip as described in any one of the embodiments of the present disclosure.
[0013] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0015] Figure 1 It is a structural diagram of a data path in a non-sparse mode provided by the related art;
[0016] Figure 2 It is a structural diagram of a data path in a sparse mode provided by the related art;
[0017] Figure 3 is a schematic diagram of an index-based selector circuit provided by the related art;
[0018] Figure 4 It is a schematic diagram of a universal data selection circuit compatible with sparse mode and non-sparse mode provided by the related art;
[0019] Figure 5 is a structural diagram of a general data preprocessing circuit provided according to an embodiment of the present disclosure;
[0020] Figure 6 is a structural diagram of a second selection submodule based on a 2:4 sparse structure provided according to an embodiment of the present disclosure;
[0021] Figure 7 is a structural diagram of a general data preprocessing circuit based on a 2:4 sparse structure provided according to an embodiment of the present disclosure;
[0022] Figure 8 is a structural diagram of a computing array provided according to an embodiment of the present disclosure;
[0023] Fig. 9is a structural diagram of another computing array provided according to an embodiment of the present disclosure;
[0024] Fig.10 This is a schematic diagram of data sorting after sparse data is selected by a general data preprocessing circuit applicable to an embodiment of the present disclosure;
[0025] Fig.11 This is a schematic diagram of data sorting after data selection is performed by a general data preprocessing circuit for non-sparse data applicable to an embodiment of the present disclosure;
[0026] Fig.12 is a structural diagram of another computing array provided by an embodiment of the present disclosure;
[0027] Fig.13 is a structural diagram of an AI chip provided by an embodiment of the present disclosure;
[0028] Fig.14 is a structural diagram of another AI chip provided by an embodiment of the present disclosure;
[0029] Fig.15 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0031] In order to facilitate understanding of the technical solutions of the embodiments of the present disclosure, the relevant technologies involved in the embodiments of the present disclosure are briefly described below.
[0032] With the continuous deepening of AI big model technology, more and more neural network computing scenarios are involved. In the neural network computing scenario, matrix multiplication is a very commonly used operation. For example, the input of each model processing layer of the AI big model is the activation transmitted by the previous layer. The model processing layer needs to multiply the activation with the weight of the layer, and continue to transmit the calculation result to the next model processing layer until the final output result is obtained. Therefore, it is often necessary to use the dedicated computing unit provided by the AI acceleration chip (also directly referred to as the AI chip) to implement the above multiplication calculation, typically, matrix multiplication calculation.
[0033] Furthermore, considering the increasing number of network parameters and computational workloads, the computational pressure on AI acceleration chips is also increasing. In order to reduce computational pressure to a certain extent, the data involved in multiplication calculations can be sparsely processed at the software level. For example, the data that contributes less to the calculation results in the network parameters (e.g., weights) can be pruned. This operation can reduce the data storage space and the actual amount of computation required without affecting the accuracy of the results.
[0034] Therefore, when the AI chip performs matrix multiplication calculations, it may encounter the situation of processing sparse data or non-sparse data. At this time, before performing the multiplication calculation, it is often necessary to use a specific data path to preprocess the input data. Generally speaking, sparse data includes structured sparse data and unstructured sparse data, and unstructured sparse data is not conducive to hardware acceleration. Therefore, the general data preprocessing circuit provided in the embodiment of the present disclosure is mainly used to implement data preprocessing before matrix multiplication of structured sparse data and non-sparse data (also called dense data).
[0035] Specifically, in Figure 1 and Figure 2 The structural diagrams of the data paths in non-sparse mode and sparse mode are shown in FIG.
[0036] like Figure 1 As shown in , when the computing unit in the AI chip works in non-sparse mode (or, it can also be called dense mode), the multiplicand and multiplier of the matrix multiplication operation are both non-sparse data. At this time, the input data can be directly multiplied with the unpruned weight matrix to obtain the multiplication result. Similarly, Figure 2 As shown in the figure, when the computing unit in the AI chip works in sparse mode, the weight matrix and data involved in the multiplication calculation are pruned. Figure 2 The weight matrix in Figure 1 The weight matrix in , contains 4 matrix elements reduced from 8. Similarly, Figure 2 The input data in the sparse data is also changed from dense data to sparse data. That is, although the sparse data still has 8 elements in data form, only 4 elements are meaningful data (non-0 data), and the remaining elements are all 0. The data positions of the above 4 non-0 data correspond to the positions of the remaining matrix elements after the weight matrix is pruned. Furthermore, it is necessary to select the above 4 non-0 data from the sparse data and perform matrix multiplication calculations corresponding to the pruned weight matrix. Optionally, the above input data and weight data are both binary data.
[0037] Based on this, Figure 2The sparse data input in needs to first pass through a selector (also called a data selector), which, under the control of the index, selects the non-zero data from the input sparse data and performs multiplication calculation with the pruned weight matrix.
[0038] Further, in Figure 3 A schematic diagram of an index-based selector circuit is shown in FIG. The selector circuit is used to select non-zero data from a structured sparse data. Different types of sparse structures have different applicable selector circuits. Figure 3 In the figure, the specific architecture of the selector circuit is described by taking data selection of sparse data with a 2:4 sparse structure as an example.
[0039] The sparse data of the 2:4 sparse structure can be understood as a sparse matrix consisting of 4 elements, in which non-zero elements occupy 2 positions. Figure 2 The sparse data consisting of 8 elements in the sparse data can be understood as consisting of two sparse data of 2:4 sparse structure. Furthermore, for each sparse data of 2:4 sparse structure, two 4-to-1 data selectors are required to select two non-zero data from the input sparse data, and the specific two data selected are controlled by the index.
[0040] Specifically, since a 4-to-1 data selector is used, 4 optional positions can be determined by a 2-bit (b) selection signal. Then, for the above two 4-to-1 data selectors, a 4-bit index can be constructed. Figure 3 Taking "4'b1000" as an example, the lower 2-bit 2-bit selection signal (2'b00) is input to the first 4-to-1 data selector to select the 0th number from the 4 data from 0 to 3, and the upper 2-bit 2-bit selection signal (2'b10) is input to the second 4-to-1 data selector to select the second number from the 4 data from 0 to 3.
[0041] In general, in order to make full use of computing resources, multiple times the data needs to be input in sparse mode to maintain the same amount of data input as in non-sparse mode. For example, for sparse data with a 2:4 sparse structure, double the data needs to be input. At the same time, in order to be compatible with the data path in non-sparse mode, it is necessary to select whether to use the original non-sparse data and weights for multiplication calculations, or to use the sparse data and weights selected by the index to perform multiplication calculations, based on the input data mode.
[0042] Specifically, in Figure 4 A schematic diagram of a general data selection circuit compatible with sparse mode and non-sparse mode is shown in FIG. Figure 4As shown in the figure, when the input data mode is determined to be a sparse mode, by assigning input matching selection signals to the bottom four 2-to-1 data selectors, four data selected by the index in the sparse data (two data selected from the 0th to 3rd data, and two data selected from the 4th to 7th data) can be finally selected to perform subsequent calculations. When the input data mode is determined to be a non-sparse mode, by assigning input matching selection signals to the bottom four 2-to-1 data selectors, four non-sparse data from the 0th to the 3rd data can be finally selected to perform subsequent calculations.
[0043] In the related art, the number of computing units contained in the AI chip is generally large, and the computing parallelism of each computing unit is also large. Figure 4 The universal data selection circuit achieves compatibility with sparse mode and non-sparse mode, which will greatly increase the resource overhead of the AI chip. In a specific example, if the parallelism of each computing unit in the AI chip is 64, each computing unit can process the multiplication calculation of 64 numbers in parallel. If the sparse structure of sparse data in the sparse mode is 2:4, it is necessary to read 128 data at a time, with 4 numbers in a group, a total of 32 groups, and perform 2:4 selection in each group to select 64 data for subsequent calculations. At this time, in order to be compatible with the non-sparse mode, a total of 64 2-to-1 data selectors are required for the computing unit to select sparse data or non-sparse data.
[0044] Similarly, for each additional computing unit in the AI chip, a matching universal data selection circuit needs to be added to the computing unit, and then 64 2-to-1 data selectors will be introduced accordingly. In other words, the resource overhead brought by the use of universal data selection circuits is linearly related to the data parallelism of the computing units contained in the AI chip. The greater the data parallelism, the greater the overhead of a single universal data selection circuit. At the same time, the number of computing units in the AI chip also determines the number of universal data selection circuits. Therefore, in high-parallelism and high-computing-power scenarios, the resource overhead brought by supporting sparse mode will be very large.
[0045] Based on this, through the technical solutions of the various embodiments of the present disclosure, the above-mentioned general data selection circuit is improved to minimize the resource overhead introduced by the coefficient mode.
[0046] Figure 5 It is a structural diagram of a universal data preprocessing circuit provided according to an embodiment of the present disclosure. The embodiment of the present disclosure can be applied to the situation where data matching the input data pattern is selected through the universal data preprocessing circuit before the computing unit performs multiplication calculation on the input data.
[0047] like Figure 5As shown, the universal data preprocessing circuit specifically includes: a plurality of combination selection modules 500 , each of which includes a first selection submodule 501 and a second selection submodule 502 .
[0048] In this embodiment, the general data preprocessing circuit inputs multiple groups of input data at one time, and the number of data contained in each group of input data is the same. One group of input data is input into one combination selection module to select a set number of data from the group of input data.
[0049] For example, if the universal data preprocessing circuit includes four combination selection modules 500 , it is necessary to construct four groups of input data with the same data volume and input them into the four combination selection modules 500 respectively to perform data selection.
[0050] The first selection submodule 501 is used to select in the first index or the second index according to the input data mode, and provide the selected target index to the second selection submodule 502 in the same module.
[0051] The input data mode is used to describe the data form of data input, such as sparse data or dense data, etc. In this embodiment, data selection is mainly performed in two data forms, and thus, the input data mode can be a first data mode and a second data mode.
[0052] Optionally, the first data pattern may correspond to the first index, and the second data pattern may correspond to the second index. Accordingly, the first selection submodule 501 may use the input data pattern as a selection signal to select from the first index and the second index, and provide the selected first index or second index as a target index to the second selection submodule 502 in the same module, so that the second selection submodule 502 can perform real data selection for the target index.
[0053] In a specific example, if the input data mode is the first data mode, the first selection submodule 501 selects the first index as the target index and provides the target index to the second selection submodule 502 in the same module. If the input data mode is the second data mode, the first selection submodule 501 selects the second index as the target index and provides the target index to the second selection submodule 502 in the same module.
[0054] In this embodiment, the universal data preprocessing circuit is constructed using various types of logic gate arrays. In order to adapt to the circuit structure of the universal data preprocessing circuit, the first index, the second index and the input data are all binary data.
[0055] In an optional implementation of this embodiment, the first selection submodule 501 can be obtained by overlapping one or more data selectors of specific specifications according to the data format of the first index and the second index. Optionally, the data format of the first index and the second index is the same, for example, the binary bit width of the two is the same, for example, both are 4-bit or 8-bit data in length.
[0056] The second selection submodule 502 is used to select a target number of data from a set of input data according to the received target index.
[0057] The second selection submodule 502 is used to select a specific number of data from a set of input data using a target index as a selection signal. It is understandable that the target number is less than the total amount of data contained in a set of data. For example, a set of data contains 4 data in total, and the target number can be 1 or 2.
[0058] In this embodiment, the target index is used to describe the data position of the data to be selected in a group of data. For example, a group of data contains 4 data in total, and the above 4 data positions can be represented by 2 bits of data. Then, the data positions of the 2 data to be selected from the above 4 data can be determined by a 4-bit target index. For example, when the target index is "4'b1011", it is determined that the 2 data positions to be selected are located at the 3rd position and the 2nd position in the group of data respectively (assuming that the numbering starts from the 0th position).
[0059] It is understandable that different input data modes select different target indexes. Further, based on input data in different data forms, the adapted index can be used to select data that meets the requirements for subsequent data processing, such as multiplication calculation.
[0060] It should be emphasized that, in contrast Figure 4 It can be seen from the plan that Figure 4 When selecting data in two input data modes, the solution adopts the implementation method of adding a two-to-one data selector for each two output data in the two input data modes before the final output data. This implementation method will bring a lot of hardware resource overhead when the amount of data to be selected for output is relatively large.
[0061] However, the technical solution of the embodiment of the present disclosure only needs to select different indexes under the two input data modes, and then the input data can be selected only according to the selected index using a unified second selection submodule 502. Through the above settings, the selection of two types of input data can be compatible through a unified circuit design.
[0062] At the same time, this implementation can reduce one layer of processing logic for mode selection of each output data in each combination selection module 500. Since the bit width of the index data is generally smaller than the bit width of the output data, for example, a 2-bit index can be selected from 4 output data, and a 3-bit index can be selected from 8 output data. This method of selecting only the index can greatly save resource overhead.
[0063] It should be noted that in Figure 5 For intuitive representation, at least three combination selection modules 500 are drawn. In fact, the universal data preprocessing circuit may include only two combination selection modules 500.
[0064] The technical solution of the disclosed embodiment is to form a general data preprocessing circuit by using multiple combination selection modules. In each combination selection module, the first selection submodule uses the input data mode as the selection signal to select between two indexes, and the selected index is used as a new selection signal, which is provided together with a group of input data to the second selection submodule in the same group for data selection. Compared with the implementation of the related technology, the data selection logic can be effectively reduced, and the hardware resource overhead can be greatly simplified. This implementation is particularly suitable for application scenarios with high parallelism and high computing power.
[0065] Based on the above embodiments, the input data mode may include: sparse input data or non-sparse input data.
[0066] Correspondingly, when the input data mode is the sparse input data, the input set of data is sparse data, and when the input data mode is the non-sparse input data, the input set of data is non-sparse data.
[0067] The first index is used to describe the index position of the non-zero data to be selected in a set of sparse data, and the second index is used to describe the index position of the data to be selected in a set of non-sparse data.
[0068] In this embodiment, the general data preprocessing circuit is compatible with the input of sparse data and non-sparse data. It is understandable that when the input is sparse data, non-zero data in the input data must be selected through the first index for output.
[0069] When the input is non-sparse data, the complete non-sparse data needs to be output. The reason why non-sparse data also needs to be selected through the second index is to control the universal data preprocessing circuit to ensure the consistency of input and output data volume regardless of whether sparse data or non-sparse data is input. Furthermore, non-sparse data is generally redundant when input.
[0070] Assume that the universal data preprocessing circuit can input 16 data at a time, and when sparse data is input, the sparse structure of the sparse data is 4:2, and then the universal data preprocessing circuit can input 4 groups of sparse data of the above sparse structure at a time, and finally output 8 non-zero dense data. In order to ensure that the universal data preprocessing circuit can also meet the above input and output data requirements when inputting non-sparse data, the 8 non-sparse data can be completely copied, and the obtained 16 data can be input into the universal data preprocessing circuit. After that, the 8 non-sparse data can be completely selected through the selection of the universal data preprocessing circuit, so that the compatible processing of the above sparse data and non-sparse data can be achieved.
[0071] The complete selection of the above 8 non-sparse data {1,2,3,4,5,6,7,8} is controlled by the second index. For example, after copying the above 8 non-sparse data, the copied data of the form {1,2,3,4,5,6,7,8,1,2,3,4,5,6,7,8} is obtained. After grouping the copied data, 4 groups of data can be obtained, {1,2,3,4}, {5,6,7,8}, {1,2,3,4} and {5,6,7,8}. By being compatible with the 4-out-of-2 selection mechanism used in sparse data selection, the first two numbers can be fixedly selected from the first two data sets {1,2,3,4} and {5,6,7,8}, and the last two numbers can be fixedly selected from the last two data sets {1,2,3,4} and {5,6,7,8}, so as to finally obtain the complete {1,2,3,4,5,6,7,8}.
[0072] By applying the above-mentioned universal data preprocessing circuit to the data selection of sparse data and non-sparse data, the universal data preprocessing circuit can be widely used in the computing scenarios of AI chips, so that the AI chip only has a small hardware resource overhead regardless of performing model calculations for sparse data or non-sparse data.
[0073] On the basis of the above embodiments, in the general data preprocessing circuit, the total amount of data input at a time is n*k, and the total amount of data selected at a time is n*m, wherein:
[0074] The number of the combination selection modules is n, the target number is m, and the combination selection module is used to select m data from k data in each group when n groups of data are input; n and k are integers greater than or equal to 2, and m is a positive integer less than k.
[0075] Specifically, considering that the universal data preprocessing circuit is compatible with the input sparse data and non-sparse data, it is necessary to use multiple combination selection modules to input different selection data respectively. Furthermore, the number of combination selection modules needs to be greater than or equal to 2. Similarly, each combination selection module needs to implement a data selection function, so the amount of data input by each combination selection module at a time also needs to be greater than or equal to 2.
[0076] When the universal data preprocessing circuit in this embodiment is designed, the circuit structure of each combination selection module is the same. Furthermore, each combination selection module can be implemented as a circuit design for selecting m data from a group of k data. Furthermore, the total amount of data input to the universal data preprocessing circuit at a time can be calculated by n*k, and the total amount of data output by the universal data preprocessing circuit at a time can be directly calculated by n*m.
[0077] On the basis of the above embodiments, when the input data mode is the sparse input data, each of the combination selection modules respectively inputs a group of sparse data of m:k sparse structure;
[0078] Wherein, when the input data mode is the non-sparse input data, each combination selection module is used to input k non-sparse data selected in sequence from the copy results obtained by copying n*m non-sparse data by k / m copies.
[0079] Among them, in the m:k sparse structure, k can divide m, that is, k is an integer multiple of m. Further, when non-sparse data is copied, an integer number of copies is generally copied, for example, 2, 3 or 4 copies.
[0080] As mentioned above, each combination selection module in the embodiment of the present disclosure is used to select m data from a group of k data, and is therefore particularly suitable for selecting sparse data with an m:k sparse structure. Among them, sparse data with an m:k sparse structure means that in a group of k data, only m data are non-zero data (meaningful data). Furthermore, multiple groups of sparse data with an m:k sparse structure can be preprocessed directly through multiple combination selection modules in a universal data preprocessing circuit, respectively, to obtain all non-zero data for subsequent operations.
[0081] As mentioned above, in order to make the universal data preprocessing circuit compatible with input non-sparse data, it is necessary to send a set amount of non-sparse data redundantly. Among them, the amount of non-sparse data that really needs to be selected should be n*m, and the total amount of data input into the universal data preprocessing circuit at one time should be n*k, and then, the number of copies of the non-sparse data required is (n*k) / (n*m)=k / m.
[0082] In a specific example, if a general data preprocessing circuit is used to preprocess both sparse data and non-sparse data, when the sparse structure of the non-sparse data that the circuit can process is 2:4, then when the circuit is used to process non-sparse data, it is necessary to copy the complete sparse data twice to obtain the copied data, and then divide the copied data into n parts, each with k data, and then provide each k data to each combination selection module in turn.
[0083] By effectively controlling the data transmission method of non-sparse data based on the sparse structure of sparse data, the compatibility of the universal data preprocessing circuit with sparse data and non-sparse data can also be guaranteed, meeting the data preprocessing requirements in different scenarios.
[0084] On the basis of the above embodiments, the general data preprocessing circuit is used to provide input data for the computing unit;
[0085] The number n of the combination selection modules is associated with the computational parallelism of the computing unit and the sparse structure of the input sparse data.
[0086] As mentioned above, the universal data preprocessing circuit provided in the embodiment of the present disclosure can provide processing data for a special computing unit so that the computing unit can perform special calculations, typically, matrix multiplication calculations. Furthermore, the universal data preprocessing circuit has a one-to-one correspondence with the computing unit. Assuming that an AI chip has a total of 256 computing units, 256 universal data preprocessing circuits can be constructed accordingly. A universal data preprocessing circuit is pre-connected before a computing unit to provide the computing unit with all the data required for a single parallel processing.
[0087] Furthermore, the data that the universal data preprocessing circuit needs to select and output at one time needs to be compatible with the computational parallelism of the connected computing units.
[0088] For example, if the computing parallelism of the computing unit is 64, it means that the computing unit can perform calculations on 64 data at one time, and thus, the general data preprocessing circuit needs to output 64 data at one time. It can be seen from the above description that when the sparse structure of the sparse data that can be processed by the general data preprocessing circuit is determined, the total amount of data selected by the general data preprocessing circuit at one time is n*m.
[0089] Correspondingly, when a general data preprocessing circuit is constructed for a specific computing unit, once the computational parallelism of the computing unit and the sparse structure of the sparse data that the general data preprocessing circuit needs to process are clear, the number of combination selection modules required to be included in the general data preprocessing circuit is uniquely determined.
[0090] On the basis of the above embodiments, when the sparse structure of the input sparse data is m:k, the number of the combination selection modules n=l / m;
[0091] Wherein, l is the computational parallelism of the computing unit, and the m:k sparse structure means that in a group of sparse data containing k data, there are m non-zero data.
[0092] Through the above settings, it can be ensured that the general data preprocessing circuit is accurately adapted to the computing unit that actually performs the calculation, effectively meeting the actual parallel computing needs of the computing unit.
[0093] Based on the above embodiments, the first selection submodule includes a plurality of two-to-one data selectors.
[0094] As mentioned above, the first selection submodule is used to select between the first index and the second index. The first index and the second index are used to select a target number of data from a set of data. Accordingly, multiple two-to-one data selectors can be used to perform bitwise selection in the first index and the second index, and the combination of the selected results can be restored to obtain the above-mentioned first index or second index.
[0095] Through the above settings, a simple logic circuit can be used to achieve the desired technical effect of index selection.
[0096] For example, each bit in the first index is input to the first input end of each two-to-one data selector, and each bit in the second index is input to the second input end of each two-to-one data selector.
[0097] When it is determined according to the input data pattern that the first index needs to be selected, a selection signal of "0" can be sent to each of the two-to-one data selectors, so that each of the two-to-one data selectors can select each bit of the input of the first input terminal to restore the first index. Similarly, when it is determined according to the input data pattern that the second index needs to be selected, a selection signal of "1" can be sent to each of the two-to-one data selectors, so that each of the two-to-one data selectors can select each bit of the input of the second input terminal to restore the second index.
[0098] Correspondingly, the number of two-to-one data selectors included in the first selection submodule matches the binary length of the first index or the second index; wherein the binary length of the first index and the second index are the same.
[0099] In a specific example, if the binary length of the first index and the second index are both 4 bits, they are used to select one data from four data respectively through two 2-bit indexes. Furthermore, a first selection submodule including four two-to-one data selectors can be constructed to realize the selection of the first index and the second index.
[0100] Based on the above embodiments, the first index is determined by the data form of the sparse data input into the combination selection module;
[0101] The second index is a predetermined fixed index, and the data selected by each combination selection module through each second index together constitute complete n*m non-sparse data.
[0102] In this embodiment, when a group of data input to each combination selection module is sparse data with a set sparse structure, it is necessary to clarify the data at which position in the group of data needs to be taken based on the first index matching the group of data. For example, if a group of 2:4 sparse data in the form of {0,1,1,0} is input, the first index matching the sparse data is 4'b1001, indicating that the first and second data (numbering starting from the 0th data) need to be selected from the group of data.
[0103] That is, the sparse data and the matching first index appear in pairs, and the two have a corresponding relationship. When setting the second index, it is necessary to consider selecting a complete piece of data to output from the multiple copies of non-sparse data obtained after replication. Furthermore, the second index can be fixed based on the number of copies of non-sparse data and the actual grouping method of the replicated data.
[0104] The disclosed embodiment can enable a general data preprocessing circuit to compatibly meet the data preprocessing requirements for sparse data and non-sparse data by flexibly setting the first index and the second index according to the input data mode.
[0105] On the basis of the above embodiments, the second selection submodule includes m w-to-one data selectors, wherein w is a positive integer less than k.
[0106] Combined with the above description of the related technology, for a set of sparse data with a 2:4 sparse structure, in order to select 2 non-zero data from 4 input data, it is equivalent to using a 4-to-2 data selector to complete the above data selection operation. In actual circuit implementation, the above 4-to-2 data selector is implemented by using 2 4-to-1 data selectors. Figure 3 It can be seen that by connecting the four input data to two 4-to-1 data selectors respectively, the actual required 4-to-2 function can be achieved.
[0107] Related Art When selecting data in sparse data with an m:k sparse structure, m k-to-1 data selectors are required. When the k value contained in the sparse structure is larger, the hardware scale of the k-to-1 data selector is also larger, and thus the hardware resource overhead it brings is also larger.
[0108] In contrast, the inventors of the present disclosure analyze the data characteristics during sparse data selection and consider improving the design of the above circuit structure to further reduce resource overhead.
[0109] Take a set of sparse data with a 2:4 sparse structure as an example. Assume that the data format of the original set of data is {a0, a1, a2, a3}. Two of the above four data are removed through pruning, and a set of sparse data with a 2:4 sparse structure is obtained. However, there are only six cases of two data retained by the actual pruning process: {a0, a1}, {a0, a2}, {a0, a3}, {a1, a2}, {a1, a3}, and {a2, a3}. Each case has a corresponding first index representation. In other words, the positions of the two numbers selected by the first index in a set of four data are definitely different. In addition, the order of selecting the two numbers is from small to large, so the 4-to-1 data selector can be optimized to a 3-to-1 data selector.
[0110] Similarly, for a set of sparse data with a 4:8 sparse structure, the related technology originally requires 4 8-to-1 data selectors. However, considering that 4 data are selected from 8 data without duplication and from small to large, there are a total of (8*7*6*5) / (4*3*2*1)=70 situations. Furthermore, only 4 5-to-1 data selectors (with a total of 5*4*3*2*1=120 combinations) are needed to fully select the data of the above 70 situations.
[0111] In this way, after determining the specific m:k sparse structure, the specific value of w in the actual w-select-one data selector can be accurately calculated. It can be understood that because the above design abandons multiple selections of the same data, the value of w must be less than k. Therefore, this design effectively reduces the overhead of hardware resources while ensuring the completeness of data selection.
[0112] Correspondingly, since the dimension of the multiple-choice-one data selector is reduced in the second selection submodule, in order to ensure the completeness of data selection, the above-mentioned multiple w-choice-one data selectors need to be connected with the input group of k data according to a certain rule.
[0113] That is, on the basis of the above embodiments, in a group of k input data, w data are sequentially obtained by sliding with a sliding interval of 1 bit and are respectively connected to each of the w-to-one data selectors in the second selection submodule.
[0114] For example, for a set of sparse data {a0, a1, a2, a3} with a 2:4 sparse structure, a total of 2 3-to-1 data selectors are required. Furthermore, when the above two 3-to-1 data selectors are connected to the input data, the three input ends of the first 3-to-1 data selector are respectively connected to the first three data {a0, a1, a2}, and the three input ends of the second 3-to-1 data selector are respectively connected to the last three data {a1, a2, a3} with an interleaved 1-bit length.
[0115] Accordingly, in Figure 6 FIG. 4 shows a structural diagram of a second selection submodule based on a 2:4 sparse structure according to an embodiment of the present disclosure. Figure 6 As shown, for all data selection scenarios of the 2:4 sparse structure, two 3-to-1 data selectors can be used to achieve a 4-to-2 data selection effect, thereby effectively reducing resource overhead. Specifically, when the input index is 4'b1000, the lower two indexes 2'b00 are provided to the first 3-to-1 data selector to select the 0th data from the first three data. At the same time, the upper two indexes 2'b10 are provided to the second 3-to-1 data selector to select the second data from the last three data.
[0116] In this embodiment, in order to ensure the simplification of the data processing logic, the encoding logic of the selection signals of the two 3-to-1 data selectors is unified. That is, for the first 3-to-1 data selector, 2'b00 represents taking the 0th number, 2'b01 represents taking the 1st number, and 2'b10 represents taking the 2nd number. For the second 3-to-1 data selector, 2'b01 represents taking the 1st number, 2'b10 represents taking the 2nd number, and 2'b11 represents taking the 3rd number. Of course, those skilled in the art can select a matching index encoding method according to actual needs, and this embodiment does not limit this.
[0117] Further, in Figure 7 Graph 1 shows a structural diagram of a general data preprocessing circuit based on a 2:4 sparse structure provided according to an embodiment of the present disclosure. Figure 7 Can be used with Figure 4 The common data selection circuits provided are cross-referenced.
[0118] like Figure 7As shown, when two groups of 4-bit data are input for compatibility with the sparse mode and the non-sparse mode, a total of four 3-to-1 data selectors are required. These four 3-to-1 data selectors are used to select two 2-bit data that are arranged in ascending order and do not repeat each other from the two groups of 4-bit data. At the same time, the universal data preprocessing circuit of the embodiment of the present disclosure is designed for different input data modes ( Figure 7 In the example of sparse mode), only the index needs to be selected, without Figure 4 Similarly, a layer of mode selection logic is added to each output data, which can greatly reduce the implementation logic. In addition, since the bit width of the index is generally much smaller than the number of bits of the data, this implementation method of selecting the index can also reduce resource overhead.
[0119] exist Figure 7 In , the index represents the aforementioned first index, and 4'b0100 and 4'b1110 represent the aforementioned second index. The sparse pattern represents the input data pattern provided to the general data pre-processing circuit.
[0120] As mentioned earlier, when Figure 7 When processing non-sparse data, the general data preprocessing circuit needs to copy 4 non-sparse data to obtain 8 data in total. Figure 7 The 0th data and the 4th data are the same, the 1st data and the 5th data are the same, the 2nd data and the 6th data are the same, and the 3rd data and the 7th data are the same. Based on this, by constructing two fixed indexes, the above 4 data can be completely selected without overlap, which will not be repeated here.
[0121] Figure 8 It is a structural diagram of a computing array provided according to an embodiment of the present disclosure. The embodiment of the present disclosure can be applied to the situation where after the data to be calculated is selected through multiple general data preprocessing circuits in the computing array, multiple computing units in the computing array perform corresponding calculations on the selected data.
[0122] like Figure 8 As shown, the computing array includes a plurality of computing units 810 and a plurality of general data preprocessing circuits 820 as described in the embodiment of the present disclosure, which are respectively pre-connected to each of the computing units.
[0123] The calculation unit 810 is used to receive the data selected from multiple groups of data by the connected general data preprocessing circuit 820, and perform calculation after combining the data.
[0124] Generally speaking, the computing units in the computing array need to be arranged in a certain arrangement, and adjacent or two computing units generally have communication connections. In this embodiment, there is no restriction on the arrangement and connection of the computing units.
[0125] In this embodiment, the number of computing units 810 is consistent with the number of general data preprocessing circuits 820. That is, each computing unit 810 is configured with a corresponding general data preprocessing circuit 820 for independently selecting the data required for calculation for the computing unit 810.
[0126] The technical solution of the disclosed embodiment can effectively reduce the hardware resource overhead of the entire computing array while being compatible with multiple input data modes by applying a universal data preprocessing circuit in the computing array. It is particularly suitable for high-computing-power scenarios in which the computing array contains a large number of computing units.
[0127] Further, in Fig. 9 FIG. 4 shows a structural diagram of another computing array provided according to an embodiment of the present disclosure. Fig. 9 As shown, in the computing array, the computing units are arranged in a square matrix, multiple computing units in the same row are connected in series, and multiple computing units in the same column are connected in series, wherein:
[0128] Each of the computing units in the first row of the computing array is used to input preset weight data ( Fig. 9 Simplified as weights in the figure), each of the weight data is used to connect to the universal data preprocessing circuit ( Fig. 9 The selected data (simplified as a preprocessing module) are calculated together, and the calculated partial products are respectively transmitted to each of the calculation units in the next row;
[0129] Each of the computing units in the non-first row of the computing array is used to perform calculations based on the received partial products and the data selected by the connected universal data preprocessing circuit, and pass the calculated new partial products to each of the computing units in the next row until the complete calculation process is completed.
[0130] Among them, Fig. 9 The connection method of each computing unit in the computing array and the data flow when each computing unit jointly performs multiplication calculations are specifically clarified. Through the above settings, the efficiency of the computing array can be effectively improved, and the computing needs of high-precision and timely computing scenarios can be effectively met.
[0131] It should be further explained that when the computing array described above processes sparse data and non-sparse data, the arrangement of the data selected by the general data preprocessing circuit has certain differences.
[0132] Specifically, in Fig.10 FIG. 2 shows a schematic diagram of data sorting after a sparse data is selected by a general data preprocessing circuit according to an embodiment of the present disclosure. Fig.10 As shown, when the parallelism of the computing units in the computing array is 64, and the sparse format of the input sparse data is 2:4, each general data preprocessing circuit needs to input 128 data at a time, and a total of 16 groups of 2:4 sparse data need to be input. One group of data needs to select 2 data from 4 input data through a 4-to-2 logic unit (the embodiment of the present disclosure can be implemented by 2 3-to-1 data selectors) to finally obtain 64 data. Fig.10 It can be seen that the 64 data finally screened out are arranged according to the data arrangement of the 128 data originally input, and the only difference is that the 0-valued data in the above 128 data are removed. That is, when the general data preprocessing circuit performs data selection on the input sparse data, no data rearrangement occurs.
[0133] Further, in Fig.11 FIG. 2 shows a schematic diagram of data sorting after a general data preprocessing circuit performs data selection on a non-sparse data applicable to an embodiment of the present disclosure. Fig.11 As shown, when the general data preprocessing circuit of the above structure performs data selection on non-sparse data, it is necessary to first copy the 64 non-sparse data input to obtain 128 non-sparse data. Afterwards, every 4 data are divided into a group, and a total of 16 groups of data are similarly input into the above-mentioned 16 4-to-2 logic units. Among them, the first 8 groups of data use the second index 4'b0100 in common to select the first two data in each group of data. That is: in the group of data {0,1,2,3}, select the first two data {0,1}, in the group of data {4,5,6,7}, select the first two data {4,5}, and so on. In addition, the second index 4'b1110 is used in the last 8 groups of data to select the last two data in each group of data. That is, in the data set {0,1,2,3}, select the last two data {2,3}, and in the data set {4,5,6,7}, select the first two data {6,7}. Obviously, this selection method can definitely select all 64 data, but the order of these 64 data is inconsistent with the input data. Fig.11 shown.
[0134] In order to ensure that such reordered data will not cause errors in subsequent calculations, the weight data needs to be reordered as well. Fig.12 A structural diagram of another computing array provided in an embodiment of the present disclosure.
[0135] like Fig.12 As shown, the computing array also includes: a plurality of weight preprocessing units ( Fig.12 , which is simplified as a preprocessing module located after the input weights);
[0136] Each of the weight preprocessing units is respectively pre-connected with each of the computing units in the first row of the computing array;
[0137] The weight preprocessing unit is used to rearrange the input weight data when the computing array performs calculations matching non-sparse data, so that the rearranged weight data is compatible with the arrangement order of the non-sparse data by the connected general data preprocessing circuit.
[0138] like Fig.12 As shown, in order to adapt the calculation to the calculation unit, the weight data also needs to input 64 data. When the input data is sparse data, the weight preprocessing unit does not work, and the data sorting method of the weight data does not change. When the input data is non-sparse data, the weight preprocessing unit needs to reorder the weight data according to the reordering method of non-sparse data to ensure the accuracy of subsequent matrix multiplication calculations.
[0139] It is understandable that, since multiple computing units can reuse the same weight data, the reused weight data only needs to be pre-processed for rearrangement once, and the additional overhead caused by the rearrangement of the weight data is very small.
[0140] exist Fig.13 FIG. 2 shows a structural diagram of an AI chip provided by an embodiment of the present disclosure. Fig.13 As shown, the AI chip includes a computing array 1310 as described in any one of the embodiments of the present disclosure.
[0141] exist Fig.14 , a structural diagram of another AI chip provided by an embodiment of the present disclosure is shown in FIG. The AI chip further includes: an instruction parsing and control module 1410, a storage module 1420, and an accumulation unit 1430;
[0142] The instruction parsing and control module 1410 is connected to the computing array and the storage module 1420 respectively, the computing array is connected to the storage module 1420 and the accumulating unit 1430 respectively, and the accumulating unit 1430 is connected to the storage module 1420;
[0143] The instruction parsing and control module 1410 is used to parse the received instruction; obtain the operation data from the storage module 1420 according to the parsing result and provide it to the computing array, and send a computing instruction to the computing array;
[0144] The calculation array is used to respond to the calculation instruction, perform calculation according to the received operation data, and send the calculated multiple partial products to the accumulation unit 1430;
[0145] The accumulation unit 1430 is used to perform accumulation calculation according to the received multiple partial products to obtain a final calculation result, and send the final calculation result to the storage module 1420 for storage.
[0146] Fig.15 is a structural diagram of an electronic device provided by an embodiment of the present disclosure. Fig.15 As shown, the electronic device includes the AI chip 1510 as described in any one of the embodiments of the present disclosure.
[0147] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions provided by this disclosure can be achieved, and this document does not limit this.
[0148] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A general data preprocessing circuit, characterized in that: include: A plurality of combined selection modules, each of which includes a first selection submodule and a second selection submodule; The first selection submodule is used to select in the first index or the second index according to the input data mode, and provide the selected target index to the second selection submodule in the same module; The second selection submodule is used to select a target number of data from a set of input data according to the received target index.
2. The universal data preprocessing circuit according to claim 1, characterized in that: The input data mode includes: sparse input data or non-sparse input data; When the input data mode is the sparse input data, the input set of data is sparse data, and when the input data mode is the non-sparse input data, the input set of data is non-sparse data; The first index is used to describe the index position of the non-zero data to be selected in a set of the sparse data, and the second index is used to describe the index position of the data to be selected in a set of the non-sparse data.
3. The universal data preprocessing circuit according to claim 2, characterized in that: In the general data preprocessing circuit, the total amount of data input at a time is n*k, and the total amount of data selected at a time is n*m, where: The number of the combination selection modules is n, the target number is m, and the combination selection module is used to select m data from k data in each group when n groups of data are input; n and k are integers greater than or equal to 2, and m is a positive integer less than k.
4. The universal data preprocessing circuit according to claim 3, characterized in that: When the input data mode is the sparse input data, each of the combination selection modules respectively inputs a group of sparse data of m:k sparse structure; When the input data mode is the non-sparse input data, each combination selection module is used to input k non-sparse data selected in sequence from the copy results obtained by copying n*m non-sparse data by k / m copies.
5. The universal data preprocessing circuit according to claim 4, characterized in that: The general data preprocessing circuit is used to provide input data to the computing unit; The number n of the combination selection modules is associated with the computational parallelism of the computing unit and the sparse structure of the input sparse data.
6. The universal data preprocessing circuit according to claim 5, characterized in that: When the sparse structure of the input sparse data is m:k, the number of the combination selection modules n=l / m; Wherein, l is the computational parallelism of the computing unit, and the m:k sparse structure means that in a group of sparse data containing k data, there are m non-zero data.
7. The universal data preprocessing circuit according to any one of claims 4 to 6, characterized in that: The first selection submodule includes a plurality of two-to-one data selectors.
8. The universal data preprocessing circuit according to claim 7, characterized in that: The number of two-to-one data selectors included in the first selection submodule matches the binary length of the first index or the second index; The binary lengths of the first index and the second index are the same.
9. The universal data preprocessing circuit according to claim 8, characterized in that: The first index is determined by the data form of the sparse data input into the combination selection module; The second index is a predetermined fixed index, and the data selected by each combination selection module through each second index together constitute complete n*m non-sparse data.
10. The universal data preprocessing circuit according to any one of claims 4 to 6, characterized in that: The second selection submodule includes m w-to-one data selectors, where w is a positive integer less than k.
11. The universal data preprocessing circuit according to claim 10, characterized in that: In a group of k input data, w data are sequentially obtained by sliding with a sliding interval of 1 bit length and are respectively connected to each of the w-to-one data selectors in the second selection submodule.
12. A computing array, characterized in that: The method comprises a plurality of computing units, and a plurality of universal data preprocessing circuits according to any one of claims 1 to 11 respectively connected in front of each of the computing units; The calculation unit is used to receive the data selected from multiple groups of data by the connected general data preprocessing circuit, and perform calculation after combining the data.
13. The computing array according to claim 12, characterized in that: In the computing array, the computing units are arranged in a square matrix, multiple computing units in the same row are connected in series, and multiple computing units in the same column are connected in series, wherein: Each of the calculation units in the first row of the calculation array is used to input preset weight data, each of the weight data is used to perform calculations together with the data selected by the connected universal data preprocessing circuit, and the calculated partial products are respectively transmitted to each of the calculation units in the next row; Each of the computing units in the non-first row of the computing array is used to perform calculations based on the received partial products and the data selected by the connected universal data preprocessing circuit, and pass the calculated new partial products to each of the computing units in the next row until the complete calculation process is completed.
14. The computing array according to claim 13, characterized in that: The computing array further comprises: a plurality of weight preprocessing units; Each of the weight preprocessing units is respectively pre-connected with each of the computing units in the first row of the computing array; The weight preprocessing unit is used to rearrange the input weight data when the computing array performs calculations matching non-sparse data, so that the rearranged weight data is compatible with the arrangement order of the non-sparse data by the connected general data preprocessing circuit.
15. An artificial intelligence (AI) chip, characterized in that: Comprising a computing array as described in any one of claims 12-14.
16. The AI chip according to claim 15, characterized in that: The AI chip also includes: an instruction parsing and control module, a storage module and an accumulation unit; The instruction parsing and control module is connected to the computing array and the storage module respectively, the computing array is connected to the storage module and the accumulating unit respectively, and the accumulating unit is connected to the storage module; The instruction parsing and control module is used to parse the received instructions; obtain the operation data from the storage module according to the parsing result and provide it to the computing array, and send a computing instruction to the computing array; The calculation array is used to perform calculation according to the received operation data in response to the calculation instruction, and send the calculated multiple partial products to the accumulation unit; The accumulation unit is used to perform accumulation calculation according to the received multiple partial products to obtain a final calculation result, and send the final calculation result to the storage module for storage.
17. An electronic device, characterized in that: Comprising the AI chip as described in claim 15 or 16.