tensor core component of a processor and a processor
Patent Information
- Application Number
- CN202511120698.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-08-11
AI Technical Summary
由此,标量核心组件就需要额外执行从结果矩阵中筛选出各行的元素最大值的操作,从而,导致标量核心组件的处理效率降低
[0032]基于本申请的实施例,处理器的张量核心组件中的脉动阵列除了包括用于通过输入矩阵之间的乘法运算计算结果矩阵的多个计算单元之外,还成列部署有分别与结果矩阵的多行对应的一列筛选单元。其中,该列筛选单元可以在结果矩阵的元素从脉动阵列逐列输出的过程中,筛选出各行的元素中的元素最大值;并且,该列筛选单元还可以响应于结果矩阵的元素从脉动阵列的逐列输出完成,将筛选出的各行的元素最大值成列输出。因此,张量核心组件可以具有伴随着结果矩阵的逐列输出而实现的元素最大值筛选功能,从而,可以避免处理器的标量核心组件为了筛选结果矩阵的各行的元素最大值而额外执行筛选操作,进而可以提升处理器中的标量核心组件的处理效率。
Smart Images

Figure CN121029685B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence chips, and in particular to a tensor core component of a processor, a processor, and a method for element selection of the tensor core component of a processor. Background Technology
[0002] Processors used in artificial intelligence can include a TensorCore component for performing matrix multiplication operations, and the TensorCore component can output the result matrix obtained by multiplying any two matrices to the ScalarCore component in the processor for operator-based function operations.
[0003] In some applications of artificial intelligence, the independent variables of function operations not only need to include all elements in the result matrix, but also the maximum value of each element in each row of the result matrix. Therefore, the scalar core component needs to perform an additional operation to filter out the maximum value of each row from the result matrix, thus reducing the processing efficiency of the scalar core component.
[0004] As can be seen above, how to improve the processing efficiency of scalar core components in processors has become a technical problem that needs to be solved in the existing technology. Summary of the Invention
[0005] Embodiments of this application provide a tensor core component for a processor, a processor, and a method for filtering elements in the tensor core component for a processor.
[0006] In one embodiment of this application, a tensor core component for a processor is provided, comprising:
[0007] The first input buffer is used to buffer the first input matrix;
[0008] The second input buffer is used to cache the second input matrix;
[0009] A pulsating array includes multiple computing units arranged in an array and a row of filtering units deployed in a column;
[0010] in:
[0011] The plurality of computing units are used to collaboratively calculate the result matrix obtained by multiplying the first input matrix and the second input matrix. The plurality of computing units correspond to each matrix position of the result matrix. When the result matrix is calculated, the elements at each matrix position in the result matrix are stored in the corresponding computing units of the plurality of computing units.
[0012] The filtering unit corresponds to multiple rows of the result matrix. Elements in the plurality of calculation units are output column by column through the filtering unit. Furthermore, the filtering unit is configured as follows:
[0013] During the column-by-column output of elements in the plurality of computational units, the maximum value of each row's elements is selected; and,
[0014] In response to the completion of the column-by-column output of elements in the plurality of computing units, the maximum values of the elements in each of the selected rows are output as a column.
[0015] In some examples, optionally, the column filtering unit is specifically configured to: determine the maximum value of the elements in each row by filtering the elements in each row during the column-by-column output of the elements in the plurality of calculation units.
[0016] In some examples, optionally, the elements in the plurality of computing units are output column by column shifting from the plurality of computing units; the column filtering unit is specifically configured to: filter the elements in each row column by column according to the rhythm of column shifting to determine the maximum value of the elements in each row.
[0017] In some examples, optionally, elements in the plurality of computing units are shifted column by column to an edge computing unit located at the end of the shift in the plurality of computing units, and are shifted out column by column from the edge computing unit; the filter unit is specifically configured to: perform a peer comparison for filtering the elements in each row of each column to be shifted out in the edge computing unit.
[0018] In some examples, the column filtering unit is optionally configured to dynamically maintain the maximum value of each row of elements in the row as the maximum value of the elements already output in each row, based on the comparison results of the row-by-row element comparison.
[0019] In some examples, optionally, the column filtering unit is specifically configured to: record the elements of each row in the first column to be shifted and output in the column edge computing unit as the maximum value of the elements of each row initially filtered; starting from the second column to be output in the column edge computing unit after the first column shifted and output, compare the elements of each row in the current column to be output in the column edge computing unit with the maximum value of the elements of the recorded row; wherein, if the element of any row in the current column to be output is greater than the maximum value of the elements of the recorded row, then update the maximum value of the elements of that row in the current column to be output to the element of that row.
[0020] In some examples, optionally, each of the column filtering units includes: a dynamic recording register for caching the maximum value of the corresponding row; wherein the maximum value of the corresponding row is initially cached as the element of the corresponding row of the first column to be shifted and output in the column edge computing unit; and a comparator for shifting from the first column output to the next column to be output in the column edge computing unit, comparing the element of the corresponding row of the current column to be output with the maximum value of the element of the same row currently cached in the dynamic recording register; wherein if the element of the corresponding row of the current column to be output is greater than the maximum value of the element of the same row currently cached in the dynamic recording register, then the maximum value of the element of the same row cached in the dynamic recording register is updated to the element of the corresponding row of the current column to be output.
[0021] In some examples, optionally, each of the plurality of computing units includes a result element register, and when the result matrix is calculated, the elements at each matrix position in the result matrix are stored in the result element register of the corresponding computing unit in the plurality of computing units; the dynamic record register of each of the column of filtering units is time-division multiplexed with the result element register of the column of edge computing units.
[0022] In another embodiment of this application, a processor is provided, including a tensor core component and a scalar core component as described in the foregoing embodiments; wherein, the elements in the plurality of computing units are output to the scalar core component column by column through the column filtering unit, and the maximum values of the elements in each row filtered by the column filtering unit are output to the scalar core component in a column.
[0023] In another embodiment of this application, a method for element filtering of a tensor core component for a processor is provided. The tensor core component includes a systolic array, which includes multiple computing units arranged in an array. The multiple computing units are used to collaboratively compute a result matrix obtained by multiplying a first input matrix and a second input matrix. The multiple computing units correspond to different matrix positions in the result matrix, and when the result matrix is computed, the elements at each matrix position in the result matrix are stored in the corresponding computing units of the multiple computing units. The element filtering method includes:
[0024] During the column-by-column output of elements in the plurality of computational units, the maximum value of each row's elements is selected; and,
[0025] In response to the completion of the column-by-column output of elements in the plurality of computing units, the maximum values of the elements in each of the selected rows are output as a column.
[0026] In some examples, optionally, the filtering of the maximum value of elements in each row includes: determining the maximum value of elements in each row by filtering the elements column by column.
[0027] In some examples, optionally, the elements in the plurality of computing units are output column by column by column shifting; the step of filtering out the maximum value of the elements in each row includes: filtering the elements in each row column by column according to the rhythm of column shifting to determine the maximum value of the elements in each row.
[0028] In some examples, optionally, elements in the plurality of computing units are shifted column by column to an edge computing unit located at the end of the shift in the plurality of computing units, and are shifted out column by column from the edge computing unit; the step of determining the maximum value of elements in each row by filtering the elements in each row according to the shifting rhythm includes: performing a peer comparison for filtering the maximum value of elements in each row of elements to be shifted out in each column of the edge computing unit.
[0029] In some examples, optionally, the step of performing a peer element comparison for filtering the maximum value of each row of elements to be shifted and output in each column of the edge computing unit includes: dynamically maintaining the maximum value of each row of elements recorded as the maximum value of elements in each row of output elements according to the comparison results of the peer element comparison column by column.
[0030] In some examples, optionally, the step of dynamically maintaining the maximum value of each row of elements recorded based on the comparison results of row-by-row element comparison includes: recording the elements of each row to be shifted and output in the first column of the edge computing unit as the maximum value of each row of elements initially selected; starting from the second column to be output in the edge computing unit after the first column shifted and output, comparing the elements of each row of the current column to be output in the edge computing unit with the maximum value of the recorded row of elements; wherein, if the element of any row in the current column to be output is greater than the maximum value of the recorded row of elements, then updating the maximum value of the recorded row to the element of the current row to be output.
[0031] In another embodiment of this application, an electronic device is provided, including a processor as described in the foregoing embodiments.
[0032] Based on embodiments of this application, the systolic array in the tensor core component of the processor, in addition to including multiple computational units for calculating the result matrix through multiplication operations between input matrices, also includes a column of filtering units arranged in rows, each corresponding to a row of the result matrix. This column of filtering units can filter out the maximum value of each row's elements as the elements of the result matrix are output column by column from the systolic array; furthermore, in response to the completion of the column-by-column output of the result matrix's elements from the systolic array, this column of filtering units can output the filtered maximum values of each row in a column. Therefore, the tensor core component can have a maximum value filtering function implemented along with the column-by-column output of the result matrix, thereby avoiding the need for the processor's scalar core component to perform additional filtering operations to filter the maximum values of each row of the result matrix, and thus improving the processing efficiency of the scalar core component in the processor. Attached Figure Description
[0033] The following figures are for illustrative purposes only and do not limit the scope of this application:
[0034] Figure 1 This is an exemplary structural diagram of the basic framework of the processor in an embodiment of this application;
[0035] Figure 2 This is an exemplary structural diagram of the tensor core component of the processor in an embodiment of this application;
[0036] Figures 3 to 9 This is a schematic diagram illustrating a computational example of the tensor core component of the processor in an embodiment of this application.
[0037] Figures 10 to 12 This is a schematic diagram illustrating an example of the output of the tensor core component of the processor in an embodiment of this application.
[0038] Figure 13 This is a schematic diagram of an exemplary optimized structure of the tensor core component of the processor in an embodiment of this application;
[0039] Figures 14 to 16 The tensor core component of the processor in this application embodiment is based on, as Figure 13 The diagram shows an example of the output of the exemplary optimized structure.
[0040] Figure 17 This is an exemplary structural diagram illustrating a preferred deployment method of the filtering unit in the tensor core component of the processor according to an embodiment of this application.
[0041] Figures 18 to 20 The filtering unit in the tensor core component of the processor in this application embodiment is based on, as Figure 17 The diagram shows the internal instance structure of the preferred deployment method.
[0042] Figure 21 This is an exemplary flowchart illustrating an element filtering method for a tensor core component of a processor in an embodiment of this application. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided with reference to the accompanying drawings and embodiments.
[0044] For example, in the embodiments of this application, the processor can be any one of the following integrated circuit chips suitable for artificial intelligence: GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), NPU (Neural Network Processing Unit), DPU (Deep Learning Processing Unit), APU (Accelerated Processing Unit), and GPGPU (General-Purpose computing on Graphics Processing Units).
[0045] Figure 1 This is an exemplary structural diagram of the basic framework of the processor in an embodiment of this application. Please refer to... Figure 1 The processor may include a tensor core component 10 and a scalar core component 30.
[0046] For example, in an embodiment of this application, the tensor core component 10 can be used to perform a multiplication operation between two matrices (e.g., a first input matrix and a second input matrix input to the tensor core component 10) and obtain a result matrix through the multiplication operation.
[0047] For example, in an embodiment of this application, the scalar core component 30 can be used to perform function operations on the result matrix generated by the tensor core component 10 based on various pre-loaded operators.
[0048] Figure 2 This is an exemplary structural diagram of the tensor core component of the processor in an embodiment of this application. Please refer to... Figure 2In embodiments of this application, the tensor core component 10 may include a first input buffer 110, a second input buffer 120, and a systolic array 150. The first input buffer 110 can be used to buffer the first input array, the second input buffer 120 can be used to buffer the second input array, and the systolic array 150 can be used to compute the result matrix obtained by multiplying the first input matrix and the second input matrix.
[0049] For example, in embodiments of this application, the pulsating array 150 may include multiple computing units arranged in an array, such as... Figure 2 The m×n computational units U_i,j shown in the diagram, i represents the row of computational unit U_i,j in the array, and j represents the column of computational unit in the array, that is, 1≤i≤m and 1≤j≤n.
[0050] For example, in the embodiments of this application, multiple computing units (e.g., m×n computing units U_i,j) can be used to collaboratively compute the result matrix obtained by multiplying the first input matrix and the second input matrix. The multiple computing units (e.g., m×n computing units U_i,j) can respectively correspond to each matrix position (i,j) of the result matrix. Furthermore, when the result matrix is computed, the element E at each matrix position (i,j) in the result matrix... ij They can be generated and stored in corresponding computational units U_i and j in multiple computational units respectively.
[0051] For example, in an embodiment of this application, the pulsating array 150 may use GEMM (General Matrix Multiply) to calculate the result matrix obtained by multiplying the first input matrix and the second input matrix. In this case, multiple computing units (e.g., m×n computing units U_i,j) may each include a MAC (Multiply Accumulate) unit.
[0052] For example, in an embodiment of this application, if the systolic array 150 can use GEMM to calculate the result matrix obtained by multiplying the first input matrix and the second input matrix, then the first input buffer 110 can be arranged on one side of the row direction of the systolic array 150, the second input buffer 120 can be arranged on one side of the column direction of the systolic array 150, each calculation unit U_i,j has two inputs in the row direction and the column direction, each calculation unit U_i,j also has one output on the other side of the row direction, and the side of the systolic array 150 opposite to the first input buffer 110 in the row direction is the output side of the result matrix.
[0053] Figures 3 to 9This is a schematic diagram illustrating a computational example of the tensor core component of the processor in an embodiment of this application. Figures 3 to 9 In this example, taking the first input matrix, the second input matrix, the result input matrix, and the pulsation array 150 as all adopting a 3×3 specification, that is, m and n mentioned above are both 3, it shows the process of the pulsation array 150 using GEMM to calculate the result matrix obtained by multiplying the first input matrix and the second input matrix.
[0054] exist Figures 3 to 9 In the first input buffer 110, the first input matrix A can be represented by the following expression (1), and the second input matrix B in the second input buffer 120 can be represented by the following expression (2):
[0055]
[0056] exist Figures 3 to 9 In the first input buffer 110, the first input matrix A and the second input matrix B in the second input buffer 120 are periodically input from the two input sides in the row and column directions of the pulsating array 150 according to the element grouping method along the diagonal direction, and periodically shifted and passed in the row and column directions.
[0057] In the first cycle of the computation instance of Tensor Core Component 10, such as Figure 3 As shown, element a in the first input matrix A 11 and b of the second input matrix B 11 The data are input from the row and column directions of the pulsating array 150 to the computation unit U_1,1 respectively, and multiplied in the computation unit U_1,1 to obtain a. 11 ×b 11 .
[0058] In the second cycle of the computation instance of Tensor Core Component 10, such as Figure 4 As shown, element a in the first input matrix A 12 and b of the second input matrix B 21 The data are input from the row and column directions of the pulsating array 150 to the computation unit U_1,1 respectively, and multiplied and summed in the computation unit U_1,1 to obtain a. 11 ×b 11 +a 12 ×b 21 Meanwhile, see also Figure 4 The element a in the first input matrix A 11 The shift in the row direction of the pulsating array 150 is passed to the computing units U_1, 2, and the second input matrix B. 12 The inputs are fed into the computation units U_1,2 along the column direction of the pulsating array 150, and multiplied by the inputs in the computation units U_1,2 to obtain a.11 ×b 12 Similarly, in Figure 4 In the second input matrix B, b 11 The shifted element a in the column direction of the pulsating array 150 is transferred to the computing unit U_2,1, and the element a in the first input matrix A. 21 The data is input from the row direction of the pulsating array 150 to the computation units U_2,1, and can be multiplied in the computation units U_1,2 to obtain a. 21 ×b 11 .
[0059] In the third cycle of the computation instance of Tensor Core Component 10, such as Figure 5 As shown, element a in the first input matrix A 13 a 22 a 31 The elements b in the second input matrix B are respectively input into the calculation units U_1,1, U_2,1, and U_3,1 in the row direction. 31 b 22 b 13 Elements from the first input matrix A, already input to the systolic array 150, are shifted and passed along the row direction, while elements from the second input matrix B, already input to the systolic array 150, are shifted and passed along the column direction. Thus, the calculation unit U_1,1 can obtain element E in the result matrix by multiplying the new input elements and summing them with the previous calculation result. 11 =a 11 ×b 11 +a 12 ×b 21 +a 13 ×b 31 The calculation units U_1,2 and U_2,1 can multiply the new input elements passed by the shift and add them to the previous calculation result. In addition, the calculation units U_1,3, U_2,2 and U_3,1 can multiply the new input elements.
[0060] In the fourth cycle of the computation instance of Tensor Core Component 10, such as Figure 6 As shown, element a in the first input matrix A 32 and a 23 The elements b in the second input matrix B are respectively input into the calculation units U_2,1 and U_3,1 in the row direction. 32 and b 23Elements from the first input matrix A, already input to the systolic array 150, are shifted and passed along the row direction, while elements from the second input matrix B, already input to the systolic array 150, are shifted and passed along the column direction. Thus, the calculation units U_1 and U_2 can obtain element E in the result matrix by multiplying the new input elements and summing them with the previous calculation result. 12 =a 11 ×b 12 +a 12 ×b 22 +a 13 ×b 32 The computation unit U_2,1 can obtain the element E in the result matrix by multiplying the new input element and summing it with the result of the previous calculation. 21 =a 21 ×b 11 +a 22 ×b 21 +a 23 ×b 31 The calculation units U_1,3, U_2,2, and U_1,3 can multiply the new input elements passed by the shift and add them to the previous calculation result. Furthermore, the calculation units U_2,3 and U_3,2 can multiply the new input elements.
[0061] In the fifth cycle of the computation instance of Tensor Core Component 10, such as Figure 7 As shown, element a in the first input matrix A 33 The elements b in the second input matrix B are respectively input into the calculation units U_2,1 and U_3,1 in the row direction. 33 Elements from the first input matrix A, already input to the systolic array 150, are shifted and passed along the row direction, while elements from the second input matrix B, already input to the systolic array 150, are shifted and passed along the column direction. Thus, the calculation units U_1 and U_3 can obtain element E in the result matrix by multiplying the new input elements and summing them with the previous calculation results. 13 =a 11 ×b 13 +a 12 ×b 23 +a 13 ×b 33 The computation unit U_2,2 can obtain the element E in the result matrix by multiplying the new input element and summing it with the result of the previous calculation. 22 =a 21 ×b 12 +a 22 ×b 22 +a 23 ×b32 The computation unit U_3,1 can obtain the element E in the result matrix by multiplying the new input element and summing it with the result of the previous calculation. 31 =a 31 ×b 11 +a 32 ×b 21 +a 33 ×b 31 The calculation units U_2,3 and U_3,2 can multiply the new input element passed by the shift and add it to the previous calculation result, and the calculation unit U_3,3 can multiply the new input element.
[0062] In the sixth cycle of the computation instance of Tensor Core Component 10, such as Figure 8 As shown, the input of the first input matrix A and the second input matrix B to the systolic array 150 ends. Elements in the first input matrix A that have already been input to the systolic array 150 continue to be shifted and passed along the row direction, and elements in the second input matrix B that have already been input to the systolic array 150 continue to be shifted and passed along the column direction. Therefore, computation units U_2 and U_3 can obtain element E in the result matrix by multiplying the new input elements and accumulating the result with the previous calculation result. 23 =a 21 ×b 13 +a 22 ×b 23 +a 23 ×b 33 The computation unit U_3,2 can obtain the element E in the result matrix by multiplying the new input element and accumulating it with the result of the previous calculation. 32 =a 31 ×b 12 +a 32 ×b 22 +a 33 ×b 33 Furthermore, the computation unit U_3 can multiply the new input element passed by shift and add it to the previous calculation result.
[0063] In the seventh cycle of the computation instance of Tensor Core Component 10, such as Figure 9 As shown, the elements already input into the systolic array 150 in the first input matrix A continue to be shifted and passed along the row direction, and the elements already input into the systolic array 150 in the second input matrix B continue to be shifted and passed along the column direction. Thus, the computation units U_3,3 obtain element E in the result matrix by multiplying the new input elements and accumulating the result with the previous calculation. 33 =a 31 ×b 13 +a 32 ×b 23 +a 33 ×b33 At this point, all elements in the result matrix are determined and stored in their respective computational units U_i,j.
[0064] For example, in an embodiment of this application, when the result matrix is calculated, the element E at each matrix position in the result matrix is... ij The results are generated and stored in the corresponding computational units U_i,j of the systolic array 150, and all elements of the result matrix in the multiple computational units of the systolic array 150 can be output column by column.
[0065] For example, in an embodiment of this application, all elements of the result matrix in the plurality of computing units of the systolic array 150 can be output column by column by column shifting in the row direction of the systolic array 150, for example, output column by column to the scalar core component 30.
[0066] Figures 10 to 12 This is a schematic diagram illustrating an example of the output of the tensor core component of the processor in an embodiment of this application. Figures 10 to 12 In this example, taking a result input matrix and a systolic array 150 both of 3×3 size (i.e., m and n are both 3 as mentioned above), the process of the systolic array 150 outputting all elements of the result matrix column by column through column-by-column shifting is shown. The three columns of the result matrix are sequentially processed... Figure 10 , Figure 11 as well as Figure 12 The output instance of the tensor core component 10 shown is output by a three-shift operation, and then output column by column to the scalar core component 30.
[0067] For example, in an embodiment of this application, the elements in the result matrix output column by column from the systolic array 15 can be accessed by the scalar core component 30 via the TLR (Thread local register) and by the scalar core component 30 through the running thread.
[0068] For example, in embodiments of this application, the tensor core component 10 and scalar core component 30 of the processor can undertake the computational tasks of large models. For instance, the tensor core component 10 and scalar core component 30 can undertake the computational tasks of the self-attention layer in a large model of the Transformer architecture. The self-attention layer can capture the dependencies between various elements (such as words, tags, etc.) in the input sequence and determine the elements in the input sequence that require focused attention through dynamic weight calculation.
[0069] For example, in the embodiments of this application, if the tensor core component 10 and the scalar core component 30 undertake the computational task of the self-attention layer, the input sequence of the self-attention layer can be input into the processor in the form of tensor data, and the first input matrix and the second input matrix of the tensor core component 10 can be associated with the input sequence of the self-attention layer. For example, the first input matrix and the second input matrix can respectively include a Q (query) matrix and a K (key) matrix associated with the input sequence, and the result matrix obtained by multiplying the first input matrix and the second input matrix by the tensor core component 10 can be a V (value) matrix.
[0070] For example, in an embodiment of this application, if the tensor core component 10 and the scalar core component 30 undertake the computational task of the self-attention layer, then the function operation performed by the scalar core component 30 on the result matrix based on the operator may include a softmax function operation based on the softmax (normalized exponential function) operator. The principle of the softmax function can be represented by the following expression (3):
[0071]
[0072] In the above expression (1), softmax(z) i ) represents the result of the softmax function, z i and z j Let i and j represent the i-th and j-th elements in the input vector of the softmax function, respectively. That is, i and j in expression (3) represent the order of the elements in the input vector (i.e., different from the meaning of i and j in "U_i,j" described above), and M represents the inclusion of zi and z. j The maximum value of all elements in the input vector, i.e., M, can be expressed by the following expression (4):
[0073] M = max{z} (4)
[0074] In the above expression (4), {z} represents all elements in the input vector, and max{z} is the maximum value of all elements in the input vector.
[0075] For example, in an embodiment of this application, the input vector of the softmax function can be any row {E} of the result moment (e.g., V matrix) output by the tensor core component 10. i1 ~E inIn this case, M in expression (3) and M represented by expression (4) can be the maximum value of any row element in the result matrix output by the tensor core component 10 (i.e., the systolic array 150). Therefore, it can be understood that in the embodiments of this application, the row direction of the systolic array 150 can be the direction in which the elements of the input vector of the scalar core component 30 are arranged in the result matrix, and the column direction of the systolic array 150 is the opposite direction to the row direction. That is, in the embodiments of this application, the row and column directions of the systolic array 150 can be determined by the recognition direction of the input vector by the scalar core component 30 in the result matrix.
[0076] For example, in embodiments of this application, structural optimizations were performed on the tensor core component 10 (e.g., systolic array 150) in the processor to avoid the scalar core component 30 performing additional operations for determining the maximum value of the elements in each row of the resulting matrix.
[0077] Figure 13 This is a schematic diagram of an exemplary optimized structure of the tensor core component of the processor in an embodiment of this application. Please refer to... Figure 13 In the embodiments of this application, the systolic array 150 of the tensor core component 10 of the processor may further include a column of filtering units F_1 to F_m deployed in the column direction of the systolic array 150. The column of filtering units F_1 to F_m corresponds to multiple rows (i.e. m rows) of the result matrix. The elements of the result matrix in the multiple computing units of the systolic array 150 are output column by column through the column of filtering units F_1 to F_m, that is, output to the scalar core component 30 through the column of filtering units F_1 to F_m.
[0078] For example, in an embodiment of this application, a column of filtering units F_1 to F_m of the systolic array 150 can be deployed in a column on the output side of the systolic array 150 opposite to the first input buffer 110 in the row direction. Thus, the elements of the result matrix in the multiple computational units of the systolic array 150 can be output column by column through the filtering units F_1 to F_m. For example, a column of computational units U_1,n to U_m,n located at the end of the shift in the row direction of the systolic array 150 can be called an edge computational unit. Furthermore, the filtering units F_1 to F_m can be deployed in a column downstream of the shift of the edge computational units U_1,n to U_m,n, or integrated into the edge computational units U_1,n to U_m,n.
[0079] For example, in an embodiment of this application, a column of filtering units F_1 to F_m of the systolic array 150 can be configured to: filter out the maximum value of the elements in each row during the column-by-column output of elements in multiple calculation units (i.e., the elements of the result matrix in the multiple calculation units of the systolic array 150). That is, the filtering unit Fi corresponding to any row can be configured to: filter out the maximum value of the elements in the calculation units U_i,1 to U_i,n of the corresponding row (i.e., the elements E in the calculation units U_i,1 to U_i,n of the result matrix in that row of the systolic array 150). i1 ~E in During the column-by-column output process, the element E in that row is selected. i1 ~E in The maximum value of elements in Max{E i1 ~E in}
[0080] For example, in an embodiment of this application, a column of filtering units F_1 to F_m of the pulsating array 150 can also be configured to: in response to the completion of column-by-column output of elements in multiple computing units (i.e., elements of the result matrix in multiple computing units of the pulsating array 150), output the maximum value of each row of elements in the filtered array. That is, the filtering unit F_i corresponding to any row can be configured to: in response to the element E of U_i,1 to U_i,n in the computing unit of the corresponding row i1 ~E in (That is, the element E in the calculation unit U_i,1 to U_i,n of the result matrix in the pulsating array 150) i1 ~E in The column-by-column output is complete, and the maximum value of the element in the selected row, Max{E}, is found. i1 ~E in Output the maximum value of elements in the same row as the maximum value of the elements in the other rows.
[0081] For example, the maximum values of the elements in each filtered row can also be output in a column to the scalar core component 30, such as in the TLR of the scalar core component 30.
[0082] Based on the embodiments of this application, the systolic array 150 in the tensor core component 10 of the processor, in addition to including multiple computing units for calculating the result matrix through multiplication operations between input matrices, also includes a column of filtering units F_1 to F_m, each corresponding to a row of the result matrix. These filtering units F_1 to F_m can filter out the maximum value of each row's elements as the elements of the result matrix are output column by column from the systolic array 150; furthermore, in response to the completion of the column-by-column output of the result matrix's elements from the systolic array 150, these filtering units F_1 to F_m can also output the filtered maximum values of each row in a column. Therefore, the tensor core component 10 can have a maximum value filtering function implemented along with the column-by-column output of the result matrix, thereby avoiding the need for the scalar core component 30 of the processor to perform additional filtering operations to filter the maximum values of each row of the result matrix, and thus improving the processing efficiency of the scalar core component 30 in the processor.
[0083] For example, in the embodiments of this application, a column of filtering units F_1 to F_m of the systolic array 150 can be specifically configured as follows: during the column-by-column output of elements (i.e., elements of the result matrix in the multiple calculation units of the systolic array 150) in the multiple calculation units of the systolic array 150, the maximum value of each element in each row is determined by filtering the elements column by column. That is, the filtering unit F_i corresponding to any row can be configured as follows: the maximum value of each element in the calculation units U_i,1 to U_i,n of that row of the systolic array 150 is determined by filtering the elements column by column. i1 ~E in (That is, the element E in the calculation unit U_i,1 to U_i,n of the result matrix in the pulsating array 150) i1 ~E in During the column-by-column output of ), by using the element E of that row... i1 ~E in Perform column-by-column filtering to determine element E in the current row. i1 ~E in The maximum value of elements in Max{E i1 ~E in}
[0084] For example, in an embodiment of this application, the elements in the plurality of computing units of the systolic array 150 (i.e., the elements of the result matrix in the plurality of computing units of the systolic array 150) can be output column by column by column shifting (i.e., column by column shifting in the row direction) between the plurality of computing units. In this case, a column filtering unit F_1 to F_m of the systolic array 150 can be specifically configured to: filter the elements of each row of the result matrix column by column according to the rhythm of column shifting to determine the maximum value of the elements in each row of the result matrix. That is, the filtering unit F_i corresponding to any row can be configured to: filter the elements E in the row computing units U_i,1 to U_i,n according to the rhythm of column shifting. i1 ~E in (That is, the element E in the calculation unit U_i,1 to U_i,n of the result matrix in the pulsating array 150) i1 ~E in Perform column-by-column filtering to determine the element E of that row in the result matrix. i1 ~E in The maximum value of elements in Max{E i1 ~E in}
[0085] For example, in an embodiment of this application, elements in multiple computing units of the systolic array 150 (i.e., elements of the result matrix in multiple computing units of the systolic array 150) are shifted column-by-column to a column of edge computing units U_1,n to U_m,n, and output column-by-column from the column of edge computing units U_1,n to U_m,n. In this case, a column of filtering units F_1 to F_m of the systolic array 150 can be specifically configured to: perform a row-by-row element comparison for each row of elements to be shifted and output in each column of edge computing units U_1,n to U_m,n, so that the column-by-column filtering of elements in each row of the result matrix can be synchronized with the column-by-column shifting rhythm. That is, the filtering unit F_i corresponding to any row can be configured to: for the elements E to be shifted and output in the corresponding row of edge computing units U_i,n in a column of edge computing units U_1,n to U_m,n. ij Perform the filtering of elements with the maximum value Max{E} i1 ~E in Comparison of elements in the same category.
[0086] For example, in an embodiment of this application, the peer element comparison performed by a column filtering unit F_1 to F_m of the systolic array 150 may include a comparison between the currently output element and the maximum value among the already output peer elements. In this case, the column filtering units F_1 to F_m of the systolic array 150 can be specifically configured to: dynamically maintain the maximum value of each row of elements recorded in each row at the maximum value among the already output elements in each row, based on the comparison results of the peer element comparison performed column by column. That is, the filtering unit F_i corresponding to any row can be configured to: dynamically maintain the maximum value of the row of elements recorded in that row at the maximum value among the currently output elements in that row, based on the comparison results of the peer element comparison performed column by column for that row.
[0087] Figures 14 to 16 The tensor core component of the processor in this application embodiment is based on, as Figure 13 The diagram illustrates an example of the output of an exemplary optimized structure. Figures 14 to 16 In the example, taking the result input matrix and the pulsation array 150 as both adopting the 3×3 specification, and a column of filtering units F_1 to F_m including three (that is, m and n mentioned above are both 3), it is shown that the filtering units F_1 to F_m dynamically maintain the maximum value of each row of elements recorded in each row as the maximum value of the elements already output in each row.
[0088] Please refer to [the website / information] first. Figure 14 In the first cycle of the output instance of the tensor core component 10 in this embodiment, a column of elements E is calculated by a column of edge computing units U_1,3 to U_3,3. 13 ~E 33 After passing through a series of filtering units F_1 to F_3, the elements E in each row of the first column to be shifted and output are sent to the scalar core component 30. Furthermore, the filtering units F_1 to F_3 can also shift the elements E in each row of the first column of the edge computing units U_1,3 to U_3,3. 13 ~E 33 Let each of these be the maximum values of the elements in the initially filtered rows, i.e., the maximum value of the elements in the first row of the initial filtering, Max{E}. 13} = E 13 The maximum value of the elements in the second row is Max{E}. 23} = E 23 And the maximum value of the elements in the third row, Max{E} 33} = E 33Subsequently, starting from the first column after shifting output, and moving to the next column to be output in the edge computing unit, the filtering units F_1 to F_3 can compare the elements of each row in the current column to be output in the edge computing units U_1,3 to U_3,3 with the maximum value of the elements in the same row of the record. If any element in any row of the current column to be output is greater than the maximum value of the elements in the same row of the record, then the maximum value of that row in the record is updated to the element of that row in the current column to be output.
[0089] Please refer to [the website / information] first. Figure 15 In the second cycle of the output instance of the tensor core component 10 in this embodiment, the element E is newly shifted to a column of edge computing units U_1,3 to U_3,3. 12 ~E 32 (That is, calculated by calculation units U_1,2 to U_3,2) After passing through a column of filtering units F_1 to F_3, the output is shifted to the scalar core component 30. Furthermore, the column of filtering units F_1 to F_3 can shift the elements E of each row of the current column (i.e., the second column) to be output from the column of edge calculation units U_1,3 to U_3,3. 12 ~E 32 Each element is compared with the maximum value of the element in the same row of the current record. For example, the filter unit F_1 for any row can compare the current output element of that row (i.e., the element E to be output in the second column). i2 The maximum value of the elements in the current row of the record, Max{E} i3} = E i3 The comparison is performed, and if the element in the current column to be output (i.e., the element E in the second column to be output) is... i2 The maximum value of elements in the same row as the current record (Max{E) i3} = E i3 Then, the maximum value of the recorded element in that row is updated to the element in that row in the current column to be output. Thus, the maximum value of the element in that row can be dynamically maintained as the maximum value Max{E} among the currently output elements in that row. i2 E i3}
[0090] Please continue reading Figure 16 The third cycle of the output instance of the tensor core component 10 in this application embodiment is related to... Figure 15 Similarly, in the second cycle shown, the maximum value of each row can be dynamically maintained as the maximum value Max{E} among the currently output elements in that row. i1 E i2 E i3}
[0091] exist Figure 16Then, the filtering units F_1 to F_3 of the pulsating array 150 can select the maximum value of each element in the filtered rows (e.g., the maximum value of the element in the first row, Max{E}). 11 E 12 E 13 The maximum value of the elements in the second row is Max{E}. 21 E 22 E 23} and the maximum value of the elements in the third row, Max{E 31 E 32 E 33 The output is columnar, that is, the output is sent to the scalar core component 30 to obtain the result matrix.
[0092] Figure 17 This is an exemplary structural diagram illustrating a preferred deployment of the filtering unit in the tensor core component of the processor according to an embodiment of this application. Figure 17 In the example, a column of filtering units F_1 to F_m is integrated into a column of edge computing units U_1,n to U_m,n. Figure 17 The diagram illustrates the calculation units U_i,1 to U_i,n and the corresponding filtering unit F_i for any given row.
[0093] Please see Figure 17 In embodiments of this application, each computing unit U_i,j may include a computing core ALU_i,j and a unit register RGS_i,j. The computing core ALU_i,j is used to execute the computing tasks undertaken by the computing unit U_i,j (e.g., ...). Figures 3 to 9 The matrix operation process shown includes multiplication and addition operations. Furthermore, when any computational unit U_i,j completes the corresponding element E in the result matrix using the computational kernel ALU_i,j... ij After the calculation, the corresponding element E in the result matrix corresponding to the calculation unit U_i,j ij The elements can be stored in the unit register RGS_i,j of the computation unit U_i,j. Therefore, the unit register RGS_i,j of each computation unit U_i,j can be called the result element register during matrix operations. That is, when the result matrix is calculated, the elements at each matrix position in the result matrix are stored in the result element registers of the corresponding computation units in multiple computation units.
[0094] See also Figure 17 In embodiments of this application, the filtering unit F_i corresponding to any row may include a comparator C_i integrated in the edge computing unit U_i,n of that row, and the edge computing unit U_i,n may also include a filtering shift register Filter_i. In this case, the same element E output from the edge computing unit U_i,n of any row...i1 ~E in Each element E can be first shifted to the filter shift register Filter_i of the edge computing unit U_i,n of that row, and then shifted out from the filter shift register Filter_i of the edge computing unit U_i,n of that row. Furthermore, for each element E shifted to the filter shift register Filter_i waiting for output... ij The comparator C_i integrated in the edge computing unit U_i,n of the filtering unit F_i in this row can convert the output element E in the filtering shift register Filter_i of the edge computing unit U_i,n. ij The element is compared with the element currently stored in the cell register RGS_i,j of the edge computing unit U_i,n. Specifically, if the element to be output E in the filter shift register Filter_i of the edge computing unit U_i,n is... ij If the element in the RGS_i,j register of the edge computing unit U_i,n is greater than the element currently stored therein, then the element currently stored in the RGS_i,j register of the edge computing unit U_i,n is updated to the larger output element E in the filter shift register Filter_i. ij Therefore, the elements in the cell register RGS_i,j of the edge computing unit U_i,n can be dynamically maintained (e.g., Figures 14 to 16 The process shown is to determine the maximum value of the element currently output in that row. At this point, the cell register RGS_i,n of the edge calculation unit U_i,n can be referred to as the dynamic record register during the element-by-element output process (or the element maximum value filtering process).
[0095] That is, in the embodiments of this application, the filtering unit F_i corresponding to any row may include a dynamic recording register and a comparator C_i. The dynamic recording register can be used to cache the maximum value of the corresponding row's elements. The maximum value of the corresponding row's elements can be initially cached as the first shifted output element in the edge computing unit U_i,n of the corresponding row (i.e., the element E calculated by the edge computing unit U_i,n). inThe comparator C_i can be used to shift the output from the first column to the next column of the edge computing unit U_i,n in the corresponding row. It compares the element to be output (i.e., waiting for shift output in the filter shift register Filter_i of the edge computing unit U_i,n) in the current column of the corresponding row with the maximum value of the element in the same row currently cached in the dynamic record register. If the element to be output (i.e., waiting for shift output in the filter shift register Filter_i of the edge computing unit U_i,n) in the current column of the corresponding row is greater than the maximum value of the element in the same row currently cached in the dynamic record register (i.e., the maximum value of the element currently output in the row), then the maximum value of the element in the same row cached in the dynamic record register is updated to the element to be output in the edge computing unit of the corresponding row.
[0096] For example, in an embodiment of this application, as a preferred solution to save register resources, the dynamic record register of the filtering unit F_i corresponding to any row can be time-multiplexed with the result element register in the edge computing unit U_i,n of that row (i.e., the unit register RGSi,j of the edge computing unit U_i,n). However, it is understood that the dynamic record register of the filtering unit F_i corresponding to any row can also be set to be independent of the result element register in the edge computing unit U_i,n of that row (i.e., the unit register RGS_i,j of the edge computing unit U_i,n).
[0097] Figures 18 to 20 The filtering unit in the tensor core component of the processor in this application embodiment is based on, as Figure 17 The diagram shows the internal instance structure of the preferred deployment method. Figures 18 to 20 In this example, taking a result input matrix and pulsation array 150 both as 3×3, and a column of filtering units F_1 to F_m as three (i.e., m and n are both 3 as mentioned above), the process is shown whereby a column of filtering units F_1 to F_m dynamically maintains the maximum value of each row of elements recorded as the maximum value of the elements already output in each row. Figures 18 to 20 The diagram illustrates the calculation units U_i,1 to U_i,3 and the corresponding filtering unit F_i for any given row.
[0098] Please refer to [the website / information] first. Figure 18 In the first cycle of the output instance of the tensor core component 10 in this embodiment, the edge computing unit U_i,3 calculates the element E. i3 The elements are shifted into the filter shift register Filter_i and await output. The comparator C_i of the filter unit F_i obtains the maximum value Max{E} by comparison. i3} = E i3 As a result, the currently cached element E in the dynamic record register (e.g., the cell register RGS_i,j of the edge computing unit U_i,3) is maintained. i3 This causes the maximum value of the corresponding row's element cached in the dynamic record register to be initially cached as the first shifted output element in the edge computing unit U_i,3 of the corresponding row (i.e., the element E calculated by the edge computing unit U_i,3). i3 ).
[0099] Please see Figure 19 In the second cycle of the output instance of the tensor core component 10 in this application embodiment, the element waiting to be output in the shift register Filter_i is the computation unit U_i, and the element E generated by 2 is calculated. i2 Furthermore, the comparator C_i of the filtering unit F_i is based on the comparison element E. i2 And the maximum value E among the elements currently output in this row. i3 The result determines whether to update the currently cached element E in the dynamic record register (e.g., the cell register RGS_i,j of the edge computing unit U_i,3). i3 So that in element E i2 After the second cycle of shift output, the dynamic record register can dynamically maintain the maximum value Max{E} of the currently output elements in that row. i2 E i3}
[0100] Please see Figure 20 In the third cycle of the output instance of the tensor core component 10 in this embodiment, the elements waiting to be output in the shift register Filter_i are the computation unit U_i, and the generated element E is calculated. i1 Furthermore, the comparator C_i of the filtering unit F_i is based on the comparison element E. i1 And the maximum value of the elements currently output in this row, Max{E} i2 E i3 The result determines whether to update the currently cached element (i.e., the maximum value Max{E} among the currently output elements in the line) in the dynamic record register (e.g., the cell register RGS_i,j of the edge computing unit U_i,3). i2 E i3}), so that in element E i1 After the third cycle of shift output, the dynamic record register dynamically maintains the maximum value Max{E} of the currently output elements in that row. i1 E i2 E i3}
[0101] In another embodiment of this application, a method for element filtering of a tensor core component for a processor is provided. The tensor core component may include a systolic array, which may include multiple computing units arranged in an array. The multiple computing units may be used to collaboratively compute the result matrix obtained by multiplying a first input matrix and a second input matrix. The multiple computing units may correspond to each matrix position in the result matrix, and when the result matrix is computed, the elements at each matrix position in the result matrix are stored in the corresponding computing units of the multiple computing units.
[0102] Figure 21 This is an exemplary flowchart illustrating an element filtering method for a tensor core component of a processor in an embodiment of this application. Please refer to... Figure 21 For the aforementioned tensor core components, the element filtering method may include:
[0103] S2110: During the column-by-column output of elements in multiple computational units (i.e., the elements of the result matrix obtained from the systolic array computation of the tensor core component in multiple computational units), the maximum value of each row's elements is selected; and,
[0104] S2130: In response to the completion of the column-by-column output of the elements in multiple computational units (i.e., the elements of the result matrix obtained by the systolic array calculation of the tensor core component in multiple computational units), the maximum values of the elements in each row selected are output in columns.
[0105] Based on embodiments of this application, the element filtering method can filter out the maximum value of each row of elements during the process of outputting the elements of the result matrix column by column from the systolic array; furthermore, the element filtering method can also output the filtered maximum values of each row in a column in response to the completion of the output of the elements of the result matrix column by column from the systolic array. Therefore, the tensor core component can have the element maximum value filtering function implemented along with the output of the result matrix column by column, thereby avoiding the need for the scalar core component of the processor to perform additional filtering operations to filter the maximum values of each row of the result matrix, and thus improving the processing efficiency of the scalar core component in the processor.
[0106] For example, in an embodiment of this application, the process of filtering out the maximum value of the elements in each row in S2110 may include: in the process of outputting the elements in multiple computing units (i.e. the elements of the result matrix calculated by the systolic array of the tensor core component in multiple computing units) column by column, the maximum value of the elements in each row is determined by filtering the elements in each row column by column.
[0107] For example, in an embodiment of this application, elements in multiple computing units can be output column by column by column shifting. In this case, the process of S2110 filtering out the maximum value of the elements in each row may include: during the column output of the elements in multiple computing units (i.e., the elements of the result matrix calculated by the systolic array of the tensor core component in multiple computing units), the elements in each row are filtered column by column according to the rhythm of column shifting to determine the maximum value of the elements in each row.
[0108] For example, in an embodiment of this application, elements in multiple computing units can be shifted column by column to a column of edge computing units located at the end of the shift, and then shifted out column by column from the column of edge computing units. In this case, S2110, by filtering the elements of each row column by column according to the shifting rhythm, and determining the maximum value of the elements in each row, may include: performing a peer comparison for filtering the maximum value of the elements in each row of elements to be shifted out in each column of edge computing units.
[0109] For example, in an embodiment of this application, the peer element comparison performed in S2110 for filtering the maximum value of an element may include: dynamically maintaining the maximum value of the element in each row of records as the maximum value of the element in each row of output elements, based on the comparison results of the peer element comparison column by column.
[0110] For example, in an embodiment of this application, S2110 can specifically maintain the maximum value of each row of elements in the recorded rows as the maximum value of the elements already output in each row in the following manner:
[0111] The elements of each row to be shifted and output in the first column of an edge computing unit (i.e., the elements of each row calculated by the edge computing unit) are recorded as the maximum value of the elements in each row initially selected.
[0112] After shifting the output from the first column, the process begins by shifting to the next column to be output in the edge computing unit. Each element in each row of the current column to be output (i.e., elements calculated by other column computing units and shifted into the current edge computing unit) is compared with the maximum value of the element in the same row of the current column. If any element in any row of the current column to be output is greater than the maximum value of the element in the same row of the current column, then the maximum value of that row in the current column to be output is updated to the maximum value of that row in the current column to be output.
[0113] In another embodiment of this application, an electronic device is also provided, which may include the processor in the foregoing embodiments.
[0114] It is understood that, in the embodiments of this application, the various parts described by example may be related by "and / or". In this document, "and / or" means that the contexts connected by it may be a mutual definition of "and" or a selective definition of "or". Therefore, the various parts having an "and / or" relationship can be understood to include different combinations of situations where "and / or" between each pair of parts respectively represents a mutual definition of "and" or a selective definition of "or", and such combinations of different situations can be considered substantially equivalent to the definition of "at least one of the parts".
[0115] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A tensor core component for a processor, characterized in that, include: The first input buffer is used to buffer the first input matrix; The second input buffer is used to cache the second input matrix; A pulsating array includes multiple computing units arranged in an array and a row of filtering units deployed in a column; in: The plurality of computing units are used to collaboratively calculate the result matrix obtained by multiplying the first input matrix and the second input matrix. The plurality of computing units correspond to each matrix position of the result matrix. When the result matrix is calculated, the elements at each matrix position in the result matrix are stored in the corresponding computing units of the plurality of computing units. The filtering unit corresponds to multiple rows of the result matrix. Elements in the plurality of calculation units are output column by column through the filtering unit. The elements in the plurality of calculation units are output column by column through shifting. Furthermore, the filtering unit is configured as follows: During the column-by-column output of elements in the plurality of computing units, the elements in each row are filtered column by column according to the rhythm of column-by-column shifting to determine the maximum value of the elements in each filtered row; and, In response to the completion of the column-by-column output of elements in the plurality of computing units, the maximum values of the elements in each of the selected rows are output as a column.
2. The tensor core component according to claim 1, characterized in that, The elements in the plurality of computing units are shifted column by column to a column of edge computing units located at the last shift level in the plurality of computing units, and then shifted and output column by column from the column of edge computing units; The column filtering unit is specifically configured to: perform a peer comparison on each row of elements to be shifted and output in each column of the column edge calculation unit to filter the maximum value of the element.
3. The tensor core component according to claim 2, characterized in that, The column filtering unit is specifically configured to: based on the comparison results of the row-by-row element comparison, dynamically maintain the maximum value of each row of elements as the maximum value of the elements already output in each row.
4. The tensor core component according to claim 3, characterized in that, The filter unit is specifically configured as follows: Record the elements of each row in the first column of the edge computing unit to be shifted as the maximum value of the elements in each row initially selected; After shifting the output from the first column, the next column to be output in the edge computing unit is shifted. The elements of each row in the current column to be output in the edge computing unit are compared with the maximum value of the elements in the same row of the record. If the element of any row in the current column to be output is greater than the maximum value of the elements in the same row of the record, then the maximum value of the element in that row of the record is updated to the element of that row in the current column to be output.
5. The tensor core component according to claim 4, characterized in that, Each of the filter units in the column includes: A dynamic recording register is used to cache the maximum value of the corresponding row; wherein, the maximum value of the corresponding row is initially cached as the element of the corresponding row in the first column of the edge computing unit to be shifted and output. A comparator is used to shift the output from the first column to the next column to be output in the edge computing unit, and compare the element of the corresponding row of the current column to be output with the maximum value of the element of the same row currently cached in the dynamic record register; wherein, if the element of the corresponding row of the current column to be output is greater than the maximum value of the element of the same row currently cached in the dynamic record register, then the maximum value of the element of the same row cached in the dynamic record register is updated to the element of the corresponding row of the current column to be output.
6. The tensor core component according to claim 5, characterized in that, Each of the plurality of computing units includes a result element register, and when the result matrix is calculated, the elements at each matrix position in the result matrix are stored in the result element register of the corresponding computing unit among the plurality of computing units; The dynamic record register of each of the filtering units in the column shares the same physical register as the result element register in the edge computing unit in the column.
7. A processor, characterized in that, It includes a tensor core component as described in any one of claims 1 to 6, and a scalar core component; wherein, the elements in the plurality of computation units are output to the scalar core component column by column through the column filtering unit, and the maximum values of the elements in each row filtered by the column filtering unit are output to the scalar core component in a column.
8. A method for element filtering in a tensor core component of a processor, characterized in that, The tensor core component includes a systolic array, which includes multiple computing units arranged in an array and a column of filtering units deployed in a row. The multiple computing units are used to collaboratively calculate the result matrix obtained by multiplying the first input matrix and the second input matrix. The multiple computing units correspond to each matrix position of the result matrix, and when the result matrix is calculated, the elements at each matrix position in the result matrix are stored in the corresponding computing units of the multiple computing units. The filtering unit corresponds to multiple rows of the result matrix. Elements in the plurality of calculation units are output column by column through the filtering unit. The elements in the plurality of calculation units are output column by column through shifting. The element filtering method includes the following steps performed by the filtering unit: During the column-by-column output of elements in the plurality of computing units, the maximum value of the elements in each row is determined by filtering the elements in each row according to the rhythm of column-by-column shifting. as well as, In response to the completion of the column-by-column output of elements in the plurality of computing units, the maximum values of the elements in each of the selected rows are output as a column.
9. The element screening method according to claim 8, characterized in that, The elements in the plurality of computing units are shifted column by column to a column of edge computing units located at the last shift level in the plurality of computing units, and then shifted and output column by column from the column of edge computing units; The step of filtering elements in each row according to the rhythm of column-by-column shifting to determine the maximum value of the elements in each row includes: comparing the elements in each row to be shifted and output in each column of the edge calculation unit with the elements in the same row for filtering the maximum value.
10. The element screening method according to claim 9, characterized in that, The step of performing a peer-to-peer comparison on the elements of each row to be shifted and output in each column of the edge computing unit, used to filter the maximum value of each element, includes: Based on the comparison results of the elements in the same row column by column, the maximum value of each element in the recorded row is dynamically maintained as the maximum value of the elements already output in each row.