Tensor calculation device and method, chip, medium and equipment

By setting multiple multiplication arrays in the tensor computing device and dynamically scheduling the target multiplication array, the design problems brought about by the growth of the multiplication array scale are solved, the universality and layout friendliness of the tensor computing circuit are realized, and the difficulty of physical implementation of the chip is reduced.

CN120353432APending Publication Date: 2025-07-22BEIJING HORIZON INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510414019.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

With the increase in the demand for chip computing power, the scale of multiplication and addition arrays has increased exponentially, resulting in increased difficulty in chip design and physical implementation. Especially in application scenarios such as intelligent driving and smart cockpits, the existing technology is difficult to effectively solve the layout-friendliness and scalability problems of tensor computing.

Method used

By setting up multiple multiplication arrays in the tensor calculation device, the target multiplication array is dynamically scheduled according to the tensor and operation type to be calculated, and the target control method is configured to realize tensor calculations of different sizes and types, thereby improving the universality and layout friendliness of the circuit.

Benefits of technology

It effectively reduces the scale of the multiplication array, improves the layout friendliness and scalability of the tensor computing circuit, reduces the difficulty of physical implementation of the chip, and improves the computing efficiency and flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353432A_ABST
    Figure CN120353432A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a tensor calculation device and method, a chip, a medium and equipment, and the device comprises a memory which is used for storing a to-be-calculated first tensor and a to-be-calculated second tensor; the controller is used for reading the first tensor, the second tensor and the operation type from the memory; the calculation component comprises a plurality of multiply-add arrays; any multiply-add array is used for determining a multiply-add result of at least one pair of input values; any pair of input values comprises a first input value and a second input value; the controller is further configured to determine a target multiply-add array for current operation and a corresponding target control mode from a plurality of multiply-add arrays of the calculation component based on the first tensor, the second tensor and the operation type; and based on the target control mode, the target multiply-add array is controlled to execute operation corresponding to the operation type on the first tensor and the second tensor to obtain a tensor calculation result, so that the universality, layout friendliness and expandability of the tensor calculation circuit can be improved, and the physical implementation difficulty is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to integrated circuit technology, and in particular to a tensor calculation device, method, chip, medium, and equipment. Background Art

[0002] In application scenarios such as intelligent driving and intelligent cockpits, tensor calculations such as convolution operations and matrix multiplications often rely on multiplier-accumulator arrays in the chip to implement calculations. However, with the continuous increase in the computing power requirements of the chip, the scale of the multiplier-accumulator arrays in the chip has increased exponentially, bringing huge challenges to the design and physical implementation of the chip. Summary of the Invention

[0003] Embodiments of the present disclosure provide a tensor calculation device, method, chip, medium, and equipment to improve the layout friendliness and scalability of tensor calculation circuits.

[0004] In a first aspect of the embodiments of the present disclosure, a tensor calculation device is provided, including: a memory configured to store a first tensor and a second tensor to be calculated; a controller configured to read the first tensor and the second tensor from the memory; and determine an operation type between the first tensor and the second tensor; a calculation component including a plurality of multiplier-accumulator arrays; any one of the multiplier-accumulator arrays is used to determine a multiplication-accumulation result of at least one pair of input values; any one pair of the input values includes a first input value and a second input value; the first input value is a value in the first tensor; the second input value is a value in the second tensor; the controller is further configured to: based on the first tensor, the second tensor, and the operation type, determine a target multiplier-accumulator array for the current operation from the plurality of multiplier-accumulator arrays of the calculation component, and determine a target control mode corresponding to the target multiplier-accumulator array; based on the target control mode, control the target multiplier-accumulator array to perform an operation corresponding to the operation type on the first tensor and the second tensor to obtain a tensor calculation result.

[0005] In a second aspect of the embodiments of the present disclosure, a tensor calculation method is provided, including: obtaining a first tensor and a second tensor to be calculated; determining an operation type between the first tensor and the second tensor; based on the first tensor, the second tensor, and the operation type, determining a target multiplier-accumulator array for the current operation from the plurality of multiplier-accumulator arrays of the calculation component, and determining a target control mode corresponding to the target multiplier-accumulator array; wherein, any one of the multiplier-accumulator arrays is used to determine a multiplication-accumulation result of at least one pair of input values; any one pair of the input values includes a first input value and a second input value; the first input value is a value in the first tensor; the second input value is a value in the second tensor; based on the target control mode, controlling the target multiplier-accumulator array to perform an operation corresponding to the operation type on the first tensor and the second tensor to obtain a tensor calculation result.

[0006] In a third aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program for executing the tensor calculation method according to any one of the above embodiments of the present disclosure.

[0007] In a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing executable instructions executable by the processor; the processor for reading the executable instructions from the memory and executing the instructions to implement the tensor calculation method according to any one of the above embodiments of the present disclosure; or, the electronic device includes the tensor calculation device according to any one of the above embodiments of the present disclosure.

[0008] In a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, when the instructions in the computer program product are executed by a processor, the tensor calculation method provided in any one of the above embodiments of the present disclosure is executed.

[0009] In a sixth aspect of the embodiments of the present disclosure, there is provided a chip, including: the tensor calculation device according to any one of the above embodiments of the present disclosure.

[0010] Based on the tensor calculation device, method, chip, medium and device provided in the above embodiments of the present disclosure, by setting a plurality of multiply-accumulate arrays in a computing component, each multiply-accumulate array can be used to determine the multiply-accumulate result of at least a pair of input values, the target multiply-accumulate array for the current operation can be scheduled from the plurality of multiply-accumulate arrays according to the first tensor, the second tensor to be calculated, and the operation type, and the target control mode of the target multiply-accumulate array can be configured. Furthermore, based on the target control mode, the target multiply-accumulate array is controlled to perform an operation corresponding to the operation type on the first tensor and the second tensor to obtain a tensor result. Based on different combinations of different numbers of multiply-accumulate arrays and different control modes in the plurality of multiply-accumulate arrays, the calculation of tensors of different operation types and different sizes can be supported. On the one hand, the versatility of the tensor calculation circuit can be improved, and there is no need to design corresponding calculation circuits for the calculation of tensors of different operation types and different sizes respectively, so that the scale of the multiply-accumulate array can be effectively reduced; on the other hand, based on the multiply-accumulate array, a plurality of multiply-accumulate arrays can be flexibly distributed in the circuit layout, and the number of multiply-accumulate arrays can be flexibly increased according to different application scenarios, thereby improving the layout friendliness and scalability of the tensor calculation circuit, and thus reducing the physical implementation difficulty of the chip. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is an exemplary application scenario of the tensor calculation device provided by the present disclosure;

[0012] Figure 2 is a schematic structural diagram of the tensor calculation device provided by an exemplary embodiment of the present disclosure;

[0013] Figure 3 It is a schematic structural diagram of a computing component provided by an exemplary embodiment of the present disclosure;

[0014] Figure 4 It is a schematic structural diagram of a computing component provided by another exemplary embodiment of the present disclosure;

[0015] Figure 5 It is a schematic structural diagram of a computing component provided by still another exemplary embodiment of the present disclosure;

[0016] Figure 6 It is a schematic structural diagram of a computing component provided by still another exemplary embodiment of the present disclosure;

[0017] Figure 7 It is a schematic structural diagram of a multiply-accumulate array provided by an exemplary embodiment of the present disclosure;

[0018] Figure 8 It is a schematic structural diagram of a computing component provided by yet another exemplary embodiment of the present disclosure;

[0019] Figure 9 It is a schematic structural diagram of a computing component provided by still another exemplary embodiment of the present disclosure;

[0020] Figure 10 It is a schematic structural diagram of a multiply-accumulate unit MAC provided by an exemplary embodiment of the present disclosure;

[0021] Figure 11 It is a schematic diagram of the principle of a convolution operation provided by an exemplary embodiment of the present disclosure;

[0022] Figure 12 It is a schematic flowchart of a tensor calculation method provided by an exemplary embodiment of the present disclosure;

[0023] Figure 13 It is a schematic flowchart of a tensor calculation method provided by another exemplary embodiment of the present disclosure;

[0024] Figure 14 It is a schematic flowchart of a tensor calculation method provided by still another exemplary embodiment of the present disclosure;

[0025] Figure 15 It is a schematic flowchart of a tensor calculation method provided by still another exemplary embodiment of the present disclosure;

[0026] Figure 16 It is a schematic flowchart of a tensor calculation method provided by yet another exemplary embodiment of the present disclosure;

[0027] Figure 17 It is a structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0028] To explain the present disclosure, exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all embodiments. It should be understood that the present disclosure is not limited by the exemplary embodiments.

[0029] It should be noted that: Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present disclosure.

[0030] Overview of the present disclosure

[0031] In the process of implementing the present disclosure, the inventors found that in application scenarios such as intelligent driving and intelligent cockpits, it is usually necessary to use neural network technology to achieve environmental perception, positioning, planning and control, and various functions in the cockpit. Neural network technology usually includes a large number of tensor calculations such as convolutional operations and matrix multiplications. These calculations often rely on the multiplication-accumulation array in the chip to implement. However, with the continuous increase in the computing power requirements of the chip, the scale of the multiplication-accumulation array in the chip has increased exponentially, bringing great challenges to the design and physical implementation of the chip.

[0032] Exemplary overview

[0033] Figure 1 is an exemplary application scenario of the tensor calculation device provided by the present disclosure. As Figure 1 shown, the chip 10 in the present disclosure is, for example but not limited to, an intelligent driving chip or an intelligent cockpit chip, etc. The chip 10 may include one or more processors 11 and a tensor calculation device 12. The processor 11 may be used to run application programs for functions such as intelligent driving and intelligent cockpits. During the process of running these application programs, for tasks that need to be accelerated by the tensor calculation device, the processor 11 may send a calculation task to the tensor calculation device 12. The calculation task may include task information such as the first tensor information to be calculated, the second tensor information, and the operation type. The first tensor information may include, for example, the size and storage address of the first tensor, and the second tensor information may include, for example, the size and storage address of the second tensor. The tensor calculation device 12 of the embodiments of the present disclosure may obtain the calculation task, and based on the calculation task, load the first tensor and the second tensor into the internal memory, or the processor 11 may pre-write the first tensor and the second tensor into the memory in advance. Furthermore, the tensor calculation device 12 may, based on the first tensor, the second tensor, and the operation type, control one or more multiplication-accumulation arrays in the tensor calculation device 12 by scheduling to complete the calculation of the first tensor and the second tensor, obtain the tensor calculation result, and store the tensor calculation result in a specified storage space.

[0034] The above-mentioned processor 11 includes, for example but not limited to, one or more of a Central Processing Unit (CPU), a Neural Network Processing Unit (NPU), a Graphic Processing Unit (GPU), etc. In practical applications, the interaction between the processor 11 and the tensor calculation device 12 is not limited to the above implementation manners. This is only an exemplary description here, as long as the tensor calculation device 12 can complete corresponding calculations according to the calculation requirements corresponding to the calculation tasks.

[0035] The tensor calculation device 12 in the embodiments of the present disclosure is not limited to being applied to intelligent driving and intelligent cockpit scenarios, and can also be applied to any other scenarios involving neural network calculations or multiply-accumulate calculations. The specific application scenarios are not limited.

[0036] Exemplary device

[0037] Figure 2 It is a schematic structural diagram of a tensor calculation device provided by an exemplary embodiment of the present disclosure. The tensor calculation device provided by the embodiments of the present disclosure can be applied to an electronic device. For example, it can be applied to an in-vehicle computing platform (or in-vehicle terminal), or can be applied to a system on chip (i.e., SOC), or can be applied to a neural network processor, such as Figure 2 As shown, the tensor calculation device 20 provided by the embodiments of the present disclosure may include: a memory 21, a controller 22, and a calculation component 23.

[0038] The memory (which can be referred to as the first memory) 21 is configured to store a first tensor and a second tensor to be calculated.

[0039] The controller 22 is configured to read the first tensor and the second tensor from the memory 21; and determine the operation type of the first tensor and the second tensor.

[0040] The calculation component 23 includes a plurality of multiply-accumulate arrays 23i (i = 1, 2,..., M), where M is the number of multiply-accumulate arrays and M is a positive integer; any one multiply-accumulate array 23i is used to determine the multiply-accumulate result of at least one pair of input values; any pair of input values includes a first input value and a second input value; the first input value is a value in the first tensor; the second input value is a value in the second tensor.

[0041] The controller 22 is further configured to: determine a target multiply-accumulate array for the current operation from a plurality of multiply-accumulate arrays 23i of the computing component 23 based on the first tensor, the second tensor, and the operation type, and determine a target control mode corresponding to the target multiply-accumulate array; and control the target multiply-accumulate array to perform an operation corresponding to the operation type on the first tensor and the second tensor based on the target control mode to obtain a tensor calculation result.

[0042] In some alternative embodiments, the operation type may include operations involving the multiply-accumulate (MAC) type, such as convolution operations, matrix multiplication operations, vector dot product operations, etc. The first tensor and the second tensor may be tensors that need to perform a multiply-accumulate operation in an operation corresponding to any operation type. For example, the first tensor and the second tensor may be a feature tensor and a convolution kernel tensor of a convolution operation respectively. For another example, the first tensor and the second tensor may be a first matrix (or left matrix) and a second matrix (or right matrix) of a matrix multiplication operation respectively. For yet another example, the first tensor and the second tensor may be a first vector and a second vector of a vector dot product operation respectively, and so on.

[0043] In some alternative embodiments, the first tensor and the second tensor may be written into the memory 21 by the processor that generates the computing task, or loaded from an external memory to the memory 21 through the controller 22.

[0044] In some alternative embodiments, the controller 22 may determine the operation type corresponding to the first tensor and the second tensor from the computing task from the processor.

[0045] In some alternative embodiments, the multiply-accumulate array 23i is a basic operation unit of the embodiments of the present disclosure. Each multiply-accumulate array 23i can be used alone for a computing task in one operation cycle, or multiple multiply-accumulate arrays can be combined to be used for a computing task in one operation cycle, thereby facilitating adaptation to the requirements of different computing tasks.

[0046] In some alternative embodiments, each multiply-accumulate array 23i can complete the multiply-accumulate operation on at least one pair of input values in one operation cycle to obtain a multiply-accumulate result. For example, the number of pairs of input values for the multiply-accumulate operation completed in one operation cycle can be 1 pair, 2 pairs, 3 pairs, 4 pairs, and so on. A pair of input values can include a first input value from a first tensor and a second input value from a second tensor. For example, for a vector dot product operation, the first vector (i.e., the first tensor) is [a1, a2, a3, a4], the second vector (i.e., the second tensor) is [b1, b2, b3, b4], and if the multiply-accumulate array 23i can calculate the multiply-accumulate results of four pairs of input values in one operation cycle, then the operation a1*b1 + a2*b2 + a3*b3 + a4*b4 can be completed through one multiply-accumulate array to obtain the calculation result of the first vector and the second vector. Based on this, various operation types corresponding to tensors of various sizes can be implemented through parallel operations of multiple multiply-accumulate arrays, or through one or more multiply-accumulate arrays, combined with one or more operation cycles, various operation types corresponding to tensors of larger sizes can be completed, such as being able to complete vector dot product operations, matrix multiplication operations, convolution operations, etc. of larger sizes.

[0047] In some alternative embodiments, any multiply-accumulate array can complete the multiply-accumulate operations corresponding to each set of input data among multiple sets of input data within one operation cycle, obtaining the multiply-accumulate results corresponding to each set of input data respectively. Each set of input data includes multiple pairs of input values. For example, the multiply-accumulate array can complete J sets of multiply-accumulate operations, and each set of multiply-accumulate operations is, for example, a1*b1 + a2*b2 + a3*b3 + a4*b4 + …. Based on this, more calculations can be completed through one multiply-accumulate array. For example, J sets of vector dot product operations can be completed in parallel within one operation cycle. For another example, matrix multiplication operations or convolution operations can be completed through a smaller number of operation cycles. The multiple pairs of input values in each set of input data can include multiple first input values from the first tensor and multiple second input values from the second tensor. That is to say, multiple sets of data can be determined from the data of the first tensor and the second tensor. Each set of data includes multiple first input values from the first tensor and multiple second input values from the second tensor. Each set of data can be used as a set of input data for the multiply-accumulate array, and multiple sets of data can be used as multiple sets of input data for the multiply-accumulate array. For example, if the first tensor is the first matrix and the second tensor is the second matrix, each row of data or a part of each row of data in the first matrix, and each column of data or a part of each column of data in the second matrix can be used as a set of input data. Based on this, according to the number of sets of input data for which the multiply-accumulate array can complete the multiply-accumulate operations within one operation cycle, multiple sets of input data are determined from the first matrix and the second matrix, and the multiply-accumulate results corresponding to each set of input data are obtained through parallel operations of the multiply-accumulate array within one operation cycle. The multiply-accumulate result corresponding to each set of input data is the multiply-accumulate result of multiple pairs of input values in this set of input data. Optionally, there may be shared input data among the input data of each group. For another example, for matrix multiplication operations, any row of the first matrix (i.e., the first tensor) needs to perform multiply-accumulate operations with each column of the second matrix (i.e., the second tensor) respectively. Then, one row of data or a part of one row of data in the first matrix can be used as the shared input data corresponding to multiple sets of operations of the multiply-accumulate array, and each column of data or a part of each column of data in the second matrix can be used as the other input data for each set of operations in multiple sets of operations of the multiply-accumulate array respectively, so as to complete the multiply-accumulate operations of one row of data in the first matrix with different columns of data in the second matrix in parallel through the multiply-accumulate array. Exemplarily, the first row of data in the first matrix is [a 11 ,a 12 ,a 13 ,a 14 , and the j-th column of data in the second matrix is represented as [b 1j ,b 2j ,b 3j ,b 4j, where \(j = 1, 2, 3, 4\). The multiply-accumulate array can complete the operations on four sets of input data within one operation cycle. Each set of input data includes four pairs of input values. Then, the first row data of the first matrix can be used as the first input value in four pairs of input values in four sets of input data at the same time, and the \(j\)-th column of the second matrix is used as the second input value in four pairs of input values in the \(j\)-th set of input data. That is, the multiply-accumulate array can perform multiply-accumulate operations on four sets of input data respectively, including: a 11 *b 11 +a 12 *b 21 +a 13 *b 31 +a 14 *b 41 、a 11 *b 12 +a 12 *b 22 +a 13 *b 32 +a 14 *b 42 、a 11 *b 13 +a 12 *b 23 +a 13 *b 33 +a 14 *b 43 、a 11 *b 14 +a 12 *b 24 +a 13 *b 34 +a 14 *b 44 。

[0048] In some alternative embodiments, the controller 22 may be implemented using any device and / or logic circuit having corresponding control functions or configurable control functions. For example, the controller 22 may be implemented using a microcontroller (Microcontroller Unit, abbreviated as: MCU) or other implementable devices and / or logic circuits.

[0049] In some alternative embodiments, the target multiply-accumulate array is a multiply-accumulate array determined from multiple multiply-accumulate arrays for implementing the operation of the first tensor and the second tensor. Optionally, the number of target multiply-accumulate arrays may be one or more.

[0050] In some alternative embodiments, the controller 22 may determine a target multiply-accumulate array for the current operation from multiple multiply-accumulate arrays of the computing component 23 based on the first tensor, the second tensor, the operation type, and a pre-configured determination rule. The determination rule may be set based on the tensor size, the operation type, and the operation types supported by the multiply-accumulate arrays. The determination rule may include a rule for determining the size relationship between the input data size supported by the multiply-accumulate array and the tensor size, and / or a rule for determining the current idle state of the multiply-accumulate array, and so on. For example, in the case where the sizes of the first tensor and the second tensor are large, as many target multiply-accumulate arrays as possible that can participate in the current calculation may be selected to be able to complete the tensor calculation of the first tensor and the second tensor within a smaller number of operation cycles. Or, in the case where there are fewer available multiply-accumulate arrays, the tensor calculation of the first tensor and the second tensor may be completed through a larger number of operation cycles.

[0051] For example, for a vector dot product operation, the first tensor and the second tensor are vector A and vector B respectively, the lengths of vector A and vector B are both L1, and a multiply-accumulate array can complete the multiply-accumulate operation of L2 pairs of input values at a time. If L1 is greater than L2, then the multiply-accumulate operation of some elements of vector A and vector B can be completed through the ceiling of L1 / L2 number of multiply-accumulate arrays. Combining with an adder, the calculation of vector A and vector B can be completed in the fewest operation cycles. Here, the length of a vector refers to the number of elements in the vector. The length of a vector can also be referred to as the size, shape, etc. of the vector. For example, the vector [a1, a2, a3, a4] includes 4 elements, and the length of this vector is 4. In a vector dot product operation, each element in the vector serves as an input value of the multiply-accumulate array. For example, the elements of vector A serve as the first input value, and the elements of vector B serve as the second input value.

[0052] For another example, for matrix multiplication operations, multiple multiply-accumulate arrays can be used to perform the multiply-accumulate operations of different rows of the first matrix and different columns of the second matrix in parallel. Specifically, in one calculation method, the first matrix includes 6 rows and 8 columns, that is, each row has 8 first input values, and each multiply-accumulate array supports the multiply-accumulate operation of four pairs of input values in one operation cycle. The first input values in each row of the first matrix can be calculated in parallel with the second matrix through two multiply-accumulate arrays. That is, in one operation cycle, through one multiply-accumulate array, calculate the multiply-accumulate results of the first 4 first input values in this row of the first matrix and the corresponding second input values in each column of the second matrix, and through another multiply-accumulate array, calculate the multiply-accumulate results of the last 4 first input values in this row (that is, the other input values in this row except the first 4 first input values) and the corresponding second input values in each column of the second matrix. Based on this, the first matrix including 6 rows and 8 columns and the second matrix can be calculated in parallel through 12 multiply-accumulate arrays to obtain the tensor calculation result of the first matrix and the second matrix. In another calculation method, the first input values in each row of the first matrix can be calculated by one multiply-accumulate array in two operation cycles. That is, in one operation cycle, calculate the multiply-accumulate results of the first 4 first input values in this row of the first matrix and the corresponding second input values in each column of the second matrix, and in another operation cycle, calculate the multiply-accumulate results of the last 4 first input values in this row (that is, the other input values in this row except the first 4 first input values) and the corresponding second input values in each column of the second matrix through the multiply-accumulate array.

[0053] For another example, for convolution operations, the first tensor is represented as a feature tensor, and the second tensor is represented as a convolution kernel. Multiple multiply-accumulate arrays can be used to calculate the multiply-accumulate operations of the feature tensor and different convolution kernels respectively, or calculate the multiply-accumulate operations of the feature data at different HW positions in the feature tensor and the different weight data in the convolution kernel respectively. Among them, H in HW represents height and W represents width. For example, if the feature tensor is a tensor of size H1*W1*C1, then H1 represents the height of the feature tensor, W1 represents the width of the feature tensor, and C1 represents the depth of the feature tensor. The HW position in the feature tensor refers to the position where H = i and W = j in the feature tensor (or called the HW = ij position, HiWj position), i = 1, 2,..., H1, j = 1, 2,..., W1. The feature at each HW position is a vector of C1 dimensions, that is, it includes C1 eigenvalue along the depth direction at this HW position. Different HW positions in the feature tensor refer to positions where at least one of i and j is different. For example, the feature at the HW = 11 position and the feature at the HW = 12 position in the feature tensor are two vectors of C1 dimensions in the feature tensor.

[0054] By decomposing the above convolution operations, matrix multiplication operations, vector dot product operations, etc. into multiplication-addition operations that can be performed by a multiplication-addition array, different-sized tensors can then perform different types of operations through a multiplication-addition array on a smaller scale, which can effectively improve the versatility of the multiplication-addition array, avoid designing corresponding tensor calculation circuits for different-sized tensors and different types of calculation tasks respectively, and can effectively reduce the scale of the multiplication-addition array, thereby reducing the physical implementation difficulty of related chips.

[0055] In some optional embodiments, the target control mode corresponding to the target multiplication-addition array refers to the control logic for controlling the target multiplication-addition array to complete the operation between the first tensor and the second tensor. The target control mode may include the timing of one or more control operations. Control operations include, for example, but are not limited to, operations for controlling the input of data to each target multiplication-addition array, operations for controlling each target multiplication-addition array to perform calculations, operations for controlling the output of calculation results, and so on. Among them, the operation for controlling the input of data to each target multiplication-addition array may include, for example: the number of input data input to the target multiplication-addition array each time, the order in which at least part of the data of the first tensor and / or the second tensor are used as input data, and so on. The operation for controlling each target multiplication-addition array to perform calculations may include, for example: providing a clock signal and / or other signals for driving each target multiplication-addition array to perform calculations. The operation for controlling the output of calculation results may include, for example: the operation of reading the calculation result from the register storing the calculation result and storing the calculation result in a specified storage space; or, controlling the register storing the calculation result to shift and output the calculation result, and inputting the calculation result to a specified subsequent circuit, and so on.

[0056] In some optional embodiments, after determining the target control mode corresponding to the target multiplication-addition array, based on the target control mode, the target multiplication-addition array can be controlled to perform an operation corresponding to the operation type on the first tensor and the second tensor to obtain a tensor calculation result. For example, corresponding operations are performed according to the timing of various control operations of the target control mode. By inputting data to the target multiplication-addition array, controlling the target multiplication-addition array to calculate, determining the tensor calculation result based on the multiplication-addition result of the target multiplication-addition array, and storing the tensor calculation result in a specified storage space, etc., the tensor calculation between the first tensor and the second tensor is completed accordingly.

[0057] The tensor calculation device provided in this embodiment determines, by designing a plurality of multiply-accumulate arrays in a calculation component, that each multiply-accumulate array can be used to determine the multiply-accumulate result of at least a pair of input values. According to the first tensor and the second tensor to be calculated, and the corresponding operation type, a target multiply-accumulate array for the current operation is scheduled from the plurality of multiply-accumulate arrays, and the target control mode corresponding to the target multiply-accumulate array is configured. Furthermore, based on the target control mode, the target multiply-accumulate array is controlled to perform an operation corresponding to the operation type on the first tensor and the second tensor to obtain a tensor calculation result. Based on different combinations and different control modes corresponding to different numbers of multiply-accumulate arrays in the plurality of multiply-accumulate arrays, calculations of different operation types and tensors of different sizes can be supported. On the one hand, the tensor calculation device provided in this disclosure can improve the generality of the tensor calculation circuit, avoid designing corresponding calculation circuits for tensor calculations of different operation types and different sizes respectively, and thus can effectively reduce the scale of the multiply-accumulate array. On the other hand, based on the multiply-accumulate array, a plurality of multiply-accumulate arrays can be flexibly distributed in the circuit layout, and the number of multiply-accumulate arrays can be flexibly increased according to different application scenarios, so that the layout friendliness and scalability of the tensor calculation circuit can be effectively improved, and further the physical implementation difficulty of the chip can be reduced. That is, for a tensor calculation device including a large-scale multiply-accumulate array, since a plurality of multiply-accumulate arrays can be flexibly laid out in an integrated circuit, the physical implementation difficulty of the integrated circuit can be effectively reduced.

[0058] In some alternative embodiments, on the basis of the embodiment shown above Figure 2 the controller 22 is configured to:

[0059] Based on the first size information of the first tensor, determine the quantity information of the target multiply-accumulate array of the tensor calculation device corresponding to the current operation; based on this quantity information, determine the target multiply-accumulate array from the plurality of multiply-accumulate arrays 23i; based on the first tensor, the second tensor, and the operation type, determine the target control mode corresponding to the target multiply-accumulate array.

[0060] Wherein, the first size information of the first tensor refers to the sizes of all dimensions of the first tensor. For example, if the first tensor is a one-dimensional vector, the first size information is the vector length; if the first tensor is a two-dimensional matrix, the first size information includes the height and width of the matrix; if the first tensor is a three-dimensional tensor, the first size information is the height, width, and depth of the three-dimensional tensor. The quantity information of the target multiply-accumulate array refers to the quantity of the determined multiply-accumulate arrays to participate in the current operation. It can be understood that the quantity of the target multiply-accumulate array is less than or equal to the total quantity of the multiply-accumulate arrays.

[0061] In some alternative embodiments, the number information of the target multiply-accumulate array corresponding to the current operation can be determined based on the first size information of the first tensor and the input data information of the operations supported by each multiply-accumulate array. The input data information of the operations supported by the multiply-accumulate array includes the logarithm of the input values supported by the multiply-accumulate array for the operations. For example, for the case where the first tensor is the first matrix, according to the width (i.e., the number of columns) of the first matrix, combined with the input data information of the multiply-accumulate array, the ratio of the input data that a multiply-accumulate array can complete within one operation cycle to one row of data of the first matrix can be determined. Based on this ratio and the height (i.e., the number of rows) of the first matrix, the number of target multiply-accumulate arrays is determined. For example, for a first matrix of 6 rows and 8 columns, that is, each row of data in the matrix has 8 values (or elements, element values, etc.), if the logarithm of the input values supported by the multiply-accumulate array is 4, then the number of multiply-accumulate arrays required is 8 / 4 * 6 = 12. In this case, for the operation type where each row of data in the first matrix needs to perform multiply-accumulation with each column of data in the second matrix, it can be achieved through multiple operation cycles. For example, the first input value is a set of elements in a certain row of the first matrix. In multiple operation cycles, the second input values corresponding to the first input value in each column of the second matrix are respectively input into the multiply-accumulate array to implement the operation of one row of the first matrix and multiple columns of the second matrix. For example, if the second matrix is a matrix of 8 rows and F (F is a positive integer) columns, and each row of the 6 rows of data in the first matrix needs to perform multiply-accumulation operations with the F columns of data in the second matrix respectively, then in F operation cycles, the second input values corresponding to the first input value in each column of the second matrix are respectively input into the multiply-accumulate array to obtain the multiply-accumulate results corresponding to each multiply-accumulate array in F operation cycles. Each operation cycle corresponds to a column of the second matrix, and then the tensor calculation result of the first matrix and the second matrix is obtained based on the multiply-accumulate results of the F operation cycles. It should be noted that this is only an exemplary implementation manner for determining the number information of the target multiply-accumulate array. In practical applications, the method for determining the target multiply-accumulate array is not limited to this example. Optionally, the number information of the target multiply-accumulate array can be determined based on the size information of the second tensor and the input data information of the multiply-accumulate array. For example, the number information of the target multiply-accumulate array is determined according to the ratio of the height of the second matrix to the logarithm of the input values of the multiply-accumulate array and the width of the second matrix. For another example, the number information of the target multiply-accumulate array can be comprehensively determined by combining the first size information of the first tensor, the size information of the second tensor (which can be called the second size information), and the input data information of the multiply-accumulate array. For another example, based on the first size information of the first tensor and the second size information of the second tensor, the ratio of the amount of operations completed by a multiply-accumulate array within one operation cycle to the overall amount of operations when performing the operations of the first tensor and the second tensor is determined, and then the number information of the target multiply-accumulate array is determined according to this ratio. The specific method for determining the number information of the target multiply-accumulate array is not limited.

[0062] In some alternative embodiments, the input data information may include the number of groups of input data supported by the multiply-accumulate array for operations, the logarithm of the number of input values included in each group, etc. For the case where the first tensor is a one-dimensional vector, based on the vector length of the one-dimensional vector, the number of groups of input data supported by the multiply-accumulate array, and the logarithm of the number of input values included in each group, the amount of operations that a multiply-accumulate array can complete in one operation cycle and the proportion of the overall amount of operations for performing the vector dot product operation on the first tensor and the second tensor can be determined. Based on this proportion, the quantity information of the target multiply-accumulate array is determined; for example, if the vector length of the one-dimensional vector is 100, the number of groups of input data supported by the multiply-accumulate array is 4, and each group of input data includes 5 pairs of input values, then the number of target multiply-accumulate arrays is 100 / (4*5) = 5.

[0063] For the case where the first tensor is a two-dimensional matrix, in one implementation, based on the height and width of the first matrix, combined with the number of input data groups supported by the multiply-accumulate array and the logarithm of the number of input values included in each group, the number of rows that can be parallelly computed by the multiply-accumulate array in one operation cycle and the completion ratio of each row can be determined. That is, the number of input data groups supported by the multiply-accumulate array can be used as the number of rows that can be parallelly computed by the multiply-accumulate array, and the logarithm of the number of input values in each group can be used as the completion ratio of each row of the multiply-accumulate array. For example, if the first matrix is a 6-row and 8-column matrix, the number of input data groups supported by the multiply-accumulate array is 3, and each group includes 4 pairs of input values, it means that the number of rows that can be parallelly computed by each multiply-accumulate array in one operation cycle is 3, and the completion ratio of each row is 4 / 8 (i.e., 1 / 2). That is to say, one multiply-accumulate array can complete the multiply-accumulate operation of 4 input values (i.e., the first input values) of each row of the 3 rows of data in the first matrix and 4 second input values of one column of data in the second matrix in one operation cycle. Then the number of multiply-accumulate arrays required can be (6 / 3)*(8 / 4) = 4. 6 / 3 represents the ratio of the number of rows of the first matrix to the number of rows that can be parallelly computed by each multiply-accumulate array. That is to say, each multiply-accumulate matrix can parallelly compute 3 rows in one operation cycle, and the first matrix has 6 rows, so 6 / 3 multiply-accumulate arrays are needed to support the parallel computation of 6 rows. 8 / 4 represents the ratio of the number of columns of the first matrix (i.e., the number of first input values included in each row of data) to the number of first input values supported by each row of the multiply-accumulate array in one operation cycle. That is to say, each multiply-accumulate array can only complete the multiply-accumulate operation of 4 first input values of each row of the 3 rows of data in the first matrix and 4 corresponding second input values in the second matrix in one operation cycle. 8 / 4 (i.e., 2) multiply-accumulate arrays are needed to parallelly complete the multiply-accumulate operation of the 3 rows of data in the first matrix and each column of data in the second matrix. Specifically, through one multiply-accumulate array, the multiply-accumulate result of the first 4 first input values of each row of the 3 rows of data in the first matrix and the corresponding second input values in each column of the second matrix is calculated, and through another multiply-accumulate array, the multiply-accumulate result of the last 4 first input values of each row of these 3 rows of data (i.e., other input values in the row except the first 4 first input values) and the corresponding second input values in each column of the second matrix is completed. In another implementation, based on the height of the first matrix (i.e., the number of rows of the first matrix) and the number of input data groups supported by the multiply-accumulate array (i.e., the number of rows that can be parallelly computed), the number of multiply-accumulate arrays required can be determined. Further, for the case where the completion ratio of each row of the first matrix by the multiply-accumulate array in one operation cycle is less than 1, the complete calculation of each row of the first matrix can be achieved by increasing the number of operation cycles. For example, if the first matrix is a 6-row and 8-column matrix, the number of input data groups supported by the multiply-accumulate array is 3, and each group includes 4 pairs of input values, based on the number of rows of the first matrix and the number of input data groups supported by the multiply-accumulate array, the number of multiply-accumulate arrays required is 6 / 3 = 2.Among them, the implementation method for each row of the first matrix to complete the calculation by increasing the operation cycle can be as follows: First, use the determined 2 target multiply-accumulate arrays to calculate the multiply-accumulate results of the first 4 first input values of each row of the first matrix and the second input values corresponding to the first input values in each column of the second matrix. Then, in the increased operation cycle, use these 2 target multiply-accumulate arrays to calculate the multiply-accumulate results of the last 4 first input values of each row of the first matrix (i.e., the other input values in the row data except the first 4 first input values) and the second input values corresponding to the first input values in each column of the second matrix. Then, based on post-processing units such as adders, add the multiply-accumulate results corresponding to the first 4 first input values of each row of the first matrix and the multiply-accumulate results corresponding to the last 4 first input values to obtain the multiply-accumulate results corresponding to each row of the first matrix. It should be noted that this is only a partial exemplary implementation method for determining the quantity information of the target multiply-accumulate arrays. In practical applications, the specific method for determining the target multiply-accumulate arrays is not limited to the above method. For example, the quantity information of the target multiply-accumulate arrays can be comprehensively determined by combining the first dimension information of the first tensor, the second dimension information of the second tensor, and the input data information of the multiply-accumulate arrays. Another example is that the quantity information of the target multiply-accumulate arrays can be determined based on the second dimension information of the second tensor and the input data information of the multiply-accumulate arrays. Similar to the operations based on vectors and matrices, for the operations based on three-dimensional tensors, the quantity information of the target multiply-accumulate arrays can be determined according to the dimension information of the three-dimensional tensor and the input data information supported by the multiply-accumulate arrays, which will not be elaborated one by one here.

[0064] In some alternative embodiments, the quantity information of the target multiply-accumulate arrays can be comprehensively determined based on the first dimension information of the first tensor and in combination with the idle state of the multiply-accumulate arrays. That is, the quantity of the target multiply-accumulate arrays is restricted by the quantity of the multiply-accumulate arrays in the idle state.

[0065] In some optional embodiments, after determining the target multiply-accumulate array, based on the first tensor, the second tensor, and the operation type, in combination with the target multiply-accumulate array, the operations on the first tensor and the second tensor can be decomposed into one or more operation cycles to perform the operations based on the multiply-accumulate array. According to the result of this decomposition, the control operations and timing of the entire operation are determined to obtain the target control mode of the target multiply-accumulate array. The control operations include, for example but not limited to, the operation of inputting data to the target multiply-accumulate array, the operation of controlling the target multiply-accumulate array to perform calculations, the operation of reading the multiply-accumulate results of the target multiply-accumulate array. Optionally, the control operations may further include the operation of fusing the multiply-accumulate results of each multiply-accumulate array in each operation cycle to obtain the final tensor calculation result. For example, if the total number of pairs of input values that need to perform multiply-accumulate operations in the first tensor and the second tensor exceeds the number of pairs of input values that a multiply-accumulate array can support in one operation cycle, the multiply-accumulate operations corresponding to this group of input values can be decomposed into multiple operation cycles and / or performed by multiple multiply-accumulate arrays. It is necessary to add the multiply-accumulate results obtained after the decomposition through calculation to obtain the multiply-accumulate results corresponding to this group of input values in the first tensor and the second tensor. For example, in an operation scenario performed by a tensor calculation device provided based on the present disclosure, the first tensor is the first matrix, the second tensor is the second matrix, and it is necessary to calculate the multiply-accumulate results of each row of the first matrix and each column of the second matrix. If the number of elements in each row of the first matrix is 8, that is, it is necessary to complete the calculation of the multiply-accumulate results of 8 first input values and 8 second input values in each column of the second matrix. The number of pairs of input values that a multiply-accumulate array can support in one operation cycle is 4. The multiply-accumulate operations of each row of the first matrix and each column of the second matrix are decomposed into multiple operation cycles and / or performed by multiple multiply-accumulate arrays to obtain the multiply-accumulate results of the first 4 first input values in each row of the first matrix and the first 4 second input values in each column of the second matrix, and the multiply-accumulate results of the last 4 first input values in each row of the first matrix and the last 4 second input values in each column of the second matrix. Taking the first matrix as a 6-row and 8-column matrix, the second matrix as an 8-row and F (F is a positive integer)-column matrix, and an operation type where a multiply-accumulate array can support the multiply-accumulate operation of four pairs of input values in one operation cycle as an example, if the number of determined target multiply-accumulate arrays is 1, the number of operation cycles required for the target multiply-accumulate array to perform the multiply-accumulate operation for the multiply-accumulate operation of the first matrix and the second matrix is 6*F*(8 / 4). That is, one row of data in the first matrix and one column of data in the second matrix require 8 / 4 operation cycles. Each row of the 6 rows of data in the first matrix needs to perform the multiply-accumulate operation with the F columns of data in the second matrix respectively. Therefore, the number of operation cycles required is 6*F times 8 / 4 operation cycles. If the number of determined target multiply-accumulate arrays is 2 (for example, 8 / 4), the number of operation cycles required for the two target multiply-accumulate arrays to perform the multiply-accumulate operation in parallel for the multiply-accumulate operation of the first matrix and the second matrix is 6*F.If it is determined that the number of target multiply-accumulate arrays is 12 (e.g., 6*8 / 4), the number of operation cycles required for the multiply-accumulate operation of the first matrix and the second matrix to be performed in parallel by these 12 target multiply-accumulate arrays is F. This is only some exemplary descriptions of the decomposition method, and in practical applications, it is not limited to the decomposition method of the above example. Subsequently, the multiply-accumulate results of the first 4 first input values and the second input value after decomposition can be added to the multiply-accumulate results of the last 4 first input values and the second input value to obtain the multiply-accumulate results of 8 first input values in each row of the first matrix and 8 second input values in each column of the second matrix. Optionally, the control operation may further include an operation of writing the tensor calculation result into a specified storage space. For example, when the processor issues a calculation task, the calculation task includes the storage address information of the tensor calculation result. After obtaining the tensor calculation result, the tensor calculation device 20 writes the tensor calculation result into the storage space corresponding to the storage address information, so that the processor can read the tensor calculation result and then continue the subsequent processing.

[0066] In the embodiments of the present disclosure, by combining the size information of the first tensor and the input data information of the operations supported by the multiply-accumulate arrays, the number information of the target multiply-accumulate arrays corresponding to the current operation can be determined. Furthermore, the collaborative operation of the first tensor and the second tensor can be completed by the multiply-accumulate arrays of this number, improving the efficiency of tensor calculation. Furthermore, based on the specific situations of the first tensor, the second tensor, and the operation type, the target control method for the target multiply-accumulate arrays is determined, so as to accurately control the target multiply-accumulate arrays to complete the corresponding calculations or operations and ensure the accuracy and effectiveness of the tensor calculation results.

[0067] Figure 3 It is a schematic structural diagram of a computing component provided by an exemplary embodiment of the present disclosure.

[0068] In some alternative embodiments, such as Figure 3 shown, the multiple multiply-accumulate arrays include multiply-accumulate arrays with m rows and n columns.

[0069] The first size information of the first tensor includes the first height, the first width, and the first depth of the first tensor.

[0070] The controller 22 is specifically configured to:

[0071] In response to the first dimension information of the first tensor satisfying the preset dimension condition, based on the first height, the first width, and the first depth, determine the target number of rows and the target number of columns of the target multiply-accumulate array corresponding to the current operation, and use the target number of rows and the target number of columns as the quantity information; or, in response to the first height in the first dimension information not satisfying the preset dimension condition, determine the number of rows m of the multiply-accumulate array as the target number of rows of the target multiply-accumulate array, determine the first width as the target number of columns of the target multiply-accumulate array, and determine the target number of rows and the target number of columns as the quantity information; or, in response to the first width in the first dimension information not satisfying the preset dimension condition, determine the number of columns n of the multiply-accumulate array as the target number of columns of the target multiply-accumulate array, determine the first height as the target number of rows of the target multiply-accumulate array, and determine the target number of rows and the target number of columns as the quantity information. Herein, the quantity information refers to the quantity information corresponding to the target multiply-accumulate array participating in the current operation, and the quantity information may include the target number of rows and the target number of columns.

[0072] Herein, both m and n are positive integers, and m is greater than 1 and / or n is greater than 1. The first height, the first width, and the first depth are respectively the sizes of the three dimensions of the height, width, and depth of the first tensor. In the case where the first tensor is a one-dimensional vector, its first dimension information may be expressed as 1*1*C, that is, both the first height and the first width are 1, and the first depth is C. In the case where the first tensor is a two-dimensional matrix, its first dimension information may be expressed as 1*h*w, where h is the height of the two-dimensional matrix and w is the width of the two-dimensional matrix. Optionally, in order to adapt to a specified data structure, the two-dimensional matrix may be represented as a three-dimensional data structure of 1*h*w, then the first height of the first dimension information is 1, the first width is h, and the first depth is w. In the case where the first tensor is a three-dimensional tensor, its first dimension information may be expressed as H*W*C. That is, the first height is H, the first width is W, and the first depth is C. The preset dimension condition is a condition set based on the constraints of m and n. For example, the preset dimension condition may be that the first height of the first tensor is less than the number of rows m of the multiply-accumulate array, and the first width of the first tensor is less than the number of columns n of the multiply-accumulate array.

[0073] In some alternative embodiments, when the first dimension information satisfies the preset dimension condition, it means that it may not be necessary to utilize all of the multiply-accumulate arrays, that is, based on the first height, the first width, and the first depth of the first tensor, determine the region of the multiply-accumulate array corresponding to some rows and / or some columns from the multiply-accumulate array of m rows and n columns as the region to participate in the calculation. Therefore, the target number of rows and the target number of columns of the target multiply-accumulate array can be determined based on the first height, the first width, and the first depth.

[0074] In some alternative embodiments, for tensors of different dimensions, different methods may be adopted to determine the target number of rows and the target number of columns. For example, for the case where the first tensor is a one-dimensional vector or a tensor of size 1*1*C, the first height 1 may be determined as the target number of rows, and the target number of columns may be determined according to the first width, the first depth, and the input data information supported by each multiply-accumulate array. For example, if the input data information supported by the multiply-accumulate array includes 5 pairs of input values, the method for determining the target number of columns may be 1*C / 5. If the calculation result of 1*C / 5 is not an integer, it is rounded up, and the rounded-up integer is used as the target number of columns. Alternatively, the first width may be determined as the target number of columns. If the first depth is greater than the number of pairs of input values of the multiply-accumulate array, the operation of the entire vector is completed by expanding (i.e., increasing) the number of operation cycles. Optionally, for the case where the input data information includes multiple sets of input data, the target number of columns may be further determined in combination with the number of sets of input data.

[0075] In some alternative embodiments, for a two-dimensional matrix of h*w or a first tensor represented as a size of 1*h*w, the first width h may be determined as the target number of rows, and the target number of columns may be determined according to the first depth w and the input data information of each multiply-accumulate array. For example, if the input data information of the multiply-accumulate array includes 5 pairs of input values, the target number of columns may be w*1 / 5. Alternatively, if the input data information of the multiply-accumulate array includes 4 sets of input data, with 5 pairs of input values in each set, the target number of columns may be w*1 / (5*4). For the second tensor, its size information may be represented as 1*w*f. That is, considering the data structure of the two-dimensional matrix, the number of rows of the second matrix is the same as the number of columns w of the first matrix. For each row in the first matrix, multiply-accumulate operations need to be performed with each column in the second matrix, which can be achieved through multiple operation cycles. Alternatively, the target number of columns may be further determined in combination with the first depth of the first tensor and the size information of the second tensor. Alternatively, based on the first height, the first width, and the first depth of the first tensor, as well as the second height, the second width, and the second depth of the second tensor, the target number of rows and the target number of columns may be comprehensively determined. There is no specific limitation.

[0076] In some alternative embodiments, for the case where the first tensor is a two-dimensional matrix, it is not limited to determining the first height of the first tensor as the target number of rows. For example, the second width (i.e., the number of columns of the second tensor) of the second tensor may also be determined as the target number of rows to perform multiply-accumulate operations on each column of the second tensor with each row of the first tensor.

[0077] In some alternative embodiments, for the case where the first tensor is a three-dimensional tensor, the first height may be determined as the target number of rows, and the first width may be determined as the target number of columns. Alternatively, the first height may be determined as the target number of rows, and in combination with the first width, and the logarithm of the input value of the first depth and the multiply-accumulate array, the target number of columns may be determined. Alternatively, in combination with the first width, the first depth, the number of groups of input data of the multiply-accumulate array, and the logarithm of the input value of each group, the target number of columns may be determined, and so on. Specifically, it may be set according to actual requirements.

[0078] In some alternative embodiments, if the first height in the first dimension information does not meet the preset dimension condition, and the first width meets the preset dimension condition. For example, the first height is greater than or equal to m, and the first width is less than n, indicating that the first height exceeds m rows. Then, m is determined as the target number of rows, and the first width is determined as the target number of columns. After that, the first height of the first tensor is decomposed into multiple sub-heights, that is, the first tensor is decomposed into sub-tensors with multiple sub-heights by rows, and the operations are respectively performed through the target multiply-accumulate array. Each sub-tensor can complete the multiply-accumulate calculation through at least one operation cycle. The number of operation cycles required for each sub-tensor is determined according to the depth of the sub-tensor (i.e., the first depth of the first tensor) and the input data information supported by each multiply-accumulate array. For example, if the depth of the sub-tensor exceeds the logarithm of the input value supported by each multiply-accumulate array, the depth dimension of the sub-tensor also needs to be divided into multiple sub-depths, and the overall multiply-accumulate operation of the depth dimension is completed by increasing the number of operation cycles. For example, the depth of the sub-tensor is 8, and each multiply-accumulate array supports the multiply-accumulate operation of 4 pairs of input values in one operation cycle. The depth dimension needs to be divided into two sub-depths, and the multiply-accumulate operation of the two sub-depths is completed through two operation cycles. Optionally, a certain number of operation cycles are also required to complete the fusion operation of the multiply-accumulate results of each sub-depth. Therefore, each sub-tensor requires at least one operation cycle to complete the operation, and the number of operation cycles required for multiple sub-tensors is determined according to the number of operation cycles required for each sub-tensor. For example, if the first height is d times of m, and the number of operation cycles required for each sub-tensor is e, then d*e operation cycles are required to complete the overall multiply-accumulate operation.

[0079] In some alternative embodiments, if the first width in the first dimension information does not meet the preset dimension condition, and the first height meets the preset dimension condition, for example, the first height is less than m, and the first width is greater than or equal to n, that is, the first width exceeds the total number of columns of the multiply-accumulate array, then the first width is taken as the target number of columns, and the first height is taken as the target number of rows. The first width of the first tensor is decomposed into multiple sub-widths, and the first tensor is decomposed into sub-tensors with multiple sub-widths column by column (i.e., in the width direction). By expanding the number of operation cycles of the target multiply-accumulate array, the overall multiply-accumulate operation is completed. Similar to the above embodiment in which the first tensor is decomposed into sub-tensors with multiple sub-heights row by row, the first tensor is decomposed into sub-tensors with multiple sub-widths column by column. Each sub-tensor can complete the multiply-accumulate calculation through at least one operation cycle. The number of operation cycles required for each sub-tensor is determined according to the depth of the sub-tensor and the input data information supported by each multiply-accumulate array. The number of operation cycles required for multiple sub-tensors is determined according to the number of operation cycles required for each sub-tensor. Details are not described one by one here.

[0080] In some alternative embodiments, when both the one-dimensional vector and the two-dimensional matrix are represented as the first tensor in a three-dimensional data structure, that is, the one-dimensional vector can be represented as a three-dimensional data structure of 1*1*C, and the two-dimensional matrix can be represented as a three-dimensional data structure of 1*h*w. On the basis of meeting the preset dimension condition, the actual type of the first tensor can also be determined by combining the first height and the first width of the first tensor. The actual type 1 indicates that the first tensor is a one-dimensional vector. The target number of rows and the target number of columns of the target multiply-accumulate array can be determined according to the pre-configured methods for determining the target multiply-accumulate array corresponding to the one-dimensional vector and the two-dimensional matrix. For example, for the case where the first tensor is a one-dimensional vector, the target number of rows is not limited to 1, and the total number of target multiply-accumulate arrays can meet the calculation requirements. For example, if a total of 6 multiply-accumulate arrays are required, the target number of rows and the target number of columns can be any one of 1 row and 6 columns, 6 rows and 1 column, 2 rows and 3 columns, 3 rows and 2 columns, etc.

[0081] In the embodiments of the present disclosure, by combining the first dimension information and the preset dimension condition constraints, the target number of rows and the target number of columns of the target multiply-accumulate array are determined, avoiding the situation where the determined target number of rows and the target number of columns exceed the total number of rows and / or the total number of columns of the multiply-accumulate array, thereby ensuring the effectiveness of the target number of rows and the target number of columns of the target multiply-accumulate array.

[0082] In some alternative embodiments, the controller 22 is specifically configured to:

[0083] In response to the first dimension information not meeting the preset dimension condition, based on the quantity information of the target multiply-accumulate array, the first tensor is determined as at least one first sub-tensor; according to each first sub-tensor, the second tensor, and the operation type, the target control mode corresponding to the target multiply-accumulate array is determined.

[0084] Among them, in the case where the first dimension information does not meet the preset dimension condition, in order to ensure that the target multiply-accumulate array can complete the accurate operation of the first tensor and the second tensor, based on the quantity information of the target multiply-accumulate array, the first tensor can be determined as multiple sub-tensors (i.e., the first sub-tensors), and then according to each first sub-tensor, the second tensor, and the operation type, the control operation and timing corresponding to the target multiply-accumulate array are determined to obtain the target control method. Specifically, according to each first sub-tensor, the second tensor, and the operation type, combined with the quantity information of the target multiply-accumulate array, the target multiply-accumulate array corresponding to the data in each first sub-tensor is determined, and according to the input data information supported by each target multiply-accumulate array, the input timing of the data corresponding to each target multiply-accumulate array in each first sub-tensor as the input data (the first input data) of the target multiply-accumulate array is determined, and according to the input timing of the first input data, the input timing of the second input data corresponding to the first input data of each target multiply-accumulate array in the second tensor is determined. Combining the input timings corresponding to the first input data and the second input data respectively, the calculation timings of each target multiply-accumulate array are determined, etc. Based on various control operations and their corresponding operation timings, the target control method is determined. For example, at the t0 moment after the target multiply-accumulate array enters the working state, according to the input timing of the data, at least one pair of input values corresponding to the t0 moment (i.e., at least one pair of input values corresponding to the timing of the t0 moment) are respectively input to each target multiply-accumulate array, that is, at least one pair of input values corresponding to the t0 moment are input to each target multiply-accumulate array according to the input timing of the data. The number of pairs of input values included in the at least one pair of input values is the number of pairs of input values supported by the target multiply-accumulate array; each pair of input values includes a first input value from the first tensor and a second input value from the second tensor. According to the calculation timings of each target multiply-accumulate array, a clock signal is transmitted to each target multiply-accumulate array at the t1 moment, so that each target multiply-accumulate array performs a multiply-accumulate calculation on the first input value and the second input value input at the t0 moment to obtain the multiply-accumulate result corresponding to the t1 moment. At the t2 moment, the multiply-accumulate result corresponding to the t1 moment is cached, and at least one pair of input values corresponding to the t2 moment are respectively input to each target multiply-accumulate array according to the input timing of the data. According to the calculation timings of each target multiply-accumulate array, a clock signal is transmitted to each target multiply-accumulate array at the t3 moment, so that each target multiply-accumulate array performs a multiply-accumulate calculation on the first input value and the second input value input at the t2 moment to obtain the multiply-accumulate result corresponding to the t3 moment, and so on. Each control operation is executed according to a certain timing to control each target multiply-accumulate array to cooperate to complete the operation of the first tensor and the second tensor.

[0085] In an embodiment of the present disclosure, for a first tensor that does not meet the preset size condition, the first tensor can be split into multiple first sub-tensors, and then based on the target multiply-accumulate array, through timing control, the calculation of multiple sub-tensors is completed, so that the calculation of a tensor with a larger size can be completed based on a multiply-accumulate array with a smaller scale, further improving the utilization rate and versatility of the multiply-accumulate array, effectively reducing the scale of the multiply-accumulate array, and facilitating the physical implementation of a tensor calculation device or chip integrated with the multiply-accumulate array.

[0086] In some optional embodiments, based on any of the above embodiments, the controller 22 is specifically configured to:

[0087] Based on the first tensor, the second tensor, the operation type, and the status information of each multiply-accumulate array 23i in the calculation component 23, determine the target multiply-accumulate array for the current operation from the multiply-accumulate arrays with the status information being the idle state among the multiple multiply-accumulate arrays.

[0088] Wherein, the status information of the multiply-accumulate array 23i is information indicating whether the multiply-accumulate array 23i is in the idle state. The status information of each multiply-accumulate array 23i can be maintained in real time during the scheduling process of the multiply-accumulate array. For example, the status information of the multiply-accumulate array can be recorded in real time through a status register, or the status information of the multiply-accumulate array can be determined through other monitoring methods, which is not specifically limited.

[0089] In some optional embodiments, each multiply-accumulate array can be calculated independently, and any number of multiply-accumulate arrays can be combined for calculation, so that the multiply-accumulate array can adapt to various different calculation tasks and support the simultaneous execution of multiple calculation tasks. In this case, when it is necessary to perform an operation on the first tensor and the second tensor, some of the multiply-accumulate arrays in each multiply-accumulate array may be participating in other calculation tasks. Combining the first tensor, the second tensor, the operation type, and the status information of each multiply-accumulate array 23i, comprehensively determining the target multiply-accumulate array for the current operation can effectively ensure that the target multiply-accumulate array participates in the calculation in a timely manner, avoiding the situation where it is necessary to wait for the execution of other calculation tasks to end before participating in the calculation.

[0090] In some optional embodiments, the number and / or arrangement information of the multiply-accumulate arrays with the status information being the idle state in each multiply-accumulate array can be determined first. The arrangement information refers to the number of rows and columns of the multiply-accumulate arrays in the idle state. Combining the number and / or arrangement information of the multiply-accumulate arrays in the idle state, determining the association relationship between the first size information of the first tensor and the number and / or arrangement information of the multiply-accumulate arrays in the idle state, and determining the target multiply-accumulate array in the manner of determining the target multiply-accumulate array according to the foregoing embodiments, which will not be elaborated herein.

[0091] Embodiments of the present disclosure further determine a target multiply-accumulate array in an idle state in combination with the status information of the multiply-accumulate array, which can ensure the availability of the target multiply-accumulate array, so as to be able to perform calculations in a timely manner, avoid waiting for the target multiply-accumulate array to execute other calculation tasks, and ensure the real-time performance of the current calculation task.

[0092] Figure 4 It is a schematic structural diagram of a computing component provided by another exemplary embodiment of the present disclosure. Figure 5 It is a schematic structural diagram of a computing component provided by still another exemplary embodiment of the present disclosure.

[0093] In some optional embodiments, on the basis of any of the above embodiments, as Figure 4 and Figure 5 shown, any multiply-accumulate array 23i includes an input register buf and a multiply-accumulate unit MAC. See Figure 4 , i = 1, 2, …, M, where M is the number of multiply-accumulate arrays and M is a positive integer; or see Figure 5 , i = 11, 12, …, 1n, 21, 22, …, 2n, …, mn, where m is the number of rows of the multiply-accumulate array and n is the number of columns of the multiply-accumulate array, and both m and n are positive integers.

[0094] The input register buf is configured to store at least a first input value.

[0095] The controller 23 is specifically configured to:

[0096] Write each first element value in the first tensor into the input register buf in the corresponding target multiply-accumulate array according to the first control timing in the target control mode as the first input value corresponding to the target multiply-accumulate unit in the target multiply-accumulate array; determine the second input value in the second tensor corresponding to each first input value based on the second tensor configuration information in the target control mode; control the target multiply-accumulate unit to perform a multiply-accumulate operation on the first input value in the input register and the corresponding second input value to obtain a multiply-accumulate operation result; and determine a tensor calculation result based on the multiply-accumulate operation results of the target multiply-accumulate units in each target multiply-accumulate array.

[0097] Among them, the input register buf can be any type of register, such as a general register, a shift register, etc. The input register buf can store at least a first input value (or the first input data in the input data), and the number of first input values that can be stored can be set according to the requirements of the multiply-accumulate unit MAC. For example, if the multiply-accumulate unit MAC can complete the multiply-accumulate operation of 5 pairs of input values at a time, the input register can store at least 5 first input values, and the second input value corresponding to each first input value is controlled by the controller 22 to be input as a weight value during the calculation process.

[0098] In some alternative embodiments, the input register buf is electrically connected to the multiply-accumulate unit MAC and provides a first input value to the multiply-accumulate unit MAC.

[0099] In some alternative embodiments, the input register buf is configured to store at least a pair of input values, and a pair of input values includes a first input value and a second input value.

[0100] In some alternative embodiments, the first control timing in the target control mode is the control timing for writing the first input value to the input register buf of the target multiply-accumulate array. For example, at the 1st clock edge (i.e., the operation cycle), according to the first control timing, the first input value corresponding to the current operation cycle is written to the input register buf corresponding to each target multiply-accumulate array. After completing one or more calculations based on the first input value, the first input value corresponding to the next operation cycle is written to the input register buf according to the first control timing to perform the next calculation, and so on. According to the first control timing in the target control mode and the timing of other control operations, the calculation process of the target multiply-accumulate array is controlled. Each set of first input values may need to perform multiply-accumulate operations with multiple sets of second input values respectively. Therefore, after the first input value is written to the input register buf, it may participate in one or more calculations before performing the relevant operations corresponding to the next set of first input values. The multiply-accumulate unit in the target multiply-accumulate array is referred to as the target multiply-accumulate unit. The second tensor configuration information in the target control mode is the second input value and the corresponding timing determined from the second tensor according to the input timing of the first input value, so that based on the second tensor configuration information, the second input value corresponding to each first input value can be determined. Furthermore, the target multiply-accumulate unit can be controlled to perform a multiply-accumulate operation on the first input value in the input register buf and the corresponding second input value to obtain a multiply-accumulate operation result (or called a multiply-accumulate result). According to various control operations and control timings in the target control mode, the target multiply-accumulate array is controlled to complete each operation and obtain each multiply-accumulate result.

[0101] In some alternative embodiments, each time the target multiply-accumulate array completes a multiply-accumulate operation, the multiply-accumulate result can be cached for determining the final tensor calculation result.

[0102] In some alternative embodiments, the controller 22 determines the tensor calculation result based on the multiplication and addition operation results of the target multiplication and addition units in each target multiplication and addition array. For example, the controller 22 may perform post-processing on the multiplication and addition operation results of each target multiplication and addition unit to obtain the tensor calculation result. The post-processing may include addition operations, data combination according to the data structure of the tensor calculation result, etc. The addition operation is used to add the multiplication and addition operation results of the multiplication and addition operations with larger sizes decomposed into multiple multiplication and addition operations with smaller sizes to obtain the multiplication and addition operation result with a larger size. For example, both the first tensor and the second tensor are vectors. During the multiplication and addition calculation performed by the target multiplication and addition array, the vector dot product operation of the first tensor with a larger size and the second tensor is decomposed into multiple sub-operations of the first sub-tensor and the second sub-tensor with smaller sizes. The multiplication and addition results of each sub-operation need to be added to obtain the operation result of the vector dot product operation of the first tensor with a larger size and the second tensor.

[0103] In some alternative embodiments, the computing component 23 may further include an addition operation unit. For the case where the multiplication and addition results of multiple target multiplication and addition arrays or the multiplication and addition results of multiple operation cycles need to be further added, the addition operation can be performed based on the addition operation unit to obtain the multiplication and addition result that meets the requirements. For example, for a vector dot product operation with a relatively long vector length, if it is decomposed into operations of multiple target multiplication and addition arrays or operations of multiple operation cycles, the multiplication and addition results of multiple target multiplication and addition arrays need to be added, or the multiplication and addition results of multiple operation cycles need to be added to obtain the final vector dot product operation result. Then, the addition operation of each multiplication and addition result can be implemented through the addition operation unit. That is to say, the controller 22 can control the addition operation unit to perform post-processing on each multiplication and addition result to determine the tensor calculation result. The target control method may include the control operation of this post-processing and the corresponding operation timing.

[0104] In the embodiments of the present disclosure, the first input value is stored in the input register buf and provided to the multiplication and addition unit MAC, enabling the first input value to participate in at least one calculation of the multiplication and addition unit. Furthermore, the first input value is written into the input register through the first control timing in the target control method, providing an effective first input value for the multiplication and addition unit of the target multiplication and addition array. Combining with the second tensor configuration information, a corresponding second input value is provided for the multiplication and addition unit, enabling the target multiplication and addition unit to perform a multiplication and addition operation on the first input value and the second input value to determine the tensor calculation result, thereby achieving accurate control of the target multiplication and addition unit and ensuring the effectiveness of the tensor calculation result.

[0105] Figure 6 It is a schematic structural diagram of a computing component provided by another exemplary embodiment of the present disclosure.

[0106] In some alternative embodiments, based on any of the above embodiments, such as Figure 6As shown, the number of target multiply-accumulate arrays is multiple; any multiply-accumulate array includes an input register buf and a multiply-accumulate unit MAC.

[0107] The input register buf is configured to store at least a first input value; the input register buf is a shift register; the input registers buf of each multiply-accumulate array 23i are electrically connected in a preset arrangement.

[0108] The controller 22 is specifically configured to:

[0109] Write each first element value in the first tensor into each input register buf at the first specified position in each target multiply-accumulate array in sequence according to the first control timing, control each input register buf to shift in the first direction, and shift the first element value to the input register buf corresponding to the first element value; or, write each first element value into each input register buf at the second specified position in each target multiply-accumulate array in sequence according to the first control timing, control each input register buf to shift in the second direction, and shift the first element value to the input register buf corresponding to the first element value.

[0110] Among them, the shift register may include at least one of a left shift register, a right shift register, and a bidirectional shift register, and can be specifically set according to actual requirements. The preset arrangement can be determined according to the arrangement of multiple multiply-accumulate arrays. For example, if the arrangement of multiple multiply-accumulate arrays is m rows and n columns, the preset arrangement of each input register may include electrically connecting the input registers in each row of the multiply-accumulate arrays, or electrically connecting the input registers in each column of the multiply-accumulate arrays, or electrically connecting the input registers in each row and electrically connecting the input registers in each column. The specific connection method is not limited.

[0111] In some alternative embodiments, the first designated position may be an edge row in the target multiply-accumulate array, such as the first row or the last row; the second designated position may be an edge column in the target multiply-accumulate array, such as the first column or the last column. The first direction and the second direction are the shiftable directions of the input register buf. For example, the first direction is the column-wise (or longitudinal) direction from the first row to the last row or from the last row to the first row, and the second direction is the row-wise (or transverse) direction from the first column to the last column or from the last column to the first column. The specific first designated position and the first direction, and the second designated position and the second direction can be determined according to the movable directions of the respective input registers. Taking one row of the target multiply-accumulate array as an example, if the input register buf is a left-shift register, the second designated position is the last column (i.e., the rightmost column) of the target multiply-accumulate array, and the second direction is the direction from right to left along the row. For one column of the target multiply-accumulate array, a left-shift register means shifting from top to bottom along the column, so the first designated position is the first row of the target multiply-accumulate array, and the first direction is the direction from top to bottom along the column. For example, if the target multiply-accumulate array includes Figure 6 all the multiply-accumulate arrays therein, then the rightmost column is the nth column (including the multiply-accumulate arrays 231n to 23mn), and the first row of the target multiply-accumulate array is Figure 6 the first row of the multiply-accumulate arrays 2311 to 231n in

[0112] In some alternative embodiments, the input registers of each row may be circularly shifted or non-circularly shifted. Circular shift means that the first input register of each row is also electrically connected to the last register of the same row, so that during the shifting process, the input value in the first input register can be shifted to the last input register, and / or the input value in the last input register can be shifted to the first input register.

[0113] In some alternative embodiments, the multiply-accumulate results of the respective multiply-accumulate units MAC can be cached in the corresponding registers, and the controller 22 can read the multiply-accumulate results of the respective multiply-accumulate units from the registers. The register can be a register inside the computing component 23 or a register outside the computing component, and is not specifically limited.

[0114] In some alternative embodiments, during the process of shifting and injecting the first input value, data multiplexing can be achieved. For example, the second element value of the second tensor can be written as the second input value into the target multiply-accumulate array. For the case where the first input value in a group needs to perform multiply-accumulate operations with the second input values in multiple groups, after the first input value is written into the input registers buf at the first specified position (such as the first row), it first serves as the first input value of the multiply-accumulate unit at the first specified position and performs operations with the second input value in the first group, and then is shifted to the next row and serves as the first input value of the target multiply-accumulate array in the next row to perform operations with the second input value in the second group in the next row, and so on, to achieve data multiplexing and reduce the number of data transfers.

[0115] In the embodiments of the present disclosure, by using a shift register to implement the data injection of the input registers in each row or each column, the wiring complexity between the input registers and the outside can be effectively reduced, and the multiplexing of the input values can be easily implemented through the shift register, effectively reducing the number of data transfers.

[0116] Figure 7 It is a schematic structural diagram of a multiply-accumulate array provided by an exemplary embodiment of the present disclosure.

[0117] In some alternative embodiments, based on any of the above embodiments, as Figure 7 shown, any multiply-accumulate array 23i includes an input register buf1, a multiply-accumulate unit MAC, and an output register buf2.

[0118] The input register buf1 is configured to store the first input value in at least one pair of input values.

[0119] The multiply-accumulate unit MAC is configured to perform a multiply-accumulate calculation on at least one pair of input values to obtain the multiply-accumulate result of the at least one pair of input values.

[0120] The output register buf2 is configured to store the multiply-accumulate result obtained by the multiply-accumulate unit.

[0121] Among them, the input register buf1 and the multiply-accumulate unit MAC can refer to the input register buf in the foregoing embodiments and will not be elaborated here. The output register buf2 can be any register. For example, the output register buf2 can include at least one of a general register, a shift register, etc. The shift register can include at least one of a left shift register, a right shift register, and a bidirectional shift register.

[0122] In some alternative embodiments, after each multiply-accumulate operation is completed, the multiply-accumulate unit MAC can cache the multiply-accumulate operation result into the output register buf2. The controller 22 can read the multiply-accumulate operation result from the output register buf2, and then determine the tensor calculation result based on the multiply-accumulate operation results of each target multiply-accumulate array in each operation cycle.

[0123] In an embodiment of the present disclosure, by outputting a register to cache the multiplication and addition result of the multiplication and addition unit, it is convenient for the controller to read the multiplication and addition result, thereby providing an accurate and effective data reference for determining the tensor calculation result.

[0124] Figure 8 It is a schematic structural diagram of a computing component provided by another exemplary embodiment of the present disclosure.

[0125] In some optional embodiments, on the basis of the above embodiments, as Figure 8 shown, the output register buf2 is a shift register; the number of target multiplication and addition arrays is multiple.

[0126] The output registers buf2 of each multiplication and addition array are electrically connected.

[0127] The controller 22 is further configured to: after the multiplication and addition result obtained by the multiplication and addition unit MAC in the target multiplication and addition array is written into the output register buf2, control each output register buf2 to shift and output in a preset direction.

[0128] Among them, the output registers buf2 of each multiplication and addition array can be electrically connected according to rows and / or columns, so as to shift and output the multiplication and addition results of each row or each column. Figure 8 Taking the electrical connection of the output registers buf2 of each column as an example, the actual application is not limited to this connection method.

[0129] In some optional embodiments, for the case where the number of target multiplication and addition arrays is less than or equal to the total number of multiplication and addition arrays, they can all be shifted and output according to the preset direction of the output register buf2. For example, the target multiplication and addition array includes Figure 8 the partial multiplication and addition arrays of 2 rows and 3 columns in, for example, the multiplication and addition arrays 2311 to 2313 and the multiplication and addition arrays 2321 to 2323. After one operation is completed, they can be shifted and output through the output register buf2.

[0130] In an embodiment of the present disclosure, by using a shift register as the output register, the multiplication and addition result of the multiplication and addition unit can be shifted and output, further reducing the wiring complexity, thereby reducing the physical implementation difficulty.

[0131] In some optional embodiments, on the basis of any of the above embodiments, the operation type is a convolution operation.

[0132] The multiple multiplication and addition arrays 23i include multiplication and addition arrays of m rows and n columns; any multiplication and addition array 23i includes an input register buf and a multiplication and addition unit MAC; the input registers buf are electrically connected according to a preset arrangement, see Figure 6 shown.

[0133] The target multiply-accumulate array includes a multiply-accumulate array with s rows and t columns, that is, a multiply-accumulate array with s rows and t columns is determined from multiple multiply-accumulate arrays of the computing component as the target multiply-accumulate array; s is less than or equal to m, and t is less than or equal to n.

[0134] The controller 22 is specifically configured to:

[0135] Iteratively execute: According to the second control timing in the target control mode, control each multiply-accumulate unit in each target multiply-accumulate array to perform a multiply-accumulate operation on the current first input value in the input register buf and the corresponding current second input value in the second tensor, to obtain the first multiply-accumulate result of the current cycle (i.e., the current operation cycle), and according to the shift direction corresponding to the current cycle in the target control mode, control the shift of each input register buf in the target multiply-accumulate array, so as to shift the first input value in one input register buf to another input register buf; in response to the shift times of each input register satisfying the end condition, determine the tensor calculation result based on the first multiply-accumulate results obtained in each iteration.

[0136] Among them, for convolution operations, the first tensor can be a feature tensor, and the second tensor can be a convolution kernel. The second tensor can include one or more convolution kernels, and each convolution kernel is represented as a three-dimensional tensor. The second control timing in the target control mode is the timing for controlling the calculation of the multiply-accumulate units in the target multiply-accumulate array. The current first input value and the current second input value are initially the first batch of data for performing the multiply-accumulate operation in the first tensor and the second tensor respectively, and the initial first input value and / or the second input value can be filled into each input register through the shift function of the input register. Then, iteratively execute the above process of multiply-accumulate operation, shift, multiply-accumulate operation, and shift. The shift direction corresponding to the current cycle can be one of shifting down along the column, shifting up along the column, shifting left along the row, and shifting right along the row, which is specifically determined according to the second control timing. The end condition is the shift times threshold, which can be obtained in the target control mode.

[0137] In some alternative embodiments, the input registers buf1 of each multiply-accumulate array are electrically connected according to a preset arrangement; the output registers buf2 of each multiply-accumulate array are electrically connected. Each input register buf1 is a shift register. Each output register buf2 is a shift register. During the calculation process, the multiplexing of the input values can be realized through the shift of the input register buf1, and after the multiply-accumulate result obtained by the multiply-accumulate unit is written into the output register buf2, it is output through shifting.

[0138] Figure 9 It is a schematic structural diagram of a computing component provided by another exemplary embodiment of the present disclosure. As Figure 9As shown, taking a 4-row and 3-column multiply-accumulate array as an example, the input registers buf1 of each column are electrically connected, and the input registers buf1 of each row are electrically connected. Taking the multiply-accumulate array 2312 as an example, the arrow from the input register buf1 to the multiply-accumulate array 2311 indicates an electrical connection to the input register buf1 of the multiply-accumulate array 2311 and can be shifted in the direction of the arrow (i.e., the direction to the left along the row). The arrow pointing to the multiply-accumulate array 23j3 (j = 1, 2, 3, 4) on the rightmost side represents the write end of the first input value, which can correspond to the height (H) direction of the first tensor to be calculated. That is, the input values at multiple height positions (such as four height positions h1, h2, h3, and h4) in the height direction of the first tensor can be written into the input register buf1 of the rightmost column from the right. Each height position can correspond to writing multiple first input values, or multiple groups of first input values, and each group includes multiple first input values. The input arrow above the multiply-accumulate array 231i represents the write end of the first input value, which can correspond to the width (W) direction of the first tensor. That is, the input values at multiple width positions (such as three width positions w1, w2, and w3) in the width direction of the first tensor can be written into the input register buf1 of the first row from above. In practical applications, if the input register buf1 is a right-shift shift register, the input end in the height direction can also be in the first column (such as the left side in the figure), and the input end in the width direction can also be in the last row (such as the lower side in the figure). After writing the first input value through the write end, the first input value can be shifted to the next input register through row-wise or column-wise shifting, and a new batch of first input values can be written at the write end, and so on, to fill each multiply-accumulate array. Or new input data can be filled while calculating according to the control timing, and the control operations and corresponding operation timings are specifically determined according to the operation properties of the convolution operation.

[0139] In some alternative embodiments, Figure 10 is a schematic structural diagram of a multiply-accumulate unit MAC provided by an exemplary embodiment of the present disclosure. As Figure 10 shown, each multiply-accumulate array can parallelly complete the multiply-accumulate operations of 4 groups of input data in one operation cycle. Each group of input data can include 4 pairs of input values. For convolution operations, Weight ki represents the weight value in the i-th convolution kernel (each convolution kernel inputs 4 weight values per beat as the second input value), and Feature c1-c4 represents 4 values in the depth direction of the feature tensor (as the first input value). The multiply-accumulate unit MAC includes 4 groups of multiply-accumulate sub-units. Each group of multiply-accumulate sub-units includes four multipliers and one adder. The 4 multipliers are used to complete the multiplication operations of 4 pairs of input values to obtain 4 products, and the adder is used to complete the addition operation of the 4 products to obtain the multiply-accumulate results of 4 pairs of input values. A multiply-accumulate unit can obtain the multiply-accumulate results corresponding to 4 groups of input data respectively. It should be noted that, Figure 10This is only an exemplary structural diagram of the multiply-accumulate array. In actual applications, the number of multiply-accumulate subunits and the number of input value pairs for each multiply-accumulate subunit can be set according to actual needs, and are not limited to those shown in the figure.

[0140] Exemplarily, Figure 11 is a schematic diagram of the principle of convolution operation provided by an exemplary embodiment of the present disclosure. As Figure 11 shown, the H*W*C Feature (i.e., the feature tensor) represents the first tensor, and the R*S*C Weight kp represents the pth convolution kernel, where p = 1, 2, …, K, and K represents the number of convolution kernels. The convolution result (Output) of P*Q*K is obtained through convolution operation. Using the tensor calculation device 20 of the embodiment of the present disclosure, taking the calculation component 23 Figure 9 as an example, and taking the multiply-accumulate unit Figure 10 as an example. In the case of using the 4-row and 3-column multiply-accumulate array shown in Figure 9 as the target multiply-accumulate array, it can be obtained from Figure 9The rightmost shown fills in the input data, which can correspond to the input data at 4 height positions at a time. For example, the input data at four height positions h1, h2, h3, and h4. Each height position corresponds to a row of the target multiply-accumulate array. The four groups of first input values of each multiply-accumulate array all use 4 values c1 to c4 along the depth (C) direction at the same width position. Each column of the target multiply-accumulate arrays respectively corresponds to different width positions (such as w1, w2, w3) of the Feature. The four groups of second input values respectively correspond to the weight values of 4 different convolutional kernels. The weight values at different positions of different convolutional kernels are respectively used as the second input values of each target multiply-accumulate array. The c1 to c4 at different height and different width positions (hiwj) of the Feature are respectively filled into each input register as the first input values. Since the c1 to c4 at each position hiwj (i = 1, 2,..., H, j = 1, 2,..., W) of the Feature need to perform multiply-accumulate operations with the c1 to c4 values at each risj (i = 1, 2,..., R, j = 1, 2,..., S) position of each convolutional kernel, therefore, the first input values of each batch are shifted through the input register to achieve multiply-accumulate operations with the weight values at different positions. Exemplarily, the second input value in the multiply-accumulate array 2311 includes c1 to c4 at the r1s1 position of each of the 4 convolutional kernels, the second input value in the multiply-accumulate array 2312 includes c1 to c4 at the r1s2 position of each of the 4 convolutional kernels, the second input value in the multiply-accumulate array 2321 includes c1 to c4 at the r2s1 position of each of the 4 convolutional kernels,..., the c1 to c4 at the four positions h1w1, h2w1, h3w1, and h4w1 of the Feature are first written into the input register in the rightmost column, and are filled into the first column by left shift. At the same time as each left shift, the c1 to c4 at the four positions h1w2, h2w2, h3w2, and h4w2 can be written into the rightmost column, and are filled into the second column by left shift, and so on, to fill the input registers of each target multiply-accumulate array with the corresponding first input values. Then, control each target multiply-accumulate array to perform multiply-accumulate operations on the current first input value and the current second input value to obtain the first multiply-accumulate result of the current cycle. Then, according to the shift direction (such as left shift or down shift) corresponding to the current cycle, shift the first input values in each input register so that the first input values in each input register are shifted to the input register of another target multiply-accumulate array to perform multiply-accumulate operations with the second input value of another target multiply-accumulate array. Through the iterative operations of operation → shift → operation → shift, the reuse of the first input value is realized. The specific number of shift times and the shift direction each time are determined according to the properties of the convolutional operation, and will not be elaborated one by one.It should be noted that the above is only an exemplary embodiment. In actual applications, it is not limited to the above embodiment. For example, the second input value can also be multiplexed by shifting through the input register, that is, the first input value in each multiply-accumulate array remains unchanged, and the second input value is shifted to implement the multiply-accumulate operation between the first input value and different second input values, which can be specifically set according to actual needs and will not be elaborated here. Optionally, it can also be from. Figure 9 The first input data is poured into the target multiply-accumulate array shown from above, shifted downward through the input register buf1, and poured into the corresponding input register. The specific principle is similar to that of pouring from the right side. For example, first write c1 to c4 at positions h1w1, h1w2, and h1w3 into the input registers of the first row, and then shift downward and write h2w1, h2w2, and h2w3, and so on, which will not be elaborated here. After completing the operation of a batch of first input data, the calculation of the next batch of input data can be performed, such as the multiply-accumulate operation of c5 to c8 at each position. The calculation process is the same as the above process and will not be elaborated here.

[0141] In some alternative embodiments, for the convolution operation, the first tensor is a feature tensor, and the first tensor may include N feature tensors of H*W*C, that is, the batch size of the feature tensor. The target multiply-accumulate array and the corresponding target control method can be further determined in combination with N.

[0142] In some alternative embodiments, for each multiply-accumulate array, in combination with Figure 10 As shown, by broadcasting Feature c1 to c4 to each multiply-accumulate subunit, each multiply-accumulate subunit shares a set of first input data. In actual applications, it is not limited to broadcasting the first input data. The Feature and Weight can also be exchanged. For example, c1 to c4 of Weight are broadcast to Features of different batches. That is, in the case of having multiple Features of H*W*C, a set of second input data can be broadcast to each multiply-accumulate subunit and perform multiply-accumulate operations with the first input data of different Features respectively, which is not specifically limited.

[0143] In some alternative embodiments, each column in the multiply-accumulate array of m rows and n columns can be used as a group, that is, multiple multiply-accumulate arrays include n groups, and each group includes m multiply-accumulate arrays.

[0144] It should be noted that the structural diagram of the computing component 23 shown in the embodiments of the present disclosure is only a logical structural representation. In the physical layout, each multiply-accumulate array can be distributed at any position in the circuit and electrically connected according to the logical relationship, and it is not limited to arranging each multiply-accumulate array centrally.

[0145] Embodiments of the present disclosure achieve multiplexing of input data through the shift function of the input register, reduce data movement, and improve computing efficiency. In addition, for the RS dimension of Weight and higher-dimensional convolutions, they can be unfolded in time, so that general convolution processing of any size can be achieved, and each dimension can be configured. Therefore, the structure of the computing component in the embodiments of the present disclosure has good scalability and flexibility. At the same time, it also supports setting any dimension in N*H*W*C*K to 1, where N represents the batch dimension of Feature, and K represents the number dimension of convolution kernels. By unfolding in time and completing the calculation through multiple operation cycles, the device in the embodiments of the present disclosure can adapt to tensor calculations with arbitrary dimension combinations and for different scenarios. Moreover, the computing component in the device of the embodiments of the present disclosure uses a multiply-accumulate array as the basic operation unit. During the chip layout design phase, it can be effectively sliced and extended row by row or column by column on a two-dimensional layout. The input and output data are sliced into smaller bandwidths, decomposing the bandwidth requirements and reducing the difficulty of chip layout implementation. It can also decompose the input and output data in the KC direction of the convolution kernel to meet different requirements. Each multiply-accumulate array is relatively independent, and Feature or Weight can be broadcast within the multiply-accumulate array, effectively reducing the wiring complexity. In addition, since there are multiple instances at each level, for example, the computing component as a whole includes multiple columns of multiply-accumulate arrays, and each column includes multiple multiply-accumulate arrays, the layout reuse rate can be effectively improved. For example, a layout can be made at any level and multiple layouts can be obtained by replication. This hierarchical and modular design brings great possibilities for the implementation of technologies such as soft start and clock skew design. Soft start means starting different circuit modules in batches to relieve the voltage pulse during initial startup. Clock skew design means that each circuit module can operate on different clocks, that is, different modules do not need to act together at the same clock edge to balance power consumption.

[0146] The above embodiments of the present disclosure can be implemented separately or in any combination without conflict according to actual needs. The specific settings are not limited by the present disclosure.

[0147] Exemplary method

[0148] Figure 12 is a schematic flowchart of a tensor calculation method provided by an exemplary embodiment of the present disclosure. The method of this embodiment can be implemented through the corresponding device embodiment above, such as Figure 12 As shown, the method of the embodiments of the present disclosure may include the following steps:

[0149] Step 510, obtain a first tensor and a second tensor to be calculated.

[0150] Step 520, determine the operation type of the first tensor and the second tensor.

[0151] Step 530: Based on the first tensor, the second tensor, and the operation type, determine the target multiply-accumulate array for the current operation from multiple multiply-accumulate arrays of the computing component, and determine the target control mode corresponding to the target multiply-accumulate array.

[0152] Wherein, any multiply-accumulate array is used to determine the multiply-accumulate result of at least one pair of input values; any pair of input values includes a first input value and a second input value; the first input value is a value in the first tensor; the second input value is a value in the second tensor.

[0153] Step 540: Based on the target control mode, control the target multiply-accumulate array to perform the operation corresponding to the operation type on the first tensor and the second tensor, and obtain the tensor calculation result.

[0154] Figure 13 It is a schematic flowchart of a tensor calculation method provided by another exemplary embodiment of the present disclosure.

[0155] In some optional embodiments, based on the above Figure 12 shown embodiment, as Figure 13 shown,

[0156] Step 530 of determining the target multiply-accumulate array for the current operation from multiple multiply-accumulate arrays of the computing component based on the first tensor, the second tensor, and the operation type, and determining the target control mode corresponding to the target multiply-accumulate array may include:

[0157] Step 5310: Based on the first dimension information of the first tensor, determine the quantity information of the target multiply-accumulate array corresponding to the current operation.

[0158] Step 5320: Based on the quantity information, determine the target multiply-accumulate array from multiple multiply-accumulate arrays.

[0159] Step 5330: Based on the first tensor, the second tensor, and the operation type, determine the target control mode corresponding to the target multiply-accumulate array.

[0160] In some optional embodiments, the multiple multiply-accumulate arrays include a multiply-accumulate array with m rows and n columns; the first dimension information includes a first height, a first width, and a first depth.

[0161] Step 5310 of determining the quantity information of the target multiply-accumulate array corresponding to the current operation based on the first dimension information of the first tensor may include:

[0162] In response to the first size information meeting the preset size condition, based on the first height, the first width, and the first depth, determine the target number of rows and the target number of columns of the target multiply-accumulate array corresponding to the current operation as the quantity information; or, in response to the first height in the first size information not meeting the preset size condition, determine m as the target number of rows of the target multiply-accumulate array, determine the first width as the target number of columns of the target multiply-accumulate array, and determine the target number of rows and the target number of columns as the quantity information; or, in response to the first width in the first size information not meeting the preset size condition, determine n as the target number of columns of the target multiply-accumulate array, determine the first height as the target number of rows of the target multiply-accumulate array, and determine the target number of rows and the target number of columns as the quantity information.

[0163] In some alternative embodiments, determining the target control mode corresponding to the target multiply-accumulate array based on the first tensor, the second tensor, and the operation type in step 5330 may include:

[0164] In response to the first size information not meeting the preset size condition, based on the quantity information of the target multiply-accumulate array, determine the first tensor as at least one first sub-tensor; determine the target control mode corresponding to the target multiply-accumulate array according to each first sub-tensor, the second tensor, and the operation type.

[0165] In some alternative embodiments, determining the target multiply-accumulate array for the current operation from multiple multiply-accumulate arrays of the computing component based on the first tensor, the second tensor, and the operation type in step 530 may include:

[0166] Based on the first tensor, the second tensor, the operation type, and the status information of each multiply-accumulate array in the computing component, determine the target multiply-accumulate array for the current operation from the multiply-accumulate arrays with the status information being the idle state among the multiple multiply-accumulate arrays.

[0167] Figure 14 It is a flowchart of the tensor calculation method provided by another exemplary embodiment of the present disclosure.

[0168] In some alternative embodiments, on the basis of any of the above embodiments, any multiply-accumulate array includes an input register and a multiply-accumulate unit; the input register is used to store at least the first input value.

[0169] As Figure 14 shown, controlling the target multiply-accumulate array to perform the operation corresponding to the operation type on the first tensor and the second tensor based on the target control mode in step 540 to obtain the tensor calculation result may include:

[0170] Step 5410, write each first element value in the first tensor into the corresponding input register in the target multiply-accumulate array according to the first control timing in the target control mode as the first input value of the target multiply-accumulate unit in the target multiply-accumulate array.

[0171] Step 5420: Based on the second tensor configuration information in the target control mode, determine the second input value corresponding to each first input value in the second tensor.

[0172] Step 5430: Control the target multiply-accumulate unit to perform a multiply-accumulate operation on the first input value in the input register and the corresponding second input value to obtain a multiply-accumulate operation result.

[0173] Step 5440: Based on the multiply-accumulate operation results of the target multiply-accumulate units in each target multiply-accumulate array, determine the tensor calculation result.

[0174] Figure 15 It is a schematic flowchart of a tensor calculation method provided by another exemplary embodiment of the present disclosure.

[0175] In some optional embodiments, based on any of the above embodiments, the number of target multiply-accumulate arrays is multiple; any multiply-accumulate array includes an input register and a multiply-accumulate unit. The input register is used to store at least the first input value; the input register is a shift register; the input registers of each multiply-accumulate array are electrically connected in a preset arrangement.

[0176] As Figure 15 shown, the step 540 of controlling the target multiply-accumulate array to perform an operation corresponding to the operation type on the first tensor and the second tensor based on the target control mode to obtain a tensor calculation result may include:

[0177] Step 5401: Write each first element value in the first tensor into each input register at the first specified position in each target multiply-accumulate array in sequence according to the first control timing in the target control mode, control each input register to shift in the first direction, and shift the first element value to the input register corresponding to the first element value as the first input value; or write each first element value into each input register at the second specified position in each target multiply-accumulate array in sequence according to the first control timing, control each input register to shift in the second direction, and shift the first element value to the input register corresponding to the first element value as the first input value.

[0178] Step 5402: Based on the second tensor configuration information in the target control mode, determine the second input value corresponding to each first input value in the second tensor.

[0179] Step 5403: Control the target multiply-accumulate unit to perform a multiply-accumulate operation on the first input value in the input register and the corresponding second input value to obtain a multiply-accumulate operation result.

[0180] Step 5404: Based on the multiply-accumulate operation results of the target multiply-accumulate units in each target multiply-accumulate array, determine the tensor calculation result.

[0181] In some alternative embodiments, any multiply-accumulate array includes an input register, a multiply-accumulate unit, and an output register. The input register is configured to store a first input value of at least one pair of input values. The multiply-accumulate unit is configured to perform a multiply-accumulate calculation on at least one pair of input values to obtain a multiply-accumulate result of the at least one pair of input values. The output register is configured to store the multiply-accumulate result obtained by the multiply-accumulate unit performing the multiply-accumulate calculation.

[0182] In some alternative embodiments, the output register is a shift register; the number of target multiply-accumulate arrays is multiple; and the output registers of each multiply-accumulate array are electrically connected.

[0183] The method according to an embodiment of the present disclosure further includes: after writing the multiply-accumulate result obtained by the multiply-accumulate unit in the target multiply-accumulate array into the output register, controlling each output register to shift and output in a preset direction.

[0184] Figure 16 It is a schematic flowchart of a tensor calculation method provided by another exemplary embodiment of the present disclosure.

[0185] In some alternative embodiments, based on any of the above embodiments, the operation type is a convolution operation; the multiple multiply-accumulate arrays include multiply-accumulate arrays arranged in m rows and n columns; any multiply-accumulate array includes an input register and a multiply-accumulate unit; the input registers are electrically connected in a preset arrangement; the target multiply-accumulate array includes a target multiply-accumulate array of s rows and t columns in the multiply-accumulate arrays of m rows and n columns; s is less than or equal to m, and t is less than or equal to n.

[0186] As Figure 16 shown, step 540 of controlling the target multiply-accumulate array to perform an operation corresponding to the operation type on the first tensor and the second tensor based on the target control method to obtain a tensor calculation result may include:

[0187] Step 54a0, iteratively execute: according to the second control timing in the target control method, control each multiply-accumulate unit in each target multiply-accumulate array to perform a multiply-accumulate operation on the current first input value in the input register and the corresponding current second input value in the second tensor to obtain a first multiply-accumulate result of the current cycle, and control each input register in the target multiply-accumulate array to shift according to the shift direction corresponding to the current cycle in the target control method, so as to shift the first input value in one input register to another input register.

[0188] Step 54b0, in response to the shift times of each input register satisfying an end condition, determine the tensor calculation result based on the first multiply-accumulate results of each cycle.

[0189] For the beneficial technical effects corresponding to the exemplary embodiments of the present method, reference may be made to the corresponding beneficial technical effects in the above exemplary device part, which will not be elaborated herein.

[0190] Any tensor calculation method provided by an embodiment of the present disclosure can be executed by any suitable electronic device with data processing capabilities, including but not limited to: electronic devices such as terminal devices and servers. Alternatively, any tensor calculation method provided by an embodiment of the present disclosure can be executed by a processor. For example, the processor executes any tensor calculation method mentioned in an embodiment of the present disclosure by calling corresponding instructions stored in a memory. This will not be elaborated further below.

[0191] Exemplary electronic device

[0192] Figure 17 FIG. is a structural diagram of an electronic device provided by an embodiment of the present disclosure, including at least one processor 91 and a memory (which can be referred to as a second memory) 92.

[0193] The processor 91 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 90 to execute desired functions.

[0194] The memory 92 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 91 can run one or more computer program instructions to implement the methods of various embodiments of the present disclosure above and / or other desired functions.

[0195] In one example, the electronic device 90 may further include: an input device 93 and an output device 94, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0196] The input device 93 may further include, for example, a touch screen, a microphone, various sensors, and so on. The sensors may include, for example, an image sensor (such as a camera, a webcam, etc.), a lidar, a millimeter-wave radar, an ultrasonic radar, a positioning sensor, a pressure sensor, an air quality sensor, a temperature sensor, etc. The image sensor, lidar, millimeter-wave radar, ultrasonic radar, etc. can be used for the perception of the surrounding environment, that is, to detect the static and dynamic objects in the surrounding environment. The static and dynamic objects may include, for example, static objects such as lane lines, curbs, arrows, signs, trees, buildings, etc., and dynamic objects such as surrounding vehicles, pedestrians, cyclists, etc. The positioning sensor is used to realize the positioning of the movable device where the electronic device is located (such as a vehicle, a robot, etc.). The positioning sensor may include, for example, an Inertial Measurement Unit (IMU for short), a Global Positioning System (GPS for short), etc. The pressure sensor can be used to detect the seat pressure. The temperature sensor can be used to detect the temperature inside the vehicle cockpit. The air quality sensor can be used to detect the air quality inside the vehicle cockpit.

[0197] The output device 94 can output various information to the outside, which may include, for example, a display, a speaker, a communication network, and the remote output devices connected thereto, and so on.

[0198] Of course, for simplicity, Figure 17 only some of the components related to the present disclosure in the electronic device 90 are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device 90 may further include any other appropriate components.

[0199] In some optional embodiments, the electronic device 90 may include the tensor calculation device provided in any of the above embodiments.

[0200] The embodiments of the present disclosure further provide a chip, which may include the tensor calculation device 20 provided in any of the above embodiments.

[0201] Exemplary computer program product and computer-readable storage medium

[0202] In addition to the above methods and devices, the embodiments of the present disclosure may further provide a computer program product, including computer program instructions, which when run by a processor cause the processor to execute the steps in the methods of various embodiments of the present disclosure described in the above "Exemplary Method" section.

[0203] A computer program product may write program code for performing the operations of the embodiments of the present disclosure in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on a user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0204] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the methods of various embodiments of the present disclosure described in the above "Exemplary Method" section.

[0205] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium, for example but not limited to, includes systems, devices or components of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0206] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that they are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for the purposes of illustration and easy understanding, rather than limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0207] Those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure is also intended to include these changes and modifications.

Claims

1. A tensor calculation device, comprising: a memory configured to store a first tensor and a second tensor to be calculated; a controller configured to read the first tensor and the second tensor from the memory; and determine an operation type of the first tensor and the second tensor; a calculation component including a plurality of multiply-accumulate arrays; any one of the multiply-accumulate arrays is used to determine a multiply-accumulate result of at least a pair of input values; any pair of the input values includes a first input value and a second input value; the first input value is a value in the first tensor; the second input value is a value in the second tensor; the controller is further configured to: based on the first tensor, the second tensor, and the operation type, determine a target multiply-accumulate array for a current operation from the plurality of multiply-accumulate arrays of the calculation component, and determine a target control mode corresponding to the target multiply-accumulate array; based on the target control mode, control the target multiply-accumulate array to perform an operation corresponding to the operation type on the first tensor and the second tensor to obtain a tensor calculation result.

2. The device according to claim 1, wherein, The controller is configured to: based on first dimension information of the first tensor, determine quantity information of the target multiply-accumulate array corresponding to the current operation; based on the quantity information, determine the target multiply-accumulate array from the plurality of multiply-accumulate arrays; based on the first tensor, the second tensor, and the operation type, determine the target control mode corresponding to the target multiply-accumulate array.

3. The device according to claim 2, wherein The plurality of multiply-accumulate arrays include multiply-accumulate arrays of m rows and n columns; the first dimension information includes a first height, a first width, and a first depth of the first tensor; the controller is specifically configured to: in response to the first dimension information satisfying a preset dimension condition, based on the first height, the first width, and the first depth, determine a target number of rows and a target number of columns of the target multiply-accumulate array corresponding to the current operation as the quantity information; or, in response to the first height in the first dimension information not satisfying the preset dimension condition, determine m as the target number of rows of the target multiply-accumulate array, determine the first width as the target number of columns of the target multiply-accumulate array, and determine the target number of rows and the target number of columns as the quantity information; or, in response to the first width in the first dimension information not satisfying the preset dimension condition, determine n as the target number of columns of the target multiply-accumulate array, determine the first height as the target number of rows of the target multiply-accumulate array, and determine the target number of rows and the target number of columns as the quantity information.

4. The device according to claim 2, wherein The controller is specifically configured to: in response to the first dimension information not satisfying the preset dimension condition, based on the quantity information of the target multiply-accumulate array, determine the first tensor as at least one first sub-tensor; according to each of the first sub-tensors, the second tensor, and the operation type, determine the target control mode corresponding to the target multiply-accumulate array.

5. The device according to claim 1, wherein Any one of the multiply-accumulate arrays includes an input register and a multiply-accumulate unit; the input register is configured to store at least the first input value; the controller is specifically configured to: Write each first element value in the first tensor into the input register in the corresponding target multiply-accumulate array according to the first control timing in the target control mode as the first input value corresponding to the target multiply-accumulate unit in the target multiply-accumulate array; Determine the second input value corresponding to each first input value in the second tensor based on the second tensor configuration information in the target control mode; Control the target multiply-accumulate unit to perform a multiply-accumulate operation on the first input value in the input register and the corresponding second input value to obtain a multiply-accumulate operation result; Determine the tensor calculation result based on the multiply-accumulate operation results of the target multiply-accumulate units in each target multiply-accumulate array.

6. The apparatus according to claim 1, wherein The number of target multiply-accumulate arrays is multiple; any one of the multiply-accumulate arrays includes an input register and a multiply-accumulate unit; The input register is configured to store at least the first input value; the input register is a shift register; the input registers of each multiply-accumulate array are electrically connected according to a preset arrangement; The controller is specifically configured as: Write each first element value in the first tensor into the input registers corresponding to the first specified positions in each target multiply-accumulate array in sequence according to the first control timing in the target control mode, control each input register to shift in the first direction, and shift the first element value to the input register corresponding to the first element value; Or, Write each first element value into the input registers at the second specified positions in each target multiply-accumulate array in sequence according to the first control timing, control each input register to shift in the second direction, and shift the first element value to the input register corresponding to the first element value.

7. The device according to claim 1, wherein Any one of the multiply-accumulate arrays includes an input register, a multiply-accumulate unit, and an output register; The input register is configured to store the first input value in the at least one pair of input values; The multiply-accumulate unit is configured to perform a multiply-accumulate calculation on the at least one pair of input values to obtain the multiply-accumulate result of the at least one pair of input values; The output register is configured to store the multiply-accumulate result obtained by the multiply-accumulate unit performing the multiply-accumulate calculation.

8. The apparatus according to claim 7, wherein The output register is a shift register; the number of target multiply-accumulate arrays is multiple; The output registers of each multiply-accumulate array are electrically connected; The controller is further configured to: after the multiply-accumulate result obtained by the multiply-accumulate unit in the target multiply-accumulate array is written into the output register, control each output register to shift and output in a preset direction.

9. The device according to claim 1, wherein The operation type is a convolution operation; The multiple multiply-accumulate arrays include a multiply-accumulate array of m rows and n columns; any one of the multiply-accumulate arrays includes an input register and a multiply-accumulate unit; the input registers are electrically connected according to a preset arrangement; The target multiply-accumulate array includes a multiply-accumulate array of s rows and t columns; s is less than or equal to m, and t is less than or equal to n; The controller is specifically configured as: Iterative execution: According to the second control timing in the target control mode, control each multiplication and addition unit in each target multiplication and addition array to perform a multiplication and addition operation on the current first input value in the input register and the corresponding current second input value in the second tensor, to obtain the first multiplication and addition result of the current cycle. According to the shift direction corresponding to the current cycle in the target control mode, control the input registers in the target multiplication and addition array to shift, so as to shift the first input value in one input register to another input register; In response to the shift times of each input register satisfying the end condition, determine the tensor calculation result based on the first multiplication and addition results obtained in each iteration.

10. The device according to any one of claims 1-9, wherein, The controller is specifically configured to: Based on the first tensor, the second tensor, the operation type, and the state information of each multiplication and addition array in the computing component, determine the target multiplication and addition array for the current operation from the multiplication and addition arrays with the state information being the idle state among the multiple multiplication and addition arrays.

11. A tensor calculation method, comprising: Obtain a first tensor and a second tensor to be calculated; Determine the operation type of the first tensor and the second tensor; Based on the first tensor, the second tensor, and the operation type, determine a target multiplication and addition array for the current operation from multiple multiplication and addition arrays of a computing component, and determine a target control mode corresponding to the target multiplication and addition array; wherein, any one of the multiplication and addition arrays is used to determine the multiplication and addition result of at least one pair of input values; any pair of the input values includes a first input value and a second input value; the first input value is a value in the first tensor; the second input value is a value in the second tensor; Based on the target control mode, control the target multiplication and addition array to perform the operation corresponding to the operation type on the first tensor and the second tensor, to obtain a tensor calculation result.

12. A chip, comprising: The tensor calculation device according to any one of claims 1-10.

13. A computer-readable storage medium, the storage medium stores a computer program, and the computer program is used to execute the method described in claim 11 above.

14. An electronic device, the electronic device includes: A processor; A memory for storing executable instructions of the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the method described in claim 11 above; or, The electronic device includes the tensor calculation device according to any one of claims 1-10.

Citation Information

Cited By

  • Output resident multiply-add array device and related equipment

    CN121433611A