Data processing device and electronic equipment
By reordering and reconstructing the data flow in the pulse array, the problems of hardware resource consumption and startup delay caused by the delay unit are solved, and more efficient matrix multiplication calculation is achieved.
Patent Information
- Application Number
- CN202511166135.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-20
AI Technical Summary
In pulse array implementations, the extensive use of delay units leads to severe hardware resource consumption and excessive startup latency, limiting its deployment and computational efficiency in resource-constrained devices.
By reordering the second initial matrix, a second matrix to be processed is generated, and the data flow between computing units in the pulse array is redesigned to eliminate the use of delay units and ensure the correctness of data timing.
It reduces hardware resource consumption and power consumption, lowers startup latency, and improves the feasibility and computational efficiency of deploying pulse arrays in resource-constrained devices.
Smart Images

Figure CN120670374B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, in particular to a data processing device and electronic equipment. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, deep learning models have been widely applied in image recognition, natural language processing, speech recognition and other fields. In these application scenarios, matrix multiplication calculation, as the core computing operation of deep learning models, directly affects the performance of the entire system.
[0003] In order to accelerate matrix multiplication calculation, the industry generally uses special hardware accelerators to replace general-purpose processors for calculation. As a classic parallel computing architecture, pulse array has been widely used in the design of matrix multiplication accelerator. However, in the related technology pulse array implementation, a large number of delay units need to be added at the data input end to ensure the correctness of the data timing. The large use of these delay units brings the problem of hardware resource consumption, especially in resource-constrained devices, the additional hardware overhead seriously limits the actual deployment and application effect of the pulse array. SUMMARY
[0004] The present application provides a data processing device and electronic equipment, which can solve the problem of large use of delay units and hardware resource consumption.
[0005] According to an aspect of the present application, a data device is provided, comprising: a cache unit configured to store a first to-be-processed matrix and a second to-be-processed matrix, the first to-be-processed matrix comprising a plurality of first to-be-processed elements, and the second to-be-processed matrix comprising a plurality of to-be-processed vector data, the to-be-processed vector data comprising a plurality of second to-be-processed elements; a pulse array comprising a plurality of pulse sub-arrays, each pulse sub-array comprising a plurality of computing units, each computing unit configured to: determine a target computing result of the computing unit according to a received first to-be-processed element and a second to-be-processed element corresponding to the computing unit; and provide the first to-be-processed element to a subsequent hardware unit corresponding to the computing unit in a target external pulse sub-array, the target external pulse sub-array being determined by cyclically moving a first moving number of positions in a first direction from a source pulse sub-array in which the computing unit is located among the plurality of pulse sub-arrays; wherein the second to-be-processed matrix comprises a plurality of to-be-processed vector data, and the to-be-processed vector data in the second to-be-processed matrix is obtained by: cyclically moving a second moving number of positions in a second direction for a plurality of elements in initial vector data of a second initial matrix, the second moving number of positions being determined according to position information of the initial vector data in the second initial matrix and a preset value.
[0006] According to another aspect of the present application, an electronic device is provided, comprising: a memory; and the data processing device provided by the present application.
[0007] The data processing apparatus and the electronic device provided by the embodiment of the present application obtain the second to-be-processed matrix by cyclically moving a plurality of elements in the initial vector data of the second initial matrix along the second direction by the second moving bit number, realize the preprocessing rearrangement of the input data, and make the rearranged second to-be-processed matrix match the data processing timing sequence of the calculation unit in the pulse array. The second moving bit number is determined according to the position information of the initial vector data in the second initial matrix and a preset value. Therefore, the second to-be-processed matrix obtained after the data rearrangement ensures that each second to-be-processed element can reach the corresponding calculation unit in the correct clock cycle, thereby eliminating the delay unit that must be set in the pulse array of the related art to ensure the correctness of the data timing.
[0008] Meanwhile, by redesigning the data flow direction between the calculation units in the pulse array, the calculation unit can provide the first to-be-processed element to the subsequent hardware unit for the calculation unit in the target external pulse sub-array, wherein the target external pulse sub-array is determined by cyclically moving the first moving bit number along the first direction from the source pulse sub-array. The reconstructed data propagation path cooperates with the rearranged second to-be-processed matrix, avoids setting the delay unit, and ensures the correctness of the matrix multiplication calculation.
[0009] Therefore, the embodiment of the present application ensures the calculation correctness, reduces the hardware resources required for the implementation of the pulse array, especially eliminates the use of a large number of delay units, reduces the hardware cost and power consumption, improves the deployment feasibility and application effect of the pulse array in the resource-limited device, and at the same time, due to the reduction of the waiting time of the data in the delay unit, the start-up delay of the entire matrix multiplication calculation is also reduced, and the calculation efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above content and other purposes, features and advantages of the present application will be more clearly understood through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:
[0011] FIG. 1A is a data flow schematic diagram of the pulse array in the related art in the 0th clock cycle;
[0012] FIG. 1B is a data flow schematic diagram of the pulse array in the related art in the 1st clock cycle;
[0013] FIG. 1C is a data flow schematic diagram of the pulse array in the related art in the 2nd clock cycle;
[0014] FIG. 1D is a data flow schematic diagram of the pulse array in the related art in the 3rd clock cycle;
[0015] FIG. 2A structural schematic diagram of a data processing device provided by an embodiment of the present application is shown in FIG. 1.
[0016] FIG. 3 A schematic diagram of a second initial matrix data reordering operation provided by an embodiment of the present application is shown in FIG. 4.
[0017] FIG. 4A A data flow schematic diagram of a pulse array in a first clock cycle provided by an embodiment of the present application is shown in FIG. 5.
[0018] FIG. 4B A data flow schematic diagram of a pulse array in a second clock cycle provided by an embodiment of the present application is shown in FIG. 6.
[0019] FIG. 4C A data flow schematic diagram of a pulse array in a third clock cycle provided by an embodiment of the present application is shown in FIG. 7.
[0020] FIG. 5 A structural schematic diagram of another data processing device provided by an embodiment of the present application is shown in FIG. 8.
[0021] FIG. 6 A structural schematic diagram of still another data processing device provided by an embodiment of the present application is shown in FIG. 9.
[0022] FIG. 7 A structural schematic diagram of a computing unit provided by an embodiment of the present application is shown in FIG. 10.
[0023] FIG. 8 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 11. DETAILED DESCRIPTION
[0024] Embodiments of the present application will be described below with reference to the accompanying drawings. It should be understood, however, that the description that follows is merely exemplary and is not intended to limit the scope of the application. In the following detailed description of embodiments of the present application, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, it will be apparent to one skilled in the art that one or more embodiments of the present application can be practiced without these specific details. In other instances, well-known structures and functions have not been described in detail in order to avoid obscuring aspects of the present application.
[0025] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present application. As used herein, the term "includes" and tautological expressions thereof, such as "including," "includes," "include," "contains," "containing," and so forth, shall be read expansively and without limitation. The terms "comprising," "comprise" and / or "comprises" and tautological expressions thereof (for example, "comprising a," "comprises an," etc.) shall be interpreted as per the term "including" and tautological expressions thereof.
[0026] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0027] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0028] Before introducing the embodiments of the present invention, the implementation methods of pulse arrays in related technologies will be introduced first.
[0029] Please refer to FIGS. 1A-1D , FIGS. 1A-1D This is a schematic diagram of the data flow of a pulse array in a related technology.
[0030] Please refer to FIG. 1A , FIG. 1A This is a schematic diagram of the data flow of a pulse array in the 0th clock cycle in related technologies; for example... FIG. 1A As shown, the pulse array in the related technology adopts a regular two-dimensional grid topology, where each processing unit 110 can transmit data with its adjacent processing units 110 through a fixed data channel. Taking the matrix multiplication operation between matrix A 120 and matrix B 130 using the pulse array as an example, based on the data channels between different processing units 110, the data flow of matrix A 120 can be: input from the left boundary of the pulse array, passed sequentially from left to right among the processing units 110 in the same row through the horizontal data channel, and each processing unit 110, after receiving an element of matrix A 120, will pass that element to the next processing unit 110 on the right side of the same row. Simultaneously, the data flow of matrix B 130 can be: input from the top boundary of the pulse array, passed sequentially from top to bottom among the processing units 110 in the same column through the vertical data channel, so that all processing units 110 in that column can receive each element in a column of matrix B 130. It can be understood that... FIGS. 1A-1D The positions of each element in matrix A120 and matrix B130 are determined in order to perform correct matrix multiplication calculations based on multiple elements of matrices A and B respectively.
[0031] During the calculation process, each processing unit 110 receives elements of matrix A 120 from the left and elements of matrix B 130 from the top. After performing multiplication and accumulation calculations, the calculation results are transmitted to the adjacent processing unit 110 below through a vertically downward data channel.
[0032] Based on the data flow of the A matrix 120 and the data flow of the B matrix 130, one or more delay units 140 can be set for the systolic array to achieve that the different column data of the B matrix 130 need to be input into the systolic array at different time points, so as to ensure that the corresponding elements of the A matrix 120 and the B matrix 130 can be calculated at the correct processing unit 110 and at the correct time point.
[0033] Please refer to FIG. 1B , FIG. 1B for the data flow diagram of the systolic array in the related art in the first clock cycle; as FIG. 1B shown, in the first clock cycle, the systolic array in the related art starts to perform the first step operation of the matrix multiplication calculation. Specifically, the element a11 of the first row and the first column of the A matrix 120 is input from the left boundary of the systolic array, and at the same time, reaches the processing unit 110 of the first row and the first column. At this time, the processing unit 110 receives the pre-loaded element b11 of the B matrix 130 and the corresponding element c11 of the C matrix, performs the multiplication and accumulation operation, and obtains the calculation result a11x b11 + c11 of the processing unit 110 of the first row and the first column in the first clock cycle.
[0034] Due to the effect of the delay unit 140, in the first clock cycle, the second column data b12, b22, b32 and the third column data b13, b23, b33 of the B matrix 130 have not been input into the corresponding processing unit 110. It is ensured that the element a11 of the A matrix 120 can be calculated at the correct time point with the corresponding element of the B matrix 130, but at the same time, it also means that the parallelism of the systolic array in the initial stage is limited, and only the processing unit 110 of the first column participates in the actual multiplication and accumulation calculation.
[0035] Please refer to FIG. 1C , FIG. 1C for the data flow diagram of the systolic array in the related art in the second clock cycle; as FIG. 1CAs shown, in the second clock cycle, element a21 of the second row and first column of matrix A 120 is input from the left boundary of the pulse array and reaches processing unit 110 in the first row and first column. It performs a multiplication-accumulation operation with the pre-loaded element b11 of matrix B 130 and the corresponding element c21 of matrix C, obtaining the calculation result a21×b11+c21 of processing unit 110 in the second clock cycle. Simultaneously, in processing unit 110 in the first row and second column, due to the control of delay unit 140, data b12 of the second column of matrix B 130 begins to be input in this clock cycle. This processing unit 110 receives element a11 propagated from the left and the newly input element b12, and together with the corresponding element c12 of matrix C, performs a multiplication-accumulation operation, obtaining the calculation result a11×b12+c12 of processing unit 110 in the first row and second column in the second clock cycle.
[0036] The processing unit 110 in the first row and first column can propagate the calculation result a11×b11+c11 of the first clock cycle down to the processing unit 110 in the second row and first column. The processing unit 110 in the second row and first column performs a multiply-accumulate calculation based on the received calculation result a11×b11+c11 propagated from above, the element a12 input from the left, and the pre-loaded element b21 of the B matrix 130, to obtain the calculation result a11×b11+c11+a12×b21 of the processing unit 110 in the second clock cycle.
[0037] Please refer to FIG. 1D , FIG. 1D This is a schematic diagram of the data flow of a pulse array in the third clock cycle in related technologies; for example... FIG. 1D As shown, in the third clock cycle, element a31 of the third row and first column of matrix A 120 is input from the left boundary of the pulse array and reaches processing unit 110 in the first row and first column. It performs a multiplication-accumulation operation with the pre-loaded element b11 of matrix B 130 and the corresponding element c31 of matrix C, obtaining the calculation result a31×b11+c31 of processing unit 110 in the third clock cycle. It can be understood that the method of obtaining the calculation result a21×b12+c22 of processing unit 110 in the third clock cycle is the same as or similar to the method of obtaining the calculation result a11×b12+c12 of processing unit 110 in the second clock cycle, and will not be elaborated further here.
[0038] For the processing unit 110 in the first row and the third column, due to the delay control of 2 clock cycles provided by the delay unit 140, the data b13 in the third column of the B matrix 130 starts to be input. The processing unit 110 in the first row and the third column receives the element a11 propagated from the left side and the newly input element b13, as well as the corresponding element c13 of the C matrix, performs a multiply-accumulate operation, and obtains the calculation result a11×b13+c13 of the processing unit 110 in the first row and the third column at the third clock cycle.
[0039] It can be understood that the way of obtaining the calculation result a21×b11+c21+a22×b21 of the processing unit 110 in the second row and the first column at the third clock cycle and the way of obtaining the calculation result a11×b12+c12+a12×b22 of the processing unit 110 in the second row and the second column at the third clock cycle are the same as or similar to the way of obtaining the calculation result a11×b11+c11+a21×b21 of the processing unit 110 in the second row and the first column at the second clock cycle, and details are not repeated herein.
[0040] The processing unit 110 in the second row and the first column can propagate the calculation result a11×b11+c11+a21×b21 at the second clock cycle to the processing unit 110 in the third row and the first column. The processing unit 110 in the third row and the first column performs a multiply-accumulate operation according to the calculation result a11×b11+c11+a21×b21 propagated from the upper side, the element a13 input from the left side, and the element b31 of the B matrix 130 pre-loaded, and obtains the calculation result a11×b11+c11+a21×b21+a13×b31 of the processing unit 110 in the third row and the first column at the third clock cycle.
[0041] It is worth noting that at the third clock cycle, the pulse array has not yet reached a fully parallel state, and some processing units 110 have not yet received valid input data, so they are still in a waiting state, which reflects the influence of the delay unit 140 on the overall calculation efficiency of the pulse array in the related art during the start-up stage.
[0042] Based on the above analysis of the related art, it can be found that the implementation manner of the pulse array in the related art has technical problems of serious hardware resource consumption and high start-up delay, which will be described below.
[0043] The use of a large number of delay units 140 leads to serious hardware resource consumption. In order to realize the timing control of the data of different columns of the B matrix 130, the pulse array in the related art needs to deploy a large number of delay units 140 on the data input path, and each delay unit 140 needs to occupy a memory device, a clock control circuit and corresponding control logic. In a resource-constrained embedded device or edge computing device, the delay unit 140 occupies a large amount of valuable chip area and power consumption budget, and competes with the computing unit or processing unit 110 in the allocation of hardware resources, resulting in low chip area utilization efficiency. Under the same hardware resource constraint, the existence of the delay unit 140 limits the number of deployable processing units 110, thereby restricting the parallel computing capability and overall performance of the pulse array.
[0044] Meanwhile, the existence of the delay unit 140 increases the start-up delay of the pulse array. Since the data of different columns of the B matrix 130 needs to pass through different numbers of delay units 140 to reach the corresponding processing unit 110, the entire pulse array must wait for the data on the longest delay path to be ready before starting effective calculation. This waiting time linearly increases with the increase of the size of the pulse array, resulting in idle of computing resources in the start-up stage, and reducing the overall computing efficiency. Especially in application scenarios that require frequent start-up of matrix multiplication calculation, the cumulative effect of start-up delay has a significant impact on system performance, limiting the actual deployment effect of the pulse array in a resource-constrained environment.
[0045] Based on the above problems, an embodiment of the present application provides a data processing apparatus.
[0046] FIG. 2 A structural schematic diagram of a data processing apparatus provided by an embodiment of the present application.
[0047] As shown in FIG. 2 The data processing apparatus 200 includes a cache unit 210 and a pulse array 220.
[0048] The cache unit 210 is configured to store a first to-be-processed matrix and a second to-be-processed matrix. The first to-be-processed matrix includes a plurality of first to-be-processed elements, and the second to-be-processed matrix includes a plurality of to-be-processed vector data. The to-be-processed vector data can include a plurality of second to-be-processed elements.
[0049] The pulse array 220 includes a plurality of pulse sub-arrays, and each pulse sub-array includes a plurality of computing units. As shown in FIG. 2As shown, the pulse array 220 can include a pulse sub-array 221, a pulse sub-array 222, and a pulse sub-array 223. The pulse sub-arrays can include a plurality of computing units. For example, the pulse sub-array 221 can include a computing unit 2211 and a computing unit 2212. The pulse sub-array 222 can include a computing unit 2221 and a computing unit 2222. The pulse sub-array 223 can include a computing unit 2231 and a computing unit 2232.
[0050] One or more of the plurality of computing units is configured to determine a target computing result of the computing unit according to the received first to-be-processed element and a second to-be-processed element for the computing unit, and provide the first to-be-processed element to a subsequent hardware unit for the computing unit in a target external pulse sub-array, which is determined by moving a first moving bit number in a first direction from a source pulse sub-array where the computing unit is located among the plurality of pulse sub-arrays.
[0051] The to-be-processed vector data in the second to-be-processed matrix is obtained by moving a plurality of elements in initial vector data of the second initial matrix in a second direction by a second moving bit number, which is determined according to position information of the initial vector data in the second initial matrix and a preset value.
[0052] For the cache unit 210 in the embodiment of the present application, it can be a high-speed storage device located between the processor and the pulse array 220, and in the embodiment of the present application, it can be understood as a storage module on a chip with fast read and write capability, which can temporarily store input data and intermediate result data participating in matrix multiplication calculation.
[0053] For example, the cache unit 210 can be a static random-access memory (SRAM), a block random access memory (BRAM), or other types of on-chip cache storage devices.
[0054] In the embodiment of the present application, the cache unit 210 can store the first to-be-processed matrix and the second to-be-processed matrix. The first to-be-processed matrix refers to a matrix that serves as a multiplicand in matrix multiplication calculation, which contains a plurality of first to-be-processed elements, each of which can be a numerical value at a position in the matrix, which is propagated through a data channel of the pulse array 220 in the calculation process. For example, the first to-be-processed matrix can be the A matrix.
[0055] The second to-be-processed matrix refers to a matrix used as a multiplier in matrix multiplication calculation, which is stored in the cache unit 210 after a pre-processing rearrangement operation and contains a plurality of second to-be-processed elements, each of which represents a value at a position in the rearranged matrix. In the matrix multiplication calculation process, the first to-be-processed element and the second to-be-processed element are specific values in different matrices participating in the multiplication and accumulation calculation, and the matrix operation is completed through the cooperative work of the calculation units in the systolic array 220.
[0056] To solve the problem of excessive hardware resource consumption and high start-up delay caused by the large number of delay units required by the systolic array 220 in the related art, an embodiment of the present application obtains a second to-be-processed matrix by performing a rearrangement operation on a second initial matrix. For example, the second initial matrix can be the B matrix described above.
[0057] For the second initial matrix in the embodiment of the present application, it can be an original input matrix used as a multiplier in matrix multiplication calculation, which can be understood in the embodiment of the present application as a weight matrix or a parameter matrix of a deep learning model, used for multiplication operation with the first to-be-processed matrix. The second initial matrix organizes data in a row-column storage manner, where each initial vector data can be a column of the matrix, containing elements at different row positions in the column.
[0058] The second moving bit number refers to the step value of the cyclic movement of the elements in the initial vector data in the second direction during data rearrangement. The second moving bit number can be determined based on the position information of the initial vector data in the second initial matrix and a preset value. The position information of the initial vector data in the second initial matrix can refer to the column serial number of the initial vector data in the matrix, used to represent the position of the data in the original matrix structure. The column serial number is counted from 1, and the position information corresponding to the initial vector data in the i-th column can be i.
[0059] For example, the second moving bit number can be the column serial number of the initial vector data minus the preset value. Different initial vector data will obtain different moving bit numbers, thereby realizing the rearrangement of the entire matrix data distribution, so that the rearranged data can match the data processing timing of the systolic array 220.
[0060] Correspondingly, the preset value can be set to 1 in the embodiment of the present application as a reference value for calculating the second moving bit number.
[0061] For example, the i-th initial vector data can be the i-th column of the second initial matrix, and the second moving bit number of the i-th initial vector data can be i-1. That is, in the second initial matrix, the moving bit number of the data in the first column can be 0 (no movement), the moving bit number of the data in the second column can be 1, the moving bit number of the data in the third column can be 2, and so on. i can be an integer greater than or equal to 1.
[0062] In combination FIG. 2 As shown in FIG. 2, for the pulse sub-array in the embodiment of the present application, a plurality of computing units in a column of the pulse array 220 can be included.
[0063] The organization structure of the pulse sub-array has the following characteristics: the jth pulse sub-array processes the data elements in the jth column of the first to-be-processed matrix. j can be an integer greater than or equal to 1. The computing units in the same pulse sub-array can receive the first to-be-processed elements, but process the second to-be-processed elements from different columns of the second to-be-processed matrix, respectively. Through this organization manner, the pulse sub-array can receive and process a plurality of first to-be-processed elements of the first to-be-processed matrix, and realize the multiplication and accumulation calculation of the plurality of first to-be-processed elements and the corresponding elements of the second to-be-processed matrix.
[0064] For the computing unit in the embodiment of the present application, it refers to a basic processing module in the pulse sub-array for performing a specific multiplication and accumulation operation. In the embodiment of the present application, the computing unit can be understood as a special hardware unit integrating a multiplier, an accumulator, and data transmission control logic. One or more computing units can be used to determine an element in a position of a result matrix of matrix multiplication, can receive a first to-be-processed element and a second to-be-processed element, perform a multiplication and accumulation calculation operation, and transmit the elements received in the calculation process and the intermediate calculation results generated. In addition, the pulse sub-array where the computing unit is located can be the source pulse sub-array of the computing unit. For example, the source pulse sub-array of the computing unit 2211 can be the pulse sub-array 221. The source pulse sub-array of the computing unit 2222 can be the pulse sub-array 222. The source pulse sub-array of the computing unit 2231 can be the pulse sub-array 223.
[0065] Different from the first to-be-processed elements only horizontally propagating within the same row in the related art, the computing unit in the embodiment of the present application adopts a reconstructed data propagation path. In the transmission process of the first to-be-processed element, the computing unit can continue to propagate the received first to-be-processed element downward for use by subsequent computing units, but the propagation manner is changed: the computing unit provides the first to-be-processed element to the subsequent hardware unit in the target external pulse sub-array for the computing unit. The above data transmission manner across the pulse sub-arrays matches the rearranged second to-be-processed matrix, and ensures the timing correctness of the matrix multiplication calculation.
[0066] The target external pulse subarray in this embodiment of the invention can be the target pulse subarray to which the current computing unit needs to send the first element to be processed during data transmission. The target external pulse subarray can be determined based on the position of the computing unit in the source pulse subarray and predefined movement rules. In this way, a data propagation topology structure different from that of pulse arrays in related technologies is established, which can receive the first element to be processed transmitted from the computing unit in the source pulse subarray and realize cross-pulse subarray data flow control.
[0067] The target external pulse subarray of the computing unit can be determined by cyclically moving the source pulse subarray along a first direction by a first number of moves. In this embodiment of the invention, the first direction refers to the direction of movement when indexing positions between pulse subarrays in the pulse array 220. In this embodiment, it can be understood as a directional reference for determining the target position of data propagation, typically pointing to the horizontal direction of the pulse array 220. Specifically, it manifests as left-right position movement within an array structure composed of multiple pulse subarrays, used to instruct the computing unit 2211 to transfer the first element to be processed to the correct target external pulse subarray. For example, when rearranging the second initial matrix, if the upward direction from the second row to the first row is taken as the second direction, the first direction can be the rightward direction from the second column to the first column.
[0068] The first shift number in this embodiment of the invention is related to the column position of the calculation unit 2211 in the source pulse subarray, ensuring the correctness of the matrix multiplication operation. The first shift number can be preset, for example, and can be 1.
[0069] In this embodiment of the invention, the cyclic movement refers to the circular movement method used in determining the position of the pulse subarray. That is, when the source pulse subarray is the boundary of the pulse array 220, the target pulse subarray can be determined from the other end of the array based on the cyclic movement method. For example, taking a pulse array 220 comprising M pulse subarrays, cyclic movement along a first direction being to the right, and the first movement position being 1 as an example, if the source pulse subarray is the first column of the pulse array 220, the target external pulse subarray can be the Mth column of the pulse array. M can be an integer greater than 1. FIG. 2 As shown, for pulse subarray 221, the target external pulse subarray determined by cyclically shifting 1 position along the first direction can be pulse subarray 223. For pulse subarray 222, the target external pulse subarray determined by cyclically shifting 1 position along the first direction can be pulse subarray 221. For pulse subarray 223, the target external pulse subarray determined by cyclically shifting 1 position along the first direction can be pulse subarray 222.
[0070] The combination operation of moving the first moving bit number in the first direction is performed, i.e., in the pulse array 220 containing a plurality of pulse sub-arrays, starting from the position index of the source pulse sub-array, moving the first moving bit number of positions in the first direction, and processing the boundary condition in a circular moving manner, to finally determine the specific position of the target external pulse sub-array.
[0071] For the post-hardware unit of the computing unit in the embodiment of the application, it refers to a hardware module located after the current computing unit in the data propagation link, used to receive the first to-be-processed element transmitted by the computing unit. In the embodiment of the application, it can be understood as a downstream receiving unit of a data stream, and the post-hardware unit of the computing unit is located in the target external pulse sub-array, which can receive the data output of the current computing unit. For example, as described above, the target external pulse sub-array of the pulse sub-array 221 can be the pulse sub-array 223. For the computing unit 2211 in the pulse sub-array 221, the post-hardware unit can be the computing unit 2232 in the pulse sub-array 223.
[0072] It should be noted that there is a correlation between the determination principle of moving the first moving bit number in the first direction and moving the second moving bit number in the second direction. By coordinating data preprocessing and propagation path reconstruction, the pulse array 220 calculation without delay units is realized. For example, the first direction and the second direction are perpendicular.
[0073] Specifically, since the to-be-processed vector data in the second to-be-processed matrix is obtained by moving the initial vector data of the second initial matrix in the second direction by the second moving bit number, the position of the second to-be-processed element is offset relative to the original position. The embodiment of the application obtains the second to-be-processed matrix by preprocessing and rearranging the second initial matrix, changes the distribution position of the second to-be-processed element in the matrix, and accordingly needs to adjust the propagation path of the first to-be-processed element in the pulse array 220 to ensure the correct matching of the data timing.
[0074] It can be understood that the above description of the application takes moving in the first direction to the right and moving in the second direction to the up as an example. However, the application is not limited thereto. In order to ensure that the first to-be-processed element can perform the multiply-accumulate calculation with the corresponding second to-be-processed element at the correct time, the data propagation path of the computing unit in the pulse array 220 is adapted to such position offset. Therefore, the corresponding relationship between the first moving bit number and the second moving bit number can be established to ensure that the rearranged second to-be-processed matrix and the reconstructed first to-be-processed element propagation path can realize timing synchronization, so as to eliminate the delay unit while ensuring the correctness of the matrix multiplication calculation.
[0075] In analogy to the implementation of the 3x3 pulse array in the related art, the element b11, the element b21, and the element b31 in the first column of the B matrix are directly input at the 0th clock cycle, the element b12, the element b22, and the element b32 in the second column need to be input after being delayed by 1 clock cycle, and the element b13, the element b23, and the element b33 in the third column need to be input after being delayed by 2 clock cycles. This time-sharing input manner needs to configure 3 delay units. When the first row data a11 of the A matrix is input from the left side, it needs to meet b11 at the 1st clock cycle, meet the delayed b12 at the 2nd clock cycle, and meet the delayed b13 at the 3rd clock cycle, and the effective calculation of the entire array needs to be delayed by 2 clock cycles.
[0076] In addition, the plurality of computing units in the pulse sub-array can respectively receive a plurality of second to-be-processed elements in a column of the second to-be-processed matrix. The data flow direction of the second to-be-processed elements can be a third direction. For example, the computing unit 2211 can provide the to-be-processed elements in the first column of the second to-be-processed matrix to the computing unit 2212. The third direction is the direction from the computing unit 2211 to the computing unit 2212.
[0077] In addition, in the pulse sub-array, the target calculation result generated by the preceding computing unit can be provided to the following computing unit. For example, the target calculation result of the computing unit 2211 can be provided to the computing unit 2212.
[0078] According to the embodiment of the application, by moving a plurality of elements in the initial vector data of the second initial matrix by a second moving bit number in a second direction to obtain the second to-be-processed matrix, the pre-processing and rearrangement of the input data are realized, so that the rearranged second to-be-processed matrix can match the data processing timing of the computing units in the pulse array 220. The second moving bit number is determined according to the position information of the initial vector data in the second initial matrix and a preset value. Therefore, the second to-be-processed matrix obtained after data rearrangement ensures that each second to-be-processed element can reach the corresponding computing unit at the correct clock cycle, thereby eliminating the delay unit that must be set in the pulse array in the related art to ensure the correctness of the data timing.
[0079] At the same time, by redesigning the data flow direction between the computing units in the pulse array 220, the computing unit can provide the first to-be-processed element to the following hardware unit in the target external pulse sub-array for the computing unit 2211, wherein the target external pulse sub-array is determined by moving a first moving bit number in a first direction from a source pulse sub-array. This reconstructed data propagation path cooperates with the rearranged second to-be-processed matrix, and ensures the correctness of the matrix multiplication calculation without setting a delay unit.
[0080] Therefore, the embodiments of the present application ensure the calculation correctness, reduce the hardware resources required for implementing the pulse array 220, especially eliminate the use of a large number of delay units, reduce the hardware cost and power consumption, improve the deployment feasibility and application effect of the pulse array 220 in the resource-limited device, and at the same time, due to the reduction of the waiting time of data in the delay unit, the start-up delay of the entire matrix multiplication calculation is also reduced, and the calculation efficiency is improved.
[0081] It can be understood that the data processing device 200 of the present application is described above, and the following will be described in combination with FIG. 3 The second to-be-processed matrix of the present application is further described.
[0082] FIG. 3 A schematic diagram of a second initial matrix data reordering operation provided by the embodiments of the present application.
[0083] As FIG. 3 shown, the original second initial matrix B310 contains 3 initial vector data, each of which can be a column data of the second initial matrix and contains 3 elements. The 1st initial vector data [b11, b21, b31] includes 3 elements, the 2nd initial vector data [b12, b22, b32] includes 3 elements, and the 3rd initial vector data [b13, b23, b33] includes 3 elements. According to the calculation rule of the second moving bit number, the second moving bit number of the 1st initial vector data is 0, the moving bit number of the 2nd initial vector data is 1, and the moving bit number of the 3rd initial vector data is 2.
[0084] For the 1st initial vector data, since the moving bit number is 0, the original order can be kept unchanged, and the 1st to-be-processed vector data obtained after rearrangement is still [b11, b21, b31];
[0085] For the 2nd initial vector data, the operation of moving 1 bit up in a cycle is performed, the element b12 is moved to the 3rd row after rearrangement, the element b22 is moved to the 1st row after rearrangement, and the element b32 is moved to the 2nd row after rearrangement, and the 2nd to-be-processed vector data obtained is [b22, b32, b12];
[0086] For the 3rd initial vector data, the operation of moving 2 bits up in a cycle is performed, the element b13 is moved to the 2nd row after rearrangement, the element b23 is moved to the 3rd row after rearrangement, and the element b33 is moved to the 1st row after rearrangement, and the 3rd to-be-processed vector data obtained is [b33, b13, b23].
[0087] After the rearrangement operation is completed, the 1st to-be-processed matrix A310 obtained can be as FIG. 3The second to-be-processed matrix B320 is shown, wherein the to-be-processed vector data of each column is the result of the corresponding initial vector data after a specific bit cyclic shift.
[0088] It can be understood that the second to-be-processed matrix of the present application is described above, and the following will be described in combination with FIGS. 4A-4C The matrix multiplication of the present application is further described.
[0089] FIGS. 4A-4C A data flow diagram of a pulse array provided for an embodiment of the present application is shown.
[0090] As FIGS. 4A-4C shown, the pulse array 220 can include a plurality of pulse sub-arrays. The pulse sub-array can include a head calculation unit, one or more intermediate calculation units, and an output unit. For example, the pulse sub-array 221 described above can include a calculation unit 2211, a calculation unit 2212, and a calculation unit 2213. The calculation unit 2211 can serve as a head calculation unit. The calculation unit 2212 can serve as an intermediate calculation unit. The calculation unit 2213 can serve as an output unit. The following will be described in combination with FIG. 4A The data flow of the pulse array in the first clock cycle of the current round is described.
[0091] In some embodiments, the head calculation unit can perform the following operation: multiplying the first to-be-processed element for the head calculation unit received from the cache unit and the second to-be-processed element for the head calculation unit to obtain the multiplication operation result of the head calculation unit. According to the multiplication operation result of the head calculation unit, the target calculation result of the head calculation unit can be determined. Before the current round of the pulse array, the calculation unit can pre-load the second to-be-processed element for the calculation unit. That is, in the current round, the second to-be-processed element for the calculation unit has been stored in the temporary storage module of the calculation unit.
[0092] Please refer to FIG. 4A , FIG. 4A A data flow diagram of a pulse array in the first clock cycle provided for an embodiment of the present application is shown.
[0093] As FIG. 4A shown, the embodiment of the present application first performs data rearrangement processing on the B matrix. In the first clock cycle, the first row data of the first to-be-processed matrix includes the element a11, the element a12, and the element a13, which are simultaneously input into the calculation unit 2211, the calculation unit 2221, and the calculation unit 2231 (i.e., the first calculation unit of each pulse sub-array). At the same time, as FIG. 3The first row data of the second to-be-processed matrix B320 shown can include an element b11, an element b22, and an element b33, which can be respectively preloaded to the computing unit 2211, the computing unit 2221, and the computing unit 2231. The computing unit 2211 can perform a multiplication operation to obtain a multiplication operation result a11 x b11. The computing unit 2221 can perform a multiplication operation to obtain a multiplication operation result a12 x b22. The computing unit 2231 can perform a multiplication operation to obtain a multiplication operation result a13 x b33. After the computation is completed, the element a11 can be transmitted to the computing unit 2232 in the second row and the third column through the first channel, the element a12 can be transmitted to the computing unit 2212 in the second row and the first column, and the element a13 can be transmitted to the computing unit 2222 in the second row and the second column. It can be understood that the three computing units in the first row of the pulse array 220 can be respectively used as head computing units of three pulse subarrays. It can also be understood that, for each computing unit in the first row, the post-processing hardware unit can be configured for the computing unit.
[0094] In addition, the multiplication operation result of the head computing unit can be used as a target operation result of the head computing unit. The head computing unit can also be configured to provide the second to-be-processed element to the post-processing computing unit of the source pulse subarray and provide the target operation result of the head computing unit to the post-processing hardware unit. That is, the elements of the second to-be-processed matrix and the intermediate calculation result can be transmitted within the source pulse subarray. For example, the element b11 can be transmitted to the computing unit 2221 in the second row and the first column through the second channel, the element b22 can be transmitted to the computing unit 2222 in the second row and the second column, and the element b33 can be transmitted to the computing unit 2232 in the second row and the third column. The target operation result of the computing unit 2211 can be transmitted to the computing unit 2212 through the third channel. The target operation result of the computing unit 2221 can be transmitted to the computing unit 2222 through the third channel. The target operation result of the computing unit 2231 can be transmitted to the computing unit 2232 through the third channel. It can be understood that three data transmission channels can be provided for the computing unit, the first to-be-processed element can be transmitted through the first channel, the second to-be-processed element can be transmitted through the second channel, and the target operation result can be transmitted through the third channel.
[0095] It can be understood that the above describes the calculation of the first clock cycle of the current round by the head computing unit, and the calculation of the second clock cycle of the current round will be described below.
[0096] In some embodiments, the computing unit can also be configured to provide the second to-be-processed element for the post-processing hardware unit in the source pulse subarray to the post-processing hardware unit in the source pulse subarray.
[0097] Please refer to FIG. 4B , FIG. 4BThis is a schematic diagram of the data flow of a pulse array in the second clock cycle, provided as an embodiment of the present invention.
[0098] like FIG. 4B As shown, in the second clock cycle, the data in the second row of the first matrix to be processed includes elements a21, a22, and a23, which can be input into the calculation units 2211, 2221, and 2231 in the first row of the pulse array 220, respectively. Calculation unit 2211 can perform multiplication operations to obtain the result a21 × b11. Calculation unit 2221 can perform multiplication operations to obtain the result a22 × b22. Calculation unit 2231 can perform multiplication operations to obtain the result a23 × b33. After the calculation is completed, element a21 can be passed to the calculation unit 2232 in the second row and third column through the first channel; element a22 can be passed to the calculation unit 2212 in the second row and first column; and element a23 can be passed to the calculation unit 2222 in the second row and second column.
[0099] In some embodiments, to implement the basic computational logic of matrix multiplication, the computation unit can also perform multiply-accumulate operations and transfer the computation result in the pulse array 220. The computation unit can perform the following operations to determine its target computation result based on the received first element to be processed and the second element to be processed for the computation unit: multiplying the received first element to be processed and the second element to be processed for the computation unit to obtain the multiplication result of the computation unit; determining the target computation result of the computation unit based on the multiplication result of the computation unit and the previous computation result provided by the preceding computation unit in the source pulse subarray; and providing the target computation result of the computation unit to the subsequent hardware unit in the source pulse subarray.
[0100] The rearranged second matrix to be processed, B320, may include elements b21, b32, and b13, which can be preloaded into computation units 2212, 2222, and 2232, respectively. Since the pulse array 220 employs a pipelined processing method, the calculation results need to be sequentially transferred and accumulated among multiple computation units. Therefore, in each clock cycle, the computation unit can determine its target calculation result based on its own multiplication results and the preceding calculation results provided by the preceding computation unit in the source pulse subarray.
[0101] like FIG. 4BAs shown, in the second clock cycle, the first column computing unit 2212 in the second row receives the element a12 passed through the first channel and the target calculation result of the computing unit 2211 in the first clock cycle passed through the third channel. The computing unit 2212 performs multiplication operation to obtain the multiplication result a12×b21 of the computing unit 2212 in the second clock cycle, and accumulates the multiplication result a12×b21 of the computing unit 2212 in the second clock cycle with the calculation result a11×b11 from the computing unit 2211 in the first clock cycle above to obtain the target calculation result a11×b11+a12×b21 of the computing unit 2212 in the second clock cycle. It can be understood that the way to obtain the target calculation result a12×b22+a13×b32 of the second column computing unit 2222 in the second clock cycle and the way to obtain the target calculation result a13×b33+a11×b13 of the third column computing unit 2232 in the second clock cycle are the same as or similar to the way to obtain the target calculation result a11×b11+a12×b21 of the computing unit 2212 in the second clock cycle, which will not be described herein again. It can also be understood that the computing unit 2211 can be the preceding computing unit of the computing unit 2212 in the source pulse subarray. For the computing unit 2212, the target calculation result provided by the computing unit 2211 can be the preceding calculation result. Due to the pre-adjustment of the data rearrangement operation, each computing unit can receive the matching matrix elements at the correct time for multiplication and accumulation operation.
[0102] After the computing units in the second row complete the calculation, the received first to-be-processed elements and the target calculation results obtained in the second clock cycle are passed to the computing units in the third row through the corresponding channels, respectively. For example, in the second clock cycle, the element a11 can be passed to the second column computing unit 2223 in the third row through the first channel, the element a12 can be passed to the third column computing unit 2233 in the third row, and the element a13 can be passed to the first column computing unit 2213 in the third row. The target calculation result of the computing unit 2212 can be passed to the computing unit 2213 through the third channel. The target calculation result of the computing unit 2222 can be passed to the computing unit 2223 through the third channel. The target calculation result of the computing unit 2232 can be passed to the computing unit 2233 through the third channel.
[0103] The calculation unit can obtain the target calculation result containing more complete accumulated information by accumulating the multiplication result and the previous calculation result. The target calculation result in the embodiment of the present application can represent the calculation of the matrix multiplication which has been completed at present. In order to ensure that the matrix multiplication calculation can continue in the pulse array 220, the calculation unit needs to provide the target calculation result of the calculation unit to the subsequent hardware unit in the source pulse subarray, so as to ensure that the calculation result can continue to propagate according to the predetermined data flow path. For example, the target calculation result a11×b11+a12×b21 contains the accumulated multiplication information of the first row of the first to-be-processed matrix and the first column of the second to-be-processed matrix, and represents the partial calculation progress of the first row and the first column of the result matrix.
[0104] According to the operation process of the multiplication calculation, the accumulation operation and the result transmission, the calculation unit 2211 realizes the processing logic of a single calculation position in the matrix multiplication according to the embodiment of the present application, and realizes the effective coordination with the overall data flow of the pulse array 220 on the basis of ensuring the operation accuracy.
[0105] It can be understood that the calculation of the second clock cycle of the current round is described above, and the calculation of the third clock cycle of the current round will be described below.
[0106] Please refer to FIG. 4C , FIG. 4C A data flow schematic diagram of the pulse array provided by the embodiment of the present application in the third clock cycle.
[0107] As shown in FIG. 4C , in the third clock cycle, the third row data of the first to-be-processed matrix includes the element a31, the element a32 and the element a33, which can be input into the calculation unit 2211, the calculation unit 2221 and the calculation unit 2231 of the first row of the pulse array 220 respectively. The calculation unit 2211 can perform multiplication operation to obtain the multiplication result a31×b11. The calculation unit 2221 can perform multiplication operation to obtain the multiplication result a32×b22. The calculation unit 2231 can perform multiplication operation to obtain the multiplication result a33×b33. After the calculation is completed, the element a31 can be transmitted to the calculation unit 2232 of the second row and the third column through the first channel, the element a32 can be transmitted to the calculation unit 2212 of the second row and the first column, and the element a33 can be transmitted to the calculation unit 2222 of the second row and the second column. It can be understood that for each intermediate calculation unit of the second row, the subsequent hardware unit can be an output unit.
[0108] In addition, in the third clock cycle, the first column of the second row of the calculation unit 2212 receives the element a22 passed through the first channel and the target calculation result of the calculation unit 2211 in the second clock cycle passed through the third channel. The calculation unit 2212 performs multiplication to obtain the multiplication result a22×b21 of the calculation unit 2212 in the third clock cycle. The calculation unit 2212 adds the multiplication result a22×b21 of the calculation unit 2212 in the second clock cycle and the calculation result a21×b11 from the above calculation unit 2211 in the second clock cycle to obtain the target calculation result a21×b11+a22×b21 of the calculation unit 2212 in the third clock cycle. It can be understood that the way of obtaining the target calculation result a22×b22+a23×b32 of the second column of the second row of the calculation unit 2222 in the third clock cycle and the way of obtaining the target calculation result a23×b33+a21×b13 of the third column of the second row of the calculation unit 2232 in the third clock cycle are the same as or similar to the way of obtaining the target calculation result a31×b11+a22×b21 of the calculation unit 2212 in the third clock cycle, and the present application will not be described here.
[0109] On the basis of the above-mentioned embodiments, as an optional embodiment, in order to complete the final calculation of the matrix multiplication calculation and realize the effective output of the calculation result, the pulse subarray further comprises an output unit; the output unit is used for multiplying the received first to-be-processed element and the second to-be-processed element for the output unit to obtain the multiplication result of the output unit; adding the multiplication result of the output unit and the previous calculation result provided by the previous calculation unit in the source pulse subarray to obtain the output result of the output unit. The output result is stored to the cache unit 210.
[0110] For the output unit in the embodiments of the present application, it refers to a hardware module located at the end of each pulse subarray 221 and outputting the final result.
[0111] For example, the rearranged second to-be-processed matrix B320 can include element b31, element b12, and element b23, which can be preloaded into the computing unit 2213, the computing unit 2223, and the computing unit 2233, respectively. The computing unit 2213 in the 3rd row and the 1st column receives the element a13 provided by the computing unit 2222 and the computing result of the computing unit 2212 in the 2nd clock cycle. The computing unit 2213 can perform a multiplication operation to obtain the multiplication result a13×b31 of the computing unit 2213 in the 3rd clock cycle, and add the multiplication result a13×b31 of the computing unit 2213 in the 3rd clock cycle to the target computing result of the computing unit 2212 in the 2nd clock cycle to obtain the output result a11×b11+a12×b21+a13×b31 of the computing unit 2213 in the 3rd clock cycle, so as to complete the final calculation of the corresponding position in the matrix multiplication. It can be understood that the manner of obtaining the output result a12×b22+a13×b32+a11×b12 of the computing unit 2223 in the 3rd row and the 2nd column in the 3rd clock cycle and the manner of obtaining the target computing result a13×b33+a11×b13+a12×b23 of the computing unit 2233 in the 3rd row and the 3rd column in the 3rd clock cycle are the same as or similar to the manner of obtaining the target computing result a11×b11+a12×b21+a13×b13 of the computing unit 2212 in the 2nd clock cycle, and details are not described herein again. At this time, the three computing units in the 3rd row have output the data in the first row of the result matrix of the matrix multiplication calculation at the end of the 3rd clock cycle. It can be understood that the computing unit 2213, the computing unit 2223, and the computing unit 2233 can serve as three output units.
[0112] To ensure that the computing result can be effectively utilized by the system and support subsequent data processing operations, the output unit needs to store the output result to the cache unit 210. By storing the final matrix multiplication result in the cache unit 210, the buffer storage of the computing result is realized. It can also be understood that the manner of obtaining the data in the second row and the third row of the result matrix of the matrix multiplication calculation is the same as or similar to the manner of obtaining the data in the first row of the result matrix of the matrix multiplication calculation, and details are not described herein again.
[0113] According to the embodiment of the present application, by arranging the output unit in the pulse subarray and explicitly defining the computing and storage functions thereof, a complete closed loop of the matrix multiplication calculation is realized, which not only guarantees the continuity of the data flow and the accuracy of the computing result in the computing process, but also realizes the effective preservation of the computing result by storing the result to the cache unit 210.
[0114] On the basis of the above-mentioned embodiments, as another optional embodiment, in order to ensure that the second to-be-processed elements in the pulse array 220 can be correctly propagated to the corresponding calculation units according to the rearranged data distribution, the calculation unit can also undertake the transfer function of the second to-be-processed elements.
[0115] Specifically, the calculation unit is configured to provide, for the second to-be-processed element of the subsequent hardware unit in the source pulse subarray, the second to-be-processed element to the subsequent hardware unit in the source pulse subarray.
[0116] For the second to-be-processed element of the subsequent hardware unit in the source pulse subarray in the embodiments of the present application, it refers to the second to-be-processed element received by the calculation unit in the current calculation period, which is not used for self-computation but needs to be transferred to the next calculation unit 1 in the source pulse subarray.
[0117] Since the data elements in the rearranged second to-be-processed matrix need to be propagated between different calculation units of the pulse array 220 according to a specific timing and path, each calculation unit needs to be responsible for continuing to pass down the received second to-be-processed element while performing its own multiply-accumulate calculation, so as to ensure that the subsequent calculation unit can obtain the corresponding calculation data at the correct time.
[0118] In the specific transfer operation process, the calculation unit provides the received second to-be-processed element to the subsequent hardware unit in the source pulse subarray through the vertically downward data channel. The second to-be-processed element propagation in the embodiments of the present application needs to match the data distribution of the rearranged second to-be-processed matrix, so as to ensure that each second to-be-processed element can reach the target calculation unit 2211 for multiply-accumulate calculation according to the predetermined timing.
[0119] Exemplarily, the scenario of the aforementioned 3x3 matrix multiplication is continued to be used for description. In the first clock cycle, the first row and first column calculation unit 2211 receives the element b11 for multiply computation with the element a11, and at the same time, the calculation unit 2211 can also pass the element b11 to the second row and first column calculation unit 2212, so that the second row and first column calculation unit 2211 can receive the element b11 from above.
[0120] According to the embodiment of the present application, it can be understood that the element b21, the element b32 and the element b13 are not the second to-be-processed elements of the calculation unit 2211, the calculation unit 2221 and the calculation unit 2231, and can be delivered to the respective subsequent hardware units of the calculation unit 2211, the calculation unit 2221 and the calculation unit 2231 without participating in the calculation. According to the embodiment of the present application, the delivery function of the second to-be-processed elements is borne by the calculation unit 2211, and the ordered propagation of the second to-be-processed elements in the pulse array 220 is realized, and a complete data flow is formed with the rearranged second to-be-processed matrix and the reconstructed first to-be-processed element propagation path.
[0121] It can be understood that the element a11, the element a12, the element a13, the element a21, the element a22, the element a23, the element a31, the element a32 and the element a33 in the above embodiment can be the first to-be-processed elements respectively. The element b11, the element b12, the element b13, the element b21, the element b22, the element b23, the element b31, the element b32 and the element b33 in the above embodiment can be the second to-be-processed elements respectively.
[0122] It can be understood from the comparison of the data flows that, in terms of hardware resource consumption, the embodiment of the present application completely eliminates the use of delay units; in terms of start-up delay, the present application almost completely eliminates the start-up delay. For example, for a 3x3 pulse array, the use of 9 delay units can be eliminated, and the start-up delay is reduced from 2 clock cycles to 0 clock cycle, and the data can start effective calculation in the first clock cycle, and the result of each calculation unit 2211 in the first row can be obtained in the third clock cycle. It can be understood that the above describes the present application by taking a 3x3 pulse array and two 3x3 matrices as examples. However, the present application is not limited thereto, and the size of the pulse array can be flexibly set according to the area of the chip. The size of the matrix that can be processed by the pulse array can also be different from the size of the pulse array.
[0123] On the basis of the above embodiment, as an optional embodiment, in order to support more complete matrix operation functions and improve the performance of the pulse array 220 in actual application, the embodiment of the present application extends the calculation capability to process the matrix multiplication and addition operation containing a bias term.
[0124] Specifically, the cache unit 210 is further configured to store a third to-be-processed matrix, and the third to-be-processed matrix comprises a plurality of third to-be-processed elements.
[0125] The head computation unit of the pulse sub-array is further configured to perform the following operations to determine a target computation result of the head computation unit according to the received first to-be-processed element and a second to-be-processed element for the computation unit: multiplying the first to-be-processed element for the head computation unit received from the cache unit and the second to-be-processed element for the head computation unit to obtain a multiplication result of the head computation unit; and adding the third to-be-processed element for the head computation unit received from the cache unit and the multiplication result of the head computation unit to obtain the target computation result of the head computation unit.
[0126] For the third to-be-processed matrix in the embodiment of the application, it refers to an input matrix used as a bias term or an initial accumulation value in the matrix multiplication and addition operation, which contains a plurality of third to-be-processed elements and realizes a linear transformation function by adding the matrix multiplication result.
[0127] For the head computation unit in the embodiment of the application, it refers to a computation unit located in the first row in each pulse sub-array, which is located at the starting position in the matrix multiplication and addition computation link. For example, the computation unit 2211 of the pulse sub-array 221 can be a head computation unit. The computation unit 2221 of the pulse sub-array 222 can be a head computation unit. The computation unit 2231 of the pulse sub-array 223 can be a head computation unit.
[0128] For example, it is assumed that the matrix multiplication and addition operation is Y=A×B+C, A can be the first to-be-processed matrix, B can be the second to-be-processed matrix, and C can be the third to-be-processed matrix. Y can be the matrix multiplication and addition operation result matrix. In the first clock cycle, the computation unit 2211 in the first row receives the first to-be-processed element a11, the second to-be-processed element b11 and the third to-be-processed element c11 from the cache unit 210. The head computation unit 2211 first performs multiplication to obtain the multiplication result a11×b11, and then adds the multiplication result a11×b11 and c11 to obtain the target computation result a11×b11+c11. The target computation result is transmitted to the computation unit 2212 in the first row and the second column, so that the computation unit 2212 performs multiplication according to the received first to-be-processed element a12 and the second to-be-processed element b22 to obtain the multiplication result a12×b22, and adds the multiplication result a12×b22 and the previous computation result from the head computation unit 2211 to obtain the target computation result a11×b11+c11+a12×b22 of the computation unit 2212, and so on until the computation of the whole row is completed.
[0129] According to the embodiment of the application, the processing function of the third to-be-processed element is integrated in the head computation unit, the direct support of the matrix multiplication and addition operation of the pulse array to the matrix containing the bias term is realized, and the application range of the pulse array 220 is expanded.
[0130] As an optional embodiment based on the above-mentioned embodiments, the first to-be-processed matrix represents feature data of the neural network model, the second to-be-processed matrix represents a weight parameter of the neural network model, and the third to-be-processed matrix represents a bias term of a neural network layer in the neural network model.
[0131] In the neural network inference calculation process, the core operation is linear transformation between neural network layers, which is usually represented as matrix multiplication calculation of feature data and a weight parameter, and then a bias term is added to obtain input features of the next layer.
[0132] In the above calculation process, the feature data carried by the first to-be-processed matrix dynamically changes with different input samples in the inference process, representing an input feature vector currently to be processed or an output feature of the previous layer of the neural network.
[0133] The weight parameter carried by the second to-be-processed matrix is a fixed parameter learned in the training phase of the neural network model, which remains unchanged and is often reused in the inference process, and can be resident data.
[0134] The bias term carried by the third to-be-processed matrix is also a fixed parameter determined in the training phase, which is used for linear offset adjustment of the matrix multiplication result.
[0135] Since the weight parameter has the resident characteristic in the neural network inference process, the embodiment of the present application pre-processes the weight parameter through data reordering operation, and pre-loads the second to-be-processed matrix after reordering into each calculation unit of the pulse array 220, thereby realizing pre-loading of the weight parameter.
[0136] The above pre-loading method can maximize the parallel computing capability of the pulse array 220, and avoid the overhead of reloading the weight parameter at each inference calculation. When new feature data is input, the calculation unit 2211 in the pulse array 220 can directly use the pre-loaded weight parameter for calculation, thereby improving the calculation efficiency of the neural network inference.
[0137] According to the embodiment of the present application, by combining the matrix multiplication calculation capability of the pulse array 220 with the data characteristics of the neural network inference, the optimization for the neural network application is realized, which not only guarantees the calculation accuracy, but also fully utilizes the advantage of weight parameter reuse, thereby improving the calculation efficiency and resource utilization of the pulse array 220 in the neural network inference task.
[0138] FIG. 5 Another structure schematic diagram of a data processing device provided by the embodiment of the present application.
[0139] On the basis of the above-mentioned embodiments, as an optional embodiment, the plurality of to-be-processed vector data is I, and the data processing apparatus 200 further comprises a processor 230 configured to perform the following operations: cyclically moving a plurality of elements in the initial vector data of the second initial matrix by a second moving bit number along a second direction, to obtain a second to-be-processed matrix; and storing the second to-be-processed matrix to the cache unit 210. I can be an integer greater than 1.
[0140] For the processor 230 in the embodiments of the present application, it refers to a general-purpose computing processing unit that cooperates with the pulse array 220 and is responsible for performing data preprocessing operations and system control functions. The processor 230 performs a reordering process on the input matrix data by executing a software algorithm, and provides preprocessed data that meets the data processing timing requirements of the pulse array 220. The processor 230 can be a processor supporting various instruction sets, such as a processor supporting a Reduced Instruction Set Computer (RISC).
[0141] After the processor 230 completes the data rearrangement operation, the rearranged data is stored in the cache unit 210, and the system can ensure that the pulse array 220 can quickly access the required second to-be-processed elements when performing matrix multiplication calculation, avoiding the influence of external memory access delay on calculation efficiency. The cache unit 210 serves as a data buffer between the processor 230 and the pulse array 220, supporting batch data write operations of the processor 230 and meeting the data access timing requirements of the pulse array 220. The second initial matrix and the second to-be-processed matrix are respectively shown in FIG. 3B and FIG. 3C, which will not be described herein again. FIG. 3
[0142] According to the embodiments of the present application, the processor 230 stores the rearranged second to-be-processed matrix in the cache unit 210 in advance, so that the pulse array 220 can directly and continuously read the data that has been adjusted in timing from the cache, avoiding data waiting and timing adjustment delay in the calculation process, and improving the overall calculation efficiency of the matrix multiplication calculation.
[0143] FIG. 6 Another data processing apparatus provided by the embodiments of the present application.
[0144] In actual matrix multiplication calculation applications, especially in deep learning inference and scientific computing scenarios, the system often needs to continuously perform a large number of matrix multiplication calculations, and the second to-be-processed matrix used in each operation can come from different data batches or calculation layers. If the processor 230 needs to wait for the pulse array 220 to complete the current calculation before starting to prepare the data required for the next round of calculation, the pulse array 220 will be idle during data preparation, which will seriously affect the overall calculation efficiency.
[0145] To solve the above problems, on the basis of the above embodiment, as an optional embodiment, as shown in FIG. 6 The data processing apparatus 200 further includes a first storage module 240 for storing the multiple state information and address information of the second to-be-processed matrix sent by the processor 230 and the running state information of the pulse array 220; the pulse array 220 is further configured to read the second to-be-processed matrix from the cache unit 210 according to the multiple state information and address information of the second to-be-processed matrix; and the processor 230 is further configured to generate the second to-be-processed matrix required for the next round of calculation of the pulse array 220 according to the running state information of the pulse array 220.
[0146] For the first storage module 240 in the embodiment of the application, it refers to a hardware module located between the processor 230 and the pulse array 220 for realizing state information exchange and data access control. The module can coordinate the working timing of the processor 230 and the pulse array 220 through multiple control registers.
[0147] The first storage module 240 stores the multiple state information of the second to-be-processed matrix, so that the pulse array 220 can accurately determine whether there is available rearranged data and the preparation state of the data. Meanwhile, the first storage module 240 also stores the address information of the second to-be-processed matrix in the cache unit 210, so that the pulse array 220 can directly locate and access the required data, avoiding the additional overhead of data searching and address calculation.
[0148] In addition, the first storage module 240 is also responsible for storing the running state information of the pulse array 220. The running state information reflects at least one of the current calculation state, the calculation progress, and the calculation completion state information of the pulse array 220.
[0149] By storing the running state information in the first storage module 240, the processor 230 can know the working state of the pulse array 220 in real time, and decide whether to start preparing the data required for the next round of calculation, thereby realizing the timing coordination between the processor 230 and the pulse array 220.
[0150] Specifically, after the processor 230 completes the rearrangement operation on the second initial matrix and stores the second to-be-processed matrix to the cache unit 210, the processor 230 updates the state information in the first storage module 240 to indicate that the new rearranged data is ready, and writes the specific address information of the data in the cache unit 210 to the first storage module 240.
[0151] The pulse array 220 judges whether the new second to-be-processed matrix is available for reading by monitoring the state information in the first storage module 240. When the state information indicates that the new data is available, the pulse array 220 reads the second to-be-processed matrix directly from the specified position of the cache unit 210 according to the address information stored in the first storage module 240, avoiding the uncertain delay of data searching and transmission.
[0152] Meanwhile, the pulse array 220 updates the running state information in the first storage module 240 to indicate that the matrix multiplication calculation is being performed when starting the calculation. Through the running state information, the processor 230 can determine the working state of the pulse array 220, and start to prepare the second to-be-processed matrix required for the next round of calculation in parallel during the execution of the current calculation by the pulse array 220.
[0153] According to the embodiment of the present application, the state synchronization realized by introducing the first storage module 240 eliminates the waiting time between the processor 230 and the pulse array 220, so that the data required for the next round of calculation is ready while the pulse array 220 completes the current calculation, and the new calculation task can be started immediately, realizing the parallelization of data preparation and calculation execution, avoiding the idle waiting of hardware resources, and improving the efficiency in the continuous matrix multiplication calculation scenario.
[0154] In an actual matrix multiplication calculation system, the pulse array 220 needs to process matrix data of different sizes, and these data may come from different calculation tasks or data batches. If the pulse array 220 reads data only according to a simple data ready signal, it may encounter problems such as incomplete data, address error or data volume mismatch, resulting in incorrect calculation results or system abnormalities.
[0155] In view of the above problems, on the basis of the above embodiment, as an optional embodiment, the first storage module 240 can be further configured to store the address information of the second to-be-processed matrix in the cache unit 210. FIG. 6The pulse array 220 is also configured to read the second to-be-processed matrix from the cache unit 210 according to the plurality of state information of the second to-be-processed matrix, address information of the second to-be-processed matrix in the cache unit 210, and data quantity of the second to-be-processed matrix. The first storage module 240 includes: a rearranged data state register 241 configured to store at least one of the plurality of state information of the second to-be-processed matrix sent by the processor 230; the plurality of state information includes first state information and second state information, the first state information is used to indicate whether the processor 230 generates the second to-be-processed matrix, and the second state information is used to indicate whether the processor 230 stores the second to-be-processed matrix to the cache unit 210; an address synchronization register 242 configured to store the address information of the second to-be-processed matrix in the cache unit 210; and a length synchronization register 243 configured to store the data quantity information of the second to-be-processed matrix.
[0156] The rearranged data state register 241 is configured to store at least one of the plurality of state information of the second to-be-processed matrix sent by the processor 230, and the plurality of state information includes first state information and second state information. The first state information is used to indicate whether the processor 230 generates the second to-be-processed matrix, that is, whether the processor 230 has completed the rearrangement operation on the second initial matrix and generates the second to-be-processed matrix meeting the data processing timing requirement of the pulse array 220. The second state information is used to indicate whether the processor 230 stores the second to-be-processed matrix to the cache unit 210, so as to ensure that the rearranged data is transferred from the working memory of the processor 230 to the cache unit 210 accessible by the pulse array 220. By distinguishing the data generation state and the data storage state, the system can more accurately control the data access timing, and avoid that the pulse array 220 performs the reading operation when the data is not completely ready.
[0157] The address synchronization register 242 is configured to store the address information of the second to-be-processed matrix in the cache unit 210, and the address information indicates the starting position of the rearranged data in the cache unit 210. Since the cache unit 210 usually uses a continuous address space to store the matrix data, the address information in the address synchronization register 242 provides an entry point for the data access of the pulse array 220, so that the pulse array 220 can directly locate the target data without performing the address search or data scanning operation.
[0158] The length synchronization register 243 is configured to store the data quantity information of the second to-be-processed matrix, and the data quantity information indicates the total number of data elements that need to be read from the cache unit 210. By explicitly indicating the data quantity information before the data reading, the pulse array 220 can verify the integrity of the reading operation, ensure that all necessary matrix elements are correctly read, and avoid reading invalid contents beyond the data boundary.
[0159] In the specific data reading process, the pulse array 220 first checks the state information in the rearranged data state register 241, verifies that the first state information and the second state information are both in the ready state, and confirms that the processor 230 has completed data rearrangement and successfully stored to the cache unit 210. Subsequently, the pulse array 220 obtains the starting address of the second to-be-processed matrix in the cache unit 210 from the address synchronization register 242, and obtains the data amount information to be read from the length synchronization register 243. Based on these information, the pulse array 220 performs a cache access operation of reading a specified data amount starting from a specified starting address, and loads all elements of the second to-be-processed matrix into the internal storage unit of the pulse array 220. It can be understood that the data amount information can represent the amount of data that has been divided, so that the size of the second to-be-processed matrix is less than or equal to the size of the pulse array.
[0160] Exemplarily, it is assumed that the processor 230 completes the rearrangement operation of a 4x4 second initial matrix, generates a second to-be-processed matrix containing 16 elements and stores it in the continuous space starting from the 0x1000 address of the cache unit 210. The processor 230 first sets the first state information of the rearranged data state register 241 to 1, indicating that the data generation is completed, sets the second state information to 1, indicating that the data storage is completed, and at the same time sets the address synchronization register 242 to 0x1000 and the length synchronization register 243 to 16. After the pulse array 220 detects the change of the state register, it reads the contents of the address synchronization register 242 and the length synchronization register 243, and confirms that 16 data elements need to be read starting from the 0x1000 address. The pulse array 220 then performs a cache access operation to read the 16 data elements starting from the 0x1000 address in turn, and completes the loading process of the second to-be-processed matrix.
[0161] According to the embodiment of the present application, by setting the state register, the address register and the length register in the first storage module 240, the overall control and verification of the access process of the second to-be-processed matrix are realized, and it is ensured that the pulse array 220 can accurately and completely obtain the data required for calculation, and the adaptability of the system in processing matrices of different sizes is improved.
[0162] On the basis of the above embodiment, as an optional embodiment, in order to establish the bidirectional state synchronization between the processor 230 and the pulse array 220, as shown in FIG. 6 The first storage module 240 further includes an array state register 244 for storing the running state information of the pulse array 220.
[0163] When the pulse array 220 starts to perform the matrix multiplication calculation, the array state register 244 is updated to the "calculation in progress" state, and the processor 230 detects the state change and can start preparing the data rearrangement work required for the next round of calculation.
[0164] When the pulse array 220 completes the calculation and stores the result to the cache unit 210, the array state register 244 is updated to the "calculation completed" state, and the processor 230 receives the state change and can trigger the next round of data transmission and calculation start process.
[0165] According to the embodiment of the application, the bidirectional state communication established by the array state register 244 realizes the real-time synchronization of the working state of the processor 230 and the pulse array 220, and effectively improves the cooperative efficiency of the continuous matrix multiplication calculation.
[0166] In the actual continuous matrix multiplication calculation scene, the pulse array 220 usually needs to process multiple rounds of continuous calculation tasks, and different second to-be-processed matrix data needs to be used for each round of calculation. If the calculation unit 2211 is only configured with a single storage space to save the second to-be-processed element, the data required for the next round of calculation cannot be received and stored at the same time during the current round of calculation execution, resulting in the pulse array 220 having to wait for the loading and distribution of the next round of data after completing the current calculation, forming idle waiting time of the calculation resource.
[0167] FIG. 7 A structural schematic diagram of a calculation unit provided in an embodiment of the application is shown.
[0168] In view of the above problems, on the basis of the above embodiment, as an optional embodiment, as shown in FIG. 7 The calculation unit 2200 includes a plurality of temporary storage modules for storing the second to-be-processed elements for the calculation unit 2200 in different rounds. The calculation unit 2200 can be any one of the calculation unit 2211, the calculation unit 2212, the calculation unit 2213, the calculation unit 2221, the calculation unit 2222, the calculation unit 2223, the calculation unit 2231, the calculation unit 2232, and the calculation unit 2233.
[0169] In this embodiment of the invention, the temporary storage module refers to an independent storage unit within the computing unit 2211 used for temporarily storing the second element to be processed. Each temporary storage module has independent read / write control logic and data retention capabilities. By configuring multiple temporary storage modules in the computing unit, the computing unit can simultaneously store the second element to be processed from different computing rounds, achieving time reuse and spatial isolation of data storage. The multiple temporary storage modules may include temporary storage module 711 and temporary storage module 712. It can be understood that temporary storage module 711 can serve as the first temporary storage module, and temporary storage module 712 can serve as the second temporary storage module.
[0170] Specifically, when the pulse array 220 performs the first round of matrix multiplication calculation, the calculation unit 2211 stores the second element to be processed required in the current round in the first temporary storage module 711, and reads data from the temporary storage module 711 to participate in the multiplication and accumulation operation. For example, the elements b11 to b33 of the current round can be stored in the temporary storage modules 711 of different calculation units respectively.
[0171] While the first round of calculations is underway, the processor 230 has already completed the rearrangement operation of the second matrix to be processed required for the second round of calculations, and the pulse array 220 begins to receive the second element to be processed in the second round. At this time, the calculation unit 2200 stores the received second element to be processed belonging to the second round of calculations in the second temporary storage module 712, while the data in the first temporary storage module 711 continues to be used for the calculation operation of the current round.
[0172] Once the first round of calculation is complete, the computing unit 2200 immediately switches its data source and reads the second element to be processed from the second temporary storage module 712 to begin a new calculation task. Simultaneously, the first temporary storage module 711 can receive data for the third round of calculation for preloading. Through the alternating use of multiple temporary storage modules, the computing unit 2200 achieves seamless switching between consecutive calculation rounds, eliminating data loading waiting time.
[0173] According to an embodiment of the present invention, by configuring multiple temporary storage modules in the computing unit 2211 and realizing parallel storage and alternating use of data in different rounds, the data waiting problem in continuous matrix multiplication calculation is effectively solved, enabling the pulse array 220 to prepare the data required for the next round of calculation while completing the current calculation, realizing the parallelization of calculation execution and data preloading, and improving the hardware utilization and overall computing efficiency in continuous computing scenarios.
[0174] Based on the above embodiments, as an optional embodiment, such as... FIG. 7As shown, each computing unit 2200 further comprises a second storage module 720, configured to store at least one of a plurality of preloading indication values corresponding to the plurality of temporary storage modules, the preloading indication value being configured to indicate that the computing unit 2200 is to store a second to-be-processed element in a next round for the computing unit 2200 to the temporary storage module corresponding to the preloading indication value.
[0175] For the second storage module 720 in the embodiment of the present application, it refers to a storage unit inside the computing unit 2200 for storing usage state information and access control information of the plurality of temporary storage modules, and the module realizes the storage function of data storage and reading operation by maintaining control information corresponding to the number of temporary storage modules.
[0176] Specifically, the preloading indication value indicates the working mode of the corresponding temporary storage module through binary state coding. The plurality of preloading indication values comprises a first preloading indication value and a second preloading indication value, the first preloading indication value indicating that the temporary storage module 711 is in an idle state or has completed data service in the current round. The second preloading indication value indicates that the temporary storage module 712 is in an idle state or has completed data service in the current round. If the second storage module stores, for example, the second preloading indication value for the temporary storage module 712, it indicates that the temporary storage module 712 is currently in an idle state or has completed data service in the current round, and can receive a second to-be-processed element in the next round for preloading operation.
[0177] According to the embodiment of the present application, by storing the preloading indication value for the temporary storage module through the second storage module 720, the computing unit 2211 realizes automatic identification and correct routing of different rounds of data in the continuous computing process, avoids data storage errors and reading confusion, and guarantees the data computing accuracy of multi-round computation.
[0178] On the basis of the above embodiment, as an optional embodiment, as shown in FIG. 7 As shown, the second storage module 720 further comprises a preloading counter 721, configured to obtain the number of second to-be-processed elements received by the computing unit 2200 in the current round; and a preloading constant register 722, configured to store a preloading trigger threshold of the computing unit 2211, the preloading trigger threshold being the difference between the row sequence number of the computing unit 2211 in the pulse array 220 and a preset value. The computing unit 2200 further comprises a comparator 730, configured to, in a case where the number of second to-be-processed elements received in the current round is equal to the preloading trigger threshold, store the received second to-be-processed element consistent with the preloading trigger threshold to one of the plurality of temporary storage modules as a second to-be-processed element in a next round for the computing unit 2200.
[0179] The preloading counter 721 is used to obtain the number of the second to-be-processed elements received by the computing unit 2200 in the current round, and the preloading counter 721 is automatically reset to zero at the beginning of each computing round, and then the second to-be-processed elements flowing through the computing unit 2200 are counted in real time. For example, in the first clock cycle, the value of the preloading counter of the computing unit 2211 is 0.
[0180] The preloading constant register 722 is used to store the preloading trigger threshold of the computing unit 2200, and the threshold is preset according to the row number of the computing unit 2200 in the pulse array 220, and the specific calculation method is the difference between the row number of the computing unit 2211 and the preset value, and the preset value can be set to 1. Through this setting method, the preloading trigger threshold of the first row of computing units is 0, the preloading trigger threshold of the second row of computing units is 1, the preloading trigger threshold of the third row of computing units is 2, and so on, so as to realize the trigger matching the distribution timing of the second to-be-processed matrix data after rearrangement.
[0181] The comparator 730 can compare the number of received elements recorded by the preloading counter 721 with the preloading trigger threshold stored in the preloading constant register 722, and when the two values are equal, it indicates that the currently received second to-be-processed element is valid data for the computing unit 2200. At this time, the comparator 730 triggers the storage operation, and stores the received second to-be-processed element consistent with the preloading trigger threshold into one of the plurality of temporary storage modules as the second to-be-processed element for the computing unit 2211 in the current round. For other second to-be-processed elements that do not meet the trigger condition, the computing unit 2211 only performs the transfer operation without storage, so as to ensure that it can continue to flow to the target computing unit 2211 downstream.
[0182] For example, in the first clock cycle, the value of the preloading counter of the computing unit 2211 is 0, the value of the preloading constant register is 0, and the preloading indication value indicates that the second to-be-processed element for the computing unit 2211 in the next round can be stored in the second temporary storage module. The computing unit 2211 can receive one second to-be-processed element in the next round, which can be stored in the second temporary storage module in the computing unit 2211. Moreover, the second to-be-processed element can also be provided to the computing unit 2212, and the value of the preloading counter of the computing unit 2211 can be set to 1.
[0183] For example, in the first clock cycle, the computing unit 2211 preloads the value of the counter as 1, the value of the constant register as 0, and the value of the indication as indicating that the second element of the next round for the computing unit 2211 can be stored to the second temporary storage module. The computing unit 2211 can receive the second element of the first next round. If the value of the counter preloaded by the computing unit 2211 is different from the value of the constant register, the second element of the first next round can not be stored to the temporary storage module of the computing unit 2211, and can be provided to the computing unit 2212. In the first clock cycle, the computing unit 2212 preloads the value of the counter as 0, and the value of the constant register as 1. The computing unit 2212 can receive the second element of the second next round. If the value of the counter preloaded by the computing unit 2212 is different from the value of the constant register, the second element of the second next round can not be stored to the temporary storage module of the computing unit 2212, and can be provided to the computing unit 2213. At the end of the first clock cycle, the value of the counter preloaded by the computing unit 2211 can be set as 2, and the value of the counter preloaded by the computing unit 2212 can be set as 1.
[0184] For example, in the third clock cycle, the calculation unit 2211 preloads the counter with a value of 2 and the constant register with a value of 0. The calculation unit 2211 can receive the second to-be-processed element of the third next round. If the value of the counter preloaded by the calculation unit 2211 is different from the value of the constant register preloaded, the second to-be-processed element of the third next round can not be stored in the temporary storage module of the calculation unit 2211, or can be provided to the calculation unit 2212. In the second clock cycle, the calculation unit 2212 preloads the counter with a value of 1, the constant register with a value of 1, and the indication value with a value indicating that the second to-be-processed element of the next round for the calculation unit 2212 can be stored in the second temporary storage module. The calculation unit 2212 can receive the second to-be-processed element of the second next round. If the value of the counter preloaded by the calculation unit 2212 is the same as the value of the constant register preloaded, the second to-be-processed element of the second next round can be stored in the second temporary storage module of the calculation unit 2212, or can be provided to the calculation unit 2213. At the end of the second clock cycle, the value of the counter preloaded by the calculation unit 2211 can be set to 2, and the value of the counter preloaded by the calculation unit 2212 can be set to 1. In the third clock cycle, the calculation unit 2213 preloads the counter with a value of 0 and the constant register with a value of 2. The calculation unit 2213 can receive the second to-be-processed element of the first next round. If the value of the counter preloaded by the calculation unit 2213 is different from the value of the constant register preloaded, the second to-be-processed element of the first next round can not be stored in the temporary storage module of the calculation unit 2213. At the end of the third clock cycle, the value of the counter preloaded by the calculation unit 2211 is 3, the value of the counter preloaded by the calculation unit 2212 is 2, and the value of the counter preloaded by the calculation unit 2213 is 1.
[0185] According to the embodiment of the present application, the accurate matching mode formed by the preloaded counter 721, the preloaded constant register 722, and the comparator 730 can realize intelligent screening and accurate acquisition of the data flowing through the calculation unit 2211, ensure that each calculation unit 2211 can obtain the correct second to-be-processed element at the correct time, and avoid the problems of data mismatch and calculation error.
[0186] Based on the above embodiment, as an optional embodiment, the calculation unit 2211 can be configured to preload the counter with a value of 0 and the constant register with a value of 1 in the first clock cycle. FIG. 7As shown, the computing unit 2200 further comprises a preloading selection register 723, which is used to store a first preloading indication value or a second preloading indication value. The first preloading indication value is used to indicate that the next round of the second to-be-processed element of the computing unit is stored to the first temporary storage module, and the second preloading indication value is used to indicate that the next round of the second to-be-processed element of the computing unit is stored to the second temporary storage module. For example, the first preloading indication value can be 0, and the second preloading indication value can be 1.
[0187] The preloading selection register 723 refers to a special register inside the computing unit 2211 for controlling the selection of the temporary storage module, and the data storage target is indicated by maintaining a preloading indication bit corresponding to the temporary storage module. The preloading indication bit is encoded in binary, and each indication bit corresponds to a temporary storage module, and identifies whether the module is selected to receive the next round of data.
[0188] The computing unit is further configured to perform at least one of the following operations: in a case where the pulse array completes the computation of the current round, switching the value of the preloading indication bit from the first preloading indication value to the second preloading indication value. In a case where the pulse array completes the computation of the current round, switching the value of the preloading indication bit from the second preloading indication value to the first preloading indication value. The switching mode of the preloading selection register 723 is associated with the running state of the pulse array 220. When the pulse array 220 completes the computation of the current round, the value of the preloading selection register can be flipped. The temporary storage module 711 used for computation in the current round is switched to receive new data in the next round, and the temporary storage module 712 used for preloading in the current round is switched to provide computation data in the next round.
[0189] For example, in the first round of computation, the preloading selection register is the second preloading indication value. The second to-be-processed element of the computing unit 2200 in the second round of computation can be stored to the temporary storage module 712. After the pulse array 220 completes the first round of computation, the preloading selection register can be flipped to the first preloading indication value.
[0190] In the second round of computation, the preloading selection register is the first preloading indication value. The second to-be-processed element of the computing unit 2200 in the third round of computation can be stored to the temporary storage module 711. After the pulse array 220 completes the second round of computation, the preloading selection register can be flipped to the second preloading indication value.
[0191] According to the embodiment of the present application, by the state-driven switching of the preloading selection register 723, the automatic storage of the computing unit 2200 between consecutive computation rounds is realized, and the correctness of the data flow in multiple rounds of computation is ensured.
[0192] On the basis of the above embodiment, as an optional embodiment, as shown in FIG. 7,FIG. 7 As shown, the second storage module 720 further includes a calculation control register 724 for storing a calculation indication value, the calculation indication value being used to indicate the calculation unit 2200 to select a second to-be-processed element from the plurality of temporary storage modules for multiplication calculation.
[0193] The calculation control register 724 refers to a register in the second storage module 720 for selecting a data source for calculation, and the calculation indication value stored in the calculation control register 724 is used to indicate the calculation unit 2200 to read a second to-be-processed element from the corresponding temporary storage module. The calculation indication value and the preloading indication value form a complementary control relationship, that is, when the preloading selection register 723 indicates that one temporary storage module is used for preloading, the calculation control register 724 indicates that another temporary storage module is used for calculation, so as to avoid data conflict.
[0194] The calculation indication value in the calculation control register 724 is changed synchronously with the preloading indication value in the round switching state. When the preloading indication bit is flipped, the calculation indication value is also flipped correspondingly, so as to realize switching of the data reading target. The above-mentioned synchronous flipping operation ensures that the calculation unit 2211 always reads an element from the temporary storage module containing correct data of the current round, and meanwhile, another temporary storage module safely receives preloading data.
[0195] For example, in the first round of calculation, the value of the preloading selection register 723 is 1, indicating that the temporary storage module 712 is used for preloading, and the calculation control register 724 stores a first calculation indication value (for example, 0), indicating that data is read from the temporary storage module 711 for calculation. When the round is switched, the preloading indication bit is flipped to indicate that the temporary storage module 711 is used for preloading. In addition, when the round is switched, the value of the calculation control register is flipped to a second calculation indication value (for example, 1), indicating that data of the second round is read from the temporary storage module 712 for calculation.
[0196] According to the embodiment of the present application, the independent data source control established by the calculation control register 724 forms a complete bidirectional data flow control with the preloading selection register 723, so as to realize strict isolation and coordinated cooperation of the calculation operation and the preloading operation.
[0197] On the basis of the above-mentioned embodiment, as an optional embodiment, as shown in FIG. 7 The calculation unit 2200 further includes a multiply-accumulate calculator 750, a selector 740 for determining a selected second to-be-processed element from the plurality of temporary storage modules according to the calculation indication value stored in the calculation control register 724, and providing the selected second to-be-processed element to the multiply-accumulate calculator 750 for multiplication operation.
[0198] The selector 740 refers to a hardware module inside the computing unit 2211 for implementing data path selection, which receives a computing indication value from the computing control register 724 as a selection signal, and selects a corresponding data path from the output ends of multiple temporary storage modules according to the signal to be connected. The selector 740 realizes the dynamic connection between the temporary storage module and the multiply-accumulate calculator 750 through multiplexing logic, ensuring that the correct second to-be-processed element can be transmitted to the multiply-accumulate calculator in each computing cycle.
[0199] The multiply-accumulate calculator 750 refers to a dedicated hardware unit inside the computing unit 2211 for performing matrix multiplication basic operations, which integrates the functions of a multiplier and an accumulator, and can receive the first to-be-processed element and the second to-be-processed element to perform multiplication, and add the result to the accumulated input to obtain the final computing result.
[0200] Specifically, the selector 740 performs data selection according to the computing indication value stored in the computing control register 724. When the computing indication value is 0, the selector 740 connects the output of the first temporary storage module to the second to-be-processed element input end of the multiply-accumulate calculator 750; when the computing indication value is 1, the selector 740 connects the output of the second temporary storage module to the second to-be-processed element input end of the multiply-accumulate calculator 750.
[0201] In the above manner, the selector 740 ensures that the multiply-accumulate calculator 750 always receives the correct data required in the current round. After the selector 740 completes the data selection, the selected second to-be-processed element is provided to the multiply-accumulate calculator 750 through a dedicated data path, and the multiply-accumulate calculator 750 then performs a multiply-accumulate operation together with the received first to-be-processed element and the accumulated input.
[0202] According to the embodiment of the present application, a complete hardware path from data storage to operation execution is established, accurate data transmission between the multiple temporary storage modules and the computing unit 2200 is realized, and it is ensured that the computing unit 2200 can automatically select the correct data source and perform accurate operation in continuous matrix multiplication calculations.
[0203] Please refer to FIG. 8 , FIG. 8 for a structural schematic diagram of an electronic device provided by the embodiment of the present application.
[0204] The embodiment of the present application also discloses an electronic device 800, which comprises a memory 250 and the above-mentioned data processing apparatus.
[0205] As shown in FIG. 8 , the memory 250 is used to store input data, intermediate result data and output result data required by the data processing apparatus 200 for performing matrix multiplication calculations.
[0206] For the memory 250 in the embodiments of the present application, it refers to a hardware module providing a large-capacity data storage function in the electronic device 800, which can be understood in the embodiments of the present application as an off-chip memory device, providing a data source and a result storage space for the calculation operation of the data processing apparatus 200. The memory 250 realizes hierarchical storage and efficient access of data by cooperating with the cache unit 210 in the data processing apparatus 200.
[0207] For example, the memory 250 can be a dynamic random-access memory (DRAM), a static random-access memory (SRAM), a non-volatile memory, a solid state drive (SSD), a hard disk drive (HDD), or other types of data storage devices.
[0208] In the embodiments of the present application, the processor 230 reads the original data of the second initial matrix from the memory 250, performs a data reordering operation to generate a second to-be-processed matrix, and writes the reordered data to the cache unit 210. At the same time, the pulse array 220 reads the data of the first to-be-processed matrix and the third to-be-processed matrix from the memory 250 to the cache unit 210 according to the state information in the first storage module 240. After the data processing apparatus 200 completes the matrix multiplication calculation, the calculation result is written back from the cache unit 210 to the memory 250, realizing the storage of the calculation result.
[0209] For the electronic device 800 in the embodiments of the present application, it can be a computing processing system integrating the above-mentioned data processing apparatus 200, which can be understood in the embodiments of the present application as a computing device with matrix multiplication calculation capability, realizing the data processing function by carrying the data processing apparatus 200 provided in the embodiments of the present application. The electronic device 800 serves as a carrier platform for the data processing apparatus 200, providing necessary power supply, system control, input / output interface, and communication connection function with external systems for the data processing apparatus 200.
[0210] It should be noted that the electronic device 800 can integrate one or more data processing apparatuses 200 at the same time to meet the matrix calculation requirements of different scales and complexities. When the electronic device 800 is configured with multiple data processing apparatuses 200, parallel computing can be realized through task allocation and load balancing, and continuous calculation tasks can be efficiently executed through pipeline processing.
[0211] Exemplarily, the electronic device 800 can be a server, a personal computer, a workstation, a mobile computing device, a tablet computer, a smartphone, an embedded computing device, an edge computing device, a special-purpose computing accelerator, a graphics processing device, a digital signal processing device, or other electronic devices 800 that need to perform large-scale matrix operations.
[0212] For the server in the embodiments of the present application, it refers to a high-performance computing device deployed in a data center or a cloud computing environment, which provides matrix computation services for multiple clients by integrating the data processing apparatus 200. For the personal computer in the embodiments of the present application, it refers to a desktop or notebook computing device oriented to personal users, which supports local artificial intelligence applications by built-in data processing apparatus 200. For the mobile computing device in the embodiments of the present application, it refers to a computing device with portability features, such as a smartphone, a tablet computer, etc., which realizes computing functions on the mobile end by integrating a miniaturized data processing apparatus 200.
[0213] For the embedded computing device in the embodiments of the present application, it refers to a special-purpose computing module integrated in other systems or products, which provides matrix operation capabilities for the host system by carrying the data processing apparatus 200. For the edge computing device in the embodiments of the present application, it refers to a computing device deployed at the edge node of the network, which realizes local data processing and inference computation by configuring the data processing apparatus 200.
[0214] In actual deployment applications, the electronic device 800 can select different system configurations and integration methods according to specific application requirements and use scenarios. The electronic device 800 can integrate the data processing apparatus 200 as a coprocessor unit of the main processor, or externally extend the data processing apparatus 200 as an independent computing accelerator card. The electronic device 800 can also establish a data transmission connection with the data processing apparatus 200 through a system bus, a high-speed interconnection interface, or a special data channel, to realize efficient data interaction between the main system and the data processing apparatus 200.
[0215] In a feasible implementation manner, the electronic device 800 can also configure multiple data processing apparatuses 200 to construct a distributed computing architecture, and further improve the overall matrix computation processing capability through parallel deployment.
[0216] In another feasible implementation manner, the electronic device 800 can also dynamically configure and adjust the running mode of the data processing apparatus 200 according to the power consumption budget and performance requirements, to realize balanced optimization between computing performance and energy efficiency.
[0217] It should be noted that the specific form and configuration of the electronic device 800 described above are only illustrative, and the embodiments of the present application do not limit the specific implementation form of the electronic device. In actual application, the electronic device can select a suitable hardware platform and system architecture according to specific application scenarios, performance requirements, power consumption constraints, and cost budget and other factors.
[0218] Those skilled in the art can understand that the features described in various embodiments of the present application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present application. In particular, the features described in various embodiments of the present application can be combined and / or combined in various ways without departing from the spirit and teachings of the present application. All these combinations and / or combinations fall within the scope of the present application.
[0219] The embodiments of the present application are described above. However, these embodiments are only for illustrative purposes, and are not intended to limit the scope of the present application. Although each embodiment is described above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present application, those skilled in the art can make various alternatives and modifications, which should fall within the scope of the present application.
Claims
1. A data processing apparatus, characterized by, The method comprises the following steps: a cache unit is configured to store a first to-be-processed matrix and a second to-be-processed matrix, the first to-be-processed matrix comprises a plurality of first to-be-processed elements, and the second to-be-processed matrix comprises a plurality of to-be-processed vector data, the to-be-processed vector data comprises a plurality of second to-be-processed elements; a pulse array comprises a plurality of pulse sub-arrays, each pulse sub-array comprises a plurality of computing units, and each computing unit is configured to: determine a target calculation result of the computing unit according to a received first to-be-processed element and a second to-be-processed element for the computing unit; provide the first to-be-processed element to a subsequent hardware unit for the computing unit in a target external pulse sub-array, the target external pulse sub-array is determined by moving a first moving bit number in a first direction from a source pulse sub-array in which the computing unit is located among the plurality of pulse sub-arrays; wherein the to-be-processed vector data in the second to-be-processed matrix is obtained by: moving a plurality of elements in initial vector data of a second initial matrix in a second direction by a second moving bit number, the second moving bit number is determined according to position information of the initial vector data in the second initial matrix and a preset value; wherein the cache unit is further configured to store a third to-be-processed matrix, the third to-be-processed matrix comprises a plurality of third to-be-processed elements; the plurality of computing units in the pulse sub-array comprises a head computing unit, and the head computing unit is further configured to: multiply the first to-be-processed element for the head computing unit and the second to-be-processed element for the head computing unit received from the cache unit to obtain a multiplication result of the head computing unit; add the third to-be-processed element for the head computing unit received from the cache unit and the multiplication result of the head computing unit to obtain a target calculation result of the head computing unit.
2. The apparatus of claim 1, wherein, The computing unit is configured to: multiply the received first to-be-processed element and the second to-be-processed element for the computing unit to obtain a multiplication result of the computing unit; determine a target calculation result of the computing unit according to the multiplication result of the computing unit and a previous calculation result provided by a previous computing unit in the source pulse sub-array; provide the target calculation result of the computing unit to a subsequent hardware unit in the source pulse sub-array.
3. The apparatus of claim 1, wherein, The computing unit is configured to: provide the second to-be-processed element for the subsequent hardware unit in the source pulse sub-array to the subsequent hardware unit in the source pulse sub-array.
4. The apparatus of claim 1, wherein, The pulse sub-array further comprises an output unit, the subsequent hardware unit is a computing unit or an output unit in the pulse sub-array, the output unit is configured to: multiply the received first to-be-processed element and the second to-be-processed element for the output unit to obtain a multiplication result of the output unit; Add the multiplication result of the output unit to the previous calculation result provided by the previous calculation unit in the source pulse sub-array to obtain an output result of the output unit; Store the output result to the cache unit.
5. The apparatus of claim 1, wherein, The plurality of to-be-processed vector data is I to-be-processed vector data, The processor is further configured to perform the following operations on the I to-be-processed vector data: cyclically move a plurality of elements in initial vector data of a second initial matrix along a second direction by a second moving bit number: Cyclically move the plurality of elements in the i-th initial vector data in the second initial matrix along the second direction by i-1 bit numbers to obtain a second to-be-processed matrix; And Store the second to-be-processed matrix to the cache unit.
6. The apparatus of claim 5, wherein, The first storage module is further configured to store a plurality of state information and address information of the second to-be-processed matrix sent by the processor and running state information of the pulse array, The pulse array is further configured to read the second to-be-processed matrix from the cache unit according to the plurality of state information and address information of the second to-be-processed matrix; The processor is further configured to generate a second to-be-processed matrix required for the next round calculation of the pulse array according to the running state information of the pulse array.
7. The apparatus of claim 6, wherein, The pulse array is further configured to read the second to-be-processed matrix from the cache unit according to the plurality of state information of the second to-be-processed matrix, the address information of the second to-be-processed matrix in the cache unit, and the data quantity of the second to-be-processed matrix; The first storage module comprises: The rearrangement data state register is configured to store at least one of the plurality of state information of the second to-be-processed matrix sent by the processor; the plurality of state information comprises first state information and second state information, the first state information is used to indicate whether the processor generates the second to-be-processed matrix, and the second state information is used to indicate whether the processor stores the second to-be-processed matrix to the cache unit; The address synchronization register is configured to store the address information of the second to-be-processed matrix in the cache unit; The length synchronization register is configured to store the data quantity information of the second to-be-processed matrix; The array state register is configured to store the running state information of the pulse array.
8. The apparatus of claim 1, wherein, Each of the calculation units comprises a plurality of temporary storage modules configured to store second to-be-processed elements for the calculation unit in different rounds respectively.
9. The apparatus of claim 8, wherein, Each of the calculation units further comprises a second storage module configured to store at least one of a plurality of preloading indication values corresponding to the plurality of temporary storage modules, the preloading indication value being used to indicate that the calculation unit stores the second to-be-processed element for the calculation unit in the next round to the temporary storage module corresponding to the plurality of preloading indication values.
10. The apparatus of claim 9, wherein, The second storage module further comprises: The preloading counter is configured to obtain the number of second to-be-processed elements received by the calculation unit in the current round; The preloading constant register is configured to store a preloading trigger threshold value of the calculation unit, the preloading trigger threshold value being the difference between the row sequence number of the calculation unit in the pulse array and a preset value. The computing unit further comprises a comparator configured to: in a case that the number of the second pending elements received in the current round is equal to the preloading trigger threshold, store the received second pending elements consistent with the preloading trigger threshold to one of the plurality of staging modules as the second pending elements for the computing unit in the next round.
11. The apparatus of claim 10, wherein, The plurality of staging modules comprises a first staging module and a second staging module, and the plurality of preloading indication values comprises a first preloading indication value and a second preloading indication value, The computing unit further comprises a preloading selection register configured to store the first preloading indication value or the second preloading indication value, the first preloading indication value being configured to indicate that the second pending elements for the computing unit in the next round are stored to the first staging module, and the second preloading indication value being configured to indicate that the second pending elements for the computing unit in the next round are stored to the second staging module. The computing unit is further configured to perform at least one of the following operations: in a case that the systolic array completes the computation of the current round, switch the preloading selection register from the first preloading indication value to the second preloading indication value; in a case that the systolic array completes the computation of the current round, switch the preloading selection register from the second preloading indication value to the first preloading indication value.
12. The apparatus of claim 10, wherein, The second storage module further comprises: a computation control register configured to store a computation indication value, the computation indication value being configured to indicate that the computing unit selects the second pending elements from the plurality of staging modules for multiplication computation.
13. The apparatus of claim 12, wherein, The computing unit further comprises: a multiply-accumulate calculator; a selector configured to: determine the selected second pending elements from the plurality of staging modules according to the computation indication value stored in the computation control register; and provide the selected second pending elements to the multiply-accumulate calculator for multiplication operation.
14. An electronic device, comprising: comprise: a memory; the data processing apparatus of any one of claims 1-13.
Citation Information
Patent Citations
Systolic convolutional neural network
CN111937009A