A data processing method, apparatus, and data processing chip

By sparsely processing the weight matrix and feature map of the neural network model, and using efficient internal product multiplication structure to optimize the sparse matrix multiplication operation, the problems of low neural network computing performance and low resource utilization are solved, and more efficient data processing is achieved.

CN118333127BActive Publication Date: 2025-07-22SMARTER SILICON (SHANGHAI) TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410742115.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-07
Publication Date
2025-07-22
Estimated Expiration
2044-06-07

AI Technical Summary

Technical Problem

In the prior art, the calculation performance of neural networks is low, the storage and transmission resource requirements are high, and the resource utilization rate is low. Especially in the Transformer network, the sparsity of the weight matrix and feature maps are not effectively utilized, resulting in low computing efficiency.

Method used

By sparsely processing the weight matrix and feature map of the neural network model, effective data is compressed, the number of data objects is reduced, and the multiplication operation of the sparse matrix is optimized by using an efficient internal product multiplication structure to eliminate invalid data to improve computing performance and resource utilization.

Benefits of technology

It effectively improves the computing performance of neural networks, reduces the storage and transmission resource requirements, improves the system resource utilization and computing efficiency, and ensures the accuracy of data processing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118333127B_ABST
    Figure CN118333127B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing method, apparatus, and data processing chip, belonging to the field of artificial intelligence. The data processing chip includes a plurality of available hardware processing channels and a plurality of second arithmetic components; each available hardware processing channel includes a first arithmetic component, configured to obtain corresponding first target data in at least one first target data sub-object and obtain second target data corresponding to the corresponding first target data from a second data object corresponding to the target data processing channel, and perform a first arithmetic process on the obtained first and second target data; the first target data in the at least one first target data sub-object includes all valid data of a first data object corresponding to the target data processing channel; each second arithmetic component is configured to perform a second arithmetic process on a first arithmetic result output by a corresponding available hardware processing channel; the number of first target data sub-objects is less than the number of first data sub-objects in the first data object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of artificial intelligence, and particularly relates to a data processing method, apparatus, and data processing chip. Background Art

[0002] A neural network (NN) is a complex network system formed by a large number of simple processing units (called neurons) widely interconnected. The Transformer network is a neural network based on the self-attention mechanism. Currently, in data processing based on neural networks such as the Transformer network, there are problems such as low computing performance, high resource requirements for storage and transmission, and low resource utilization. How to solve at least some of these problems has become a technical difficulty in this field. Summary of the Invention

[0003] For this reason, the present application discloses the following technical solutions:

[0004] A data processing chip includes: a plurality of available hardware processing channels and a plurality of second arithmetic components;

[0005] Each of the available hardware processing channels includes a first arithmetic component, configured to obtain corresponding first target data in at least one first target data sub-object. The first target data in the at least one first target data sub-object includes all valid data of each first data sub-object in a first data object corresponding to a target data processing channel. Each first target data corresponds to corresponding position information, configured to indicate the original position of the first target data in the first data object; and configured to obtain second target data corresponding to the corresponding first target data from a second data object corresponding to the target data processing channel according to the position information corresponding to the corresponding first target data; and perform a first arithmetic process on the obtained first target data and second target data to obtain a first arithmetic result;

[0006] Each second arithmetic component is configured to perform a second arithmetic process on the first arithmetic result output by the corresponding available hardware processing channel to obtain a second arithmetic result;

[0007] Wherein, the number of the first target data sub-objects is less than the number of the first data sub-objects in the first data object.

[0008] Optionally, the first data object includes a first data matrix having a plurality of data to be processed;

[0009] The first data sub-object is a column in the first data matrix;

[0010] The at least one first target data sub-object includes: at least one first target column containing valid data obtained by moving the valid data in the corresponding column of the first data matrix to the position of the invalid data in other columns outside the corresponding column of the first data matrix; the data in the same column of the first data matrix is in the same first target column after the movement, and the data in the same row is in different first target columns after the movement, and the number of the first target columns is less than the number of columns included in the first data matrix.

[0011] The second data object includes a second data matrix having a plurality of data to be processed.

[0012] Optionally, the first arithmetic processing includes a multiplication operation, and the multiplication operation includes the multiplication operation involved in multiplying the first data matrix and the second data matrix; the second arithmetic processing includes an accumulation operation.

[0013] When performing the first arithmetic processing on the obtained first target data and second target data, the first arithmetic component is specifically configured to: perform a multiplication operation on the first target data in the obtained corresponding first target column and the corresponding second target data to obtain a multiplication operation result.

[0014] When performing the second arithmetic processing on the first arithmetic result output by the corresponding available hardware processing channel, the second arithmetic component is specifically configured to: perform an accumulation process on the multiplication operation result output by the corresponding available hardware processing channel to obtain an accumulation operation result.

[0015] Optionally, the data processing chip further includes a routing component.

[0016] Each of the available hardware processing channels further includes an independent storage component for storing the first target data allocated to the corresponding available hardware processing channel.

[0017] The routing component is configured to send the multiplication operation result output by the available hardware processing channel to the corresponding second arithmetic component for accumulation processing according to the position information corresponding to the first target data and the second target data in the available hardware processing channel.

[0018] Wherein, the position information corresponding to the second target data is used to indicate the position of the second target data in the second data object; the multiplication operation results corresponding to the first target data originally in the same row of the first data matrix are sent to the same second arithmetic component for accumulation processing.

[0019] A data processing method includes:

[0020] Obtain at least one first target data sub-object containing first target data; the first target data in the at least one first target data sub-object includes at least all valid data of each first data sub-object in the first data object corresponding to the target data processing channel; each first target data corresponds to corresponding position information, which is used to indicate the original position of the first target data in the first data object;

[0021] According to the position information corresponding to each first target data in the at least one first target data sub-object, obtain second target data corresponding to each first target data from the second data object corresponding to the target data processing channel;

[0022] Perform data processing on each first target data and the corresponding second target data;

[0023] Wherein, the number of the first target data sub-objects is less than the number of the first data sub-objects in the first data object.

[0024] Optionally, the method for forming the first target data sub-object includes:

[0025] Obtain the first data object corresponding to the target data processing channel in the model;

[0026] Move the valid data of the corresponding first data sub-object in the first data object to the position where the invalid data of other first data sub-objects is located, so as to reduce the number of first data sub-objects included in the first data object;

[0027] Wherein, the at least one first target data sub-object includes: at least the first data sub-object containing valid data obtained after the movement; the other first data sub-objects include the first data sub-objects other than the corresponding first data sub-object in the first data object.

[0028] Optionally, the first data object includes a first data matrix having a plurality of data to be processed, and the first data sub-object is a column in the first data matrix;

[0029] The moving the valid data of the corresponding first data sub-object in the first data object to the position where the invalid data of other first data sub-objects is located includes:

[0030] Move the valid data of the corresponding column in the first data matrix to the position where the invalid data of other columns other than the corresponding column in the first data matrix is located, so as to reduce the number of columns included in the first data matrix;

[0031] Among them, the at least one first target data sub-object includes: a first target column that contains at least valid data after movement; data in the same column of the first data matrix is in the same first target column after movement, and data in the same row is in different first target columns after movement.

[0032] Optionally, the method for forming the first target data sub-object further includes:

[0033] Performing data order adjustment processing on the data in each of the first target columns, so that the valid data originally belonging to the same column in the first data matrix is continuously arranged in the first target column where it is located after movement.

[0034] Optionally, the second data object includes a second data matrix having a plurality of data to be processed;

[0035] The obtaining of the second target data corresponding to each of the first target data from the second data object corresponding to the target data processing channel according to the position information corresponding to each of the first target data in the at least one first target data sub-object includes:

[0036] Obtaining the second target data corresponding to each of the first target data from the second target column of the second data matrix according to the position information of each of the first target data in each of the first target columns; the second target column is the column currently to be processed in the second data matrix.

[0037] Optionally, the data processing of each of the first target data and the corresponding second target data includes:

[0038] Allocating the corresponding number of first target data to each available hardware processing channel according to the number of available hardware processing channels for parallel processing, and using the available hardware processing channels to perform data processing on the allocated first target data and the corresponding second target data;

[0039] Among them, the number of available hardware processing channels required for processing each of the first target data is less than the number of available hardware processing channels required for processing each of the data to be processed in the first data object.

[0040] Optionally, the valid data originally belonging to the same column in the first data matrix is continuously arranged in the first target column where it is located after movement; the data processing includes multiply-accumulate processing, and the multiplication operation in the multiply-accumulate processing includes the multiplication operation involved in multiplying the first data matrix and the second data matrix;

[0041] The data processing of each of the first target data and the corresponding second target data includes:

[0042] All the first target data in the same group in each first target column are respectively allocated to consecutive available hardware processing channels in the channel array; the first target data in the same group includes the first target data belonging to the same original column in the first data matrix within the first target column.

[0043] According to the position information corresponding to the first target data in the same group, determine the second target data corresponding to the first target data in the same group; the first target data in the same group corresponds to the same second target data.

[0044] Allocate the determined second target data to the available hardware processing channels where the first target data in the corresponding group are respectively located, so that the first arithmetic component provided by the available hardware processing channel performs a multiplication operation on the allocated first target data and second target data to obtain a multiplication operation result.

[0045] According to the position information corresponding to each first target data, send the multiplication operation results corresponding to the first target data belonging to the same original row in the first data matrix to the same second arithmetic component for accumulation processing to obtain an accumulation operation result.

[0046] Wherein, the position information corresponding to the second target data is used to indicate the position of the second target data in the second data object.

[0047] A data processing device, comprising:

[0048] A first acquisition module, configured to acquire at least one first target data sub-object including first target data; the first target data in the at least one first target data sub-object at least includes all valid data of each first data sub-object in the first data object corresponding to the target data processing channel; each first target data corresponds to corresponding position information, which is used to indicate the original position of the first target data in the first data object.

[0049] A second acquisition module, configured to acquire, according to the position information respectively corresponding to each first target data in the at least one first target data sub-object, second target data corresponding to each first target data from the second data object corresponding to the target data processing channel.

[0050] A data processing module, configured to perform data processing on each first target data and the corresponding second target data.

[0051] Wherein, the number of the first target data sub-objects is less than the number of the first data sub-objects in the first data object.

[0052] An electronic device, at least comprising:

[0053] A memory for storing a computer instruction set;

[0054] A processor for implementing the data processing method provided in any one of the above by executing the computer instruction set.

[0055] A readable storage medium storing a computer instruction set, the computer instruction set being used to be called and executed by a processor to implement the data processing method described in any one of the above.

[0056] In addition, a computer program product is also provided, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the data processing method described in any one of the above is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0058] Figure 1 is an example of the sparse weight matrix and the sparse feature map provided by the present application;

[0059] Figure 2 is a flowchart of a data processing method provided by the present application;

[0060] Figure 3 is an example of matrix multiplication operation provided by the present application;

[0061] Figure 4 is a flowchart of a method for forming a first target data sub-object provided by the present application;

[0062] Figure 5 is an example of compressing a first data object in the row direction provided by the present application;

[0063] Figure 6 is an example of compressing a first data object in the column direction provided by the present application;

[0064] Figure 7 is an example of a first target column obtained after compressing a first data object in the row and column directions provided by the present application;

[0065] Figure 8 is another flowchart of a data processing method provided by the present application;

[0066] Figure 9 is a schematic diagram of the internal structure of a PE provided by the present application;

[0067] Figure 10 It is a schematic flowchart of data processing for each first target data and the corresponding second target data provided by this application;

[0068] Figure 11 It is a schematic diagram of associating valid second target data with each group by index provided by this application;

[0069] Figure 12 It is a mapping connection diagram between the PE array and the accumulator provided by this application;

[0070] Figure 13 It is a composition structure diagram of the data processing device provided by this application;

[0071] Figure 14 It is a composition structure diagram of the data processing chip provided by this application;

[0072] Figure 15 It is a composition structure diagram of the electronic device provided by this application. Specific embodiments

[0073] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0074] Currently, in data processing based on neural network models, there are a series of problems such as low computing performance, high resource requirements for storage and transmission, and low resource utilization.

[0075] Specifically, in neural network models, quantization and pruning operations are often performed on network layer weights, resulting in a large number of 0 values in the weight matrix. At the same time, due to the ReLU (activation) operation, a large number of 0 values also appear in the feature map. For example, in Figure 1 the weight matrix and feature map shown, there are a large number of 0 values (where the gray blocks represent non-0 values and the white blocks represent 0 values). This phenomenon of a large number of 0 values in such a network is called sparsification. In particular, in the Transformer network, due to the local correlation of tokens (a token refers to the smallest unit with independent semantics, each token represents an independent unit, has a certain semantic meaning, and can be processed by the model), 0 values (sparsity) are more common.

[0076] The applicant has found that the operations in a neural network mainly include multiplication and addition (for example, the multiplication and addition operations involved in the matrix multiplication of the weight matrix in a Transformer network). A value of 0 makes no contribution to the final calculation result. If only the valid values are transmitted and stored during data transmission and storage, the bandwidth required for transmission and storage can be greatly reduced. If the 0 values are skipped during data calculation, the computing performance of the system can be greatly improved, and the resource utilization rate of the system can be enhanced.

[0077] However, the current related hardware responsible for data processing in a neural network model, such as related commercial chips, does not support weight sparsification processing. The 0-value weights still participate in the processing and occupy computing time. Therefore, how to improve the computing performance during the processing based on the sparsification characteristics of the weights in the model network, reduce the data storage and bandwidth requirements, and enhance the resource utilization rate and operation efficiency has become a difficult point.

[0078] Based on this, the present application provides a data processing method, apparatus, and data processing chip, mainly aiming at the matrix multiplication operation of the weight matrix in neural networks such as Transformer. By utilizing the characteristic that the weight matrix is fixed and known, an efficient inner product multiplication structure is proposed to optimize the multiplication operation of the sparse matrix, so as to solve various problems existing in the known technology.

[0079] The data processing method, apparatus, and data processing chip provided by the present application can be applied to, but are not limited to, electronic devices such as personal computers or servers.

[0080] See Figure 2 the flowchart of the data processing method shown, the provided data processing method at least includes the following processing steps:

[0081] Step 201, obtain at least one first target data sub-object including first target data; the first target data in the at least one first target data sub-object at least includes all valid data of each first data sub-object in the first data object corresponding to the target data processing channel; each first target data corresponds to corresponding position information, which is used to indicate the original position of the first target data in the first data object.

[0082] Wherein, the number of the first target data sub-objects is less than the number of the first data sub-objects in the first data object.

[0083] The method provided by the present application can be, but is not limited to, applicable to various fields such as natural language processing, image processing, video processing, speech recognition, and industrial detection (such as equipment defect detection).

[0084] The embodiments of the present application mainly take the data processing of a neural network model (such as a neural network model based on Transformer) as an example to illustrate the solution.

[0085] The target data processing channel can be, but is not limited to, the input channel of the network layer of a neural network model. For example, for image processing based on a neural network model, the target data processing channel can specifically include any one or more of the R, G, and B primary color input channels, texture input channel, and semantic input channel of the model network layer.

[0086] The valid data in the first data object specifically refers to the data contained in the first data object that is valuable for the data processing of the first data object, while the data contained in the first data object that is not valuable for its data processing is regarded as the invalid data or non - valid data of the first data object.

[0087] Optionally, the first data object is a data matrix including multiple data to be processed, and each first data sub - object contained in the first data object is each column in the data matrix. For the data processing scenario of a neural network model, the first data object can specifically be the weight matrix corresponding to the corresponding input channel of the network layer of the neural network model. Each first data sub - object in the first data object can be each column contained in the weight matrix. The valid data in the first data sub - objects contained in the first data object is the non - zero value in the columns contained in the weight matrix. Since the non - zero value contributes to the operation of the model network, the non - zero weights are regarded as the valid data in the first data sub - objects contained in the first data object. Correspondingly, since the zero value does not contribute to the operation of the model network at all, the zero - valued weights in the weight matrix are regarded as invalid data.

[0088] In order to improve the computing performance of the system, reduce the storage and transmission bandwidth required by the system, and improve the resource utilization rate of the system, the embodiments of the present application propose to compress the data in the first data object (such as the weight matrix of the network layer of a neural network model) according to the sparse characteristics of the first data object, so as to reduce the number of first data sub - objects contained in the first data object, thereby optimizing the data processing of the first data object (such as optimizing the matrix multiplication operation of a sparse matrix), and solving various problems existing in the known technology based on this technical idea.

[0089] Among them, for the case where the first data object is a data matrix including multiple data to be processed, and each first data sub-object included in the first data object is each column in the data matrix, the data in the first data object is compressed, which at least includes compressing the data matrix of the first data object in the row direction. By compressing the data matrix of the first data object in the row direction, all the valid data in each original column of the data matrix of the first data object is aggregated into a part of the columns in the original columns, so as to reduce the number of the first data sub-objects included in the first data object, so as to optimize the data processing of the first data object.

[0090] More specifically, by making the valid data in the corresponding original column of the data matrix of the first data object occupy the positions of the invalid data such as the 0-value weight in other original columns (one or more columns other than the corresponding original column) of the data matrix, all the valid data of the data matrix is aggregated into a part of the columns of the data matrix, so that the corresponding part of the columns of the data matrix at least includes valid data, while the other corresponding part of the columns does not include any valid data (that is, all are invalid data), so that the columns that do not include any valid data can be directly removed, realizing the compression of the data matrix of the first data object and reducing the number of the first data sub-objects (columns) included therein.

[0091] The at least one first target data sub-object is the first data sub-object that at least includes valid data obtained after compressing the data in each first data sub-object included in the first data object based on the above technical idea. For example, the columns that at least include valid data obtained after aggregating all the valid data in each original column of the data matrix, and the first data sub-objects (columns) that do not include any valid data are removed and no longer participate in the subsequent data processing of the first data object.

[0092] It should be emphasized that in order to achieve effective compression, it is required that after compressing the data in the first data object, the number of the obtained first target data sub-objects is less than the number of the first data sub-objects in the first data object; and the first target data in the at least one obtained first target data sub-object after compression at least includes all the valid data of each first data sub-object in the first data object. For example, based on the above technical idea of compression processing, the number of columns that at least include valid data obtained after aggregating the valid data of the data matrix is less than the number of each original column included in the data matrix; and each data in the columns that at least include valid data obtained after compression at least includes all the valid data of each original column of the data matrix. In addition to including all the valid data, it may also include a part of the invalid data, and of course, it may not include any invalid data, depending on the actual situation.

[0093] That is to say, through the above compression processing, at least part of the invalid data in the first data object is removed, and all valid data is retained, so as to reduce the data processing amount of the first data object while avoiding affecting the data processing result of the first data object and ensuring the accuracy of the data processing result.

[0094] For the case where the first data object is the weight matrix of the network layer of the neural network model, in practical applications, when the model training is completed, the weight matrix of the model network layer can be compressed based on the above idea, and the at least one first target data sub-object obtained by the compression processing (such as each column containing at least valid data) can be stored. Subsequently, when it is necessary to use the model for data processing, the data of the stored at least one first target data sub-object can be directly read and the required processing can be performed on it. However, this is not limited to this. In other embodiments, the weight matrix of each network layer of the model can also be compressed in real time when it is necessary to use the model for data processing, and the required processing can be performed on the at least one first target data sub-object obtained after compression, which can be determined according to the actual application requirements.

[0095] Each first target data corresponds to corresponding position information, which is used to indicate the original position of the first target data in the first data object. The position information corresponding to the first target data can specifically include a row index and a column index, which are respectively used to indicate the original row and the original column where the first target data is located in the first data object.

[0096] Step 202: Obtain second target data corresponding to each of the first target data from the second data object corresponding to the target data processing channel according to the position information corresponding to each of the first target data in the at least one first target data sub-object.

[0097] Optionally, the second data object can also be a data matrix including a plurality of data to be processed. For the data processing scenario of the neural network model, the second data object can specifically be the feature map corresponding to the corresponding input channel of the network layer of the neural network model, and the second target data can be the corresponding feature value in the feature map.

[0098] The second data object can also include valid data and invalid data. The valid data of the second data object specifically refers to the data contained in the second data object that contributes to the data processing of the second data object, while the data contained in the second data object that does not contribute to its data processing is regarded as the non-valid data or invalid data of the second data object. Taking the second data object as the feature map as an example, based on the characteristic of whether the data operation of the feature map by the feature value contributes, the non-0 feature values in the feature map are determined as the valid data of the feature map, and the 0-valued feature values are regarded as invalid data.

[0099] Among them, the feature map can be, but is not limited to, various types of data to be processed such as images and voices, depending on the specific application scenario.

[0100] For each piece of data to be processed in the first data object, there is corresponding data to be processed in the second data object, which is used to match into a pair of data to be processed between the first data object and the second data object, so as to perform the required data processing on it (the pair of data to be processed), such as performing a multiplication operation on the two pieces of data to be processed included in the pair of data to be processed, and accumulating the multiplication results of corresponding different pairs of data to be processed, etc.

[0101] Whether a piece of data to be processed in the first data object matches a piece of data to be processed in the second data object (that is, whether it should be matched into a corresponding pair of data to be processed) depends on the positions of these two pieces of data to be processed in their respective data objects. The data at the matching positions between the first data object and the second data object are correspondingly the matching data to be processed. Further, the matching positions between the first data object and the second data object specifically depend on the data processing rules for the first data object and the second data object.

[0102] Based on this, for each first target data in the at least one first target data sub-object (essentially the corresponding data to be processed in the first data object), the second target data corresponding to this first target data can be obtained from the second data object corresponding to the target data processing channel according to the position information corresponding to the first target data.

[0103] Among them, specifically, according to the data processing rules for the first data object and the second data object, the position information in the second data object that matches the position information corresponding to the first target data can be determined, and the data to be processed at the position indicated by this matching position information in the second data object is obtained as the second target data corresponding to this first target data. More specifically, for the case where the first and second data objects are data matrices, according to the data processing rules for the first data object and the second data object, the row index and column index in the second data object that match the row index and column index corresponding to the first target data can be determined, and the data to be processed at the row and column positions indicated by this matching row index and column index in the second data object is obtained as the second target data corresponding to this first target data.

[0104] The data processing of the first data object and the second data object in the embodiments of this application mainly refers to the matrix multiplication operation of the first data object and the second data object. For example, the matrix multiplication operation of a weight matrix and a feature map. And by utilizing the characteristic that the weight matrix is fixed and known, this application proposes an efficient inner product multiplication structure to optimize the multiplication operation of sparse matrices. That is to say, the data processing rule of the first data object and the second data object in this application is specifically based on the matrix multiplication operation rule of the inner product.

[0105] The matrix multiplication operation based on the inner product can be applied to, but is not limited to, the matrix multiplication processing of the weight matrix and the feature map in the network layer of a large language model (LLM). For example, the matrix multiplication processing of the weight matrix and the feature map in the network layer of a large language model based on the Transformer network, etc.

[0106] For the matrix multiplication operation based on the inner product, it is necessary to multiply the rows of the data matrix of the first data object with the columns in the data matrix of the second data object. Specifically, it is necessary to correspond the data in each row of the data matrix of the first data object with the data in each column of the data matrix of the second data object in sequence one by one, and perform a multiplication operation on the data pairs formed by the corresponding data, and accumulate the results of each multiplication operation obtained by multiplying the data in the same row of the first data object with the data in the same column of the second data object.

[0107] Among them, for the rows in the first data object, the sequence is from left to right, and for the columns in the second data object, the sequence is from top to bottom.

[0108] It is easy to understand that the above matrix multiplication operation based on the inner product essentially requires matching the data with the same column index in the row currently participating in the operation (i.e., the row to be processed) of the first data object with the data with the same row index in the column currently participating in the operation (i.e., the column to be processed) of the second data object into data pairs to be processed. For example Figure 3 In the example, assume that the first data object is matrix A (Matrix A), the second data object is matrix B (Matrix B), and assume that the row currently participating in the operation in matrix A is the first row of matrix A, and the column currently participating in the operation in matrix B is the first column of matrix B. Then, in the inner product operation of the first row of matrix A and the first column of matrix B, it is required to match the data with the same column index in the first row of matrix A and the data with the same row index in the first column of matrix B into data pairs to be processed. For example, the data pairs represented by (11, 1), (41, 4), and (71, 7) in this example. In Matrix A and Matrix B, the white boxes represent invalid data, and the gray boxes represent valid data.

[0109] Based on this, for the matrix multiplication operation based on inner product, in this step, specifically according to the column index corresponding to the first target data in the first data object, the data at the row position corresponding to this column index is obtained from the corresponding column currently participating in the operation in the second data object as the second target data corresponding to the first target data. This second target data is also the data to be processed that matches the first target data and is used to form a data pair to be processed with the first target data to participate in subsequent data processing.

[0110] It should be noted that the data processing of the weight matrix and the feature map in the neural network mainly includes two types: convolution operation and matrix multiplication operation (such as the matrix multiplication operation based on inner product in this application). Currently popular large language models, such as those based on the Transformer network, use matrix multiplication operation for the operation of the weight matrix and the feature map. In practical applications, the convolution operation can also be converted into the form of matrix multiplication operation through corresponding conversion rules, and the data processing method provided in the embodiments of this application can be used to implement it accordingly.

[0111] The neural network model can specifically perform one-dimensional convolution, two-dimensional convolution, or three-dimensional convolution on the feature map without limitation, which can be determined according to actual needs. For example, for a one-dimensional convolution kernel with a size of 1x3, one-dimensional convolution can be performed on a 1x3 feature map based on a 1x3 weight matrix. For a two-dimensional convolution kernel with a size of 3x3, two-dimensional convolution can be performed on a 3x3 feature map based on a 3x3 weight matrix accordingly.

[0112] Step 203: Perform data processing on each of the first target data and the corresponding second target data.

[0113] Optionally, the data processing performed on each of the first target data and the corresponding second target data may include multiply-accumulate processing. That is, first perform a multiplication operation on the currently processed first target data and its corresponding second data object, and then accumulate the corresponding multiplication results, but not limited to this. When applying this application, the data processing performed can be determined according to actual application requirements.

[0114] Among them, for the above-mentioned matrix multiplication operation based on inner product, after performing the multiplication operation on each first target data and its corresponding second target data, for the corresponding column currently participating in the processing of the second data object, specifically, the multiplication results corresponding to the first target data belonging to the same original row in the first data object can be accumulated, and the accumulated result is used as one of the result data in the data processing results of the first data object and the second data object.

[0115] In summary, the data processing method provided by the embodiments of the present application compresses the data included in the first data object based on the sparsity characteristics of the data in the first data object. By compressing each first data sub-object in the first data object into at least one first target data sub-object and making the number of the first target data sub-objects less than the number of the first data sub-objects, the data processing amount of the first data object is effectively reduced, thereby improving the computing performance of the system, reducing the resource requirements for storage, transmission, and operation, and improving the system resource utilization rate and operation efficiency. For application scenarios such as natural language processing, image processing, video processing, speech recognition, and industrial inspection, the processing efficiency of various applications such as natural language processing, image processing, and speech recognition can be correspondingly improved, and the system resource utilization rate can be improved.

[0116] At the same time, since the first target data in the at least one first target data sub-object obtained after compression includes all the valid data of each first data sub-object in the first data object, the compression process only eliminates at least part of the invalid data in the first data object and retains all the valid data, without affecting the data processing result of the first data object, ensuring the accuracy of the data processing result. And since the essence of the present application is to eliminate at least part of the invalid data in the first data object based on soft processing, that is, before sending the first data object to the hardware for processing, at least part of the invalid data has been removed. Therefore, the data processing method provided by the present application is still applicable to related hardware that currently does not support weight sparsity processing (0-value weights still participate in processing and occupy computing time), such as related commercial chips.

[0117] In an alternative embodiment, referring to Figure 4 the flowchart of the method for forming the first target data sub-object shown, based on the compression idea described above, the data in the first data object is compressed to form the method for forming the first target data sub-object, including:

[0118] Step 401, obtain the first data object corresponding to the target data processing channel in the model.

[0119] The model here is a neural network model, which can specifically but not limited to be a large language model based on the Transformer network. The target data processing channel can be but not limited to the input channel of the network layer of the neural network model.

[0120] In this embodiment, the first data object is a first data matrix with multiple data to be processed, and the first data sub-object in the first data object is a column in the first data matrix. For the data processing scenario of the neural network model, the first data object can specifically be the weight matrix corresponding to the input channel of the model network layer.

[0121] This step can obtain the weight matrix corresponding to the input channels of the neural network model's network layer at the moment when the training of the neural network model is completed, as the first data object, so as to combine with subsequent steps to implement the compression processing of the data in the first data object; however, it is not limited to this. It can also obtain the weight matrix corresponding to the input channels of the neural network model's network layer in real time when it is necessary to use the model to perform data processing after the training of the neural network model is completed, as the first data object, so as to perform the required data compression processing on it.

[0122] Step 402: Move the valid data of the corresponding first data sub-object in the first data object to the position where the invalid data of other first data sub-objects is located, so as to reduce the number of first data sub-objects included in the first data object.

[0123] In this embodiment, by moving the valid data of the corresponding first data sub-object in the first data object to the position where the invalid data of other first data sub-objects is located, the aggregation of the valid data in the first data object is realized, and the valid data is aggregated from the original respective first data sub-objects into a partial number of first data sub-objects, so as to reduce the number of first data sub-objects included in the first data object.

[0124] Among them, the at least one first target data sub-object includes: the first data sub-object that contains at least valid data obtained after the movement; the other first data sub-objects include the first data sub-objects other than the corresponding first data sub-objects in the first data object.

[0125] For the case where the first data object is a first data matrix such as a weight matrix, at least compress the data in the first data matrix in the row direction, that is, specifically move the valid data of the corresponding column in the first data matrix to the position where the invalid data of other columns other than the corresponding column in the first data matrix is located, so as to reduce the number of columns included in the first data matrix. In this case, the at least one first target data sub-object correspondingly includes: the at least one first target column that contains at least valid data obtained after the movement.

[0126] When the valid data of the corresponding column in the first data matrix is moved to the position of the invalid data of other columns other than the corresponding column in the first data matrix, optionally, the valid data of the corresponding column in the first data matrix can be moved to the position of the invalid data of other columns on the left side of the corresponding column in the first data matrix based on the left compression method, so that the valid data of the corresponding column occupies the position of the invalid data such as the 0-value weight of the other columns on the left side of the corresponding column; but not limited to this, the valid data of the corresponding column in the first data matrix can also be moved to the position of the invalid data of other columns on the right side of the corresponding column in the first data matrix based on the right compression method, so that the valid data of the corresponding column occupies the position of the invalid data such as the 0-value weight of the other columns on the right side of the corresponding column. Through the above-mentioned left compression or right compression method, the valid data is gathered into some columns of the first data matrix, so that the some columns at least contain valid data, and other columns other than the some columns do not contain valid data.

[0127] Preferably, data in the same column in the first data matrix are located in the same first target column after the movement is completed, and data in the same row are located in different first target columns after the movement is completed.

[0128] The following are examples:

[0129] See also Figure 5 , assuming that the first data object is Figure 5 In the sparse matrix A (Matrix A), the first data sub-object in the first data object is the column in the matrix A. The matrix A includes 8 columns of data, and the column indexes are 1, 2, 3, ... 8 from left to right. The blank squares represent the 0-valued data in the matrix A, that is, invalid data, and the non-blank squares, that is, the gray boxes, represent the non-0-valued data in the matrix A, that is, valid data. When compressing the matrix A, for example, Figure 5 As shown, based on the leftward compression method in the row direction, the valid data "25" in the second column and the valid data "38" in the third column are moved to the invalid data position in the first column, the valid data "52", "55" in the fifth column and the valid data "69" in the sixth column are moved to the invalid data position in the fourth column, and the valid data "86", "89" in the eighth column are moved to the invalid data position in the seventh column. After the movement, the first data sub-object containing at least valid data is as shown in FIG. Figure 5 As shown, specifically, it is the first data sub-object obtained by adding valid data of other columns to the original 1st, 4th and 7th columns. In the embodiment of the present application, the first data sub-object obtained after the move and containing at least valid data is called the first target data sub-object, and the number of the first target data sub-objects is less than the number of the original first data sub-objects contained in the first data object. For example, in this example, the number is 3, which is less than the number of the original columns contained in the matrix A, which is 8.

[0130] In practical applications, this example can also adopt a rightward compression method to achieve the compression of the data in matrix A in the row direction, and there is no restriction on this. Different compression methods, leftward or rightward, usually result in different compression results. For example, if the rightward compression method is adopted for matrix A, the valid data "52" and "55" in its 5th column may be compressed into the last column, i.e., the 8th column. It should be noted that although different compression methods, leftward or rightward, will lead to different compression results for matrix A, when subsequent data processing is performed on the data in matrix A, the data processing is executed on the first target data according to the position information corresponding to each first target data after compression (used to indicate the original position of the first target data in matrix A), and the position information corresponding to each first target data is fixed and unchanged, thus not affecting the data processing result of matrix A, and both methods can ensure the accuracy of the data processing result of matrix A.

[0131] The first target data sub-object in this example is also the first target column described above. After the movement, the data in the original same column in matrix A is in the same first target column, and the data in the original same row is in different first target columns. For example, the data "52" and "55" in the original 5th column are in the same target column after the movement, the data "86" and "89" in the original 8th column are also in the same target column after the movement, and the data "11", "41", and "71" in the original first row are in different first target columns after the movement.

[0132] In this embodiment, by compressing the data in the first data object, the data in the first data object is aggregated from the original respective first data sub-objects into a partial number of first data sub-objects, reducing the number of first data sub-objects included in the first data object, correspondingly reducing the amount of data to be processed in the first data object, thereby improving the computing performance of the system, reducing the resource requirements for storage, transmission, and operation, etc., and improving the system resource utilization rate.

[0133] In addition, for the case where the first data object is a first data matrix, in this embodiment, by controlling the data in the same column of the first data matrix to be in the same first target column after the movement and the data in the same row to be in different first target columns after the movement, it is convenient to reduce the data selection logic and the hardware wiring complexity when subsequent hardware is used to perform data processing on the compressed data (the first target data in the at least one first target data sub-object / first target column).

[0134] In an alternative embodiment, the method for forming the first target data sub-object may further include the following processing: performing a data order adjustment process on the data in each of the first target columns, so that the valid data originally belonging to the same column in the first data matrix are continuously arranged in the first target column after being moved.

[0135] This embodiment proposes a technical idea of further compressing in the column direction for the compression result obtained by compressing the first data matrix in the row direction in the previous embodiment.

[0136] Optionally, specifically, the invalid data in each of the first target columns may be first removed. After that, for the data in each of the first target columns, a data order adjustment process based on position movement is performed according to the corresponding column index. Through the invalid data removal and order adjustment processes, the valid data originally belonging to the same column in the first data matrix are continuously arranged in the first target column after being moved, thereby achieving further compression of the first data matrix in the column direction.

[0137] Such as Figure 6 In the example of , for the three first target columns obtained by compressing matrix A in the row direction, the 0-value data therein are first removed, resulting in three first target columns that do not contain any 0-value data. On this basis, further, according to the column index of each data in the first target column, a sequential adjustment based on position movement is performed on the data, so as to finally obtain the first target columns in which the valid data originally belonging to the same column are continuously arranged. For example, the data "11", "14", and "17" in the first first target column are the data of the original first column in matrix A. After completing the sequential adjustment based on movement, these three data are continuously arranged in the first target column where they are located.

[0138] In practical applications, it is not limited to the above column direction compression method of first removing invalid data and then performing order adjustment. In other embodiments, based on the above continuous arrangement target (the valid data originally belonging to the same column in the first data matrix are continuously arranged in the first target column after being moved), the data in each of the first target columns may be first sequentially adjusted based on position movement according to the corresponding column index, and then the invalid data such as 0 values existing in the first target column after the order adjustment are removed. By combining the data order adjustment and invalid data removal, the above continuous arrangement target is achieved, so that the valid data originally belonging to the same column in the first data matrix are continuously arranged in the first target column after being moved.

[0139] In this embodiment, based on the row - direction compression result of the first data matrix, further compression is performed in the column direction, and invalid data in the first target columns obtained after row - direction compression of the first data matrix is removed. This can further reduce the data processing volume for the first data object, correspondingly improve the computing performance of the system, reduce the resource requirements for storage, transmission, and operation, and improve the system resource utilization rate and operation efficiency.

[0140] In addition, when performing data compression in the column direction in this embodiment, by arranging the valid data originally belonging to the same column in the first data matrix continuously in the first target columns after movement, it is further convenient for subsequent data processing using hardware on the compressed data (the first target data in each first target column obtained after row - and - column - direction compression, which only includes valid data). This can reduce the data selection logic and the complexity of hardware wiring. Additionally, the index quantity of data in the first target columns can be reduced. For example, for multiple valid data originally belonging to the same column and arranged continuously in the first target column, only record all the position information (such as column index and row index) of the first valid data among the multiple valid data and record the number of the multiple valid data. For the other data except the first valid data among the multiple valid data, only the row index needs to be recorded without recording the column index, thus reducing the index quantity of data in the first target columns and saving storage space. And for the above characteristics, for multiple valid data originally belonging to the same column and arranged continuously in the first target column, the multiple valid data of the entire column (the original same column) can be read and input into the corresponding channels and quantities at one time according to the number within a continuous beat, so as to complete the calculation of a column (the original column) of data within one beat.

[0141] In an optional embodiment, matching the first data object which can be a first data matrix including multiple data to be processed (such as the weight matrix in a neural network), the second data object corresponding to the target data processing channel can be a second data matrix including multiple data to be processed, for example, specifically the feature map corresponding to the input channel of a neural network.

[0142] Step 202 in the data processing method provided in this application, that is, obtaining the second target data corresponding to each of the first target data from the second data object corresponding to the target data processing channel according to the position information corresponding to each of the first target data in the at least one first target data sub - object, can be correspondingly implemented as:

[0143] Obtaining the second target data corresponding to each of the first target data from the second target columns of the second data matrix according to the position information of each of the first target data in each of the first target columns; the second target columns are the columns currently to be processed in the second data matrix.

[0144] Among them, for the matrix multiplication operation based on inner product in the embodiments of the present application, since the essence is to require matching the data with the same column index in the current row participating in the operation in the first data object / first data matrix and the row index in the current column participating in the operation in the second data object / second data matrix into data pairs to be processed. Based on this, when obtaining the second target data corresponding to each of the first target data from the second target column of the second data matrix according to the position information of each first target data in each first target column, specifically, the data at the row position corresponding to the column index can be obtained from the second target column currently to be processed in the second data matrix according to the column index corresponding to the first target data in the first data matrix, and used as the second target data corresponding to the first target data. This second target data is also the data to be processed that matches the first target data and is used to form a data pair to be processed with the first target data to participate in subsequent data processing.

[0145] For example, continuing with Figure 3 the matrices A and B in the example, for the three first target columns in Figure 7 obtained by compressing matrix A in the row and column directions, assuming that the second target column currently to be processed in matrix B is the first column of matrix B. Taking Figure 7 the first first target column in as an example, for the data "11", "14", "17" in the first first target column, the data at the row index "1" (i.e., the data represented by "1" in the second target column) can be obtained from the second target column according to their corresponding column index "1" as the second target data corresponding to "11", "14", "17"; similarly, for the data "25" in this first target column, the data at the row index "2" (i.e., the data represented by "2" in the second target column) can be obtained from the second target column according to its corresponding column index "2" as the second target data corresponding to "25"; for the data "38" in this first target column, the data at the row index "3" (i.e., the data represented by "3" in the second target column) can be obtained from the second target column according to its corresponding column index "3" as the second target data corresponding to "38".

[0146] In this embodiment, by obtaining the second target data corresponding to each of the first target data from the second target column of the second data matrix according to the position information of each first target data in each first target column, it is convenient to accurately locate and obtain the second target data that matches the first target data from the second data object, so as to form a data pair to be processed with the first target data to participate in subsequent data processing, ensuring the efficiency and accuracy of data processing for the first data object and the second data object.

[0147] In an alternative embodiment, refer to Figure 8The flowchart of the data processing method shown. Step 203 in the data processing method provided by this application, that is, performing data processing on each of the first target data and the corresponding second target data, can be implemented as follows:

[0148] Step 801: According to the number of available hardware processing channels that can be processed in parallel, allocate the corresponding number of first target data to each available hardware processing channel, and use the available hardware processing channels to perform data processing on the allocated first target data and the corresponding second target data.

[0149] An available hardware processing channel is a hardware-based computing channel in the system of an electronic device such as a personal computer or a server that is currently not occupied and can be scheduled to perform the required operations on data. For example, a computing channel composed of an arithmetic unit and registers. Each channel can include the required number of arithmetic units and / or registers, and can also include other required hardware.

[0150] Optionally, in the embodiments of this application, the available hardware processing channel at least includes a first arithmetic component. The first arithmetic component can be a multiplier that can be used to perform multiplication operations on data. In addition, it can also include registers for storing data.

[0151] Furthermore, each available hardware processing channel can be a PE (processing element). Refer to Figure 9 The schematic internal structure diagram of the PE shown. Each PE includes a register (Reg) and a multiplier (Mul), which are respectively used for storing data and performing multiplication operations on data.

[0152] For the case where the first data object is a first data matrix and the second data object is a second data matrix, in this step 801, specifically, according to the number of available hardware processing channels that can be processed in parallel, the corresponding number of first target data in the first target columns obtained by compressing the first data matrix (compressing in the row direction, or compressing in both the row and column directions) are allocated to each available hardware processing channel in a one-to-one manner. For example, for the three first target columns obtained by compressing the above matrix A in both the row and column directions, assuming the current number of available PEs is 5, then in sequence, first allocate the 5 first target data included in the first first target column among the three first target columns to 5 PEs in a one-to-one manner, and store them in the registers of the allocated PEs.

[0153] In practical applications, optionally, when the neural network model is known, appropriate hardware resource configuration can be performed on the neural network model. For example, the number of hardware channels to be used for model data processing can be configured as the number of rows of the weight matrix corresponding to the input channels of the neural network model layer, or configured as the number of valid data included in the column with the largest amount of valid data in each column of the weight matrix, etc.

[0154] The first target data corresponds to corresponding position information, such as having a row index and a column index, which is used to indicate the original positions such as the original row and the original column where the first target data is located in the first data object (the first data matrix). After allocating the corresponding number of first target data to each available PE, the data on the row corresponding to the column index of the first target data allocated to each PE can be further obtained from the second target column currently to be processed in the second data object (the second data matrix), and used as the second target data corresponding to the first target object and sent to the PE. The PE performs a multiplication operation on the obtained first target data and the second target data based on its multiplier.

[0155] The hardware used for data processing of the first target data and the corresponding second target data, in addition to the available hardware processing channels such as the PE, may also include a plurality of second arithmetic components and at least one routing component. The second arithmetic components can be accumulators capable of performing accumulation processing (sum operation) on data. The routing component is used to transmit the multiplication operation results of the first arithmetic components in the available hardware processing channels to the corresponding second arithmetic components according to the row index of the first target data, so as to accumulate the multiplication operation results with the same row index of the corresponding first target data for the second target column currently participating in the processing in the second data matrix, thereby realizing the accumulation of the multiplication operation results corresponding to each data (the first target data) in the same row of the first data matrix for the second target column currently participating in the processing in the second data matrix.

[0156] The number of second arithmetic components is not less than the number of rows of the first data matrix. For example, for the example where the first data matrix is matrix A, 9 accumulators can be set in total, corresponding one by one to the original 9 rows of data of the first data matrix. After the multiplier of each PE performs a multiplication operation on the obtained first target data and the corresponding second target data, it sends the multiplication operation result and the row index of the first target data to the routing component. The routing component routes the received multiplication operation result to the corresponding accumulator based on the row index of the first target data, so that the multiplication operation result obtained by each accumulator is the multiplication operation result of the first target data in the same row of the first data matrix for the second target column currently participating in the processing in the second data matrix, and correspondingly realizes the accumulation of the multiplication operation results corresponding to each data (the first target data) in the same row of the first data matrix in the accumulator, which is consistent with the matrix multiplication operation rule based on the inner product and meets the requirements of the matrix multiplication operation based on the inner product.

[0157] Among them, since each of the first target data participating in data processing is the data in at least one first target data sub-object obtained by compressing a first data object, compared with the first data object, at least part of the invalid data in the first data object is removed, so that the removed invalid data can be avoided from participating in the operation, reducing the data processing amount of the first data object. Based on this, the number of available hardware processing channels required to process each of the first target data is less than the number of available hardware processing channels required to process each piece of data to be processed in the first data object, which can correspondingly improve the computing performance of the system, reduce the resource requirements such as storage, transmission, and operation, and improve the system resource utilization rate and operation efficiency. At the same time, the designed hardware structure and the way of using the hardware conform to the matrix multiplication operation rule based on inner product, meeting the requirements of the matrix multiplication operation based on inner product, and can ensure the accuracy of the matrix multiplication operation result based on inner product in this application.

[0158] In an alternative embodiment, the valid data in the original same column in the first data matrix / first data object are continuously arranged in the first target column after being moved. That is to say, in this embodiment, the first target column is the column obtained by compressing the first data matrix / first data object in the row direction and the column direction and performing sequential adjustment (adjusting the valid data in the original same column to be continuously arranged).

[0159] The data processing of the first target data and the corresponding second target data in this embodiment also includes multiply-accumulate processing. The multiplication operation in the multiply-accumulate processing includes the multiplication operation involved in multiplying the first data matrix and the second data matrix.

[0160] In this embodiment, referring to Figure 10 the flowchart of, step 203 in the method provided in this application, that is, processing each of the first target data and the corresponding second target data, can be specifically implemented as:

[0161] Step 1001: Allocate the first target data in the same group in each first target column to consecutive available hardware processing channels in the channel array; the first target data in the same group includes the first target data belonging to the original same column in the first data matrix within the first target column.

[0162] The channel array is an array formed by each available hardware processing channel, such as a PE array formed by each PE.

[0163] In this embodiment, according to the characteristic that the valid data in the original same column in the first data matrix / the first data object are continuously arranged in the first target column after being moved, the data in the first target column are grouped. Specifically, the first target data belonging to the original same column in the first data matrix in the first target column are grouped into one group, and the first target data belonging to the original different columns in the first data matrix in the first target column are correspondingly grouped into different groups.

[0164] The data included in each group are continuously arranged in the first target column where they are located. According to this characteristic, when allocating the first target data to the available hardware processing channels, specifically, the first target data in the same group in each first target column are respectively allocated to the continuous available hardware processing channels in the channel array. For example, the first target data in the same group are respectively allocated to the continuous PEs in the PE array, and are specifically stored in the registers included in the allocated PEs.

[0165] Step 1002: Determine the second target data corresponding to the first target data in the same group according to the position information corresponding to the first target data in the same group.

[0166] The column indices corresponding to the first target data in the same group are the same. Combining with the characteristics of the matrix multiplication operation based on the inner product, it can be known that the first target data in the same group correspond to the same second target data in the second data matrix.

[0167] In this step, specifically, according to the column index corresponding to the first target data in the same group, the data at the row position corresponding to this column index is determined from the second target column currently to be processed in the second data matrix as the second target data corresponding to the first target data in this group.

[0168] For example, for the first target column among the three first target columns obtained by compressing the matrix A in the above text in the row and column directions, taking the first group in this column as an example, for the data "11", "14", "17" included in the first group in this column, specifically, according to the column index "1" of "11", "14", "17", the data at the row indicated by the index "1" is determined from the second target column currently to be processed in the second data matrix as the second target data corresponding to "11", "14", "17". For example, the data at row 1 is determined from the first column currently to be processed in the matrix B as the second target data corresponding to "11", "14", "17".

[0169] Step 1003: Allocate the determined second target data to the available hardware processing channels where the first target data in the corresponding group are respectively located, so that the first arithmetic component provided by the available hardware processing channel performs a multiplication operation on the allocated first target data and the second target data to obtain a multiplication operation result.

[0170] The first arithmetic component can be a multiplier included in the PE, such asFigure 9 Multiplier Mul in

[0171] After grouping the first target data in each first target column and determining the same second target data corresponding to each piece of the first target data in each group, the same second target data can be obtained, and the obtained same second target data can be allocated to the available hardware processing channels where the respective first target data in the corresponding group are located. On this basis, the available hardware processing channels can perform a multiplication operation on the obtained first target data and second target data based on their first arithmetic components.

[0172] For example, read the same second target data corresponding to each piece of the first target data "11", "14", "17" included in the first group in the above-mentioned first first target column, and allocate the read second target data to the PEs where "11", "14", "17" included in this group are located respectively. The multipliers in the PEs where "11", "14", "17" are located respectively perform a multiplication operation on the obtained first target data and second target data.

[0173] Optionally, for the case where the second target data is invalid data (such as a 0-valued eigenvalue), the reading of the second target data can be directly skipped. Correspondingly, it is not necessary to input the second target data of this invalid data into the corresponding PE, and further, it is not necessary to perform an operation on the first target data currently allocated in this PE. That is, the multiplication operation result corresponding to the first target data currently allocated in this PE can be directly regarded as an empty result, which is consistent with the feature that invalid data does not contribute to the operation result in matrix multiplication operation and will not affect the overall matrix multiplication operation result.

[0174] Step 1004: According to the position information corresponding to each first target data, send the multiplication operation results corresponding to the first target data belonging to the same original row in the first data matrix to the same second arithmetic component for accumulation processing to obtain an accumulation operation result.

[0175] Among them, the position information corresponding to the second target data is used to indicate the position where the second target data is located in the second data object.

[0176] The second arithmetic component can be an accumulator capable of performing accumulation processing (sum operation) on data. The number of second arithmetic components is not less than the number of rows of the first data matrix. For example, for the example where the first data matrix is the above-mentioned matrix A, a total of 9 accumulators can be set, corresponding one by one to the original 9 rows of data of matrix A. Of course, in practical applications, it is not limited to this, and more than 9 accumulators can also be set.

[0177] After the multiplication operation on the allocated first target data and second target data is completed in the corresponding available hardware processing channel, such as a PE, each PE sends the multiplication result and the row index of the corresponding first target data to the routing component. Based on the row index of the first target data, the routing component routes the received multiplication result to the corresponding accumulator, so that each accumulator obtains the multiplication result of the first target data in the same row of the first data matrix for the currently to-be-processed second target column of the second data matrix. Accordingly, the multiplication results corresponding to each data (first target data) in the same row of the first data matrix are accumulated within the accumulator, which is consistent with the matrix multiplication operation rule based on the inner product and meets the requirements of the matrix multiplication operation based on the inner product.

[0178] The following is an example:

[0179] Continuing Figure 3 from the example in, for the three first target columns obtained by compressing matrix A in the row direction and the column direction, as Figure 11 shown, assuming that the number of currently available PEs in the PE array is not less than the total number of first target data in the three first target columns, the first target data in each of the three first target columns can be directly allocated one-to-one to different PEs. Specifically, the first target data belonging to the same original column of the first data matrix in each first target column can be allocated to consecutive PEs in the PE array in a grouped form, and the data of the currently to-be-processed second target column in matrix B (such as Figure 3 the first column data shown for matrix B in) is simplified, the invalid data is removed, and the valid second target data is associated with each group according to the index. Specifically, reference can be made to Figure 11 shown.

[0180] After that, the second target data corresponding to each group can be read according to the established association, and the read second target data can be directly allocated to the PEs where the respective first target data in the corresponding group are located. Each PE uses its multiplier to perform a multiplication operation on the allocated first target data and second target data. For this example, in one clock cycle, the multiplication operations on each first target data and its corresponding second target data in the three first target columns can be completed based on the parallel processing method. For matrix B, a column of data in matrix B can be sent to the corresponding PE in one clock cycle, so all operations on a column of data in matrix B can be processed in one clock cycle. The first target data (such as the compressed weight data) obtained after matrix A is compressed is stored in the PEs in advance, and the data of matrix B is directly connected to the PE array, and all PEs calculate the results in one clock cycle. Specifically, the data of the column of matrix B currently participating in the operation can be divided into multiple groups, and the data of each group is only multiplied by the data of the same first target column obtained after matrix A is compressed, and will not cross to different columns, thereby reducing the data selection logic and the complexity of the hardware wiring. The invalid data in matrix B is discarded.

[0181] Among them, one clock cycle in the embodiments of the present application specifically refers to that a PE performs a multiplication operation on the obtained pair of data to be processed.

[0182] See further Figure 12 , each PE is connected to nine accumulators adder1 - adder9 through a routing component crossbar. The multiplication results of the PEs in the same column will be sent to different accumulators through the routing component crossbar without conflict. If there are x columns (the number of columns of PEs in the PE array, which is also the number of columns of the first target column) in the PE array performing operations simultaneously, then each accumulator has at most x inputs.

[0183] Assume that the current second target column to be processed in the second data matrix is the nth column of the matrix, the first target data obtained in the PE is W ij , and the second target data is B mn (i.e., the corresponding data in the nth column of the second data matrix). Among them, i and j respectively represent the row index and column index of the first target data, m and n respectively represent the row index and column index of the second target data, and i, j, m, and n are all integers not less than 1. If the number of accumulators set is 9, numbered 1 - 9 respectively, then the multiplication result of W ij and B mn in the PE will be based on the first target data W ijThe row index i is sent into the accumulator numbered i. Similarly, for the current second target column to be processed in the second data matrix, the multiplication results of all the first target data with the row index i will be sent into the accumulator numbered i based on the row index i, so that the multiplication results of the same row (i.e., the i-th row) in the first data matrix and the data in the n-th column of the second data matrix are accumulated in this accumulator.

[0184] In the embodiment of the present application, for the matrix multiplication operation based on the inner product, essentially, the data in each row of the first data matrix and the data in each column of the second data matrix are sequentially and correspondingly multiplied one by one, and then the multiplication results are accumulated to obtain the multiply-accumulate result. The result of the matrix multiplication operation based on the inner product is still a matrix. The multiply-accumulate result of the same row in the first data matrix and the same column in the second data matrix is a value in the final matrix multiplication operation result (also a matrix). Based on this, in practical applications, optionally, an accumulator can also be set for each row-column combination formed by the data in each row of the first data matrix and the data in each column of the second data matrix. In this way, when the PE multiplies the obtained first target data W ij and the second target data B mn after multiplication, the multiplication result can be sent into the accumulator at the position R ij corresponding to the row index i of the first target data W mn and the column index n of the second target data B in . The accumulator at the position R in is specifically the accumulator corresponding to the row-column combination formed by the i-th row in the first data matrix and the n-th column in the second data matrix. Based on this routing method, for the current second target column to be processed in the second data matrix, the multiplication results of the data in the same row of the first data matrix are accumulated in the corresponding accumulator, meeting the requirements of the matrix multiplication operation.

[0185] In this embodiment, by virtue of the characteristic that the valid data originally in the same column of the first data matrix / first data object are continuously arranged in the first target column after compression, the data in the first target column are grouped, which is convenient for determining the same second target data corresponding to each first target data in the same group in the first target column, and sending the determined second target data to the available hardware processing channels such as the PEs where each first target data in the same group is located based on one read operation, thereby further simplifying the data selection logic, reducing the data reading volume and bandwidth requirements, reducing the hardware wiring complexity, improving the data operation efficiency, and reducing the memory access of the partial sum (multiply-accumulate result) compared with the outer product method.

[0186] Corresponding to the above data processing method, the embodiment of the present application further provides a data processing device, the composition structure of which is as Figure 13 shown, including:

[0187] A first acquisition module 1301, configured to acquire at least one first target data sub-object including first target data; the first target data in the at least one first target data sub-object at least includes all valid data of each first data sub-object in a first data object corresponding to a target data processing channel; each first target data corresponds to corresponding position information, which is used to indicate the original position of the first target data in the first data object.

[0188] A second acquisition module 1302, configured to acquire second target data corresponding to each of the first target data from a second data object corresponding to the target data processing channel according to the position information corresponding to each of the first target data in the at least one first target data sub-object.

[0189] A data processing module 1303, configured to perform data processing on each of the first target data and the corresponding second target data.

[0190] Wherein, the number of the first target data sub-objects is less than the number of the first data sub-objects in the first data object.

[0191] In an optional implementation manner, the device further includes a preprocessing device for forming the first target data sub-object based on preprocessing. The process of the preprocessing device forming the first target data sub-object includes:

[0192] Acquire a first data object corresponding to a target data processing channel in the model.

[0193] Move the valid data of the corresponding first data sub-object in the first data object to the position where the invalid data of other first data sub-objects is located, so as to reduce the number of first data sub-objects included in the first data object.

[0194] Wherein, the at least one first target data sub-object includes: at least one first data sub-object containing valid data obtained after the movement; the other first data sub-objects include first data sub-objects other than the corresponding first data sub-object in the first data object.

[0195] In an optional implementation manner, the first data object includes a first data matrix having a plurality of data to be processed, and the first data sub-object is a column in the first data matrix.

[0196] When the preprocessing device moves the valid data of the corresponding first data sub-object in the first data object to the position where the invalid data of other first data sub-objects is located, it is specifically configured to:

[0197] Move the valid data in the corresponding columns of the first data matrix to the positions of the invalid data in other columns outside the corresponding columns in the first data matrix, so as to reduce the number of columns included in the first data matrix;

[0198] Wherein, the at least one first target data sub-object includes: a first target column that at least contains valid data obtained after the movement; the data in the same column of the first data matrix is in the same first target column after the movement, and the data in the same row is in different first target columns after the movement.

[0199] In an alternative embodiment, the process of the preprocessing device forming the first target data sub-object further includes:

[0200] Perform data order adjustment processing on the data in each of the first target columns, so that the valid data originally belonging to the same column in the first data matrix is continuously arranged in the first target column where it is located after the movement.

[0201] In an alternative embodiment, the second data object includes a second data matrix having a plurality of data to be processed;

[0202] The second acquisition module 1302 is specifically configured to: obtain second target data corresponding to each of the first target data from the second target column of the second data matrix according to the position information of each first target data in each first target column; the second target column is the column currently to be processed in the second data matrix.

[0203] In an alternative embodiment, the data processing module 1303 is specifically configured to:

[0204] Allocate a corresponding number of first target data to each available hardware processing channel according to the number of available hardware processing channels for parallel processing, and use the available hardware processing channels to perform data processing on the allocated first target data and the corresponding second target data;

[0205] Wherein, the number of available hardware processing channels required to process each of the first target data is less than the number of available hardware processing channels required to process each of the data to be processed in the first data object.

[0206] In an alternative embodiment, the valid data originally belonging to the same column in the first data matrix is continuously arranged in the first target column where it is located after the movement; the data processing includes multiply-accumulate processing, and the multiplication operation in the multiply-accumulate processing includes the multiplication operation involved in multiplying the first data matrix and the second data matrix;

[0207] The data processing module 1303 is specifically configured to:

[0208] All the first target data in the same group in each first target column are respectively allocated to consecutive available hardware processing channels in the channel array; the first target data in the same group includes the first target data within the first target column that belongs to the same original column in the first data matrix.

[0209] According to the position information corresponding to the first target data in the same group, determine the second target data corresponding to the first target data in the same group; the first target data in the same group corresponds to the same second target data.

[0210] Allocate the determined second target data to the available hardware processing channels where the first target data in the corresponding group is located respectively, so that the first arithmetic component provided by the available hardware processing channel performs a multiplication operation on the allocated first target data and second target data to obtain a multiplication operation result.

[0211] According to the position information corresponding to each first target data, send the multiplication operation results corresponding to the first target data that belongs to the same original row in the first data matrix to the same second arithmetic component for accumulation processing to obtain an accumulation operation result.

[0212] Among them, the position information corresponding to the second target data is used to indicate the position of the second target data in the second data object.

[0213] The embodiments of the present application further provide a data processing chip. Refer to Figure 14 the shown composition structure diagram. This data processing chip includes a plurality of available hardware processing channels 1401 and a plurality of second arithmetic components 1402.

[0214] Among them, each of the available hardware processing channels 1401 includes a first arithmetic component 1403, which is used to obtain the corresponding first target data in at least one first target data sub-object. The first target data in the at least one first target data sub-object includes all valid data of each first data sub-object in the first data object corresponding to the target data processing channel. Each first target data corresponds to corresponding position information, which is used to indicate the original position of the first target data in the first data object; and is used to obtain the second target data corresponding to the corresponding first target data from the second data object corresponding to the target data processing channel according to the position information corresponding to the corresponding first target data; and perform a first arithmetic process on the obtained first target data and second target data to obtain a first arithmetic result.

[0215] Each second arithmetic component 1402 is used to perform a second arithmetic process on the first arithmetic result output by the corresponding available hardware processing channel to obtain a second arithmetic result.

[0216] Among them, the number of the first target data sub-objects is less than the number of the first data sub-objects in the first data object.

[0217] In an alternative embodiment, the first data object includes a first data matrix having a plurality of data to be processed;

[0218] The first data sub-objects are columns in the first data matrix;

[0219] The at least one first target data sub-object includes: a first target column containing at least valid data obtained by moving the valid data in the corresponding column of the first data matrix to the position of the invalid data in other columns outside the corresponding column of the first data matrix; the data in the same column of the first data matrix is in the same first target column after the movement, and the data in the same row is in different first target columns after the movement, and the number of the first target columns is less than the number of columns included in the first data matrix;

[0220] The second data object includes a second data matrix having a plurality of data to be processed.

[0221] In an alternative embodiment, the first arithmetic process includes a multiplication operation, and the multiplication operation includes the multiplication operation involved in multiplying the first data matrix and the second data matrix; the second arithmetic process includes an accumulation operation;

[0222] When the first arithmetic component 1403 performs the first arithmetic process on the obtained first target data and second target data, it is specifically configured to: perform a multiplication operation on the first target data in the obtained corresponding first target column and the corresponding second target data to obtain a multiplication operation result;

[0223] When the second arithmetic component 1402 performs the second arithmetic process on the first arithmetic result output by the corresponding available hardware processing channel, it is specifically configured to: perform an accumulation process on the multiplication operation result output by the corresponding available hardware processing channel to obtain an accumulation operation result.

[0224] In an alternative embodiment, the data processing chip further includes a routing component;

[0225] Each of the available hardware processing channels 1401 further includes an independent storage component for storing the first target data allocated to the corresponding available hardware processing channel;

[0226] The routing component is configured to send the multiplication operation result output by the available hardware processing channel to the corresponding second arithmetic component for accumulation processing according to the position information corresponding to the first target data and the second target data in the available hardware processing channel;

[0227] Among them, the position information corresponding to the second target data is used to indicate the position of the second target data in the second data object; the multiplication operation results corresponding to the first target data that originally belong to the same row in the first data matrix are sent to the same second operation component for accumulation processing.

[0228] In an alternative embodiment, the available hardware processing channel 1401 is a PE, the second operation component 1402 is an accumulator, the first operation component 1403 is a multiplier in the PE, and the storage component is a register.

[0229] The data processing chip provided in this embodiment corresponds to the data processing methods disclosed in the above method embodiments. It is used to implement the data processing methods disclosed in the above method embodiments based on the hardware structure of the data processing chip and the functions of its various components. For the more detailed functions of the various components in this data processing chip and the process of implementing data processing based on the various components of this data processing chip, reference may specifically be made to the descriptions in the above method embodiments, which will not be elaborated here.

[0230] This application embodiment also discloses an electronic device. The composition structure of the electronic device is as Figure 15 shown, and at least includes:

[0231] A memory 10 for storing a computer instruction set;

[0232] The computer instruction set can be implemented in the form of a computer program.

[0233] A processor 20 for implementing the data processing method disclosed in any of the above method embodiments by executing the computer instruction set.

[0234] The processor 20 can be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a neural network processor (NPU), a deep learning processor (DPU), or other programmable logic devices, etc.

[0235] The electronic device is equipped with a display device and / or has a display interface and can externally connect to a display device.

[0236] Optionally, the electronic device further includes a camera component and / or is connected to an external camera component.

[0237] In addition, the electronic device may further include components such as a communication interface and a communication bus. The memory, the processor, and the communication interface complete communication with each other through the communication bus.

[0238] The communication interface is used for communication between an electronic device and other devices. The communication bus can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0239] In addition, an embodiment of the present application further provides a readable storage medium, on which a computer instruction set is stored. The computer instruction set is used to be called and executed by a processor to implement the data processing method provided in any of the above method embodiments.

[0240] Moreover, a computer program product is also provided, including a computer program / instructions. When the computer program / instructions are executed by a processor, the data processing method provided in any of the above method embodiments is implemented.

[0241] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other.

[0242] For the convenience of description, when describing the above system or device, it is divided into various modules or units according to functions for description. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0243] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.

[0244] Finally, it should also be noted that in this text, relational terms such as first, second, third, and fourth are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the said element.

[0245] The above are only the preferred embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A data processing chip, comprising: Multiple available hardware processing channels and multiple second arithmetic components; Each of the available hardware processing channels includes a first arithmetic component; The first arithmetic component is configured to obtain first target data in at least one first target column column by column, each first target data corresponding to corresponding position information for indicating the original position of the first target data in the first data matrix; And configured to obtain second target data corresponding to the corresponding first target data from a second data matrix with multiple data to be processed corresponding to the target data processing channel according to the position information corresponding to the corresponding first target data; and perform a first arithmetic process on the obtained first target data and second target data to obtain a first arithmetic result; Wherein, the first target data in the at least one first target column includes all valid data in each column of a first data matrix with multiple data to be processed corresponding to the target data processing channel; the first target column is obtained by directly compressing the first data matrix in the row direction at least, including: moving the valid data in the corresponding column of the first data matrix to the position of the invalid data in other columns outside the corresponding column of the first data matrix, wherein, the data in the same column of the first data matrix are in the same first target column after the movement, and the data in the same row are in different first target columns after the movement, and the number of the first target columns is less than the number of columns included in the first data matrix; Each second arithmetic component is configured to perform a second arithmetic process on the first arithmetic result output by the corresponding available hardware processing channel to obtain a second arithmetic result.

2. The data processing chip according to claim 1, wherein the first arithmetic process includes a multiplication operation, and the multiplication operation includes the multiplication operation involved in multiplying the first data matrix and the second data matrix; the second arithmetic process includes an accumulation operation; When performing the first arithmetic process on the obtained first target data and second target data, the first arithmetic component is specifically configured to: perform a multiplication operation on the first target data in the obtained corresponding first target column and the corresponding second target data to obtain a multiplication operation result; When performing the second arithmetic process on the first arithmetic result output by the corresponding available hardware processing channel, the second arithmetic component is specifically configured to: perform an accumulation process on the multiplication operation result output by the corresponding available hardware processing channel to obtain an accumulation operation result.

3. The data processing chip according to claim 2, further comprising a routing component; Each of the available hardware processing channels further includes an independent storage component for storing the first target data allocated to the corresponding available hardware processing channel; The routing component is configured to send the multiplication operation result output by the available hardware processing channel to the corresponding second arithmetic component for accumulation processing according to the position information corresponding to the first target data and the second target data in the available hardware processing channel; Among them, The position information corresponding to the second target data is used to indicate the position of the second target data in the second data matrix; the multiplication operation results corresponding to the first target data that originally belong to the same row in the first data matrix are sent to the same second operation component for accumulation processing.

4. A data processing method, comprising: Obtaining at least one first target column containing first target data column by column; Each first target data corresponds to corresponding position information, which is used to indicate the original position of the first target data in the first data matrix; According to the position information corresponding to each first target data in the at least one first target column, obtaining second target data corresponding to each first target data from a second data matrix with multiple data to be processed corresponding to the target data processing channel; Performing data processing on each first target data and the corresponding second target data; Wherein, the first target data in the at least one first target column includes all valid data in each column of a first data matrix with multiple data to be processed corresponding to the target data processing channel; the first target column is obtained by directly compressing the first data matrix at least in the row direction, including: moving the valid data in the corresponding column of the first data matrix to the positions of the invalid data in other columns outside the corresponding column of the first data matrix, wherein the data in the same column of the first data matrix are in the same first target column after the movement, and the data in the same row are in different first target columns after the movement, and the number of the first target columns is less than the number of columns included in the first data matrix.

5. The method for forming the first target column according to claim 4, further comprising: Performing data order adjustment processing on the data in each first target column, so that the valid data that originally belonged to the same column in the first data matrix are continuously arranged in the first target column where they are located after the movement.

6. The data processing method according to claim 4 or 5, wherein the obtaining the second target data corresponding to each first target data from a second data matrix with multiple data to be processed corresponding to the target data processing channel according to the position information corresponding to each first target data in the at least one first target column includes: Obtaining the second target data corresponding to each first target data from the second target column of the second data matrix according to the position information of each first target data in each first target column; The second target column is the column to be currently processed in the second data matrix.

7. The data processing method according to claim 4, wherein the performing data processing on each first target data and the corresponding second target data includes: Allocating the corresponding number of first target data to each available hardware processing channel according to the number of available hardware processing channels for parallel processing, and using the available hardware processing channels to perform data processing on the allocated first target data and the corresponding second target data; Among them, the number of available hardware processing channels required to process each of the first target data is less than the number of available hardware processing channels required to process each piece of data to be processed in the first data matrix.

8. The data processing method according to claim 6, wherein the valid data originally belonging to the same column in the first data matrix are continuously arranged in the first target column after being moved; the data processing includes multiply-accumulate processing, and the multiplication operation in the multiply-accumulate processing includes the multiplication operation involved in multiplying the first data matrix and the second data matrix; The data processing of each of the first target data and the corresponding second target data includes: Allocating the first target data in the same group in each first target column to consecutive available hardware processing channels in the channel array; the first target data in the same group includes the first target data belonging to the same original column in the first data matrix in the first target column; Determining the second target data corresponding to the first target data in the same group according to the position information corresponding to the first target data in the same group; the first target data in the same group corresponds to the same second target data; Allocating the determined second target data to the available hardware processing channels where the first target data in the corresponding group are located respectively, so that the first arithmetic component provided by the available hardware processing channel performs a multiplication operation on the allocated first target data and second target data to obtain a multiplication operation result; Sending the multiplication operation results corresponding to the first target data belonging to the same original row in the first data matrix to the same second arithmetic component for accumulation processing to obtain an accumulation operation result according to the position information corresponding to each first target data; Among them, the position information corresponding to the second target data is used to indicate the position of the second target data in the second data object.

9. A data processing device, comprising: A first acquisition module, configured to acquire at least one first target column containing first target data column by column; Each first target data corresponds to corresponding position information, which is used to indicate the original position of the first target data in the first data matrix; A second acquisition module, configured to acquire second target data corresponding to each of the first target data from the second data matrix with multiple pieces of data to be processed corresponding to the target data processing channels according to the position information corresponding to each of the first target data in the at least one first target column; A data processing module, configured to perform data processing on each of the first target data and the corresponding second target data; Among them, the first target data in the at least one first target column includes all valid data of each column in the first data matrix corresponding to the target data processing channel and having a plurality of data to be processed; the first target column is obtained by directly compressing the first data matrix at least in the row direction, including: moving the valid data of the corresponding column in the first data matrix to the position of the invalid data in other columns outside the corresponding column in the first data matrix, wherein the data in the same column of the first data matrix are in the same first target column after the movement, and the data in the same row are in different first target columns after the movement, and the number of the first target columns is less than the number of columns included in the first data matrix.

Citation Information

Patent Citations

  • Data processing method and device and hardware accelerator

    CN110929854A

  • ReRAM-based weight matrix processing method and device

    CN114187944A

  • Systems and methods for accelerating neural networks using unified sparse tensor kernels

    CN116724318A