Sparse matrix multiplication acceleration hardware, recommendation system acceleration method, and ai chip

By optimizing the sparse matrix multiplication acceleration hardware and reducing the ineffective transfer and calculation of zero-value elements in sparse matrices, the problems of high computational latency and high resource consumption in sparse matrix multiplication are solved, and a low-latency, low-cost user recommendation system is realized.

CN120610681BActive Publication Date: 2025-10-17SHANGHAI YUNSUI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511100096.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-10-17
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

The existing technology has a large number of invalid zero-value calculations in sparse matrix multiplication calculations, which leads to high calculation delays and large resource consumption, affecting user experience.

Method used

Design sparse matrix multiplication acceleration hardware, including data loading unit, multiplication calculation unit group, addition calculation unit group and normalization acceleration unit group, optimize sparse matrix multiplication calculation through hardware and reduce the transportation and calculation of zero-value elements.

Benefits of technology

This enables fast calculation of sparse matrix multiplication, reduces latency and resource consumption, and improves the user experience of the recommendation system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120610681B_ABST
    Figure CN120610681B_ABST
Patent Text Reader

Abstract

The application discloses sparse matrix multiplication acceleration hardware, a recommendation system acceleration method and an AI chip. A data loading unit in the hardware establishes a data carrying task for dense data in a sparse matrix multiplication task, carries the dense data to a register, a normalization acceleration unit group sends each left and right matrix block to a multiplication calculation unit group after performing a pre-multiplication activation operation according to the left and right matrix blocks obtained from the register, the multiplication calculation unit group is used for selecting a multiplication calculation unit to perform multiplication calculation according to the data size of the left and right matrix blocks, and providing a calculation result to an addition calculation unit group to perform addition calculation, and the normalization acceleration unit group is also used for performing normalization calculation and / or activation operation according to the addition calculation result when the post-multiplication normalization and / or post-multiplication activation operation is needed, and obtaining a final result. The technical scheme of the embodiment can realize fast sparse matrix multiplication calculation under limited calculation resources.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of AI (Artificial Intelligence) chip, in particular to a sparse matrix multiplication acceleration hardware, a recommendation system acceleration method and an AI chip. BACKGROUND

[0002] With the continuous development of Internet technology, various recommendation systems are also emerging. The essence of the recommendation system is to meet the personalized needs of users through content filtering. At present, modern industrial-level recommendation systems have borrowed from traditional NLP (Natural Language Processing) language models and have entered the era of deep learning. When implementing a recommendation system based on a large language model, sparse matrix multiplication is a core operation that needs to be frequently used.

[0003] In related technologies, when performing sparse matrix multiplication on a large language model, data is often processed in a sparse manner at the software level. When using hardware to actually multiply matrices, the diluted matrix needs to be converted into a formal dense matrix by zero padding, and then matrix multiplication is performed. For example, by padding the longest sequence, the sparse matrix is padded with zeros to a fixed tensor shape matrix, and then a fixed tensor shape multiplication is performed using a hardware calculation unit.

[0004] The inventors found during the implementation of the present application that due to the complex structure and large number of parameters of the large language model, a large number of invalid zero value calculations are introduced in the frequently used matrix multiplication, which further increases the calculation delay of the matrix multiplication and the consumption of the calculation resources, greatly increasing the calculation cost of the matrix multiplication, which seriously affects the user experience in the user recommendation scenario. SUMMARY

[0005] The embodiments of the present application provide a sparse matrix multiplication acceleration hardware, a recommendation system acceleration method and an AI chip, which creatively construct an acceleration hardware suitable for sparse matrix multiplication, and realize fast calculation of sparse matrix multiplication under limited calculation resources.

[0006] According to an aspect of the embodiments of the present application, a sparse matrix multiplication acceleration hardware is provided, which includes a data loading unit, a register, a multiplication calculation unit group, an addition calculation unit group and a normalization acceleration unit group. The multiplication calculation unit group includes a plurality of independent multiplication calculation units, each multiplication calculation unit being used to implement multiplication calculation of a plurality of data scales. The addition calculation unit group includes a plurality of independent addition calculation units. The normalization acceleration unit group includes a plurality of independent activation calculation units and normalization calculation units, wherein:

[0007] The data loading unit is configured to establish a data transfer task for each dense data in the sparse matrix multiplication task, and transfer each dense data to the register by performing a data transfer operation matched with the data transfer task.

[0008] The normalization acceleration unit group is configured to call at least one activation calculation unit to perform an activation operation according to each left and right matrix block respectively obtained from the register when the activation before multiplication is required, and send the activated left and right matrix blocks to the multiplication calculation unit group.

[0009] The multiplication calculation unit group is configured to select at least one multiplication calculation unit matched with the data size of each left and right matrix block obtained from the register or the normalization acceleration unit group to perform multiplication calculation, and provide the multiplication calculation result to the addition calculation unit group.

[0010] The addition calculation unit group is configured to call at least one addition calculation unit to perform addition calculation according to each multiplication calculation result received, and update the register according to the addition calculation result.

[0011] The normalization acceleration unit group is further configured to call at least one normalization calculation unit to perform normalization calculation and / or call at least one activation calculation unit to perform the activation operation according to the updated data obtained from the register when the post-multiplication normalization and / or post-multiplication activation operation is required, to obtain the final result.

[0012] According to another aspect of the embodiments of the present application, a recommendation system acceleration method is also provided, which comprises:

[0013] According to the received user recommendation request, a user feature matrix is constructed.

[0014] The user feature matrix is input into a pre-trained recommendation system model to perform a user recommendation result generation process.

[0015] The sparse matrix multiplication calculation requirement is constructed by the recommendation system model according to the user feature matrix and the model parameters of the recommendation system model, and is sent to the sparse matrix multiplication acceleration hardware as described in any of the embodiments of the present application, so that the sparse matrix multiplication acceleration hardware performs the matched sparse matrix multiplication acceleration operation.

[0016] The user recommendation result output after the recommendation system model calculation is fed back to the user.

[0017] According to another aspect of the embodiments of the present application, an AI chip is also provided, which comprises the sparse matrix multiplication acceleration hardware as described in any of the embodiments of the present application.

[0018] The technical scheme of the embodiment of the application is characterized in that, in the data loading unit of the sparse matrix multiplication acceleration hardware, only the dense data in the sparse matrix multiplication task is subjected to data transfer operation, in the multiplication calculation unit group, only the dense left and right matrix blocks are subjected to adaptive matrix multiplication calculation, and in the normalization acceleration unit group, the normalization and activation operations are implemented in a pure hardware manner, thus a new type of sparse matrix multiplication acceleration hardware that can be effectively applied to sparse matrix multiplication is creatively provided, based on which the invalid transfer and calculation of zero value elements in the sparse matrix can be reduced to the greatest extent, the fast calculation of the sparse matrix multiplication can be realized under limited calculation resources, and further, by applying the above acceleration hardware in the recommendation system scenario based on a large language model, a user recommendation system with low latency, low cost and low resource consumption can be realized, and the user experience is effectively improved.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the application, nor is it intended to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0021] Figure 1 is a structural schematic diagram of a sparse matrix multiplication acceleration hardware provided according to an embodiment of the application;

[0022] Figure 2 is a structural schematic diagram of another sparse matrix multiplication acceleration hardware provided according to an embodiment of the application;

[0023] Figure 3 is a structural schematic diagram of a data loading unit suitable for an embodiment of the application;

[0024] Figure 4 is a whole flowchart when the sparse matrix multiplication acceleration hardware provided by an embodiment of the application implements a sparse matrix multiplication task;

[0025] Figure 5 is a flowchart of a recommendation system acceleration method provided according to an embodiment of the application;

[0026] Figure 6 is a structural diagram of an AI chip provided according to an embodiment of the application. DETAILED DESCRIPTION

[0027] In order to make the person skilled in the art better understand the technical scheme of the present application, the technical scheme in the embodiments of the present application will be described clearly and completely below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the scope of protection of the present application.

[0028] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0029] Figure 1 A structural diagram of sparse matrix multiplication acceleration hardware is provided for the embodiments of the present application. The present embodiment can be applied to the case of implementing matrix multiplication calculation between sparse matrices or between a sparse matrix and a dense matrix by a pure hardware manner, and is particularly suitable for use in a recommendation system scenario based on a large language model. Correspondingly, as shown in Figure 1 The acceleration hardware includes:

[0030] The data loading unit 110, the register 120, the multiplication calculation unit group 130, the addition calculation unit group 140 and the normalization acceleration unit group 150, the multiplication calculation unit group 130 includes a plurality of independent multiplication calculation units 1301, each multiplication calculation unit 1301 is used to implement multiplication calculation of multiple data scales, the addition calculation unit group 140 includes a plurality of independent addition calculation units 1401, and the normalization acceleration unit group 150 includes a plurality of independent activation calculation units 1501 and normalization calculation units 1502.

[0031] Among them, the direction of the arrow drawn in Figure 1 represents the flow direction of data, specifically:

[0032] The data loading unit 110 is used to establish a data carrying task for the dense data in the sparse matrix multiplication task, and carries each dense data to the register 120 by executing a data carrying operation matched with the data carrying task.

[0033] wherein the sparse matrix multiplication task can be understood as a task in which one or both of the two matrices required to perform the multiplication calculation are sparse matrices. For example, for a sparse matrix multiplication task that requires the calculation of matrix A * matrix B, it is indicated that one or both of matrix A (which can also be referred to as the left operand matrix) and matrix B (which can also be referred to as the right operand matrix) are sparse matrices. Specifically, a sparse matrix can be understood as a matrix in which the number of zero elements is much greater than the number of non-zero elements.

[0034] Generally, the sparse matrix multiplication task is issued by the software layer to the acceleration hardware. For example, in a recommendation system scenario, it is required to multiply a user feature matrix with a specific parameter weight matrix in a large language model. However, the user feature matrix is often a large sparse matrix. Alternatively, if a Dropout operation is involved in the training process of the large language model, the parameter weight matrix can also be understood as a specific sparse matrix. Dropout is a regularization technique for neural networks that randomly "drops out" (temporarily disables) a portion of neurons (for example, 50% or 60%) during the training phase, forcing the model to not rely on specific neurons, thereby improving the generalization ability.

[0035] In the present embodiment, the sparse matrix multiplication task can be automatically generated by the sparse matrix multiplication acceleration hardware according to the information of the left and right operand matrices issued by the software layer. For example, if the matrix scale of the left and right operand matrices required to perform the matrix multiplication calculation is small, only one sparse matrix multiplication task can be generated. If the matrix scale of the left and right operand matrices required to perform the matrix multiplication calculation is large, the left and right operand matrices can be divided into blocks to generate multiple sparse matrix multiplication tasks, which are not limited by the present embodiment.

[0036] That is, the multiplication calculation required to be performed by the software layer has left and right operand matrices adapted thereto, and each sparse matrix multiplication task also has left and right operand matrices adapted thereto, which are the same or different.

[0037] It can be understood that after the large matrix required to perform the multiplication calculation is divided into blocks to construct the sparse matrix multiplication task, if the left operand matrix or the right operand matrix in a sparse matrix multiplication task is a full 0 matrix, the sparse matrix multiplication task can be directly filtered out.

[0038] In the related art, a certain compression technology can be used in the software level to remove a large number of 0 values in the sparse matrix, and then the obtained compressed matrix is stored efficiently. However, when the multiplication calculation hardware actually performs matrix multiplication calculation on the sparse matrix, the compressed matrix is first restored to the original sparse matrix by the software, and then the sparse matrix with a large amount of 0 value data is transported to the calculation hardware for storage. The multiplication calculation hardware takes the sparse matrix as the operation object and performs multiplication calculation on the sparse matrix with a large amount of 0 value data. It can be understood that the implementation manner of the related art introduces a large amount of invalid 0 value data transportation and a large amount of invalid 0 value data calculation.

[0039] In the embodiment, the data loading unit 110 is improved in hardware to transport only the dense data in the sparse matrix multiplication task to the sparse matrix multiplication acceleration hardware, so as to greatly improve the data transportation efficiency.

[0040] The dense data can be understood as the non-zero data in the left and right operand matrices matched with the sparse matrix multiplication task. In the embodiment, it is considered that the left or right operand matrix in the sparse matrix multiplication task may have non-zero data concentrated in some specific position region, and the other regions except the regions may be all 0 data. Further, the data loading unit 110 continues to perform data blocking of the left and right operand matrices in a finer granularity blocking dimension, and performs data transportation in the unit of the data block of the left and right operand matrices in the finer granularity.

[0041] For example, the left operation data block in the data block of a certain left and right operand matrix is all 0 data, or the right operation data block is all 0, and then the data transportation task for the data block of the left and right operand matrix can be directly filtered out. For another example, if a data transportation task needs to transport 128 data for the left operand, and the first 64 data are all 0, the 0 value data to be transported can be filtered out by modifying the task boundary of the data transportation task (taking the 65th data as the starting point of data transportation).

[0042] Through the above setting, the 0 value data in the large-scale left and right operand matrix in the sparse matrix multiplication task dimension can be effectively filtered out, and invalid 0 value data transportation is avoided.

[0043] In the embodiment, the data loading unit 110 is used to transport the data defined by the data transportation task from the outside of the sparse matrix multiplication acceleration hardware to the on-chip register 120 of the sparse matrix multiplication acceleration hardware, so as to be used for subsequent matrix multiplication calculation. That is, the essence of the data loading unit 110 is to load the left and right operand matrices required for the sparse matrix multiplication calculation in the form of a smaller granularity rectangular block to the on-chip register based on the rectangular loading shape.

[0044] The normalization acceleration unit group 150 is configured to, when activation needs to be performed before multiplication, call at least one activation calculation unit 1501 to perform activation operation according to each left and right matrix block obtained from the register 120 respectively, and send the activated left and right matrix blocks to the multiplication calculation unit group 130.

[0045] In the embodiment, the sparse matrix multiplication acceleration hardware is mainly used to realize calculation acceleration in various matrix multiplication scenarios. However, some specific matrix multiplication scenarios may need to perform activation operation on the matrix used for multiplication calculation first. For example, before performing attention mechanism-based multiplication calculation, the attention weight matrix needs to be subjected to activation operation of a certain type first, or in some model layers containing gating mechanism, the left and right operands both need to be subjected to activation operation of a certain type before matrix multiplication calculation.

[0046] In the embodiment, by constructing the activation calculation unit 1501 in the form of full hardware, the activation operation can be performed without software assistance, and the efficient data carrying result of the data loading unit 110 can be directly reused to further improve the calculation efficiency of the sparse matrix multiplication calculation.

[0047] Typically, the activation calculation unit 1501 can be specifically a SiLU hardware calculation unit, which is specifically configured to implement SiLU activation calculation on input data. The calculation formula of SiLU activation calculation is as follows:

[0048] ;

[0049] wherein x is an input value, is a Sigmoid function, which is used to map the input to the interval (0, 1).

[0050] The multiplication calculation unit group 130 is configured to select at least one multiplication calculation unit 1301 matched with the data size of each left and right matrix block obtained from the register 120 or the normalization acceleration unit group 150 to perform multiplication calculation, and provide the multiplication calculation result to the addition calculation unit group 140.

[0051] In the related art, the multiplication calculation module implemented by hardware generally also includes multiple multiplication calculation units. However, the data scale that each multiplication calculation unit can calculate is the same. The data scale can be understood as the data dimension of the left and right operand matrices, for example, the data dimension of the left operand matrix is 64*128, and the data dimension of the right operand matrix is 128*64, and the like. It can be understood that the implementation manner of the related art that the fixed multiplication calculation unit calculates the data of the fixed data scale, even if the 0-value data is effectively filtered out in the data loading unit 110, in order to adapt to the fixed data scale of the multiplication calculation unit, the 0-value data filtered out needs to be refilled into the left and right operand matrices with more dense data forms to adapt to the fixed data scale of the multiplication calculation unit, and then a lot of multiplication calculations of 0-value data are increased.

[0052] Based on this, each embodiment of the present application sets multiple multiplication calculation units 1301 with different data scales in the multiplication calculation unit group 130, and sets multiple multiplication calculation units 1301 with different data scales under each data scale. According to the actual data scale of each left and right operand matrix after the 0-value data is filtered out, the adaptive multiplication calculation unit 1301 is selected to implement calculation, and then the invalid multiplication calculation of 0-value data is further reduced on the basis of effectively reducing the 0-value data to be transported.

[0053] Among them, the data scale of the multiplication calculation unit 1301 can be adaptively designed according to the hardware parameters such as the data bit width of the smallest calculation unit, the parallel calculation capability, and the storage and bandwidth limitation, and the present embodiment does not limit this.

[0054] Among them, the multiplication calculation unit 1301 with different data scales can process left and right operand matrices with different data scales.

[0055] The addition calculation unit group 140 is used to call at least one addition calculation unit 1401 to implement addition calculation according to the received multiplication calculation results, and update the register 120 according to the addition calculation results.

[0056] Among them, the addition calculation unit group 140 is used to perform addition (or accumulation) calculation according to the multiplication calculation results calculated by one or more multiplication calculation units 1301 to obtain the final matrix multiplication calculation result. The number of addition calculation units 1401 included in the addition calculation unit group 140 is multiple, so as to realize parallel addition calculation.

[0057] Considering that the computing power consumption of addition calculation is much smaller than that of multiplication calculation, different addition calculation units 1401 can be set to the same data scale.

[0058] The normalization acceleration unit group 150 is also configured to, when multiplication post-normalization and / or multiplication post-activation operations are required, invoke at least one normalization calculation unit 1502 to perform normalization calculation and / or at least one activation calculation unit 1501 to perform activation operation according to the updated data obtained from the register 120, to obtain final results.

[0059] As a continuation of the previous example, in some other matrix multiplication scenarios, it is required to perform normalization operation or activation operation or both normalization operation and activation operation on the multiplication calculation results after the matrix multiplication calculation is completed. For example, multiplication post-normalization operation is required in the Softmax normalization scenario based on the attention mechanism, or multiplication post-activation operation is required in the nonlinear transformation scenario in the feedforward neural network.

[0060] In the prior art, the layer normalization (i.e., Layernormal) calculation is mainly implemented by pure software or a combination of software and hardware. The implementation of the layer normalization calculation needs to rely on sample expectation and sample variance, and therefore, the calculation is highly dependent on the output row data and has low calculation efficiency. With the continuous development of technology, a new type of normalization function, i.e., a dynamic tanh calculation function (DynamicTanh, DyT), has appeared. DyT can achieve the same effect as the layer normalization in most scenarios, and can even be superior to the calculation effect of the layer normalization, but does not need to calculate the mean and variance at all.

[0061] Further, in the embodiment, a tanh hardware calculation unit completely implemented by pure hardware is creatively proposed as the normalization calculation unit 1502. By arranging the tanh hardware calculation unit on the sparse matrix multiplication acceleration hardware, the data transfer result can be directly reused, the dynamic tanh is performed on the dense data required by the matrix multiplication calculation, and the normalization operation requirements of various matrix multiplication scenarios are met.

[0062] Correspondingly, in an optional implementation of the embodiment, the normalization calculation unit is specifically a tanh hardware calculation unit. The tanh hardware calculation unit is configured to perform dynamic tanh normalization calculation on input data.

[0063] The calculation formula of the dynamic tanh normalization calculation is as follows:

[0064] ;

[0065] x is an input value, is a standard hyperbolic tangent function, is a scaling factor, is a temperature coefficient, is a center offset, is a translation term.

[0066] In this embodiment, the normalization acceleration unit group 150 also contains a plurality of activation calculation units 1501 and a plurality of normalization calculation units 1502 to meet the parallelized activation operation and normalization operation requirements.

[0067] The technical scheme of the embodiment of the application, by only carrying out data transfer operations on the dense data in the sparse matrix multiplication task in the data loading unit of the sparse matrix multiplication acceleration hardware, only executing the adaptive matrix multiplication calculation on the dense left and right matrix blocks in the multiplication calculation unit group, and realizing the normalization and activation operations in the normalization acceleration unit group in a pure hardware manner, creatively provides a new type of acceleration hardware that can be effectively applied to sparse matrix multiplication. Based on the acceleration hardware, the invalid transfer and calculation of zero elements in the sparse matrix can be reduced to the maximum extent, and fast calculation of sparse matrix multiplication can be realized under limited computing resources. Furthermore, by applying the above acceleration hardware in the recommendation system scenario based on a large language model, a user recommendation system with low latency, low cost and low resource consumption can be realized, effectively improving the user experience.

[0068] Figure 2 Another structural diagram of the sparse matrix multiplication acceleration hardware provided by the embodiment of the application is provided. In this embodiment, the constituent elements in the sparse matrix multiplication acceleration hardware are further refined. Accordingly, as shown in Figure 2 The acceleration hardware further includes a user configuration unit 210, a sparse matrix multiplication task management unit 220, a task queue 230 and a sparse matrix multiplication task scheduling unit 240, wherein:

[0069] The user configuration unit 210 is configured to receive first user configuration information, wherein the first user configuration information includes various description parameters of the left and right matrices matched with the sparse matrix multiplication calculation requirement, indication information of whether to perform the pre-multiplication activation operation, indication information of whether to perform the post-multiplication normalization operation and indication information of whether to perform the post-multiplication activation operation.

[0070] In a specific example, the first user configuration information can specifically include:

[0071] Left matrix block (left operand required to perform multiplication calculation) base address, left matrix block shape participating in calculation, right matrix block base address, right matrix block shape participating in calculation, output base address, output base address based on dense representation vertex coordinates, left current calculation rectangular block left vertex row and column offset, left input rectangular width and height, right current calculation rectangular block left vertex row and column offset, right input rectangular width and height, data precision, whether to perform SiLu activation, whether to perform upper triangular mask, whether the output needs to perform dropout (the output needs to perform dropout will be combined with the mask matrix), output mask matrix and dropout matrix combined base address, offset (length is batch), whether the normalization parameter is a parameter or a scalar (configured as true or false, configured as true needs to perform data transfer, tanh or sigmod operation (0: tanh 1: sigmod), normalization rectangular data base address (empty when the scalar is empty), and normalization scalar parameter configuration (default is 0).

[0072] The sparse matrix multiplication task management unit 220 is configured to construct at least one sparse matrix multiplication task according to the first user configuration information in the user configuration unit 210, and send each sparse matrix multiplication task to the task queue 230.

[0073] In this embodiment, after obtaining the first user configuration information, the vertex storage positions of the left and right operands of the matrix multiplication required to be executed, the data shape, and whether the activation or normalization operation needs to be performed before and after multiplication and other information can be obtained. Based on the above information, one or more sparse matrix multiplication tasks can be constructed accordingly, and each sparse matrix multiplication task can further include the storage position and data shape (dimension) of the left and right operands required to be calculated in the task.

[0074] Correspondingly, after constructing one or more sparse matrix multiplication tasks, the sparse matrix multiplication task management unit 220 can send each sparse matrix multiplication task to the task queue 230, so that the sparse matrix multiplication task scheduling unit 240 schedules and executes each sparse matrix multiplication task constructed in the task queue 230.

[0075] In an optional embodiment of this embodiment, the sparse matrix multiplication task management unit 220 can be specifically used for:

[0076] After the calculation result matrix structure is constructed according to the first user configuration information and the left and right matrixes matching the sparse matrix multiplication calculation requirement, at least one candidate sparse matrix multiplication task is constructed according to the calculation result matrix structure, and the candidate sparse matrix multiplication tasks with all-0 elements as the calculation result are filtered out in each candidate sparse matrix multiplication task, so as to obtain each sparse matrix multiplication task.

[0077] In this embodiment, the specific calculation result matrix structure can be determined according to each item of information of the left and right operands corresponding to the most initial sparse matrix multiplication calculation requirement. For example, when the left operand matrix is 128*64 and the right operand is 64*128, the calculation result matrix structure is 128*128. Correspondingly, after the calculation result matrix structure is obtained, when it is determined that the data size of the calculation result matrix is relatively large, the data block method (for example, divided into four data blocks) can be used to determine each item of information of the left and right operands corresponding to each calculation result matrix block. Then, the candidate sparse matrix multiplication task corresponding to each calculation result matrix block can be determined.

[0078] Further, after each candidate sparse matrix multiplication task is determined, the left and right operand matrices required by each candidate sparse matrix multiplication task can be quickly scanned from the corresponding storage area. If it is scanned that the matrix elements of the left or right operand matrix of the candidate sparse matrix multiplication task A are all-0, it can be determined that the calculation result of the candidate sparse matrix multiplication task A is all-0. Therefore, the candidate sparse matrix multiplication task A can be directly filtered out, so as to effectively reduce the subsequent data transfer amount and data calculation amount, and improve the calculation efficiency of the sparse matrix multiplication.

[0079] The sparse matrix multiplication task scheduling unit 240 is configured to sequentially obtain the sparse matrix multiplication tasks from the task queue 230, and call each unit and / or unit group in the sparse matrix multiplication acceleration hardware to execute each sparse matrix multiplication task.

[0080] In this embodiment, the sparse matrix multiplication task scheduling unit 240 obtains each sparse matrix multiplication task stored in the task queue 230, triggers the data loading unit, one or more multiplication calculation units in the multiplication calculation unit group, one or more addition calculation units in the addition calculation unit group, one or more activation calculation units in the normalization acceleration unit group, and one or more normalization calculation units in the normalization acceleration unit group in the sparse matrix multiplication acceleration hardware through the serial and / or parallel scheduling mode, so that the above hardware units jointly execute each sparse matrix multiplication task.

[0081] The technical solution of the embodiment can construct a sparse matrix multiplication task by arranging a sparse matrix multiplication task management unit in the sparse matrix multiplication acceleration hardware, and can delete a sparse matrix multiplication task with all 0 calculation results in real time during the task construction process, further avoiding data transfer and dot multiplication of most invalid 0 value elements, and further improving the computing power resources of the chip.

[0082] To further describe the embodiments of the present application in detail, Figure 3 a structural diagram of a data loading unit suitable for the embodiments of the present application is shown. The embodiment is refined on the basis of the above-mentioned embodiments, and specifically further refines the structure and function of the data loading unit.

[0083] Correspondingly, as shown in Figure 3 , the data loading unit specifically includes a data transfer configuration subunit 310, a subtask management subunit 320, a subtask queue 330, a subtask scheduling executor 340, a plurality of hardware subtask execution threads 350, and a data transfer cache 360, wherein:

[0084] The data transfer configuration subunit 310 is configured to receive second user configuration information, wherein the second user configuration information includes description information of each data block of each left and right data block required for a single sparse matrix multiplication task to perform data transfer.

[0085] The second user configuration information can also be issued to the sparse matrix multiplication hardware by the upper software. The second user configuration information specifically includes left matrix shape, right matrix shape, left matrix transfer vertex coordinates, right matrix transfer vertex coordinates, width and height of the left matrix data block, and width and height of the right matrix data block.

[0086] Specifically, the left matrix shape and the right matrix shape can be understood as the data dimension information of the left and right operation operand matrices required to be completely transferred by a single sparse matrix multiplication task. The left matrix transfer vertex coordinates and the right matrix transfer vertex coordinates can be understood as the storage location of the vertex (generally, the top-left vertex of the matrix) of the above-mentioned left and right operation operand matrices in the external storage space. The width and height of the left matrix data block and the width and height of the right matrix data block can be understood as the dimension information of the left and right data blocks transferred during a single data transfer, which is pre-set.

[0087] It can be understood that if only the left matrix shape, the left matrix transfer vertex coordinates, and the width and height of the left matrix data block are configured in the data transfer configuration subunit 310, or only the right matrix shape, the right matrix transfer vertex coordinates, and the width and height of the right matrix data block are configured, data transfer of a single rectangular data can be realized.

[0088] The subtask management subunit 320 is configured to establish a plurality of data carrying tasks matched with the sparse matrix multiplication task according to the first user configuration information and the second user configuration information, and send each data carrying task to the subtask queue 330.

[0089] In this embodiment, the first user configuration information in the user configuration unit and the second user configuration information in the data carrying configuration subunit 310 are combined to construct a plurality of data carrying tasks with finer granularity for each sparse matrix multiplication task, and send the data carrying tasks to the subtask queue 330 for execution.

[0090] In the process of creating data carrying tasks by the subtask management subunit 320, the data carrying task configuration parameters can be further detected, for example, the data carrying task should not exceed the physical boundary of the entire sparse matrix multiplication task, or when the left and right operands in the sparse matrix multiplication task are the same matrix, a plurality of data carrying tasks can be merged accordingly when constructing the data carrying tasks.

[0091] Optionally, the vertex coordinates of the left and right data blocks to be carried, and the width and height of the left and right data blocks to be carried can be explicitly defined in each data carrying task.

[0092] In an optional implementation of this embodiment, the subtask management subunit 320 can be specifically configured to:

[0093] According to the first user configuration information and the second user configuration information, a plurality of candidate data carrying tasks matched with the sparse matrix multiplication task are initially established; the candidate data carrying tasks with all 0 data are filtered out, and the data carrying length of the candidate data carrying tasks with partial all 0 data is updated to obtain each data carrying task.

[0094] Similarly, the subtask management subunit 320 can first initially construct a plurality of candidate data carrying tasks, and then quickly scan the left and right data blocks to be carried by each candidate data carrying task from the external storage area. If it is scanned that all the data in a candidate data carrying task A is 0, the candidate data carrying task A can be filtered out accordingly. Alternatively, if it is scanned that a plurality of continuous data at the beginning or the end of a candidate data carrying task B are 0, the data length (for example, the width and height of the left and right data blocks to be carried) in the candidate data carrying task B can be updated accordingly.

[0095] Further, after updating the data carrying length of the plurality of candidate data carrying tasks, a plurality of candidate data carrying tasks with non-full load carrying can be further selected for task merging to obtain full load data carrying tasks, so as to further improve the data carrying efficiency.

[0096] The sub-task scheduling executor 340 is configured to sequentially obtain each data carrying task from the sub-task queue 330 and provide the at least one hardware sub-task execution thread 350.

[0097] On the basis of the above embodiments, the data loading unit can further include a task state management table 370.

[0098] Correspondingly, the sub-task scheduling executor 340 can be further configured to:

[0099] update the task state of each data carrying task recorded in the task state management table 370.

[0100] That is, the task state management table 370 records the task state of each data carrying task. For example, the task identification (ID) of a sparse matrix multiplication task (which can also be referred to as a main task), the task list of all data carrying tasks (which can also be referred to as sub-tasks) matched with the sparse matrix multiplication task, the state of each sub-task, the hardware sub-task execution thread 350 allocated to each sub-task, and the main task state.

[0101] It can be understood that the task state of the data carrying task maintained in the task state management table 370 is dynamically updated, and the task state management table 370 can be updated by the sub-task scheduling executor 340 at each task scheduling.

[0102] The sub-task scheduling executor 340 can sequentially obtain one data carrying task from the sub-task queue 330 and allocate the obtained data carrying task to an idle hardware sub-task execution thread 350 to perform a matched data carrying operation.

[0103] The hardware sub-task execution thread 350 is configured to perform a matched data carrying operation for the received data carrying task.

[0104] In this embodiment, the number of hardware sub-task execution threads 350 configured in the data loading unit can be preset according to the actual hardware resource condition, for example, can be 4, 6 or 8, and the embodiment does not limit this. Through the above setting, parallel operation of multiple data carrying tasks can be realized in a parallel pipeline manner.

[0105] Alternatively, each hardware sub-task execution thread 350 can also perform parallel data loading on multiple rows of data in a single data carrying task.

[0106] After each hardware sub-task execution thread 350 completes the execution of the allocated data carrying task, the execution result can be notified to the sub-task scheduling executor 340, and the sub-task scheduling executor 340 can correspondingly update the task state management table 370.

[0107] A data carrying cache area 360 is configured to cache the data to be carried.

[0108] Based on the above embodiments, the data carrying cache area 360 can specifically include a normal data cache area 3601 and a hit data cache area 3602, wherein:

[0109] The normal data cache area 3601 is configured to cache the normal data to be carried.

[0110] The hit data cache area 3602 is configured to cache the frequently hit data to be carried.

[0111] In a specific example, the normal data cache area 3601 can be a plurality of buffers, and the hit data cache area 3602 can be a whole cache. In the hit data cache area 3602, only data hit processing is performed, for example, whether the cache is hit can be determined based on the carrying base address and the length, and if so, the cache is reused to further improve the data read-write efficiency.

[0112] Based on the above embodiments, at least one of each multiplication calculation unit in the multiplication calculation unit group, each addition calculation unit in the addition calculation unit group, each normalization calculation unit in the normalization acceleration unit group, each activation calculation unit in the normalization acceleration unit group, and each hardware sub-task execution thread is configured to perform the matched operation in a parallel pipeline manner.

[0113] It is again emphasized that the data loading unit provided in the embodiments of the present application is mainly responsible for loading the double-rectangular data, i.e., loading the left and right matrix block data. Different data loading tasks are performed in multiple threads, i.e., each thread loads one row of continuous rectangular data. If multiple rows of data are continuous, the data can be loaded once or n times, thereby reducing the number of resources required for data loading, and supporting the carrying and loading of single-rectangular data, such as the data carrying of the normalization rectangular parameters which is single-rectangular.

[0114] Figure 4 is a whole flowchart of the sparse matrix multiplication acceleration hardware when implementing a sparse matrix multiplication task. As shown in the figure, the whole flow specifically includes: Figure 4

[0115] 1. The sparse matrix multiplication task scheduling unit in the sparse matrix multiplication acceleration hardware acquires the sparse matrix multiplication task based on a preset scheduling strategy.

[0116] ​2. Allocate a register for the computed result matrix (i.e., the output matrix block) of the current acquired sparse matrix multiplication task, i.e., allocate a storage space for the computed result of the sparse matrix multiplication task. If the register allocation fails, the sparse matrix multiplication task can be directly discarded, and a new sparse matrix multiplication task is acquired to perform the register allocation operation of the computed result matrix again.

[0117] 3. After the register allocation of the output matrix block is completed, start the multi-pipelined parallel operation to execute the sparse matrix multiplication task.

[0118] Specifically, first, configure a plurality of data carrying tasks matched with the sparse matrix multiplication task based on the configuration information, i.e., configure the loading information of the left and right operands of each data carrying task, and then acquire the left and right operands corresponding to each data carrying task from the register area based on the loading information of the left and right operands. If it is determined based on the configuration information that the activation operation before multiplication calculation is required, an activation unit included in the normalization acceleration unit group is called to perform the activation operation on each left and right operand, and then each left and right operand required for each small-scale multiplication operation is provided to the multiplication calculation unit group to perform multiplication calculation, and the calculation result is output to the accumulation queue for the corresponding accumulation calculation by the addition calculation unit group. If it is determined based on the configuration information that the activation operation before multiplication calculation is not required, each left and right operand required for each small-scale multiplication operation is directly provided to the multiplication calculation unit group to perform multiplication calculation, and the calculation result is output to the accumulation queue.

[0119] 4. Each addition calculation unit in the addition calculation unit group continuously consumes each accumulation calculation task in the accumulation queue.

[0120] 5. After the accumulation task in the accumulation queue is empty, if it is determined based on the configuration information that the normalization operation after multiplication calculation is required, the normalization operation is performed on the result after accumulation, and if it is determined based on the configuration information that the activation operation after multiplication is required, the activation operation is continuously performed on the normalized result to obtain the final result, and the final result is output to the outside of the sparse matrix multiplication acceleration hardware.

[0121] Figure 5 A flowchart of a recommendation system acceleration method provided by the embodiment of the present application. The embodiment can be applied to the case where the sparse matrix multiplication acceleration hardware performs calculation on the sparse matrix multiplication involved in the information recommendation process of the recommendation system. As shown in the figure, the method can include: Figure 5

[0122] S510, according to the received user recommendation request, construct a sparse user feature matrix.

[0123] ​S520, input the user feature matrix into a pre-trained recommendation system model to perform a user recommendation result generation process.

[0124] Optionally, the recommendation system model of the embodiments of the present application can be a deep learning model based on Transformer, or can also be a generative recommendation model based on a hierarchical sequence translation unit HSTU.

[0125] Specifically, the generative recommendation model based on HSTU uses dynamic tanh normalization calculation instead of layer normalization, and can directly reuse the tanh hardware calculation unit in the sparse matrix multiplication acceleration hardware of the embodiments of the present application.

[0126] S530, construct a sparse matrix multiplication calculation requirement by the recommendation system model according to the user feature matrix and the model parameters of the recommendation system model and issue it to the sparse matrix multiplication acceleration hardware as described in any one of the embodiments of the present application, so that the sparse matrix multiplication acceleration hardware performs matching sparse matrix multiplication acceleration operation.

[0127] S540, perform user feedback on the user recommendation result output after the recommendation system model calculation is completed.

[0128] The technical solution of the embodiments of the present application applies sparse matrix multiplication acceleration hardware to the recommendation system scenario where the user feature matrix has obvious sparse characteristics, and through the implementation mode of software and hardware combination scheduling, the model inference efficiency of the recommendation system model can be greatly improved. Even if a large language model with complex structure and large number of parameters is used, the model inference result can be obtained in a relatively short time, which can effectively satisfy the user's demand for the recommendation result generated by the large language model, effectively realize the function of fast model inference under effective computing resources, and effectively reduce the computing cost.

[0129] Figure 6 is a structural diagram of an AI chip according to the embodiments of the present application. As shown in Figure 6 The AI chip includes the sparse matrix multiplication acceleration hardware 610 as described in any one of the embodiments of the present application.

[0130] It should be understood that various forms of flows shown above can be used to reorder, add or delete steps. For example, the steps described in the present application can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present application can be achieved, which is not limited herein.

[0131] The above detailed description does not limit the scope of the application. Various modifications, combinations, sub-combinations and alternatives can be made to the detailed embodiment within the scope of the application. Any modification, equivalent replacement and improvement made without departing from the spirit and principle of the application shall fall within the scope of the application.

Claims

1. A sparse matrix multiplication acceleration hardware, characterized in that: The acceleration hardware includes a data loading unit, a register, a multiplication calculation unit group, an addition calculation unit group, and a normalization acceleration unit group. The multiplication calculation unit group includes multiple independent multiplication calculation units, each of which is used to implement multiplication calculations of various data sizes. The addition calculation unit group includes multiple independent addition calculation units. The normalization acceleration unit group includes multiple independent activation calculation units and normalization calculation units. A data loading unit is used to establish a data transfer task for the dense data in the sparse matrix multiplication task, and transfer each dense data to the register by executing a data transfer operation matching the data transfer task; The normalization acceleration unit group is used to call at least one activation calculation unit to perform an activation operation according to the left and right matrix blocks respectively obtained from the register when activation before multiplication is required, and send the activated left and right matrix blocks to the multiplication calculation unit group; The multiplication calculation unit group is used to select at least one matching multiplication calculation unit to perform multiplication calculation according to the data size of each left and right matrix block obtained from the register or the normalization acceleration unit group, and provide the multiplication calculation result to the addition calculation unit group; an addition calculation unit group, configured to call at least one addition calculation unit to perform addition calculation according to each received multiplication calculation result, and update a register according to the addition calculation result; The normalization acceleration unit group is also used to call at least one normalization calculation unit to perform normalization calculation and / or call at least one activation calculation unit to perform activation operation according to the updated data obtained from the register when post-multiplication normalization and / or post-multiplication activation operations are required to obtain the final result.

2. The sparse matrix multiplication acceleration hardware according to claim 1, characterized in that The acceleration hardware further includes a user configuration unit, a sparse matrix multiplication task management unit, a task queue, and a sparse matrix multiplication task scheduling unit, wherein: A user configuration unit, configured to receive first user configuration information, wherein the first user configuration information includes various descriptive parameters of left and right matrices that match sparse matrix multiplication calculation requirements, indication information of whether a pre-multiplication activation operation is required, indication information of whether a post-multiplication normalization operation is required, and indication information of whether a post-multiplication activation operation is required; a sparse matrix multiplication task management unit, configured to construct at least one sparse matrix multiplication task according to the first user configuration information in the user configuration unit, and send each sparse matrix multiplication task to a task queue; The sparse matrix multiplication task scheduling unit is used to sequentially obtain sparse matrix multiplication tasks from the task queue and call each unit and / or unit group in the sparse matrix multiplication acceleration hardware to execute each sparse matrix multiplication task.

3. The sparse matrix multiplication acceleration hardware according to claim 2, characterized in that The sparse matrix multiplication task management unit is specifically used for: After constructing the calculation result matrix structure based on the first user configuration information and the left and right matrices that match the sparse matrix multiplication calculation requirements, at least one alternative sparse matrix multiplication task is constructed based on the calculation result matrix structure, and among each alternative sparse matrix multiplication task, the alternative sparse matrix multiplication tasks whose calculation results are all 0 elements are filtered out to obtain each sparse matrix multiplication task.

4. The sparse matrix multiplication acceleration hardware according to claim 1, characterized in that The data loading unit specifically includes a data handling configuration subunit, a subtask management subunit, a subtask queue, a subtask scheduling executor, multiple hardware subtask execution threads, and a data handling buffer area, wherein: A data handling configuration subunit is configured to receive second user configuration information, wherein the second user configuration information includes description information of each left and right data block required for data handling in a single sparse matrix multiplication task; a subtask management subunit, configured to establish, based on the first user configuration information and the second user configuration information, a plurality of data handling tasks matching the sparse matrix multiplication task, and send each data handling task to a subtask queue; The subtask scheduling executor is used to sequentially obtain each data handling task from the subtask queue and provide it to at least one hardware subtask execution thread; The hardware subtask execution thread is used to execute the matching data handling operation for the received data handling task; The data transfer buffer area is used to cache the transferred data.

5. The sparse matrix multiplication acceleration hardware according to claim 4, characterized in that The subtask management subunit is specifically used to: Based on the first user configuration information and the second user configuration information, a plurality of alternative data handling tasks matching the sparse matrix multiplication task are preliminarily established; the alternative data handling tasks with all-0 data are filtered out, and the data handling length of the alternative data handling tasks with local all-0 data is updated to obtain each of the data handling tasks.

6. The sparse matrix multiplication acceleration hardware according to claim 4, characterized in that The data loading unit further includes a task status management table; Accordingly, the subtask scheduling executor is further used to: Update the task status of each data transfer task recorded in the task status management table; The data transfer buffer area specifically includes a normal data buffer area and a hit data buffer area, wherein: The general data buffer area is used to cache general data that needs to be moved; The hit data cache area is used to cache frequently hit data that needs to be moved.

7. The sparse matrix multiplication acceleration hardware according to any one of claims 1 to 6, characterized in that: The normalization calculation unit is specifically a tanh hardware calculation unit, and the activation calculation unit is a SiLU hardware calculation unit; The tanh hardware calculation unit is used to perform dynamic tanh normalization calculation on input data; The SiLU hardware computing unit is used to perform SiLU activation calculations on input data.

8. The sparse matrix multiplication acceleration hardware according to any one of claims 4 to 6, characterized in that: Each multiplication calculation unit in the multiplication calculation unit group, each addition calculation unit in the addition calculation unit group, each normalization calculation unit in the normalization acceleration unit group, each activation calculation unit in the normalization acceleration unit group, and at least one of the hardware subtask execution threads are used to perform matching operations in a parallel pipeline manner.

9. A recommendation system acceleration method, characterized in that: The method comprises: Construct a sparse user feature matrix based on the received user recommendation request; Inputting the user feature matrix into a pre-trained recommendation system model to generate user recommendation results; Constructing a sparse matrix multiplication calculation requirement according to the user feature matrix and the model parameters of the recommendation system model through the recommendation system model and sending the requirement to the sparse matrix multiplication acceleration hardware according to any one of claims 1 to 8, so that the sparse matrix multiplication acceleration hardware performs a matching sparse matrix multiplication acceleration operation; After the recommendation system model is calculated, the user recommendation results output are used for user feedback.

10. The method according to claim 9, characterized in that The recommendation system model is a deep learning model based on Transformer, or a generative recommendation model based on Hierarchical Sequence Transformation Unit (HSTU).

11. An AI chip, characterized in that: The method comprises the sparse matrix multiplication acceleration hardware as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Sparse matrix operation programming method and device based on data flow

    CN118409734A

  • Multiplication acceleration method and device of double sparse matrixes, equipment, medium and product

    CN119806639A