Computing devices, computing systems, storage media, and computer program products
Patent Information
- Application Number
- CN202610787969.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-09-11
AI Technical Summary
然而,在实际AI加速器中,尤其是在计算精度较低(例如,以 FP16 作为主要计算精度)的数据通路中,现有矩阵处理方法往往面临中间结果动态范围扩张、数值溢出、舍入误差累积以及访存开销高等问题
Smart Images

Figure CN122735901A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence chip technology, and in particular to a computing device, computing system, storage medium, and computer program product. Background Technology
[0002] With the widespread application of large models, sequence models, attention mechanisms, and structured matrix computation in AI inference and training, an increasing number of algorithms require matrix inversion, linear recursion, triangular system solving, or equivalent transformations to be performed on dedicated AI accelerators. In some application scenarios, the matrices to be processed have a strict lower triangular structure or a block lower triangular structure, such as causal structure matrices, block temporal dependency matrices, local attention matrices, and certain state propagation matrices.
[0003] Theoretically, since strictly lower triangular matrices possess the null-power property, the inverse matrix or equivalent results can be obtained using finite-term Neumann series, stepwise algebra, explicit higher-order power accumulation, or methods based on repeated squaring. However, in practical AI accelerators, especially in data paths with low computational precision (e.g., using FP16 as the primary computational precision), existing matrix processing methods often face problems such as dynamic range expansion of intermediate results, numerical overflow, accumulation of rounding errors, and high memory access overhead. Summary of the Invention
[0004] To address the aforementioned technical problems, this disclosure is proposed. Embodiments of this disclosure provide a computing device, a computing system, a storage medium, and a computer program product.
[0005] According to one aspect of the present disclosure, a computing device is provided, including: a memory and a computing unit;
[0006] The memory is used to store the first matrix output by the previous computing unit; the first matrix is a structured matrix; the previous computing unit is the computing unit corresponding to the previous operator in the operators corresponding to the computing unit in the model inference;
[0007] The computing unit is used to perform matrix multiplication and vector processing on the first matrix to obtain a target matrix; the target matrix is the inverse of the first matrix.
[0008] Optionally, the computing unit includes:
[0009] The matrix multiplication unit is used to perform matrix multiplication on the matrix to be processed and store the resulting candidate matrix into a temporary cache area in the computing unit; the matrix to be processed is the first matrix or a partial or vector unit of the first matrix, which is a locally updated matrix obtained by processing the vector unit.
[0010] The vector unit is used to iteratively update the candidate matrix in the temporary buffer multiple times to obtain the local update matrix corresponding to each iteration, until all regions in the candidate matrix are updated to obtain the second matrix; the target matrix is obtained based on the second matrix.
[0011] Optionally, the matrix multiplication unit is specifically used to use the first matrix as the matrix to be updated during the first iteration;
[0012] In non-first iterations, the local update matrix of the vector unit obtained from the temporary buffer in the previous iteration is used as the matrix to be updated; the previous iteration is the iteration before the current iteration;
[0013] Perform matrix multiplication on the matrix to be updated to obtain the candidate matrix corresponding to the current iteration and store it in the temporary cache area.
[0014] Optionally, the vector unit includes:
[0015] The data acquisition module is used to acquire the candidate matrix corresponding to the current iteration from the temporary cache area;
[0016] The element processing module is used to perform local region update on the candidate matrix according to the current iteration round corresponding to the current iteration, so as to obtain the local update matrix corresponding to the current iteration round;
[0017] A matrix adder is used to add the local update matrix to the matrix to be updated stored in the temporary cache to obtain the third matrix corresponding to the current iteration round;
[0018] A discriminator is used to judge the third matrix and determine whether to output the second matrix.
[0019] Optionally, the element processing module is specifically used to perform element-wise processing on the candidate matrix and the mask matrix corresponding to the current iteration round to obtain the local update matrix corresponding to the current iteration round and store it in the temporary cache area; the element-wise processing includes element-wise multiplication or element-wise selection.
[0020] Optionally, the mask matrix is determined based on the mask parameters pre-stored in the temporary buffer and the current iteration round; the mask parameters are determined based on the first matrix.
[0021] Optionally, the discriminator is specifically used to determine whether the third matrix has completed all region updates relative to the candidate matrix; in response to the third matrix having completed all region updates, the third matrix is used as the second matrix; in response to the third matrix not having completed all region updates, the third matrix is used as the matrix to be updated and stored in the temporary cache area.
[0022] Optionally, the vector unit further includes a scaling module, configured to perform scaling processing on the local region of the candidate matrix to be updated in the current iteration before or after the element processing module performs local region update.
[0023] Optionally, when the vector unit obtains the target matrix based on the second matrix, it is used to perform a synthesis process on the second matrix and the identity matrix to obtain the target matrix.
[0024] Optionally, the memory is an on-chip memory within the computing unit or an off-chip memory outside the computing unit.
[0025] According to another aspect of the embodiments of this disclosure, a computing system is provided, including:
[0026] At least one computing device as described in any of the above embodiments is used to perform matrix inverse calculations in model inference;
[0027] At least one computational core is used to perform computational processing in model inference and output the results of model inference.
[0028] According to another aspect of the present disclosure, a computer-readable storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the functions of the computing device described in any of the preceding embodiments.
[0029] According to another aspect of the present disclosure, a computer program product is provided, including computer program instructions that, when executed by a processor, implement the functions of the computing device described in any of the above embodiments.
[0030] The computing device provided in the above embodiments of this disclosure includes: a memory and a computing unit; the memory is used to store a first matrix output by the previous computing unit; the first matrix is a structured matrix; the previous computing unit is the computing unit corresponding to the previous operator in the operators corresponding to the computing unit in model inference; the computing unit is used to perform matrix multiplication and vector processing on the first matrix to obtain a target matrix; the target matrix is the inverse matrix of the first matrix. This embodiment relies only on matrix multiplication and vector processing to achieve matrix inversion, avoiding the intermediate value magnitude explosion caused by existing methods, reducing the risk of overflow and error accumulation, and improving numerical stability.
[0031] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0032] The accompanying drawings, which form part of this specification, illustrate embodiments of this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0033] This disclosure will become clearer with reference to the accompanying drawings and the following detailed description, wherein:
[0034] Figure 1 This is a schematic diagram of the structure of a computing device provided in an exemplary embodiment of this disclosure;
[0035] Figure 2 This is a schematic diagram of the process of a computing device for processing a correlation matrix provided in an exemplary embodiment of this disclosure;
[0036] Figure 3 This is a comparison chart of the dynamic range of intermediate results between the computing device provided in this disclosure and the prior art;
[0037] Figure 4 A block diagram of an electronic device according to an embodiment of the present disclosure is shown. Detailed Implementation
[0038] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. It is obvious that the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0039] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0040] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.
[0041] It should also be understood that in the embodiments disclosed herein, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0042] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.
[0043] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship. The data referred to in this disclosure can include unstructured data such as text, images, and videos, as well as structured data.
[0044] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.
[0045] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0046] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0047] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0048] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0049] The embodiments disclosed herein can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate together with a wide range of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems, etc.
[0050] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are executed by remote processing devices linked through communication networks. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.
[0051] Application Overview
[0052] In the process of realizing this disclosure, the inventors discovered that current mainstream AI accelerators typically include high-throughput matrix multiplication units and vector processing units, which are good at performing regular matrix multiplication and addition and element-wise operations, but have weak support for triangular predecessors or element-wise recursion with strong dependency chains and obvious irregular control flows.
[0053] Exemplary device
[0054] Figure 1 This is a schematic diagram of the structure of a computing device provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices, such as... Figure 1 As shown, the computing device provided in this embodiment includes a memory 11 and a computing unit 12.
[0055] Memory 11 is used to store the first matrix output by the previous computing unit.
[0056] The first matrix is a structured matrix, such as a strictly lower triangular matrix, a block lower triangular matrix, a block attention matrix, or other structured matrices. The previous computational unit is the computational unit corresponding to the previous operator in the operator corresponding to the computational unit in model inference. The model provided in this embodiment can be any deep learning model, such as a Large Language Model (LLM), a Vision-Language Model (VLM), etc. In this embodiment, as long as the output matrix is a structured matrix, it can be used as the previous computational unit.
[0057] In some alternative implementations, the first matrix is derived from the local dependency matrix, causal structure matrix, or other propagation matrix with a strict lower triangular structure during the attention calculation process.
[0058] Optionally, in this embodiment, the memory 11 is an on-chip memory within the computing unit 12 or an off-chip memory outside the computing unit 12.
[0059] In some optional embodiments, before inputting the first matrix into the computing unit, it can be processed according to the batch, head, chunk, or block dimensions, depending on the actual application scenario. This processing may include, but is not limited to, at least one of the following: rearrangement, grouping, or flattening, to form a data layout suitable for AI accelerator execution. In this embodiment, rearrangement simply changes the position of the dimensions while keeping the data content unchanged. By changing the order of the dimensions, the first matrix is arranged continuously in memory, conforming to the accelerator's reading habits. Grouping divides the first matrix into multiple blocks along any dimension, and these blocks can be executed in parallel by multiple computing cores. Flattening merges other dimensions into a single large dimension, retaining only one computing dimension, allowing data to directly enter the computing unit for computation.
[0060] The calculation unit 12 is used to perform matrix multiplication and vector processing on the first matrix to obtain the target matrix; the target matrix is the inverse of the first matrix.
[0061] For example, when the first matrix is represented as A, the target matrix is... Or its equivalent result.
[0062] The computing device provided in this embodiment can be part of an AI accelerator. The computing unit provided in this embodiment is the computing unit in the AI accelerator. The computing unit includes matrix multiplication units (Cube / PE array), vector units, on-chip cache SRAM and other structures. The memory can be part of the storage system of the AI accelerator.
[0063] This embodiment does not require any changes to the hardware structure of the accelerator. It only utilizes the inherent hardware structure of the accelerator to achieve the inversion of the first matrix through different execution orders. It also overcomes the problems of dynamic range expansion of intermediate results, numerical overflow, error accumulation, low utilization of matrix multiplication units, and high on-chip cache pressure that exist in the prior art when performing structured matrix inversion calculations on AI accelerators with limited computing power (e.g., FP16 (16-bit half-precision floating-point) AI accelerators).
[0064] This embodiment is particularly applicable to AI accelerators that include at least one of the following hardware features: support for FP16 data format; high-throughput matrix multiplication unit; vector processing unit; on-chip cache or scratchpad; and support for element-wise masking, selective write-back, scaling, and format conversion operations.
[0065] The computing device provided in the above embodiments of this disclosure includes: a memory and a computing unit; the memory is used to store a first matrix output by the previous computing unit; the first matrix is a structured matrix; the previous computing unit is the computing unit corresponding to the previous operator in the operators corresponding to the computing unit in model inference; the computing unit is used to perform matrix multiplication and vector processing on the first matrix to obtain a target matrix; the target matrix is the inverse matrix of the first matrix. This embodiment relies only on matrix multiplication and vector processing to achieve matrix inversion, avoiding the intermediate value magnitude explosion caused by existing methods, reducing the risk of overflow and error accumulation, and improving numerical stability.
[0066] In some optional embodiments, the computing unit 12 includes:
[0067] The matrix multiplication unit (GEMM core) is used to perform matrix multiplication on the matrix to be processed and store the resulting candidate matrix in a temporary buffer in the computation unit.
[0068] The matrix to be processed is the first matrix or a partial or vector unit of the first matrix, which is a locally updated matrix obtained by processing it.
[0069] In this embodiment, the temporary cache area is a portion of the on-chip cache SRAM in the computing unit.
[0070] Optionally, the calculation unit 12 provided in this embodiment can perform the inversion operation on the first matrix at once, or decompose the first matrix into multiple blocks, perform the inversion operation on one block as a matrix each time, and finally concatenate the results of the inversion operations on multiple blocks to obtain the inversion result matrix corresponding to the first matrix.
[0071] In this embodiment, matrix multiplication is performed by a matrix multiplication unit, which is the reason for the high throughput advantage of AI accelerators in regular matrix multiplication; and by storing candidate matrices in a temporary cache instead of writing them back in full immediately, external storage traffic and cache usage are reduced.
[0072] The vector unit is used to iteratively update the candidate matrix in the temporary buffer multiple times to obtain the local update matrix corresponding to each iteration, until all regions in the candidate matrix are updated to obtain the second matrix; the target matrix is obtained based on the second matrix.
[0073] In this embodiment, the vector unit restricts the update of each iteration to a local range, thereby significantly suppressing the overall expansion of the dynamic range of intermediate results.
[0074] Alternatively, the target matrix can be obtained based on the second matrix by performing a synthesis process on the second matrix and the identity matrix.
[0075] The Vector Unit in the AI accelerator is a parallel element-wise computation engine within the computational kernel, responsible for all non-matrix multiplication operators such as nonlinear activation, matrix addition and subtraction, element-wise multiplication (element-wise selection), numerical transformation, and conditional selection.
[0076] In this embodiment, the inherent function of the vector unit is utilized to perform multiple iterations, updating only a local region of the candidate matrix each time. This effectively reduces the risks of overflow, rounding errors, and outlier propagation in accelerators with lower computing power. Furthermore, the matrix addition and subtraction functions of the vector unit are used to synthesize the second matrix, after all regions have been updated, with the identity matrix to obtain the target matrix, which is equivalent to the stable inverse computation result of the target operation.
[0077] In some optional embodiments, the matrix multiplication unit (GEMM core) is specifically used to take the first matrix as the matrix to be updated during the first iteration; and during subsequent iterations, to take the local update matrix of the vector unit obtained from the temporary buffer as the matrix to be updated in the previous iteration.
[0078] The previous iteration refers to the iteration preceding the current iteration; for example, if the current iteration is the 3rd iteration, the previous iteration is the 2nd iteration.
[0079] Perform matrix multiplication on the matrix to be updated to obtain the candidate matrix corresponding to this iteration and store it in a temporary buffer.
[0080] In this embodiment, the matrix multiplication unit (GEMM core) is used to perform matrix multiplication on two matrices to be updated. For example, in an optional example, the matrix multiplication unit determines the candidate matrix using the following formula (1):
[0081] C = X × X Formula (1)
[0082] Where C represents the candidate matrix, X represents the matrix to be updated, and × represents matrix multiplication. Additionally, in some alternative embodiments, when the first matrix is split into multiple parts for processing, in this embodiment, the matrix to be updated is a part of the first matrix during the first iteration.
[0083] In some alternative embodiments, the vector unit includes:
[0084] The data acquisition module is used to obtain the candidate matrix corresponding to the current iteration from the temporary cache.
[0085] In this embodiment, since the matrix multiplication unit generates a new candidate matrix based on the current number of iterations, the candidate matrix needs to be retrieved from the temporary cache during each iteration.
[0086] The element processing module is used to perform local region updates on the candidate matrix according to the current iteration round, and obtain the local update matrix corresponding to the current iteration round.
[0087] In this embodiment, each iteration targets different local regions within the candidate matrix. Therefore, it is necessary to determine the region to be updated for each iteration before performing the update to avoid updating the same region multiple times or missing regions. For example, when updating a row in the candidate matrix, the first iteration updates the first row, the second iteration updates the second row, and so on, until all regions have been updated. Alternatively, when updating a block in the candidate matrix, all blocks in the candidate matrix can be sorted, and the corresponding block can be locally updated according to the sorting until all blocks in the candidate matrix have been updated.
[0088] The matrix adder is used to add the local update matrix to the matrix to be updated stored in the temporary buffer to obtain the third matrix corresponding to the current iteration round.
[0089] This embodiment utilizes the matrix addition and subtraction function of the vector unit to accumulate the local update matrix into the corresponding target region of the matrix to be updated, thereby obtaining the third matrix of local iteration. Since each iteration only updates a single target row or target block, while the rest of the region remains unchanged, the intermediate results will not expand synchronously across the entire matrix range as in the full high-order power expansion, thus keeping the intermediate dynamic range under control and preventing the explosion of intermediate value amplitude.
[0090] The discriminator is used to evaluate the third matrix and determine whether to output the second matrix.
[0091] In this embodiment, after each update of a region of the candidate matrix, it is necessary to determine whether the update of all regions of the candidate matrix has been completed after this iteration. Optionally, this determination can be achieved by comparing the third matrix with the candidate matrix across the entire region. Alternatively, when the region update is executed according to the target row order, target block order, or other predetermined scheduling order, it can be determined whether the third matrix obtained in this iteration corresponds to the last one in the target row order, target block order, or other predetermined scheduling order (e.g., the last row in the target row order, the last block in the target block order, etc.) in the corresponding order. This determines whether the candidate matrix has been updated. After the update is completed, the second matrix can be output.
[0092] Optionally, the discriminator is specifically used to determine whether the relative candidate matrix in the third matrix has completed the update of all regions. In response to the fact that the third matrix has completed the update of all regions, the third matrix is used as the second matrix; in response to the fact that the third matrix has not completed the update of all regions, the third matrix is stored as the matrix to be updated in the temporary buffer.
[0093] If the third matrix has not been updated in all regions, it is necessary to continue iterating to update other regions in the candidate matrix. At this time, the third matrix is stored in a temporary buffer as the matrix to be updated, so that the matrix multiplication unit can retrieve the matrix to be updated from the temporary buffer to determine the candidate matrix corresponding to the next iteration. This process is repeated until all regions are updated and the second matrix is output.
[0094] In some optional embodiments, the element processing module is specifically used to perform element-wise processing on the candidate matrix and the mask matrix corresponding to the current iteration round to obtain the local update matrix corresponding to the current iteration round and store it in a temporary cache area.
[0095] Element-by-element processing includes element-by-element multiplication or element-by-element selection.
[0096] In this embodiment, the vector processing unit is invoked to generate, load, or select a mask corresponding to the current iteration. This mask can be a row mask, column mask, block mask, or a combination thereof. The mask is determined based on the method of updating the region in each iteration of the inversion process (the mask form corresponds to the form of each updated region; for example, if each updated region is a row, the corresponding mask is a row mask). By using the mask, only data corresponding to the current target updated region is retained, reducing invalid calculations and invalid write-backs, thus achieving the technical effect of preventing intermediate value overflow.
[0097] This disclosure does not directly perform element-wise pre-generation in the traditional CPU style, but instead achieves rule execution for AI accelerators through "candidate matrix multiplication + local mask submission".
[0098] Optionally, the mask matrix is determined based on the mask parameters pre-stored in the temporary buffer and the current iteration round.
[0099] The mask parameters are determined based on the first matrix. Optionally, the mask parameters are determined based on the first matrix and the size of the region updated in each iteration.
[0100] Optionally, the mask matrix is a matrix that retains 1s only in the region updated in the current iteration round, and sets all other regions to 0. The current iteration round determines which region of the first matrix to retain data based on the mask parameters. For example, if the first matrix is a 10×10 matrix, the corresponding mask matrix is also a 10×10 matrix. The update method is to update one row at a time. When the current iteration round is the 5th round, only the 5th row of the mask matrix has a value, and all other positions are set to 0. The mask matrix ensures that the local update matrix after element-wise processing based on the mask matrix only retains the data of the region that needs to be updated in this iteration, while other regions are cleaned up, not committed, or not written back, reducing the processing load of intermediate values.
[0101] In one implementation, the mask matrix is dynamically generated by the vector unit based on the current iteration round, without the need to pre-store the complete dense mask matrix, thereby reducing storage overhead.
[0102] In some optional embodiments, the vector unit further includes a scaling module for performing scaling processing on the local region of the candidate matrix to be updated in the current iteration before or after the element processing module performs local region update.
[0103] Optionally, scaling may include, but is not limited to, at least one of the following: 1. Calculating statistics for the target row or block, such as maximum absolute value, mean square value, norm, or preset threshold comparison value; 2. Determining a scaling factor based on the statistics; 3. Scaling the candidate update amount before updating; 4. Performing normalization, cropping, descaling, or format adjustment on the target area after updating. Adding row-by-row or block-by-block scaling, cropping, normalization, or mixed precision accumulation strategies during the local update stage reduces the risk of low-precision accelerator overflow.
[0104] This embodiment further reduces the risk of overflow and rounding errors under low-power accelerators by performing scaling processing on local areas.
[0105] Figure 2 This is a schematic diagram illustrating the process of a computing device processing a correlation matrix according to an exemplary embodiment of this disclosure. Figure 2 As shown, the processing procedure in this embodiment includes:
[0106] Step 201: Obtain the first matrix A.
[0107] Optionally, the first matrix A may be processed according to the batch, head, chunk, or block dimensions. The processing may include, but is not limited to, at least one of the following: rearrangement, grouping, or flattening to form a data layout suitable for execution by the AI accelerator. These processing methods do not change the data content and structure of the first matrix, that is, the processed first matrix A is still a structured matrix.
[0108] Step 202: Initialize the first matrix A to obtain the matrix X to be updated, and record the identity matrix I.
[0109] The first matrix A is initialized as a structured representation, denoted as X. This structured initialization stores data according to matrix rules (strict lower triangular, block lower triangular, block attention, etc. in this application), storing only key parameters and not the entire dataset. This results in smaller storage, faster computation, and better compatibility with accelerators with lower computing power. Simultaneously, an identity matrix I or a unit block matrix is pre-configured for final reconstruction. The output format.
[0110] In some alternative embodiments, matrix X is stored in FP16 data format in an on-chip cache or off-chip memory and loaded in blocks into a local cache near the computing array during each iteration.
[0111] Step 203: Perform matrix multiplication on the current matrix X to be updated through the matrix multiplication unit in the calculation unit to generate a candidate matrix C.
[0112] For example, in an alternative embodiment, the candidate matrix C is determined by the method provided by the above formula (1).
[0113] This step is primarily performed by the matrix multiplication unit (GEMM core), leveraging the high throughput advantage of the AI accelerator in regular matrix multiplication. Furthermore, candidate matrices can be stored in an on-chip temporary buffer instead of being immediately written back in full.
[0114] Step 204: Local update of subsequent matrix C is achieved through mask matrix.
[0115] Optionally, a vector unit is invoked to generate, load, or select a mask matrix M corresponding to the current round. This mask matrix M can be a row mask, column mask, block mask, or a combination thereof, and is used to retain only the data corresponding to the currently updated target region.
[0116] For example, by performing element-wise filtering on the candidate matrix C and the mask matrix M (which has the same size as the candidate matrix C), the local update amount U can be obtained by referring to the following formula (2):
[0117] U = C ⊙ M Formula (2)
[0118] Here, ⊙ represents element-wise multiplication or element-wise selection. Through this step, only the data in the target row or target block is preserved, while the remaining non-target areas are zeroed out, not committed, or not written back.
[0119] Step 205: Perform local cumulative update to obtain the third matrix.
[0120] The local update amount U is accumulated into the corresponding target region of the matrix X to be updated, resulting in a new matrix to be updated (the third matrix). In some optional examples, local accumulation can be achieved based on the following formula (3):
[0121] X'← X + U Formula (3)
[0122] Where X' is the third matrix.
[0123] Since only a single target row or target block is updated in each round, while the rest of the region remains unchanged, the intermediate results do not expand synchronously across the entire matrix as in a full higher-order power expansion, thus keeping the intermediate dynamic range under control.
[0124] Step 206: Determine whether the third matrix has completed the update of all regions; if the update of all regions has been completed, output the third matrix as the second matrix; otherwise, store the third matrix as the matrix to be updated and return to step 203.
[0125] Step 207: Matrix composition and output.
[0126] Optionally, the second matrix X' is combined (accumulated) with the identity matrix I to obtain the target matrix Y, which is the inverse of the first matrix. For example, in an optional example, the matrix combination can be achieved based on the following formula (4):
[0127] Y = I + X formula (4)
[0128] Where Y is Or a stable inverse computation result equivalent to the target computation.
[0129] Figure 3 This is a comparison chart of the dynamic range of intermediate results between the computing device provided in this disclosure and the prior art. For example... Figure 3 As shown, the specific implementation process of matrix inversion based on Neumann series expansion includes: since strictly lower triangular matrices satisfy the null terminator property, they can be... Expand into a finite series I + A + + … + This method is simple in form, but if multiple powers are explicitly accumulated in hardware, it can easily introduce large intermediate tensors and cause the amplitude of intermediate values to increase rapidly. The specific implementation process of matrix inversion based on repeated squaring acceleration methods includes reducing the number of rounds of series expansion and improving theoretical parallelism by using squaring recursion or grouping recursion. However, on low-precision AI accelerators, such methods often form high-order terms with large amplitudes in the intermediate stages, causing FP16 overflow, numerical instability, and increased intermediate cache usage. The computing device provided in this disclosure has a slower increase in the amplitude of intermediate results and a lower peak value, making it more suitable for execution in FP16 environments.
[0130] This disclosure also provides a computing system, including:
[0131] At least one computing device as provided in any of the above embodiments is used to perform matrix inverse calculations in model inference.
[0132] At least one computational core is used to perform computational processing in model inference and output the results of model inference.
[0133] Any of the computing devices provided in the embodiments of this disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers. Alternatively, any of the computing devices provided in the embodiments of this disclosure can be executed by an accelerator. Further details will not be provided below.
[0134] Exemplary electronic devices
[0135] Below, for reference Figure 4 This describes an electronic device according to embodiments of the present disclosure. The electronic device may be either or both of a first device and a second device, or a standalone device independent of them, which may communicate with the first device and the second device to receive acquired input signals from them.
[0136] Figure 4 A block diagram of an electronic device according to an embodiment of the present disclosure is shown.
[0137] like Figure 4 As shown, the electronic device includes one or more processors and memory.
[0138] A processor can be a central processing unit (CPU) or other form of processing unit with data processing and / or instruction execution capabilities, and can control other components in an electronic device to perform desired functions.
[0139] The memory can store one or more computer program products, and the memory can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program products can be stored on the computer-readable storage medium, and the processor can run the computer program products to implement the operation of the computing device of the various embodiments of this disclosure described above and / or other desired functions.
[0140] In one example, the electronic device may also include input devices and output devices, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0141] In addition, the input device may also include, for example, a keyboard, a mouse, etc.
[0142] This output device can output various information to the outside, including determined distance information, direction information, etc. The output device may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0143] Of course, for the sake of simplicity, Figure 4 Only some of the components of the electronic device relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device may include any other suitable components depending on the specific application.
[0144] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform operations of the computing device according to various embodiments of this disclosure as described in the foregoing portions of this specification.
[0145] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0146] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform operations of the computing device according to various embodiments of this disclosure as described in the foregoing portion of this specification.
[0147] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0148] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0149] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0150] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0151] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.
[0152] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0153] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0154] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A computing device, comprising: include: Memory and computing units; The memory is used to store the first matrix output by the previous computing unit; The first matrix is a structured matrix; The previous calculation unit is the calculation unit corresponding to the previous operator in the operator corresponding to the calculation unit in the model inference; The computing unit is used to perform matrix multiplication and vector processing on the first matrix to obtain a target matrix; the target matrix is the inverse of the first matrix.
2. The computing device of claim 1, wherein, The computing unit includes: The matrix multiplication unit is used to perform matrix multiplication on the matrix to be processed and store the resulting candidate matrix into a temporary cache area in the computing unit; the matrix to be processed is the first matrix or a partial or vector unit of the first matrix, which is a locally updated matrix obtained by processing the vector unit. The vector unit is used to iteratively update the candidate matrix in the temporary buffer multiple times to obtain the local update matrix corresponding to each iteration, until all regions in the candidate matrix are updated to obtain the second matrix; the target matrix is obtained based on the second matrix.
3. The computing device of claim 2, wherein, The matrix multiplication unit is specifically used to use the first matrix as the matrix to be updated during the first iteration. In non-first iterations, the local update matrix of the vector unit obtained from the temporary buffer in the previous iteration is used as the matrix to be updated; the previous iteration is the iteration before the current iteration; Perform matrix multiplication on the matrix to be updated to obtain the candidate matrix corresponding to the current iteration and store it in the temporary cache area.
4. The computing device of claim 2 or 3, wherein, The vector unit includes: The data acquisition module is used to acquire the candidate matrix corresponding to the current iteration from the temporary cache area; The element processing module is used to perform local region update on the candidate matrix according to the current iteration round corresponding to the current iteration, so as to obtain the local update matrix corresponding to the current iteration round; A matrix adder is used to add the local update matrix to the matrix to be updated stored in the temporary cache to obtain the third matrix corresponding to the current iteration round; A discriminator is used to judge the third matrix and determine whether to output the second matrix.
5. The computing device of claim 4, wherein, The element processing module is specifically used to perform element-wise processing on the candidate matrix and the mask matrix corresponding to the current iteration round to obtain the local update matrix corresponding to the current iteration round and store it in the temporary cache area; the element-wise processing includes element-wise multiplication or element-wise selection.
6. The computing device of claim 5, wherein, The mask matrix is determined based on the mask parameters pre-stored in the temporary buffer and the current iteration round; The mask parameters are determined based on the first matrix.
7. The computing device of any of claims 4-6, wherein, The discriminator is specifically used to determine whether the third matrix has completed all region updates relative to the candidate matrix, and in response to the third matrix having completed all region updates, the third matrix is used as the second matrix. In response to the fact that the third matrix has not completed updating all regions, the third matrix is stored as the matrix to be updated in the temporary cache area.
8. The computing device of any of claims 4-7, wherein, The vector unit further includes a scaling module, which is used to perform scaling processing on the local region of the candidate matrix to be updated in the current iteration before or after the element processing module performs local region update.
9. The computing device of any of claims 2-8, wherein, When the vector unit obtains the target matrix based on the second matrix, it is used to perform a synthesis process on the second matrix and the identity matrix to obtain the target matrix.
10. The computing device of any of claims 1-9, wherein, The memory is either on-chip memory within the computing unit or off-chip memory outside the computing unit.
11. A computing system, characterized in that, include: At least one computing device as described in any one of claims 1-10, for performing matrix inverse calculations in model inference; At least one computational core is used to perform computational processing in model inference and output the results of model inference.
12. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the functions of the computing device described in any one of claims 1-10.
13. A computer program product comprising computer program instructions, characterized in that, When executed by a processor, the computer program instructions implement the functions of the computing device as described in any one of claims 1-10.