Tensor core, processor, data processing method, electronic device, and storage medium

By reusing dot product units of different precisions in the tensor kernel to perform matrix multiplication operations with scaling factors, the problem of increased hardware area in the prior art is solved, hardware cost and power consumption are reduced, and system performance and chip integration are improved.

CN120872284BActive Publication Date: 2025-11-28SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites -1 Cited by

Patent Information

Application Number
CN202511299659.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-11-28
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Existing technologies require additional multipliers and adders inside the tensor kernel for matrix multiplication operations using scaling factors, which increases hardware area, hardware cost, and power consumption.

Method used

By reusing multiplication units of different precisions in the tensor kernel to perform matrix multiplication operations using scaling factors, the additional hardware area cost is reduced, existing hardware resources are fully utilized, the multiplication and accumulation modules are eliminated, and matrix multiplication operations are performed using hardware paths of multiplication units across precisions.

Benefits of technology

This reduces hardware area costs, optimizes system performance, lowers hardware costs and power consumption, and increases chip integration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872284B_ABST
    Figure CN120872284B_ABST
Patent Text Reader

Abstract

A tensor core, a processor, a data processing method, an electronic device and a storage medium are applied to the field of tensor processing. The tensor core comprises a first point multiplication unit and a scaling factor matrix multiplication processing module, the scaling factor matrix multiplication processing module comprises a second point multiplication unit, the first point multiplication unit and the second point multiplication unit support point multiplication operations of different floating point precisions, the tensor core is configured to receive a first tensor, a second tensor, a scaling factor of the first tensor, a scaling factor of the second tensor, an offset term of the first tensor and an offset term of the second tensor, perform a matrix multiplication operation using a scaling factor and an offset term by using the first point multiplication unit and the scaling factor matrix multiplication processing module, and obtain a matrix multiplication operation result of the first tensor and the second tensor. By multiplexing the point multiplication units of different precisions in the tensor core to perform the matrix multiplication operation using the scaling factor, the additional hardware area cost is reduced, and the existing hardware resources are fully multiplexed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to a tensor core, a processor, a data processing method, an electronic device and a non-transitory computer-readable storage medium. BACKGROUND

[0002] Floating point quantization refers to format conversion of a high-precision floating point number format (such as FP16, BF16, FP32, etc.) to a low-precision floating point number format (such as FP4, FP8, etc.). The low-precision floating point number format can greatly improve the calculation efficiency, especially the general matrix multiplication (GEMM) operation required by an artificial intelligence (AI) model. However, low-precision matrix multiplication will cause precision loss, thereby reducing the prediction and generation effect of the AI model. Therefore, in the low-precision calculation process of the AI model, a corresponding scaling factor is usually calculated from the current data in the process of floating point quantization to scale the expression range of the data, thereby reducing the precision loss of the calculation result.

[0003] At present, for the matrix multiplication operation using the scaling factor, additional multipliers and adders need to be added inside the tensor core, which brings an additional hardware area cost. SUMMARY

[0004] The present application provides a tensor core, including a first point multiplication unit, a scaling factor matrix multiplication processing module and a cross-precision point multiplication unit hardware path, the scaling factor matrix multiplication processing module includes a second point multiplication unit, the first point multiplication unit and the second point multiplication unit are configured to support point multiplication operations of different floating point precisions, the tensor core is configured to receive a first tensor, a second tensor, a scaling factor of the first tensor and a scaling factor of the second tensor, perform a matrix multiplication operation using a scaling factor by using the first point multiplication unit and the scaling factor matrix multiplication processing module, and obtain a matrix multiplication operation result of the first tensor and the second tensor, the first tensor and the second tensor are both of a first floating point precision, wherein in the process of performing the matrix multiplication operation using a scaling factor, the first point multiplication unit is configured to perform a matrix multiplication operation of the first floating point precision of the first tensor and the second tensor to obtain a first multiplication result, the cross-precision point multiplication unit hardware path is configured to transmit the first multiplication result to the scaling factor matrix multiplication processing module, and the scaling factor matrix multiplication processing module is configured to determine the matrix multiplication operation result by using the second point multiplication unit according to the first multiplication result, the scaling factor of the first tensor and the scaling factor of the second tensor.

[0005] For example, in the tensor core provided in at least one embodiment of the application, the scaling factor matrix multiplication processing module further comprises a scaling factor multiplication processing submodule, the scaling factor multiplication processing submodule is configured to determine a second multiplication result of the scaling factor of the first tensor and the scaling factor of the second tensor; the second point multiplication unit is configured to determine the matrix multiplication operation result according to the second multiplication result and the first multiplication result.

[0006] For example, in the tensor core provided in at least one embodiment of the application, the second point multiplication unit supports point multiplication operation of a second floating point precision higher than the first floating point precision.

[0007] For example, in the tensor core provided in at least one embodiment of the application, the first tensor is divided into T first tensor blocks in the column direction, each first tensor block has a shape size of m x k, the second tensor is divided into T second tensor blocks in the row direction, each second tensor block has a shape size of k x n, the scaling factor of the first tensor is divided into T first vectors in the column direction, each first vector has a shape size of m x 1, the scaling factor of the second tensor is divided into T second vectors in the row direction, each second vector has a shape size of 1 x n, the first point multiplication unit is configured to perform matrix multiplication operation between the T first tensor blocks and the T second tensor blocks to obtain T matrix multiplication intermediate results, wherein the first multiplication result comprises the T matrix multiplication intermediate results; the scaling factor multiplication processing submodule is configured to perform outer product operation between the T first vectors and the T second vectors to obtain T outer product operation results, wherein the second multiplication result comprises the T outer product operation results; the second point multiplication unit is configured to determine the matrix multiplication operation result according to the T matrix multiplication intermediate results and the T outer product operation results, wherein m, n, k, T are all positive integers.

[0008] For example, in the tensor core provided in at least one embodiment of the application, the shape size of the matrix multiplication operation result, each matrix multiplication intermediate result and each outer product operation result is m x n, a plurality of second point multiplication units are configured to perform a plurality of first operations in parallel to obtain the matrix multiplication operation result through m x n first operations, wherein the first operation comprises performing multiplication and accumulation operation of the i-th row and the j-th column elements in the T matrix multiplication intermediate results and the T outer product operation results to obtain the i-th row and the j-th column elements in the matrix multiplication operation result, wherein i is a positive integer less than or equal to m, and j is a positive integer less than or equal to n.

[0009] For example, in the tensor core provided in at least one embodiment of the application, the second point multiplication unit is configured to perform one first operation at a time, and when performing the first operation, the second point multiplication unit includes determining an i-th row and j-th column element of each of the T matrix multiplication intermediate results, to obtain T first elements; determining an i-th row and j-th column element of each of the T outer product operation results, to obtain T second elements; and performing point multiplication operations between the T first elements and the T second elements, to obtain an i-th row and j-th column element of the matrix multiplication operation result.

[0010] For example, in the tensor core provided in at least one embodiment of the application, the minimum number of point multiplications N of the single-time point multiplication operation supported by the second point multiplication unit is greater than or equal to T, N is a positive integer, and when performing the point multiplication operations of the T first elements and the T second elements, the second point multiplication unit includes performing N point multiplication calculation operations and determining the sum of operation results of the N point multiplication calculation operations, and in response to N being greater than T, performing point multiplication calculation operations of N first elements and N second elements when performing the N point multiplication calculation operations, and the N first elements include the T first elements and N-T 0s, and the N second elements include the T second elements and N-T 0s.

[0011] For example, in the tensor core provided in at least one embodiment of the application, the first floating-point precision is FP4, and the second point multiplication unit supports point multiplication operations of FP8 and higher than FP8 precision.

[0012] For example, in the tensor core provided in at least one embodiment of the application, the second point multiplication unit is configured to perform the process of determining the matrix multiplication operation result at FP16 precision.

[0013] For example, in the tensor core provided in at least one embodiment of the application, each first point multiplication unit is configured to perform a matrix multiplication operation between one first tensor block and a corresponding second tensor block, to obtain one matrix multiplication intermediate result, and a plurality of the first point multiplication units perform a plurality of the matrix multiplication operations in parallel, wherein when the first point multiplication unit performs the matrix multiplication operation of the first tensor block and the corresponding second tensor block, the first point multiplication unit includes performing m*n second operations, and each second operation includes performing point multiplication calculation operations of k pairs of elements.

[0014] In at least one embodiment of the application, a processor is provided, which includes the tensor core described in any embodiment of the application.

[0015] The data processing method provided in at least one embodiment of the present application is applied to a tensor core, wherein the tensor core includes a first point multiplication unit, a scaling factor matrix multiplication processing module, and a cross-precision point multiplication unit hardware channel, the scaling factor matrix multiplication processing module includes a second point multiplication unit, the first point multiplication unit and the second point multiplication unit are configured to support point multiplication operations of different floating-point precisions, and the data processing method includes the following steps: receiving a first tensor, a second tensor, a scaling factor of the first tensor, and a scaling factor of the second tensor; performing a scaling factor matrix multiplication operation by using the first point multiplication unit and the scaling factor matrix multiplication processing module to obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating-point precision; and performing the scaling factor matrix multiplication operation by using the first point multiplication unit and the scaling factor matrix multiplication processing module to obtain the matrix multiplication operation result of the first tensor and the second tensor includes the following steps: performing a matrix multiplication operation of the first floating-point precision of the first tensor and the second tensor by using the first point multiplication unit to obtain a first multiplication result; transmitting the first multiplication result to the scaling factor matrix multiplication processing module by using the cross-precision point multiplication unit hardware channel; and determining the matrix multiplication operation result by using the second point multiplication unit according to the first multiplication result, the scaling factor of the first tensor, and the scaling factor of the second tensor.

[0016] For example, in the data processing method provided in at least one embodiment of the present application, the scaling factor matrix multiplication processing module further includes a scaling factor multiplication processing sub-module, and determining the matrix multiplication operation result by using the second point multiplication unit according to the first multiplication result, the scaling factor of the first tensor, and the scaling factor of the second tensor includes the following steps: determining a second multiplication result of the scaling factor of the first tensor and the scaling factor of the second tensor by using the scaling factor multiplication processing sub-module; and determining the matrix multiplication operation result by using the second point multiplication unit according to the second multiplication result and the first multiplication result.

[0017] The electronic device provided in at least one embodiment of the present application includes a memory that non-transiently stores computer executable instructions, and a processor configured to run the computer executable instructions, wherein the computer executable instructions, when run by the processor, implement the data processing method according to any one of the embodiments of the present application.

[0018] The non-transient computer readable storage medium provided in at least one embodiment of the present application stores computer executable instructions, and the computer executable instructions, when executed by a processor, implement the data processing method according to any one of the embodiments of the present application.

[0019] The application performs the matrix multiplication operation using the scaling factor by multiplexing the point multiplication units with different precisions in the tensor core, reduces the additional hardware area cost, fully multiplexes the existing hardware resources, reduces the multiplier setting and does not need to additionally set the adder, reduces the hardware area cost, optimizes the system performance, reduces the hardware cost and power consumption, and improves the chip integration. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only related to some embodiments of the present disclosure, but not limited to the present disclosure.

[0021] Figure 1A It is a schematic structural diagram of a general-purpose graphics processor (GPGPU);

[0022] Figure 1A It is a schematic structural diagram of a tensor core;

[0023] Figure 2 It is a schematic structural diagram of the tensor core provided by at least one embodiment of the present disclosure;

[0024] Figure 3 It is a schematic structural diagram of the scaling factor matrix multiplication processing module 102 provided by at least one embodiment of the present disclosure;

[0025] Figure 4 It is a schematic diagram of the first tensor, the second tensor, the scaling factor of the first tensor, and the scaling factor of the second tensor provided by at least one embodiment of the present disclosure;

[0026] Figure 5 It is a schematic diagram of the first operation provided by at least one embodiment of the present disclosure;

[0027] Figure 6 It is a schematic structural diagram of the processor provided by at least one embodiment of the present disclosure;

[0028] Figure 7 It is a schematic flowchart of the data processing method provided by at least one embodiment of the present disclosure;

[0029] Figure 8 It is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure;

[0030] Figure 9 It is a schematic diagram of a non-transitory computer readable storage medium provided by at least one embodiment of the present disclosure. DETAILED DESCRIPTION

[0031] In order to make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the following will be combined with the drawings of the embodiments of the present disclosure to make a clear and complete description of the technical solutions of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without any inventive effort fall within the protection scope of the present disclosure.

[0032] Unless otherwise defined, technical terms or scientific terms used in the present disclosure shall have the ordinary meaning of the terms to a person of ordinary skill in the art to which the present disclosure belongs. The terms "first", "second" and similar terms used in the present disclosure do not denote any order, quantity or importance, but are used to distinguish different components. The terms "comprise", "contain" and similar terms mean that the elements or objects before the terms encompass the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "up", "down", "left", "right" and the like are only used to indicate relative positional relationships, and when the absolute positions of the described objects are changed, the relative positional relationships can also be changed accordingly. In order to keep the following description of the embodiments of the present disclosure clear and concise, the present disclosure omits the detailed description of some known functions and known components.

[0033] A floating point number (FP) is mainly used to represent a decimal number, and is usually composed of three parts, i.e. a sign bit, an exponent part and a mantissa part. For example, a floating point number V can be represented as follows:

[0034]

[0035] wherein the sign bit s can be 1 bit, which determines whether the floating point number V is negative or positive; M represents the mantissa part, which can include multiple bits, and is in the form of a binary fraction, defining the precision of the floating point number; E represents the exponent (also referred to as the index value), which is used to weight the floating point number, representing the position of the decimal point in the floating point number V, and defining the value range of the floating point number.

[0036] Traditional floating point numbers usually include three formats, i.e. half-precision floating point number (FP16), single-precision floating point number (FP32) and double-precision floating point number (FP64), and the exponent parts and the mantissa parts of the three formats have different bit numbers.

[0037] AI accelerators and the like have been widely used for deep learning model training. For the convolution operation commonly used in deep learning models, special optimization is made in the design of software and hardware to accelerate the calculation. For example, various data formats of floating point numbers are developed for optimization in the field of artificial intelligence or deep learning, such as BF16 (brain floating point 16, bit width of 16 bits), BF24 (brain floating point 24, bit width of 24 bits), TF32 (Tensor Float 32, bit width of 19 bits), and the like. These data formats can greatly reduce the required operation resources and power consumption, especially for matrix multiplication or convolution multiplication operations. In addition, the processor also supports some conventional floating point number types, such as half-precision floating point number (FP16, bit width of 16 bits) or single-precision floating point number (FP32, bit width of 32 bits).

[0038] Figure 1A An exemplary structure diagram of a general-purpose graphics processor (GPGPU).

[0039] As shown in Figure 1A , the general-purpose graphics processor is actually an array of programmable multi-processors. For example, the programmable multi-processors can be streaming processor clusters (SPCs), such as the streaming processor cluster 1,..., streaming processor cluster M shown in Figure 1A , where M is a positive integer greater than 1. In the general-purpose graphics processor, one streaming processor cluster processes one computing task, or multiple streaming processor clusters process one computing task. The multiple streaming processor clusters share data through a global cache or a global memory.

[0040] As shown in Figure 1A , taking the streaming processor cluster 1 as an example, one streaming processor cluster includes multiple computing units, such as the computing unit 1, the computing unit 2,..., the computing unit N shown in Figure 1A , where N is a positive integer.

[0041] A computing unit includes multiple cores (also referred to as computing cores or computing cores), each of which includes an arithmetic logic unit (ALU), a floating point calculation unit, and the like, and is used to execute specific computing tasks. In addition, the computing unit also includes a register (such as the register stack shown in Figure 1A ) and a shared memory, which are used to store source data and destination data related to the computing task in a hierarchical manner. The shared memory in one computing unit is used to share data among the cores in the computing unit.

[0042] As shown in Figure 1AAs shown, each computational unit also provides a tensor core for performing tensor-related computations, such as matrix multiplication operations using the GEMM operator. Tensors are a crucial data structure in deep learning; they are high-dimensional generalizations of scalars, vectors, and matrices. Tensor operations are commonly used in the training and inference of deep learning models, and tensor cores can accelerate matrix multiplication. Tensor cores across multiple computational units can be uniformly scheduled and controlled.

[0043] For example, each computing unit can also provide a Vector Core Engine (not shown in the figure) for performing vector-related computations, such as vector-related arithmetic and logical operations, such as accumulation, reduction, and regular addition, subtraction, multiplication, and division.

[0044] In parallel computing, computational tasks are typically executed by multiple threads. These threads are divided into multiple thread blocks before execution in a general-purpose graphics processor (or parallel computing processor), and then dispatched via a thread block distribution module. Figure 1A (Not shown in the image) Multiple thread blocks are distributed to various computation units. All threads in a thread block must be assigned to the same computation unit for execution. Simultaneously, thread blocks are broken down into minimum execution thread bundles (or simply warps), each containing a fixed number (or less than this fixed number) of threads, for example, 32 threads. Multiple thread blocks can execute in the same computation unit or in different computation units.

[0045] In each computing unit, the thread beam scheduling / distribution module ( Figure 1A (Not shown in the diagram) Thread bundles are scheduled and allocated so that multiple computing cores within the computing unit can run thread bundles. Depending on the number of computing cores in the computing unit, multiple thread bundles within a thread block can be executed concurrently or in a time-sharing manner. Multiple threads within each thread bundle execute the same instructions. Memory execution instructions are issued to shared memory within the computing unit or further issued to intermediate-level caches (such as vector operation unit caches), global caches, or global memory for read and write operations, etc.

[0046] Low-precision matrix multiplication is gradually widely used in the training and inference of AI large models, because it can bring huge performance benefits with acceptable precision loss. In GPU (Graphics Processing Unit) or GPGPU, matrix multiplication is usually executed by tensor cores in hardware. Among them, the low-precision tensor core is several times the power of the high-precision tensor core, bringing higher computing efficiency. In addition, the data volume of the low-precision tensor itself is also several times less than that of the high-precision tensor, bringing higher data transmission efficiency. Therefore, the end-to-end efficiency brought by low-precision tensor calculation is almost several times improved. For example, the tensor core power of FP4 can be 2-8 times that of FP8, 4-16 times that of FP16 / BF16, or even higher. And the data volume is 1 / 2 of FP8, 1 / 4 of FP16 / BF16.

[0047] The point multiplication unit is a hardware structure in the tensor core for performing point multiplication operations, which can efficiently calculate the point multiplication operation of each element in two vectors and accumulate the operation results of the point multiplication operation to obtain the point multiplication operation result. Through the cooperative work of multiple point multiplication units, the tensor core realizes large-scale matrix multiplication operation. For example, the point multiplication units are organized into an array to efficiently perform matrix multiplication.

[0048] In order to reduce the precision loss, the low-precision matrix multiplication introduces a scaling factor to maximize the numerical representation ability of the low-precision tensor. For example, the low-precision matrix multiplication includes matrix multiplication for low-precision (such as FP4, FP8, etc.) tensors.

[0049] For low-precision matrix multiplication, for example, D=A'×B'+C', A', B', C' and D are all high-precision tensors, and the general matrix multiplication operation using the scaling factor can be described as:

[0050]

[0051] Among them, represents element multiplication, represents outer product, represents matrix multiplication, α is the scaling factor of tensor A', β is the scaling factor of tensor B', σ is the scaling factor of tensor C', γ is the scaling factor of tensor D, quantized tensor A, quantized tensor B, quantized tensor C and quantized tensor D are low-precision tensors after floating-point quantization of tensors A', B', C' and D. For example, the data size of each parameter is as follows:

[0052] A: [m, T×k], B: [T×k, n], C: [m, n], D: [m, n], α: [m, T], β: [T, n], σ: [m, r] or [r, n], γ: [m, r] or [r, n].

[0053] wherein m, n, s, k, r are positive integers. For any row in the tensor A', every k tensor elements in the row share one scaling factor parameter, and for any column in the tensor B', every k tensor elements in the column share one scaling factor parameter.

[0054] To maintain the performance advantage brought by low-precision matrix computation, the computation of the scaling factor and the element multiplication of the low-precision tensor are placed inside the tensor core. As shown in the following equation 1, the tensor core performs one round of matrix multiplication operation of A t and B t , which includes m*n point multiplication operations, each of which includes k element point multiplication operations, and then the result needs to be multiplied by the scaling factor α t × β t , which is performed T times:

[0055] (Equation 1)

[0056] t takes 0, 1, 2,..., T-1 in turn. The data size of A t is [m, k], which represents the t+1th first tensor block in T first tensor blocks obtained by dividing the matrix A in the column direction, and the data size of B t is [k, n], which represents the t+1th second tensor block in T second tensor blocks obtained by dividing the matrix B in the row direction. The data size of α t is [m, 1], which represents the t+1th first vector in T first vectors obtained by dividing the scaling factor α of the tensor A in the column direction, and the data size of β t is [1, n], which represents the t+1th second vector in T second vectors obtained by dividing the scaling factor β of the tensor B in the column direction.

[0057] Figure 1A is a schematic structural diagram of a tensor core. As shown in Figure 1B , the multiplication operation of A t × B t in equation 1 is performed by the point multiplication unit in the tensor core, and α t × β t and (α t × β t ) (A t × B t ) need additional multipliers to perform, that is, the point multiplication and accumulation module in Figure 1B . In addition, additional adders are also needed to perform, which are also performed by Figure 1BThe point multiplication and accumulation module in the tensor core is executed. Therefore, the current matrix multiplication using a scaling factor needs to introduce more additional adders and multipliers inside the tensor core, which brings an additional hardware area cost, increases the hardware cost and power consumption, and reduces the processor performance and reliability.

[0058] The tensor core provided by at least one embodiment of the present disclosure includes a first point multiplication unit and a scaling factor matrix multiplication processing module, the scaling factor matrix multiplication processing module includes a second point multiplication unit, the first point multiplication unit and the second point multiplication unit are configured to support point multiplication operations of different floating-point precisions, the tensor core is configured to receive a first tensor, a second tensor, a scaling factor of the first tensor, and a scaling factor of the second tensor, perform a matrix multiplication operation using a scaling factor by using the first point multiplication unit and the scaling factor matrix multiplication processing module, and obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating-point precision, and in the process of performing the matrix multiplication operation using a scaling factor, the first point multiplication unit is configured to perform a matrix multiplication operation of the first floating-point precision of the first tensor and the second tensor to obtain a first multiplication result, and the scaling factor matrix multiplication processing module is configured to determine the matrix multiplication operation result by using the second point multiplication unit according to the first multiplication result, the scaling factor of the first tensor, and the scaling factor of the second tensor.

[0059] In at least one embodiment, the present disclosure performs a matrix multiplication operation using a scaling factor by multiplexing point multiplication units of different precisions in a tensor core, reduces the additional hardware area cost, fully multiplexes the existing hardware resources, reduces the setting of multipliers and does not need to additionally set adders, reduces the hardware area cost, optimizes the system performance, reduces the hardware cost and power consumption, and improves the chip integration.

[0060] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.

[0061] Figure 1B The schematic structural diagram of the tensor core provided by at least one embodiment of the present disclosure is shown.

[0062] As shown in Figure 2 , the tensor core 100 includes a first point multiplication unit 101 and a scaling factor matrix multiplication processing module 102, the scaling factor matrix multiplication processing module 102 includes a second point multiplication unit 103, and the first point multiplication unit 101 and the second point multiplication unit 103 support point multiplication operations of different floating-point precisions.

[0063] It should be noted that Figure 2 the first point multiplication unit 101 and the second point multiplication unit 103 shown are schematic, and the tensor core can include a plurality of first point multiplication units 101 and a plurality of second point multiplication units 103, and the plurality of first point multiplication units 101 and the plurality of second point multiplication units 103 can be arranged in an array.

[0064] like Figure 2 As shown, the tensor kernel 100 also includes a cross-precision multiplication unit hardware path, indicated by the black arrow, between the first dot product unit 101 and the scaling factor matrix multiplication processing module 102. This cross-precision multiplication unit hardware path is configured to transmit the first multiplication result Y to the scaling factor matrix multiplication processing module 102. If the tensor kernel includes multiple first dot product units 101, each of the multiple first dot product units 101 has its own cross-precision multiplication unit hardware path, used to transmit the dot product results calculated by each first dot product unit 101 to the scaling factor matrix multiplication processing module for subsequent processing.

[0065] For example, the first dot product unit 101 supports dot product operations with a first floating-point precision, and the second dot product unit 103 supports dot product operations with a second floating-point precision. The first floating-point precision and the second floating-point precision are different.

[0066] For example, if the precision of the first floating-point number is lower than that of the second floating-point number, the second dot product unit supports dot product operations with a precision of the second floating-point number that is higher than that of the first floating-point number.

[0067] For example, the first floating-point precision is FP4, and the second floating-point precision includes FP8 and higher floating-point precisions (e.g., FP16, FP32, etc.). For example, in this embodiment, the second dot product unit 103 supports dot product operations of multiple floating-point precisions, but does not support dot product operations of the first floating-point precision. Of course, this disclosure is not limited to this, and the second dot product unit 103 may also support only one floating-point precision dot product operation.

[0068] The first dot product unit 101 and the second dot product unit 103 operate independently and do not reuse each other when performing dot product operations with first or second floating-point precision. For example, when the first dot product unit 101 performs a dot product operation with first floating-point precision, the second dot product unit 103 will not perform a dot product operation with first floating-point precision, for example, it may be in an idle state; similarly, when the second dot product unit 103 performs a dot product operation with second floating-point precision, the first dot product unit 101 will not perform a dot product operation with second floating-point precision.

[0069] For example, the minimum number of dot products supported by the second dot product unit in a single dot product operation is N, where N is a positive integer greater than 1.

[0070] For example, taking the second dot product unit supporting FP8 precision dot product operations as an example, the minimum number of dot products supported in a single dot product operation by the second dot product unit is 32. That is, when performing a dot product operation with FP8 floating-point precision, at least 32 pairs of elements are multiplied and then summed. The result of the dot product operation is this sum. Taking the second dot product unit supporting FP16 dot product operations as an example, N is 16; taking the second dot product unit supporting FP32 dot product operations as an example, N is 8.

[0071] Similar restrictions may also exist for the first dot product unit. For example, if the first dot product unit supports FP4 dot product operations, the minimum number of dot products supported by the second dot product unit in a single dot product operation is 16.

[0072] Of course, this disclosure is not limited to this. The minimum number of dot products supported by the first dot product unit and the second dot product unit in a single dot product operation can be set according to hardware requirements.

[0073] Tensor kernel 100 is configured to receive a first tensor A, a second tensor B, a scaling factor α of the first tensor, and a scaling factor β of the second tensor. It then performs a matrix multiplication operation using the scaling factor using a first dot multiplication unit 101 and a scaling factor matrix multiplication processing module 102 to obtain the matrix multiplication result of the first tensor A and the second tensor B. In this context, both the first tensor A and the second tensor B have the first floating-point precision.

[0074] During the matrix multiplication operation using a scaling factor, the first dot multiplication unit 101 is configured to perform a matrix multiplication operation with a first floating-point precision for the first tensor A and the second tensor B, to obtain a first multiplication result Y.

[0075] The first floating-point precision is low, such as FP4. The first dot multiplication unit 101 is configured to perform low-precision matrix multiplication, which has high computing power and high data transmission efficiency.

[0076] The scaling factor matrix multiplication processing module 102 is configured to determine the matrix multiplication operation result E using the second dot multiplication unit 103 based on the first multiplication result Y, the scaling factor α of the first tensor, and the scaling factor β of the second tensor.

[0077] Figure 2 This is a schematic structural diagram of a scaling factor matrix multiplication processing module 102 provided in at least one embodiment of the present disclosure.

[0078] like Figure 3 As shown, the scaling factor matrix multiplication processing module 102 further includes a scaling factor multiplication processing submodule 104. For example, the scaling factor multiplication processing submodule 104 includes a multiplier.

[0079] The scaling factor multiplication processing sub-module 104 is configured to determine a second multiplication result γ of the scaling factor α of the first tensor and the scaling factor β of the second tensor.

[0080] The second point multiplication unit 103 is configured to determine a matrix multiplication operation result E according to the second multiplication result γ and the first multiplication result Y.

[0081] In the tensor core provided in at least one embodiment of the present disclosure, reference is made to Figure 3 and Figure 1B The point multiplication and accumulation module is cancelled, thereby reducing the settings of the multiplier and the adder, saving the accumulation of the calculation result of the point multiplication unit, and saving the hardware area and hardware overhead. In addition, by adding a passageway between point multiplication units with different precisions, that is, a cross-precision point multiplication unit hardware passageway, the second point multiplication unit can be reused to perform other point multiplication and accumulation operations in addition to the low-precision matrix multiplication operation, thereby reducing the additional hardware area cost and fully reusing the existing hardware resources.

[0082] Figure 2 The schematic diagram of the first tensor, the second tensor, the scaling factor of the first tensor, and the scaling factor of the second tensor provided in at least one embodiment of the present disclosure is shown.

[0083] As shown in Figure 4 , the first tensor A is divided into T first tensor blocks in the column direction, for example, the first tensor block A0, the first tensor block A1,..., and the first tensor block AT in Figure 4 . T-1 The shape size of each first tensor block is m×k. The second tensor is divided into T second tensor blocks in the row direction, for example, the second tensor block B0, the second tensor block B1,..., and the second tensor block BT in Figure 4 . T-1 The shape size of each second tensor block is k×n.

[0084] The scaling factor α of the first tensor is divided into T first vectors in the column direction, for example, the first vector α0, the first vector α1,..., and the first vector αT in Figure 4 . T-1 The shape size of each first vector is m×1. The scaling factor β of the second tensor is divided into T second vectors in the row direction, and the shape size of each second vector is 1×n.

[0085] For example, the matrix multiplication operation result E can be represented as:

[0086] (Formula 2)

[0087] The first point multiplication unit 101 is configured to perform a matrix multiplication operation between the T first tensor blocks and the T second tensor blocks to obtain T matrix multiplication intermediate results Y0, Y1,... Y T-1 The first multiplication result includes the T matrix multiplication intermediate results, and each matrix multiplication intermediate result has a shape size of m x n.

[0088] For example, the scaling factor multiplication processing sub-module 104 is configured to perform an outer product operation between the T first vectors and the T second vectors to obtain T outer product operation results γ0, γ1,... γ T-1 The second multiplication result includes the T outer product operation results, and each outer product operation result has a shape size of m x n.

[0089] The matrix multiplication operation result E also has a shape size of m x n.

[0090] The execution process of each hardware module when performing a specific processing operation is described in detail below.

[0091] For example, each first point multiplication unit 101 is configured to perform a matrix multiplication operation between one first tensor block and a corresponding second tensor block to obtain one matrix multiplication intermediate result. A plurality of first point multiplication units perform a plurality of matrix multiplication operations between a plurality of first tensor blocks and a plurality of second tensor blocks in parallel.

[0092] For example, one first point multiplication unit 101 is configured to perform a matrix multiplication operation A0xB0 between the first tensor block A0 and the second tensor block B0 to obtain a matrix multiplication intermediate result Y0, and another first point multiplication unit 101 is configured to perform a matrix multiplication operation A1xB1 between the first tensor block A1 and the second tensor block B1 to obtain a matrix multiplication intermediate result Y1, and so on. For example, T first point multiplication units perform T matrix multiplication operations in parallel to obtain T matrix multiplication intermediate results, i.e., Y0, Y1,... Y T-1 .

[0093] For example, when the first point multiplication unit 101 performs a matrix multiplication operation of a first tensor block and a corresponding second tensor block, it includes performing m x n second operations, where each second operation includes performing a point multiplication calculation operation of k pairs of elements.

[0094] Taking the first tensor block A0 and the second tensor block B0 as an example, when the first point multiplication unit 101 performs the matrix multiplication operation of the first tensor block A0 and the second tensor block B0 to obtain the matrix multiplication intermediate result Y0 with the data size of m x n, it includes performing m x n second operations, each of which includes performing k point multiplication calculation operations of elements and determining the sum of operation results of the k point multiplication calculation operations. Specifically, the second operation includes performing the point multiplication calculation operation between the k elements of the first row of the first tensor block A0 and the k elements of the first column of the first tensor block B0, and obtaining the sum result as the element of the 1st row and the 1st column in the matrix multiplication intermediate result Y0.

[0095] For example, k represents the number of elements sharing the same scaling factor, and the larger k is, the smaller the effect of the scaling factor is, and the lower the calculation accuracy is. T is related to k and the width of the first tensor, and the larger T is, the higher the parallelism is, but the larger the calculation overhead is. The selection of T and k can be set by considering the balance between the calculation accuracy requirement and the hardware resource overhead. For example, in an embodiment, k = 16, and according to the different width of the first tensor A, T can be set to different values, such as T = 4, 8, 16, etc., of course, the present disclosure does not make specific limitations on this.

[0096] For example, the scaling factor multiplication processing sub-module 104 includes a plurality of multipliers, which are configured to perform a plurality of outer product operations in parallel. For example, one multiplier performs an outer product operation between the first vector a0 and the second vector b0 to obtain an outer product operation result g0, another multiplier performs an outer product operation between the first vector a1 and the second vector b1 to obtain an outer product operation result g1, and so on. The second multiplication result g includes the T outer product operation results.

[0097] The second point multiplication unit is configured to determine the matrix multiplication operation result E according to the T matrix multiplication intermediate results and the T outer product operation results.

[0098] For example, a plurality of second point multiplication units are configured to perform a plurality of first operations in parallel to obtain the matrix multiplication operation result through m x n first operations.

[0099] For example, expanding the above formula 2 can obtain the following formula 3:

[0100]

[0101] wherein, represents the element of the i th row and the j th column in the outer product operation result g0, represents the element of the i th row and the j th column in the matrix multiplication intermediate result Y0, represents the element of the i th row and the j th column in the outer product operation result g1, represents the element of the i th row and the j th column in the matrix multiplication intermediate result Y1, and so on.

[0102] For example, the first operation includes performing a multiplication and accumulation operation on the element in the i-th row and j-th column of the intermediate results of T matrix multiplications and the results of T outer product operations to obtain the element in the i-th row and j-th column of the matrix multiplication operation result, where i is a positive integer less than or equal to m and j is a positive integer less than or equal to n.

[0103] For example, the second dot product unit is configured to execute a first operation at a time. When executing the first operation, the second dot product unit includes the following operations: determining the element in the i-th row and j-th column of each intermediate matrix multiplication result in the T intermediate matrix multiplication results to obtain T first elements; determining the element in the i-th row and j-th column of each outer product operation result in the T outer product operation results to obtain T second elements; and performing a dot product operation between the T elements and the T second elements to obtain the element in the i-th row and j-th column of the matrix multiplication operation result.

[0104] Specifically, referring to Formula 3, the first operation can be expressed as Formula 4 as follows:

[0105]

[0106] For example, This represents the element in the i-th row and j-th column of the result of the matrix multiplication operation. The other parameters are defined as described above and will not be repeated here.

[0107] When the second dot multiplication unit 103 performs the first operation, it determines the element in the i-th row and j-th column of each intermediate matrix multiplication result in the T intermediate matrix multiplication results, and obtains the T first elements, which are represented as follows: ; Determine the i-th row and j-th column element of each of the T outer product results to obtain T second elements, denoted as follows: Perform dot product operations on T elements and T second elements to obtain the element in the i-th row and j-th column of the matrix multiplication result. Specifically, the dot product operation here, as shown in Formula 4, includes performing dot product calculations on each first element and its corresponding second element, and summing the results of the dot product calculations as the element in the i-th row and j-th column of the matrix multiplication result. .

[0108] For example, multiple second-point multiplication units can perform the first operation described above in parallel. Each first operation can obtain one element in the matrix multiplication result. By performing m×n first operations, all elements in the matrix multiplication result E can be obtained, thus obtaining the matrix multiplication result.

[0109] In the above embodiment, the disclosure implements the accumulation and element multiplication operation in the matrix multiplication operation using the scaling factor by the second point multiplication unit, to reuse the second point multiplication unit in the idle state when performing the matrix multiplication operation of the first floating point precision, make full use of the existing hardware, reduce the setting of the multiplier, avoid the additional adder setting, and save the hardware area. Moreover, since the floating point precision of the point multiplication operation supported by the second point multiplication unit is higher than the first floating point precision, the execution of the multiplication and addition operation using the second point multiplication unit can preserve the effective data to a greater extent, and the calculation precision is not affected.

[0110] As described previously, the minimum number of point multiplications N of the single point multiplication operation supported by the second point multiplication unit is limited, and in at least one embodiment of the disclosure, when selecting the second point multiplication unit, the point multiplication unit supporting the minimum number of point multiplications N of the single point multiplication operation greater than or equal to T is selected as the second point multiplication unit, N is a positive integer, thereby maximizing the balance between hardware overhead and operator execution speed and efficiency.

[0111] For example, when the second point multiplication unit performs the point multiplication operation of T first elements and T second elements, it includes performing N point multiplication calculation operations and determining the sum of the operation results of the N point multiplication calculation operations. That is, due to the limitation of the minimum number of point multiplications, the minimum number of point multiplications of the single point multiplication operation of the second point multiplication unit is at least N, that is, at least N point multiplication calculation operations need to be performed.

[0112] Figure 4 The schematic diagram of the first operation provided by at least one embodiment of the disclosure.

[0113] As Figure 5 shown, in response to N greater than T, when performing N point multiplication calculation operations, the point multiplication calculation operations of N first elements and N second elements are performed. Among them, the N first elements include T first elements, that is, , and in addition, the N first elements also include N-T 0s; the N second elements include T second elements, that is, , and in addition, the N second elements also include N-T 0s. When the second point multiplication unit performs the point multiplication operation, each point multiplication calculation operation performs the point multiplication of two elements in the same column, and then adds the operation results of the N point multiplication calculation operations.

[0114] Therefore, when the above-mentioned first operation is performed using the second point multiplication unit, the second point multiplication unit is still made to perform N point multiplication calculation operations in the form of supplementing 0s, but the actual calculation result is the operation result of the point multiplication operation of T first elements and T second elements, thereby reusing the second point multiplication unit.

[0115] For example, in one embodiment, the first floating point precision is FP4, the first dot product unit supports dot product operation of FP4, and the second dot product unit supports dot product operation of FP8 and higher precision than FP8. For example, assume k = 16, T is set according to the width of the first tensor, for example, T = 4, 8, 16, etc.

[0116] The first dot product unit performs matrix multiplication operation of the first tensor and the second tensor in FP4 floating point precision, to obtain a first multiplication result, for example, T first dot product units perform matrix multiplication operation between T first tensor blocks and T second tensor blocks in parallel to obtain T matrix multiplication intermediate results, and the first multiplication result includes the T matrix multiplication intermediate results. The specific process is described above and will not be repeated here.

[0117] The scaling factor multiplication processing sub-module performs outer product operation between T first vectors and T second vectors, for example, T multipliers are used to perform outer product operation between T first vectors and T second vectors in parallel to obtain T outer product operation results.

[0118] The plurality of second dot product units perform a plurality of first operations in parallel, and obtain a matrix multiplication operation result through m x n first operations.

[0119] For example, after obtaining the matrix multiplication operation result E described above, if C is not equal to 0, the tensor core is further configured to add the matrix multiplication operation result E and σ ⊙ C to obtain a final matrix multiplication operation result.

[0120] For example, in one embodiment, when the second dot product unit performs dot product operation of FP8 precision, the minimum number of dot products supported by a single dot product operation is 32, when the second dot product unit performs dot product operation of FP16 precision, the minimum number of dot products supported by a single dot product operation is 16, and when the second dot product unit performs dot product operation of FP32 precision, the minimum number of dot products supported by a single dot product operation is 8.

[0121] Assume T = 4, the first operation can be performed using dot product operation of FP8, FP16 or FP32 precision. For example, when FP8 is used to perform the first operation, positions other than the positions of 4 first elements and 4 second elements are filled with 0 in the dot product operation, that is, 2 x (32-4) spare positions are filled with 0. For example, when FP16 is used to perform the first operation, positions other than the positions of 4 first elements and 4 second elements are filled with 0 in the dot product operation, and 2 x (16-4) spare positions are filled with 0. For example, when FP32 is used to perform the first operation, positions other than the positions of 4 first elements and 4 second elements are filled with 0 in the dot product operation, and 2 x (8-4) spare positions are filled with 0.

[0122] In the above embodiment, the second point multiplication unit performs the first operation in FP16 in consideration of the accuracy requirement and the calculation overhead.

[0123] If T = 16, the first operation can be performed using the point multiplication operation in FP8 and FP16 accuracy. If T = 32, the first operation can be performed using the point multiplication operation in FP8 accuracy.

[0124] In the above embodiment, the point multiplication operation of the first multiplication result and the second multiplication result and the addition operation of the operation results of T rounds of the point multiplication operation are performed by multiplexing the second point multiplication unit originally in an idle state in the tensor core. Since the second point multiplication unit is independent of the first point multiplication unit and is originally in an idle state when the first point multiplication unit performs the matrix multiplication operation in the first floating point accuracy, the hardware resource utilization rate can be improved by multiplexing the second point multiplication unit, the existing hardware resources can be fully utilized, the settings of the multipliers for performing the point multiplication operation of the first multiplication result and the second multiplication result are reduced, the settings of the adders for performing the addition operation of the operation results of T rounds of the point multiplication operation are reduced, and the hardware area cost of the additional multipliers and adders required for implementing the low-precision matrix multiplication using the scaling factor is reduced.

[0125] Figure 5 The schematic structural diagram of the processor provided by at least one embodiment of the present disclosure is shown in FIG. 2. As shown in FIG. 2, the processor 200 includes a tensor core 100. Figure 6

[0126] For example, the tensor core includes a first point multiplication unit and a scaling factor matrix multiplication processing module, the scaling factor matrix multiplication processing module includes a second point multiplication unit, and the first point multiplication unit and the second point multiplication unit are configured to support point multiplication operations in different floating point accuracies.

[0127] The tensor core is configured to receive a first tensor, a second tensor, a scaling factor of the first tensor, and a scaling factor of the second tensor, perform a matrix multiplication operation using a scaling factor using the point multiplication unit and the scaling factor matrix multiplication processing module to obtain a matrix multiplication operation result of the first tensor and the second tensor, and the first tensor and the second tensor are both in a first floating point accuracy.

[0128] In the process of performing the matrix multiplication operation using a scaling factor, the first point multiplication unit is configured to perform a matrix multiplication operation in the first floating point accuracy of the first tensor and the second tensor to obtain a first multiplication result, and the scaling factor matrix multiplication processing module is configured to determine a matrix multiplication operation result using the second point multiplication unit according to the first multiplication result, the scaling factor of the first tensor, and the scaling factor of the second tensor.

[0129] ​The specific structure and functions of the tensor core 100 can refer to the related descriptions of the tensor core 100 in the foregoing embodiments, and will not be repeated here.

[0130] For example, the processor can include a graphics processor, and the processor includes a plurality of stream processor clusters and a memory, each stream processor cluster includes a plurality of computing units, and each computing unit includes a tensor core, with reference to the structure of the graphics processor shown in Figure 6

[0131] Of course, the processor can also be other processors provided with tensor cores, such as neural network processors, digital signal processors, central processing units, etc., and the present disclosure does not make specific limitations thereto.

[0132] In at least one embodiment, the processor of the present disclosure performs a matrix multiplication operation using a scaling factor by multiplexing point multiplication units of different precisions in the tensor core, reduces the additional hardware area cost, fully multiplexes the existing hardware resources, reduces the setting of the multiplier and does not need to additionally set the adder, reduces the hardware area cost, optimizes the system performance, reduces the hardware cost and power consumption, and improves the chip integration.

[0133] In addition, it should be noted that Figure 1A The components of the processor 200 shown are exemplary only and are not limiting, and the processor 200 can also have other components according to actual application needs.

[0134] The present disclosure at least one embodiment also provides a data processing method. Figure 6 The data processing method provided by at least one embodiment of the present disclosure is shown in the schematic flowchart.

[0135] The data processing method is applied to a tensor core. The tensor core includes a first point multiplication unit and a scaling factor matrix multiplication processing module, the scaling factor matrix multiplication processing module includes a second point multiplication unit, and the first point multiplication unit and the second point multiplication unit support point multiplication operations of different floating point precisions. For more information about the tensor core, please refer to the related descriptions of the tensor core 100 in the foregoing embodiments, which will not be repeated here.

[0136] As shown in Figure 7 The data processing method provided by at least one embodiment of the present disclosure includes steps S10-S20.

[0137] In S10, a first tensor, a second tensor, a scaling factor of the first tensor, and a scaling factor of the second tensor are received.

[0138] In S20, a matrix multiplication operation using a scaling factor is performed by using the first point multiplication unit and the scaling factor matrix multiplication processing module to obtain a matrix multiplication operation result of the first tensor and the second tensor.

[0139] ​The first tensor and the second tensor are both of a first floating-point precision.

[0140] For example, the S20 can include: performing, by the first dot product unit, a matrix multiplication operation of the first tensor and the second tensor at the first floating-point precision to obtain a first multiplication result; and determining, by the second dot product unit, the matrix multiplication operation result according to the first multiplication result, a scaling factor of the first tensor, and a scaling factor of the second tensor.

[0141] For example, the scaling factor matrix multiplication processing module further includes a scaling factor multiplication processing submodule, and the scaling factor multiplication processing submodule includes a plurality of multipliers.

[0142] Determining, by the second dot product unit, the matrix multiplication operation result according to the first multiplication result, the scaling factor of the first tensor, and the scaling factor of the second tensor can include: determining, by the scaling factor multiplication processing submodule, a second multiplication result of the scaling factor of the first tensor and the scaling factor of the second tensor; and determining, by the second dot product unit, the matrix multiplication operation result according to the second multiplication result and the first multiplication result.

[0143] For example, the first tensor is divided into T first tensor blocks in a column direction, each first tensor block has a shape size of m×k, the second tensor is divided into T second tensor blocks in a row direction, each second tensor block has a shape size of k×n, the scaling factor of the first tensor is divided into T first vectors in the column direction, each first vector has a shape size of m×1, and the scaling factor of the second tensor is divided into T second vectors in the row direction, each second vector has a shape size of 1×n. More descriptions about the first tensor block, the second tensor block, the first vector, and the second vector can be referred to the foregoing content, and will not be described here.

[0144] For example, in some embodiments, performing, by the first dot product unit, the matrix multiplication operation of the first tensor and the second tensor at the first floating-point precision to obtain the first multiplication result can include: performing, by the first dot product unit, a matrix multiplication operation between the T first tensor blocks and the T second tensor blocks to obtain T matrix multiplication intermediate results, wherein the first multiplication result includes the T matrix multiplication intermediate results.

[0145] For example, in some embodiments, determining, by the scaling factor multiplication processing submodule, the second multiplication result of the scaling factor of the first tensor and the scaling factor of the second tensor can include: performing, by the scaling factor multiplication processing submodule, an outer product operation between the T first vectors and the T second vectors to obtain T outer product operation results, wherein the second multiplication result includes the T outer product operation results.

[0146] For example, in some embodiments, determining the matrix multiplication operation result by the second dot product unit according to the second multiplication result and the first multiplication result can comprise: determining the matrix multiplication operation result by the second dot product unit according to the T matrix multiplication intermediate results and the T outer product operation results.

[0147] The data processing method provided by at least one embodiment of the present disclosure performs the matrix multiplication operation using the scaling factor by multiplexing the dot product units with different precisions in the tensor core, reduces the additional hardware area cost, fully multiplexes the existing hardware resources, reduces the setting of the multiplier and does not need to additionally set the adder, reduces the hardware area cost, optimizes the system performance, reduces the hardware cost and power consumption, and improves the chip integration.

[0148] It should be noted that the first tensor and the second tensor can have different physical meanings according to different application fields.

[0149] For example, in the field of speech processing, the first tensor and the second tensor can be any parameter used, input, or generated in tasks such as feature extraction, speech enhancement, speech recognition, etc., which needs to perform a matrix multiplication operation, such as a speech feature vector, a filtering parameter, etc.

[0150] For example, in the field of image processing, the first tensor and the second tensor can be any parameter used, input, or generated in tasks such as image preprocessing, feature extraction, image segmentation, target detection, etc., which needs to perform a matrix multiplication operation, such as an image feature vector, various edge detection operators (such as Sobel operator, Canny operator, Prewitt operator, etc.), image filtering operators (such as Gaussian filtering, median filtering, bilateral filtering, etc.), morphological operators (such as erosion, dilation, opening operation, closing operation, etc.), etc.

[0151] For example, in the field of text processing, the first tensor and the second tensor can be any parameter used, input, or generated in tasks such as text classification, sentiment analysis, text generation, etc., which needs to perform a matrix multiplication operation, such as a semantic feature vector of the text, etc.

[0152] For example, in the field of video processing, the first tensor and the second tensor can be the parameters in the field of image processing as described above, or parameters used, input, or generated in the field of video processing, such as an optical flow operator (used to estimate the motion between video frames), a target tracking operator (used to track a specific target in a video), etc.

[0153] Of course, the present disclosure is not limited thereto, and for other application scenarios or fields, such as the technical fields of graphics, machine learning, data mining, signal processing, natural language processing, geographic information systems, oceanography, environmental science, etc., as long as a matrix multiplication operation is needed, the data processing method described in at least one embodiment of the present disclosure can be applied, and details are not repeated here.

[0154] Figure 7 A schematic block diagram of an electronic device is provided for an embodiment of the present disclosure.

[0155] As shown in Figure 8 , the electronic device 300 is suitable for implementing the data processing method provided by the embodiments of the present disclosure, for example. It should be noted that Figure 8 The components of the electronic device 300 shown are only exemplary and are not limiting, and the electronic device 300 can also have other components according to actual application needs.

[0156] As shown in Figure 8 , the electronic device 300 can include a processing device 301, which includes the aforementioned processor 200, for example, which can perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in the memory to achieve various functions. For example, the processing device 301 can also include a central processing unit (CPU), a tensor processor (TPU), and the like, which have instruction optimization capabilities and / or program execution capabilities. The central processing unit (CPU) can be X86, ARM, RISC-V architecture, etc. The GPU can be directly integrated into the SOC, directly integrated into the mainboard, or built into the north bridge chip of the mainboard.

[0157] For example, when the computer-readable instructions are executed by the processing device 301, one or more steps of the data processing method according to any of the above embodiments can be performed. It should be noted that the detailed description of the processing process of the data processing method can refer to the related description in the above embodiments of the data processing method.

[0158] As shown in Figure 8 , for example, the memory can include any combination of one or more computer program products, which can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. For example, the volatile memory can include random access memory (RAM) 303 and / or cache memory, etc., for example, the computer-readable instructions can be loaded from the storage device 308 into the random access memory (RAM) 303 to run the computer-readable instructions. The non-volatile memory can include read-only memory (ROM) 302, hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. Various application programs and various data can also be stored in the computer-readable storage medium, such as various data used and / or generated by the application programs, etc.

[0159] For example, the processing device 301, a read only memory (ROM) 302, and a random access memory (RAM) 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0160] Generally, the following devices can be connected to the input / output (I / O) interface 305: an input device 306 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage device 308 including, for example, a magnetic tape, a hard disk, a flash memory, and the like; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate wirelessly or wiredly with other electronic devices to exchange data. Although Figure 8 The electronic device 300 having various devices is illustrated, but it is understood that all of the illustrated devices are not required to be implemented or possessed, and the electronic device 300 can instead implement or possess more or fewer devices. For example, the processing device 301 can control other components in the electronic device 300 to perform desired functions.

[0161] For example, the electronic device can be in the form of a server for various application scenarios such as deep learning and artificial intelligence, scientific computing, graphics rendering and video editing, virtual reality and game development, cloud services, and the like.

[0162] The flowcharts and block diagrams in the attached drawings illustrate the possible implementation architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in different orders than those noted in the attached drawings. For example, two blocks that are shown in succession can actually be executed in parallel, and they can also be executed in reverse order, depending on the involved functions. It should also be noted that each block in the block diagrams and / or flowcharts, and the combination of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0163] The units described in the embodiments of the present disclosure can be implemented by software or by hardware. In some cases, the names of the units do not constitute a limitation on the units themselves.

[0164] The functionality described herein above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0165] Figure 8 A schematic diagram of a non-transitory computer-readable storage medium is provided for at least one embodiment of the present disclosure. For example, as shown in FIG. 4, a storage medium 400 can be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 401 can be non-transitorily stored in the storage medium 400. For example, when the computer-readable instructions 401 are executed by a processor, one or more steps of a data processing method according to the above description can be performed. Figure 9 Figure 9

[0166] For example, the storage medium 400 can be applied in the electronic device 300. For example, the storage medium 400 can include the storage apparatus 308 in the electronic device 300.

[0167] For example, the storage apparatus can include any combination of one or more computer program products. The computer program product can include various forms of computer-readable storage media for storing information that can be used to fully or partially train, program, or otherwise configure a processor. For example, the computer-readable storage media can include volatile memory or non-volatile memory. For example, volatile memory can include random access memory (RAM), cache memory, and / or the like, and non-volatile memory can include read only memory (ROM), hard disk, erasable programmable read only memory (EPROM), compact disc read only memory (CD-ROM), USB memory, flash memory, and / or the like. One or more computer-readable instructions can be stored on the computer-readable storage medium. The processor can execute the computer-readable instructions to implement various functions of the processor. Various application programs and various data, etc. can also be stored in the storage medium.

[0168] For example, the storage medium can include a memory card of a smart phone, a cache component of a tablet computer, a hard disk of a personal computer, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM), compact disc read only memory (CD-ROM), flash memory, or any combination of the above storage medium, and can also be other applicable storage medium.

[0169] ​The above description merely illustrates the preferred embodiments of the present disclosure and the principles of the technology applied. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0170] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.

[0171] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely illustrative of example forms of implementing the claims.

[0172] For the present disclosure, the following points need to be explained:

[0173] (1) The drawings of the embodiments of the present disclosure only involve the structures involved in the embodiments of the present disclosure, and other structures can refer to the general design.

[0174] (2) In the case of no conflict, the embodiments of the present disclosure and the features in the embodiments can be combined to obtain new embodiments.

[0175] The above description is merely a specific implementation of the present disclosure, but the protection scope of the present disclosure is not limited thereto, and the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A tensor core, comprising: The first point multiplication unit, the scaling factor matrix multiplication processing module and the cross-precision point multiplication unit hardware path, the scaling factor matrix multiplication processing module includes a second point multiplication unit, the first point multiplication unit and the second point multiplication unit are configured to support different floating point precision point multiplication operation, The tensor core is configured to receive a first tensor, a second tensor, a scaling factor of the first tensor and a scaling factor of the second tensor, perform a scaling factor matrix multiplication operation using the first point multiplication unit and the scaling factor matrix multiplication processing module, and obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating point precision, In the process of performing the scaling factor matrix multiplication operation, the first point multiplication unit is configured to perform a matrix multiplication operation of the first floating point precision of the first tensor and the second tensor to obtain a first multiplication result, The cross-precision point multiplication unit hardware path is configured to transmit the first multiplication result to the scaling factor matrix multiplication processing module, The scaling factor matrix multiplication processing module is configured to determine the matrix multiplication operation result using the second point multiplication unit according to the first multiplication result, the scaling factor of the first tensor and the scaling factor of the second tensor.

2. The tensor core of claim 1, wherein, The scaling factor matrix multiplication processing module further includes a scaling factor multiplication processing sub-module, The scaling factor multiplication processing sub-module is configured to determine a second multiplication result of the scaling factor of the first tensor and the scaling factor of the second tensor; The second point multiplication unit is configured to determine the matrix multiplication operation result according to the second multiplication result and the first multiplication result.

3. The tensor core of claim 1, wherein, The second point multiplication unit supports a second floating point precision higher than the first floating point precision.

4. The tensor core of claim 2, wherein, The first tensor is divided into T first tensor blocks in the column direction, each first tensor block has a shape size of m×k, the second tensor is divided into T second tensor blocks in the row direction, each second tensor block has a shape size of k×n, The scaling factor of the first tensor is divided into T first vectors in the column direction, each first vector has a shape size of m×1, and the scaling factor of the second tensor is divided into T second vectors in the row direction, each second vector has a shape size of 1×n, The first point multiplication unit is configured to perform a matrix multiplication operation between the T first tensor blocks and the T second tensor blocks to obtain T matrix multiplication intermediate results, wherein the first multiplication result includes the T matrix multiplication intermediate results; The scaling factor multiplication processing sub-module is configured to perform an outer product operation between the T first vectors and the T second vectors to obtain T outer product operation results, wherein the second multiplication result includes the T outer product operation results; The second point multiplication unit is configured to determine the matrix multiplication operation result according to the T matrix multiplication intermediate results and the T outer product operation results, Wherein, m, n, k, T are positive integers.

5. The tensor core of claim 4, wherein, The shape size of the matrix multiplication operation result, each matrix multiplication intermediate result and each outer product operation result is m×n. The second point multiplication units are configured to perform a plurality of first operations in parallel to obtain the matrix multiplication operation result through m*n first operations, The first operation includes performing multiplication and accumulation operation of the i-th row and j-th column element in the T matrix multiplication intermediate results and the T outer product operation results to obtain the i-th row and j-th column element in the matrix multiplication operation result, Wherein, i is a positive integer less than or equal to m, and j is a positive integer less than or equal to n.

6. The tensor core of claim 5, wherein, The second point multiplication unit is configured to perform one first operation at a time, When the second point multiplication unit performs the first operation, the following operations are included: Determine the i-th row and j-th column element of each matrix multiplication intermediate result in the T matrix multiplication intermediate results to obtain T first elements; Determine the i-th row and j-th column element of each outer product operation result in the T outer product operation results to obtain T second elements; Perform point multiplication operation between the T first elements and the T second elements to obtain the i-th row and j-th column element in the matrix multiplication operation result.

7. The tensor core of claim 6, wherein, The minimum point multiplication number N of a single point multiplication operation supported by the second point multiplication unit is greater than or equal to T, and N is a positive integer, When the second point multiplication unit performs the point multiplication operation of the T first elements and the T second elements, N point multiplication calculation operations are performed and the sum of the operation results of the N point multiplication calculation operations is determined, In response to N being greater than T, when the N point multiplication calculation operations are performed, N first elements and N second elements are subjected to point multiplication calculation operations, and the N first elements include the T first elements and N-T zeros, and the N second elements include the T second elements and N-T zeros.

8. Tensor core according to any of claims 1-7, characterized in that, The first floating point precision is FP4, and the second point multiplication unit supports point multiplication operations with FP8 and higher than FP8 precision.

9. The tensor core of claim 8, wherein, The second point multiplication unit is configured to perform the process of determining the matrix multiplication operation result with FP16 precision.

10. The tensor core of claim 4, wherein, Each first point multiplication unit is configured to perform a matrix multiplication operation between a first tensor block and a corresponding second tensor block to obtain a matrix multiplication intermediate result, A plurality of first point multiplication units perform a plurality of matrix multiplication operations in parallel, When the first point multiplication unit performs the matrix multiplication operation of the first tensor block and the corresponding second tensor block, m*n second operations are performed, and each second operation includes performing point multiplication calculation operation of k pairs of elements.

11. A processor, comprising: The tensor core includes any one of claims 1-10.

12. A data processing method applied to tensor kernels, characterized in that, The tensor core includes a first point multiplication unit, a scaling factor matrix multiplication processing module, and a cross-precision point multiplication unit hardware path, the scaling factor matrix multiplication processing module includes a second point multiplication unit, and the first point multiplication unit and the second point multiplication unit are configured to support point multiplication operations with different floating point precisions. The data processing method includes: Receiving a first tensor, a second tensor, a scaling factor of the first tensor, and a scaling factor of the second tensor; The first point multiplication unit and the scaling factor matrix multiplication processing module are used to perform a matrix multiplication operation using a scaling factor to obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating point precision. The first point multiplication unit and the scaling factor matrix multiplication processing module are used to perform a matrix multiplication operation using a scaling factor to obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating point precision. The first point multiplication unit and the scaling factor matrix multiplication processing module are used to perform a matrix multiplication operation using a scaling factor to obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating point precision. The first point multiplication unit and the scaling factor matrix multiplication processing module are used to perform a matrix multiplication operation using a scaling factor to obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating point precision. The first point multiplication unit and the scaling factor matrix multiplication processing module are used to perform a matrix multiplication operation using a scaling factor to obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating point precision.

13. The data processing method according to claim 12, characterized in that, The first point multiplication unit and the scaling factor matrix multiplication processing module are used to perform a matrix multiplication operation using a scaling factor to obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating point precision. The first point multiplication unit and the scaling factor matrix multiplication processing module are used to perform a matrix multiplication operation using a scaling factor to obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating point precision. The first point multiplication unit and the scaling factor matrix multiplication processing module are used to perform a matrix multiplication operation using a scaling factor to obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating point precision. The first point multiplication unit and the scaling factor matrix multiplication processing module are used to perform a matrix multiplication operation using a scaling factor to obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating point precision.

14. An electronic device, comprising: The first point multiplication unit and the scaling factor matrix multiplication processing module are used to perform a matrix multiplication operation using a scaling factor to obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating point precision. The first point multiplication unit and the scaling factor matrix multiplication processing module are used to perform a matrix multiplication operation using a scaling factor to obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating point precision. The first point multiplication unit and the scaling factor matrix multiplication processing module are used to perform a matrix multiplication operation using a scaling factor to obtain a matrix multiplication operation result of the first tensor and the second tensor, wherein the first tensor and the second tensor are both of a first floating point precision. ​ 15. A non-transitory computer-readable storage medium, comprising: ​ ​

Citation Information

Patent Citations

  • Tensor calculation device, data processor, tensor calculation method, and storage medium

    CN114741650A

  • Generalized acceleration of matrix multiply accumulate operations

    US20220391206A1