Methods and apparatuses for neural network optimized matrix-matrix multiplication (NNMM)
By retraining the weight coefficients of the neural network to have a uniform pattern, the unstructured sparseness problem caused by weight pruning is solved, and the prediction performance of the neural network is maintained or improved while reducing computational costs and storage needs.
Patent Information
- Application Number
- CN202110180290.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-09
- Filing Date
- 2021-02-08
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-02-08
AI Technical Summary
In the prior art, weight pruning of deep neural networks results in unstructured sparsity causing random memory access, inability to improve inference calculations, and removing larger percentages of weights will lead to degradation in prediction performance.
By retraining the weight coefficients of the neural network to have a predetermined uniform mode, the characteristics of the uniform mode are used to skip unnecessary multiplication and addition operations in the matrix multiplication operation, reducing the amount of operations, while maintaining the performance of the neural network.
It realizes that while reducing multiplication and addition operations, maintain or improve the prediction performance of neural networks, reducing computational costs and storage requirements.
Smart Images

Figure CN113282879B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Provisional Application No. 62 / 979,034, filed on February 20, 2020, and U.S. Application No. 17 / 092,925, filed on November 9, 2020, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to neural network technologies, and more particularly, to methods and apparatuses for neural network optimized matrix - matrix multiplication (NNMM), and computer - readable media. Background Art
[0004] Deep neural networks (DNNs) have achieved great success in solving various tasks such as computer vision and natural language processing. The large model capacity of deep network structures with a large number of parameters brings high prediction performance, but it also makes DNN models too expensive for practical applications, especially for mobile and on - device applications with strict limitations in storage, computing power, and energy consumption. How to reduce the cost of using DNNs in academia and industry has attracted great attention. The international standardization organization MPEG has also established a special group to address this issue. Summary of the Invention
[0005] Embodiments of this application provide methods and apparatuses for neural network optimized matrix - matrix multiplication (NNMM), and computer - readable media, aiming to solve the problems in the prior art that the unstructured sparsity in the weight matrix after weight pruning of DNNs causes random memory access, thus unable to improve inference calculation, and that removing a large percentage of weights will lead to a large decline in prediction performance.
[0006] According to an embodiment, a method for neural network optimized matrix - matrix multiplication (NNMM) includes: determining a first matrix of input coefficients; determining a second matrix of weight coefficients of a trained neural network such that the second matrix has a predetermined uniform pattern, the predetermined uniform pattern including at least two of the weight coefficients having the same value. The method further includes: performing multiplication on the determined first matrix and the determined second matrix to determine output coefficients.
[0007] An apparatus for neural network optimized matrix - matrix multiplication NNMM includes: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code to perform the method for neural network optimized matrix - matrix multiplication (NNMM) described in the embodiment.
[0008] According to an embodiment, an apparatus for neural network optimized matrix-matrix multiplication (NNMM) includes: a first determination module configured to determine a first matrix of input coefficients; a second determination module configured to determine a second matrix of weight coefficients of a trained neural network such that the second matrix has a predetermined uniform pattern, the predetermined uniform pattern including at least two of the weight coefficients having the same value; and an execution module configured to perform multiplication on the determined first matrix and the determined second matrix to determine output coefficients.
[0009] A non-transitory computer-readable medium for storing instructions that, when executed by at least one processor for neural network optimized matrix-matrix multiplication NNMM, cause the at least one processor to perform the method for neural network optimized matrix-matrix multiplication (NNMM) as described in the embodiment.
[0010] The method for NNMM provided by the embodiments of the present application can achieve a smaller number of multiplication and addition operations by determining a second matrix of weight coefficients of a neural network and making the second matrix have a predetermined uniform pattern, thereby improving inference calculation and maintaining the performance of the neural network. Description of the Drawings
[0011] Figure 1A is a schematic diagram of a general panel-matrix multiplication (GEPM, general panel-matrix multiplication) / general block-panel multiplication (GEBP, general block-panel multiplication) partitioning method and a general panel-panel multiplication (GEPP, general panel-panel multiplication) / GEBP partitioning method.
[0012] Figure 1B is a schematic diagram of another GEPM / GEBP partitioning method and another GEPP / GEBP partitioning method.
[0013] Figure 2 is a schematic diagram of an environment in which the methods, apparatuses, and systems described herein can be implemented according to an embodiment.
[0014] Figure 3 is Figure 2 a block diagram of example components of one or more devices of.
[0015] Figure 4 is a schematic diagram of a 2D tensor layout and a 3D tensor layout of a convolutional layer according to an embodiment.
[0016] Figure 5Schematic diagram of a uniform pattern of 2D sensor layout and 3D tensor layout using convolutional layers according to an embodiment.
[0017] Figure 6 Schematic diagram of a three-dimensional general matrix-matrix multiplication (GEMM3D) partition based on a uniform pattern.
[0018] Figure 7 Flowchart of a method for NNMM according to an embodiment.
[0019] Figure 8 Block diagram of an apparatus for NNMM according to an embodiment. Detailed implementation
[0020] The present disclosure relates to neural network model acceleration. More specifically, it relates to a method for neural network model acceleration based on a general matrix-matrix multiplication (GEMM) operation with a uniform pattern.
[0021] Inference operations in deep learning systems extensively use matrix multiplication, so high-performance GEMM is crucial for inference operations. Depending on the sizes of the left-hand-side (lhs) matrix and the right-hand-side (rhs) matrix, two GEMM routines (GEPP / GEBP, GEPM / GEBP) have been recognized as the best GEMM solutions in the industry over the past decade. Both methods recursively partition the lhs matrix and the rhs matrix to fully utilize the different characteristics of off-chip memory (such as DDR) and on-chip memory (such as multi-level caches) in modern computing platforms. The lhs matrix and the rhs matrix are matrices of weight coefficients of a neural network, which are referred to as the second matrix in an embodiment to distinguish them from the first matrix of input coefficients of the neural network.
[0022] As Figure 1A shown, the GEPM method partitions the lhs matrix into multiple slices, and each slice is further divided into multiple blocks with dimensions [mc, kc]. The processing scan order of the GEPM method is a raster scan order in the horizontal direction, where each slice is processed from top to bottom, and each block within a slice is read from the main memory from left to right. To generate the complete result of the matrix multiplication, the rhs matrix needs to be read from the main memory the number of times equal to the number of slices.
[0023] The GEPP method divides the lhs matrix into multiple tiles, and each tile is further divided into multiple blocks with dimensions [mc, kc]. The processing scan order of the GEPP method is the raster scan order in the vertical direction, where each tile is processed from left to right, and each block within a tile is read from the main memory from top to bottom. To generate the complete result of matrix multiplication, the result matrix needs to be read from the main memory tile number - 1 times, and the result matrix also needs to be written to the main memory tile number of times.
[0024] As Figure 1B shown, after reading the blocks of the lhs matrix [mc, kc] and the slices of the rhs matrix [kc, SIZE] from the main memory into the computer processing unit (CPU) cache, additional partitioning is performed.
[0025] The [mc, kc] lhs block is further divided into multiple [mr, kc] blocks, and then further divided into multiple [mr, 1] blocks; the [kc, SIZE] rhs block is further divided into multiple [kc, nr] blocks, and then further divided into multiple [1, nr] blocks. Finally, the [mr, 1] lhs block and the [1, nr] rhs block are loaded into the CPU registers, and the [mr, nr] result block is calculated using the CPU media access control (MAC) unit and stored in the register before unloading the [mr, nr] result block to the CPU cache.
[0026] This process is repeated until all coefficients in the lhs block and the rhs block are processed, and the [mc, SIZE] result block is generated and written to the main memory. After that, the next lhs block and rhs block are loaded into the CPU cache for processing.
[0027] In the past few years, active research has been carried out to compress large DNN models. The overall goal is to reduce the size of the model (i.e., the required storage space) and speed up the inference speed without sacrificing too much the performance of the original task (e.g., classification accuracy). Effective solutions usually require multidisciplinary knowledge in machine learning, computer architecture, hardware design, etc., and great progress has been made by using different techniques including weight pruning, weight quantization, low-rank factorization, and knowledge distillation.
[0028] Among all the efforts, weight pruning and weight quantization are the most popular directions. Weight pruning aims to remove unimportant weight coefficients and reduce redundancy in network connections.
[0029] According to the partitioning process of GEMM operations, the main / cache memory access and matrix multiplication operations can be skipped only when the coefficient values of the entire block or sub-block are zero. Although a high compression rate can be achieved with little prediction loss, the unstructured sparsity in the pruned weight matrix causes random memory access, so the unstructured weight pruning method cannot improve the inference calculation in most cases (and sometimes even makes the problem worse).
[0030] In addition, removing a large percentage of weights usually leads to a significant drop in prediction performance.
[0031] The weight pruning method aims to change more coefficients to zero values so that the blocks / sub-blocks of matrix multiplication can be skipped during GEMM operations.
[0032] A GEMM operation method based on a uniform pattern is proposed to achieve a similar multiplication skip result. The neural network retraining process is performed to generate various predefined uniform patterns. The pruning method is regarded as a special case of the uniform pattern, where the coefficient values of the uniform pattern are zero, in which case the multiplication can be completely skipped for this pattern. When the coefficient values of the uniform pattern are not zero, the multiplication cannot be completely skipped for this pattern, but the multiplication results can be shared within the block, thereby reducing the number of multiplication operations.
[0033] Figure 2 is a schematic diagram of an environment 200 in which the methods, apparatuses, and systems described herein can be implemented according to an embodiment. As Figure 2 shown, the environment 200 can include a user device 210, a platform 220, and a network 230. The devices in the environment 200 can be interconnected by a wired connection, a wireless connection, or a combination of wired and wireless connections.
[0034] The user device 210 includes one or more devices that are capable of receiving, generating, storing, processing, and / or providing information related to the platform 220. For example, the user device 210 can include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smart phone, a wireless phone, etc.), a wearable device (e.g., smart glasses or a smart watch), or a similar device. In some embodiments, the user device 210 can receive information from the platform 220 and / or send information to the platform 220.
[0035] The platform 220 includes one or more devices as described elsewhere herein. In some embodiments, the platform 220 may include a cloud server or a group of cloud servers. In some embodiments, the platform 220 may be designed to be modular such that software components can be swapped in or out. In this way, the platform 220 can be easily and / or quickly reconfigured for different uses.
[0036] In some embodiments, as shown, the platform 220 may be hosted in a cloud computing environment 222. It is noted that while the embodiments described herein describe the platform 220 as being hosted in the cloud computing environment 222, in some embodiments, the platform 220 is not cloud-based (i.e., can be implemented outside of a cloud computing environment) or may be partially cloud-based.
[0037] The cloud computing environment 222 includes an environment that hosts the platform 220. The cloud computing environment 222 may provide services such as computing, software, data access, storage, etc., which do not require an end user (e.g., the user device 210) to know the physical location and configuration of the systems and / or devices of the hosted platform 220. As shown, the cloud computing environment 222 may include a set of computing resources 224 (collectively referred to as "computing resources 224" and individually referred to as "computing resource 224").
[0038] The computing resources 224 include one or more personal computers, workstation computers, server devices, or other types of computing and / or communication devices. In some embodiments, the computing resources 224 may host the platform 220. Cloud resources may include computing instances executed in the computing resources 224, storage devices provided in the computing resources 224, data transfer devices provided by the computing resources 224, etc. In some embodiments, the computing resources 224 may communicate with other computing resources 224 via a wired connection, a wireless connection, or a combination of wired and wireless connections.
[0039] As Figure 2 Further shown, the computing resources 224 include a set of cloud resources, such as one or more applications ("APP") 224-1, one or more virtual machines ("VM") 224-2, virtualized storage ("VS") 224-3, one or more hypervisors ("HYP") 224-4, etc.
[0040] The application 224-1 includes one or more software applications, which can be provided to and / or accessed by the user device 210 and / or the platform 220. The application 224-1 can provide software applications without installing and executing them on the user device 210. For example, the application 224-1 can include software related to the platform 220 and / or any other software that can be provided through the cloud computing environment 222. In some embodiments, one application 224-1 can send / receive information to / from one or more other applications 224-1 through the virtual machine 224-2.
[0041] The virtual machine 224-2 includes a software implementation of a machine (e.g., a computer) that executes programs, similar to a physical machine. The virtual machine 224-2 can be a system virtual machine or a process virtual machine, depending on the usage and correspondence of the virtual machine 224-2 to any real machine. A system virtual machine can provide a complete system platform that supports the execution of a complete operating system (“OS”). A process virtual machine can execute a single program and can support a single process. In some embodiments, the virtual machine 224-2 can execute on behalf of a user (e.g., the user device 210) and can manage the infrastructure of the cloud computing environment 222, such as data management, synchronization, or long-term data transfer.
[0042] The virtualized storage 224-3 includes one or more storage systems and / or one or more devices that use virtualization technology within the storage system or device of the computing resource 224. In some embodiments, within the context of a storage system, the types of virtualization can include block virtualization and file virtualization. Block virtualization can refer to the abstraction (or separation) of logical storage from physical storage so that the storage system can be accessed without considering the physical storage or heterogeneous structure. The separation can allow the administrator of the storage system to flexibly manage the storage of end users. File virtualization can eliminate the dependence between the data accessed at the file level and the location of the physical storage file. This can optimize the performance of storage usage, server consolidation, and / or uninterrupted file migration.
[0043] The hypervisor 224-4 can provide hardware virtualization technology that allows multiple operating systems (e.g., “guest operating systems”) to execute simultaneously on a host computer such as the computing resource 224. The hypervisor 224-4 can provide a virtual operating platform for the guest operating systems and can manage the execution of the guest operating systems. Multiple instances of various operating systems can share the virtualized hardware resources.
[0044] Network 230 includes one or more wired and / or wireless networks. For example, network 230 can include a cellular network (e.g., a fifth-generation (5G) network, a long-term evolution (LTE) network, a third-generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, a fiber-optic based network, etc., and / or a combination of these or other types of networks.
[0045] Figure 2 The number and arrangement of the illustrated devices and networks are provided as an example. In fact, compared with Figure 2 the illustrated devices and / or networks, there can be more devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks with a different arrangement. Additionally, Figure 2 two or more of the illustrated devices can be implemented within a single device, or Figure 2 the illustrated single device can be implemented as multiple distributed devices. Additionally or alternatively, a set of devices (e.g., one or more devices) in environment 200 can perform one or more functions described as being performed by another set of devices in environment 200.
[0046] Figure 3 is Figure 2 a block diagram of example components of one or more of the devices. Device 300 can correspond to user device 210 and / or platform 220. As Figure 3 illustrated, device 300 can include a bus 310, a processor 320, a memory 330, a storage component 340, an input component 350, an output component 360, and a communication interface 370.
[0047] Bus 310 includes components that allow communication between the components of device 300. Processor 320 is implemented in hardware, firmware, or a combination of hardware and software. Processor 320 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or another type of processing component. In some implementations, processor 320 includes one or more processors that can be programmed to perform functions. Memory 330 includes random access memory (RAM), read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by processor 320.
[0048] The storage component 340 stores information and / or software related to the operation and use of the device 300. For example, the storage component 340 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cassette tape, a magnetic tape, and / or another type of non-volatile computer-readable medium, as well as corresponding drives.
[0049] The input component 350 includes components that allow the device 300 to receive information, for example, via user input, such as a touch screen display, a keyboard, a keypad, a mouse, buttons, switches, and / or a microphone. Additionally or alternatively, the input component 350 may include sensors for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, and / or an actuator). The output component 360 includes components that provide output information from the device 300, such as a display, a speaker, and / or one or more light-emitting diodes (LEDs).
[0050] The communication interface 370 includes transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter) that enable the device 300 to communicate with other devices, for example, via a wired connection, a wireless connection, or a combination of wired and wireless connections. The communication interface 370 may allow the device 300 to receive information from another device and / or provide information to another device. For example, the communication interface 370 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.
[0051] The device 300 may perform one or more of the processes described herein. In response to the processor 320 executing software instructions stored in a non-volatile computer-readable medium (e.g., the memory 330 and / or the storage component 340), the device 300 may perform these processes. A computer-readable medium is defined herein as a non-volatile memory device. A memory device includes storage space within a single physical storage device or storage space distributed across multiple physical storage devices.
[0052] The software instructions may be read into the memory 330 and / or the storage component 340 from another computer-readable medium or from another device via the communication interface 370. When executed, the software instructions stored in the memory 330 and / or the storage component 340 may cause the processor 320 to perform one or more of the processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with software instructions to perform one or more of the processes described herein. Accordingly, the embodiments described herein are not limited to any particular combination of hardware circuitry and software.
[0053] Figure 3The number and arrangement of the components shown are provided as examples. In fact, compared with the components shown in Figure 3 , device 300 may include more components, fewer components, different components, or components with a different arrangement. Additionally or alternatively, a set of components (e.g., one or more components) of device 300 may perform one or more functions described as being performed by another set of components of device 300.
[0054] Methods and apparatuses for NNMM will now be described in detail.
[0055] In an embodiment, a second matrix of weight coefficients of a neural network can be partitioned into multiple partitions. As shown in Figure 1B , the minimum matrix partition in a GEMM operation is [mr, 1] for the lhs block and [1, nr] for the rhs block, and a general MAC unit or a dedicated single instruction multiple data (SIMD) unit can be used to perform the multiplication of mr by nr.
[0056] In an embodiment, the minimum matrix partition is [mr, p] for the lhs block and [p, nr] for the rhs block. mr, nr, and p can be adjusted arbitrarily, and a uniform pattern can be defined for each combination of mr, nr, and p. The pattern can be defined in such a way that the total number of multiplications and additions is within the capacity of a general MAC unit or a dedicated SIMD unit.
[0057] Generally, a 2D matrix multiplication of [mr, nr] = [mr, p] x [p, nr] requires p multiplication operations to generate one output:
[0058]
[0059] where j = [0, mr - 1], i = [0, nr - 1] (1)
[0060] However, if n coefficients in a row of the lhs matrix share the same value, these coefficients can share a multiplication result, thus n - 1 multiplication operations can be skipped. For example, without loss of generality, if the first n coefficients in a row of the lhs matrix are the same, the following formula can be used to generate the output so as to combine n multiplication operations into one multiplication operation:
[0061]
[0062] where j = [0, mr - 1], i = [0, nr - 1] (2)
[0063] If n coefficients in a row of the lhs matrix share the same absolute value, these coefficients can also share a multiplication result, so that n - 1 multiplication operations can be skipped. For example, if the absolute values of the first n coefficients in a row of the lhs matrix are the same, the following formula can be used to generate the output so that the n multiplication operations are combined into one multiplication operation:
[0064]
[0065] where j = [0, mr - 1], i = [0, nr - 1] (3)
[0066] Therefore, in an embodiment, the neural network can be trained such that the second matrix of its weight coefficients has a predetermined uniform pattern, and the predetermined uniform pattern includes at least two of the weight coefficients having the same value. For example, after dividing the second matrix into multiple partitions, the predetermined uniform pattern includes each of the multiple partitions, and each of the multiple partitions includes the weight coefficients having the same value. For another example, after dividing the second matrix into multiple partitions, the predetermined uniform pattern only includes one of the multiple partitions, and one of the multiple partitions includes the weight coefficients having the same value. Wherein, the same value is an absolute value or a non - absolute value, and the same value is an unquantized floating - point value or a quantized integer value.
[0067] Based on the properties of matrix multiplication, several row or column reordering operations are defined:
[0068] 1) If row j0 and row j1 of the lhs matrix are the same, the calculation of O[j1, i] can be completely skipped (O[j1, i]=O[j0, i]).
[0069] 2) If multiple rows in the lhs matrix are swapped or reordered, the corresponding rows in the output are also swapped or reordered. If more uniform patterns can be found after this operation, multiple rows can be swapped or reordered.
[0070] If multiple columns in the lhs matrix are swapped or reordered, and the corresponding rows in the rhs matrix are also swapped or reordered, the output remains unchanged. If more uniform patterns can be found after this operation, multiple rows can be swapped or reordered in the lhs matrix, and multiple columns can be swapped or reordered in the rhs matrix.
[0071] 4) Similar properties can be applied to the rhs matrix block.
[0072] By swapping or reordering the rows or columns of the lhs matrix and the rhs matrix, that is, swapping or reordering multiple partitions of the second matrix partition, the number of uniform patterns in a predetermined uniform pattern can be increased, and then the neural network is retrained so that the second matrix has the predetermined uniform pattern.
[0073] To illustrate the uniform pattern, [4,4]x[4,4], [4,2]x[2,4], [4,1]x[1,4], and [2,2]x[2,2] can be used as examples; other matrix shape combinations and uniform patterns can be defined using the same principle.
[0074] As shown below, general [4,4]x[4,4] matrix multiplication uses 64 multipliers and 48 adders.
[0075] Table 1-3
[0076]
[0077]
[0078] If the uniform pattern is used to represent the lhs block, a smaller number of multiplication and addition operations will be achieved. Several uniform patterns of the [4,4] lhs block and the corresponding number of shared multipliers and adders are listed below as examples. a, b, c, d are unquantized floating-point values or quantized integer values.
[0079] Table 4
[0080]
[0081]
[0082]
[0083]
[0084] Instead of using 64 multipliers for general [4,4]x[4,4] matrix multiplication, only 4, 8, or 16 multipliers are used by using different uniform patterns for the [4,4] lhs matrix block.
[0085] As shown below, general [4,2]x[2,4] matrix multiplication uses 32 multipliers and 16 adders.
[0086] Table 5-7
[0087]
[0088] If the uniform mode is used to represent the lhs block, a smaller number of multiplications and additions are achieved. Several uniform modes of the [4, 2] lhs block and the corresponding numbers of shared multipliers and adders are listed below as examples. a, b, c, d are unquantized floating-point values or quantized integer values.
[0089] Table 8
[0090]
[0091]
[0092] Instead of using 32 multipliers for a general [4, 2] x [2, 4] matrix multiplication, by using different uniform modes for the [4, 2] lhs matrix block, only 4, 8, 12, or 16 multipliers are used.
[0093] As shown below, a general [4, 1] x [1, 4] matrix multiplication uses 16 multipliers.
[0094] Tables 9 - 11
[0095]
[0096] If the uniform mode is used to represent the lhs block, a smaller number of multiplications and additions are achieved. Several uniform modes of the [4, 1] lhs block and the corresponding numbers of shared multipliers are listed below as examples. a, b, c are unquantized floating-point values or quantized integer values.
[0097] Table 12
[0098]
[0099]
[0100] Instead of using 16 multipliers for a general [4, 1] x [1, 4] matrix multiplication, by using different uniform modes for the [4, 1] lhs matrix block, only 4, 8, or 12 multipliers are used.
[0101] As shown below, a general [2, 2] x [2, 2] matrix multiplication uses 8 multipliers and 4 adders.
[0102] Tables 13 - 15
[0103]
[0104] If the uniform mode is used to represent the lhs block, a smaller number of multiplications and additions are implemented. Several uniform modes of the [2, 2] lhs block and the corresponding numbers of shared multipliers and adders are listed below as examples. a, b, c, and d are unquantized floating-point values or quantized integer values.
[0105] Table 16
[0106]
[0107] Instead of using eight multipliers for general [2, 2] x [2, 2] matrix multiplication, different uniform modes are applied to the [2, 2] lhs matrix block, so that only two or four multipliers are used.
[0108] In an embodiment, the weight coefficients of the neural network include four-dimensional (4D) tensors. For example, the convolutional layer of the neural network is usually a 4D tensor with the shape of [R][S][K][C], where R / S is the convolutional kernel size, C is the input feature size, and K is the output feature size. Convolutional calculations usually reshape the 4D tensor [R][S][K][C] into a 2D tensor [K][CRS], where each kernel [R][S] is stored in contiguous memory. If the uniform mode is applied to this layout, most of the kernel coefficients, if not all, will be modified to the same uniform mode, resulting in a performance degradation of the large neural network that cannot be recovered even after the retraining process.
[0109] To generate more uniform modes and maintain the performance of the neural network after the retraining process, the 2D [R][S] dimension is reshaped into a 1D [RS] dimension, so that the 4D tensor [R][S][K][C] is reshaped into a 3D tensor [RS][K][C]. After reshaping the 4D tensor into a 3D tensor, the neural network is trained so that the matrix of the weight coefficients of the neural network has a predetermined uniform mode. The uniform mode development space is mainly on the [K][C] 2D plane. Optional uniform mode development operations can still be applied to the [RS] axis.
[0110] Figure 4 The 2D tensor layout 410 used in general convolutional operations and the proposed 3D tensor layout 420 are shown. Figure 5 Examples of the uniform mode using the 2D tensor layout 510 and the 3D tensor layout 520 are shown. It can be seen from the examples that more uniform modes can be developed using the 3D layout.
[0111] In one embodiment, two or more 2D planes along the RS axis are reordered to generate more uniform modes. In another embodiment, operations that do not allow 2D plane reordering are not allowed during this process.
[0112] In an embodiment, the NNMM routine is presented, which is a uniform-pattern-based GEMM3D method for neural network matrix-matrix multiplication routines.
[0113] After reading the block of the lhs matrix [mc, kc] and the slice of the rhs matrix [kc, SIZE] from the main memory into the CPU cache, additional partitioning is performed.
[0114] As Figure 6 shown, the [mc, kc] lhs block is partitioned using a uniform-pattern-based method, where the [mc, kc] lhs block is divided into multiple [mr, kc] blocks, and then into multiple [mr, p] blocks, and the [kc, SIZE] rhs block is divided into multiple [kc, nr] blocks, and then into multiple [p, nr] blocks. In this example, mr, nr, and p are all equal to 4.
[0115] Different matrix multiplication routines are defined to handle different uniform patterns of the [mr, p] lhs blocks. Depending on the requirements of a specific matrix multiplication routine, not all coefficients in the [mr, p] lhs block need to be loaded into the CPU registers.
[0116] To generate the [mr, nr] result block, the required coefficients in the [mr, p] lhs block and the [p, nr] rhs block are loaded into the CPU registers, and a specific matrix multiplication routine is used to calculate the [mr, nr] result block. Before unloading the [mr, nr] result block into the CPU cache, it is stored in the registers.
[0117] This process is repeated until all coefficients in the lhs block and the rhs block are processed, and the [mc, SIZE] result block is generated and written to the main memory. After that, the next lhs block and rhs block are loaded into the CPU cache for processing.
[0118] In one embodiment, two or more rows in the lhs matrix block are reordered to generate more uniform patterns. In another embodiment, when processing the lhs matrix block, operations that reorder rows are not allowed.
[0119] In one embodiment, two or more columns in the rhs matrix block are reordered to generate more uniform patterns. In another embodiment, when processing the rhs matrix block, operations that reorder columns are not allowed.
[0120] In one embodiment, only one uniform pattern is allowed for the lhs matrix block. In another embodiment, two or more uniform patterns can be used for the lhs matrix block.
[0121] In an embodiment, a multiplication is performed on a first matrix of input coefficients and a second matrix of weight coefficients of a neural network to determine output coefficients, and the determined output coefficients are output.
[0122] Figure 7 is a flowchart of a method 700 for NNMM according to an embodiment. In some embodiments, Figure 7 one or more processing blocks of can be executed by platform 220. In some embodiments, Figure 7 one or more processing blocks of can be executed by another device or a group of devices separate from or including platform 220, such as user device 210.
[0123] As Figure 7 shown, at step 710, method 700 includes: determining a first matrix of input coefficients.
[0124] At step 720, method 700 includes: determining a second matrix of weight coefficients of a trained neural network such that the second matrix has a predetermined uniform pattern, the predetermined uniform pattern including at least two of the weight coefficients having the same value.
[0125] At step 730, method 700 includes: performing a multiplication on the determined first matrix and the determined second matrix to determine output coefficients.
[0126] The second matrix is divided into a plurality of partitions, the predetermined uniform pattern includes each of the plurality of partitions, and each of the plurality of partitions includes the weight coefficients having the same value.
[0127] After swapping or reordering the plurality of partitions to increase the number of uniform patterns in the predetermined uniform pattern, the neural network is trained such that the second matrix has the predetermined uniform pattern.
[0128] The second matrix is divided into a plurality of partitions, the predetermined uniform pattern includes only one of the plurality of partitions, and one of the plurality of partitions includes the weight coefficients having the same value.
[0129] The same value is an absolute value or a non - absolute value, and the same value is an unquantized floating - point value or a quantized integer value.
[0130] The weight coefficients include a four - dimensional 4D tensor, and
[0131] After deforming the 4D tensor into a three - dimensional 3D tensor, the neural network is trained such that the second matrix has the predetermined uniform pattern.
[0132] The method may further include: outputting the determined output coefficients.
[0133] Although Figure 7 example blocks of method 700 are shown, in some embodiments, method 700 may include more blocks, fewer blocks, different blocks, or differently arranged blocks than those shown. Additionally or alternatively, two or more blocks of method 700 may be executed in parallel. Figure 7
[0134] Figure 8 is a schematic diagram of an apparatus 800 for NNMM according to an embodiment. As Figure 8 shown, apparatus 800 includes a first determination code 810, a second determination code 820, and an execution code 830.
[0135] The first determination code 810 is configured to cause at least one processor to determine a first matrix of input coefficients;
[0136] The second determination code 820 is configured to cause the at least one processor to determine a second matrix of weight coefficients of a trained neural network such that the second matrix has a predetermined uniform pattern, the predetermined uniform pattern including at least two of the weight coefficients having the same value; and
[0137] The execution code 830 is configured to cause the at least one processor to perform a multiplication on the determined first matrix and the determined second matrix to determine output coefficients.
[0138] The second matrix is divided into a plurality of partitions, the predetermined uniform pattern includes each of the plurality of partitions, and each of the plurality of partitions includes the weight coefficients having the same value.
[0139] After swapping or reordering the plurality of partitions to increase the number of uniform patterns in the predetermined uniform pattern, the neural network is trained such that the second matrix has the predetermined uniform pattern.
[0140] The second matrix is divided into a plurality of partitions, the predetermined uniform pattern includes only one of the plurality of partitions, and one of the plurality of partitions includes the weight coefficients having the same value.
[0141] The same value is an absolute value or a non - absolute value, and the same value is an unquantized floating - point value or a quantized integer value.
[0142] The weight coefficients include a four - dimensional 4D tensor, and after deforming the 4D tensor into a three - dimensional 3D tensor, the neural network is trained such that the second matrix has the predetermined uniform pattern.
[0143] The apparatus 800 may further include: an output code configured to cause the at least one processor to output the determined output coefficients.
[0144] According to an embodiment, there is also provided an apparatus for NNMM, including: a first determination module configured to determine a first matrix of input coefficients; a second determination module configured to determine a second matrix of weight coefficients of a trained neural network such that the second matrix has a predetermined uniform pattern, the predetermined uniform pattern including at least two of the weight coefficients having the same value; and an execution module configured to perform multiplication on the determined first matrix and the determined second matrix to determine output coefficients.
[0145] The foregoing embodiments have been provided for purposes of illustration and description, and are not intended to be exhaustive or to limit the implementations to the precise forms disclosed. Modifications and variations may be made in accordance with the above embodiments, or may be obtained from practice of the implementations.
[0146] As used herein, the term "component" is intended to be broadly construed as hardware, firmware, or a combination of hardware and software.
[0147] It is apparent that the systems and / or methods described herein may be implemented in different forms of hardware, firmware, or a combination of hardware and software. The actual specific control hardware or software code used to implement these systems and / or methods does not limit the implementations. Accordingly, the operations and performance of the systems and / or methods are described herein without reference to specific software code - it should be understood that software and hardware may be designed based on the description herein to implement the systems and / or methods.
[0148] Even if combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or not disclosed in the specification. Although each dependent claim listed below may directly depend on only one claim, the disclosure of possible implementations includes combinations of each dependent claim in the claim set with every other claim.
[0149] Unless explicitly stated, elements, acts, or instructions used herein should not be construed as critical or essential. Additionally, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." Further, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with "one or more." The term "one" or similar language is used where only one item is intended. Additionally, as used herein, the terms "has," "have," "having," etc. are intended to be open-ended terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless otherwise explicitly stated.
Claims
1. A method for optimizing matrix-matrix multiplication in a neural network, characterized in that, The method is executed by at least one processor and includes: Determining a first matrix of input coefficients; Determining a second matrix of weight coefficients of a neural network, wherein the neural network is trained to produce the second matrix having a predetermined uniform pattern, the predetermined uniform pattern includes at least two of the weight coefficients having the same value, and the second matrix is divided into a plurality of partitions; Exchanging the plurality of partitions to increase the number of uniform patterns in the predetermined uniform pattern and generating an exchanged second matrix; Performing multiplication of the first matrix and the exchanged second matrix by executing computational operations using only a first number of arithmetic units executed by the at least one processor to generate a result block; wherein the first number is less than a second number of arithmetic units based on which the at least one processor performs multiplication of the first matrix and the second matrix, and the arithmetic units include multipliers; Storing the result block in a main memory; Wherein executing the computational operations includes: based on the uniform pattern of the exchanged second matrix, performing block partitioning on the first matrix and the exchanged second matrix to obtain a plurality of [mr, p] blocks corresponding to the first matrix and a plurality of [p, nr] blocks corresponding to the exchanged second matrix; loading coefficients required for each of the [mr, p] blocks and the corresponding [p, nr] blocks into registers associated with the at least one processor, and using a specific matrix multiplication routine to calculate corresponding [mr, nr] result blocks.
2. The method according to claim 1, wherein The predetermined uniform pattern includes each of the plurality of partitions, and each of the plurality of partitions includes the weight coefficients having the same value.
3. The method according to claim 1, wherein The predetermined uniform pattern includes only one of the plurality of partitions, and one of the plurality of partitions includes the weight coefficients having the same value.
4. The method according to claim 1, wherein The same value is an absolute value or a non-absolute value.
5. The method according to any one of claims 1-4, wherein The same value is an unquantified floating-point value or a quantified integer value.
6. The method according to any one of claims 1-4, characterized in that, The weight coefficients include a four-dimensional 4D tensor, The method further includes: after deforming the 4D tensor into a three-dimensional 3D tensor, training the neural network such that the second matrix has the predetermined uniform pattern.
7. The method according to any one of claims 1 to 4, characterized in that Further includes: Outputting output coefficients based on the result block.
8. An apparatus for optimizing matrix-matrix multiplication in a neural network, characterized in that, The apparatus includes: A first determination module for determining a first matrix of input coefficients; A second determination module for determining a second matrix of weight coefficients of a neural network, wherein the neural network is trained to produce the second matrix having a predetermined uniform pattern, the predetermined uniform pattern includes at least two of the weight coefficients having the same value, and the second matrix is divided into a plurality of partitions; An exchange module for exchanging the plurality of partitions to increase the number of uniform patterns in the predetermined uniform pattern and generating an exchanged second matrix; An execution module, configured to generate a result block by performing multiplication of the first matrix and the swapped second matrix by only using computational operations performed by the at least one processor based on a first number of arithmetic units; wherein the first number is less than a second number of arithmetic units based on which the at least one processor performs multiplication of the first matrix and the second matrix, and the arithmetic units include multipliers. A storage module, configured to store the result block in a main memory. Wherein, when the execution module performs the computational operations, it includes: based on the uniform pattern of the swapped second matrix, partitioning the first matrix and the swapped second matrix into blocks, to obtain a plurality of [mr, p] blocks corresponding to the first matrix and a plurality of [p, nr] blocks corresponding to the swapped second matrix; loading coefficients required for each [mr, p] block and the corresponding [p, nr] block into registers associated with the at least one processor, and using a specific matrix multiplication routine to calculate the corresponding [mr, nr] result block.
9. A computer device, characterized in that, The computer device includes: At least one memory, configured to store program code; and At least one processor, configured to read the program code and operate according to the instructions of the program code to perform the method according to any one of claims 1-7.
10. A non - volatile computer - readable medium, characterized in that, Instructions for storage, when executed by at least one processor for neural network optimized matrix-matrix multiplication, cause the at least one processor to perform the method according to any one of claims 1-7.
Citation Information
Patent Citations
Neural network based calculation method and device
CN107402905A