Method and system for optimizing general matrix multiplication, and device and medium

By optimizing the dimensional slicing strategy in general matrix multiplication and adopting double buffering method, the problem of low computational parallelism and performance in dwarf matrix-matrix multiplication is solved, and higher parallelism and memory fetching performance are achieved.

WO2025092443A1PCT designated stage expired Publication Date: 2025-05-08SHANGHAI INFINIGENCE AI INTELLIGENT TECHNOLOGY CO LTD

Patent Information

Application Number
PCT/CN2024/125514
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-03
Filing Date
2024-10-17
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The existing general matrix multiplication sub-base has low computational parallelism and performance in dwarf matrix-matrix multiplication, making it difficult to make full use of hardware computational parallelism and improve computational memory fetch ratio.

Method used

By using the influence of the dimension slice size on parallelism and memory fetch ratio, a dimensional slice strategy is obtained, and at least one target dimension in the general matrix multiplication is sliced ​​to optimize the parallelism and memory fetch performance of matrix multiplication. The specific strategy includes using more slices when the target dimension is large, and allocating two buffers in shared memory using the double buffering method to achieve overlapping loading and computing.

Benefits of technology

The operation parallelism of matrix multiplication is improved and the memory fetch limit is reduced. The calculation parallelism of hardware is fully utilized, and the computational memory fetch ratio is improved, which improves the overall performance of matrix multiplication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024125514_08052025_PF_FP_ABST
    Figure CN2024125514_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present invention are a method and system for optimizing general matrix multiplication, and a device and a medium. The method comprises the following steps: on the basis of the influence of the size of a dimension slice on the degree of parallelism and a memory access ratio, obtaining a dimension slicing strategy; and on the basis of the dimension slicing strategy, slicing at least one target dimension in general matrix multiplication to obtain optimized dimension slices, so as to complete the optimization of matrix multiplication. In the present application, the degree of parallelism and a memory access ratio are both used as considerations for a dimension slicing strategy, and the size of a target dimension and the number of slices can be taken into comprehensive consideration, so that parallel computing units can be filled, thereby fully utilizing the degree of computing parallelism of hardware, and also increasing a computing-to-memory access ratio.
Need to check novelty before this filing date? Find Prior Art

Description

A general matrix multiplication optimization method, system, device and medium

[0001] Cross-references

[0002] This application claims priority to Chinese patent application No. 202311456255.4 filed on November 3, 2023, the entire contents of which are incorporated by reference in their entirety into this application. Technical Field

[0003] The present invention belongs to the field of deep learning technology, and specifically relates to an optimization method, system, device and medium for general matrix multiplication. Background Art

[0004] Deep learning is an emerging technology that uses complex neural networks to automate manual tasks. Computer vision is one of the main applications of deep learning, mainly performing tasks such as image classification, object detection, and semantic segmentation on digital images. These technologies have been widely used in various earth science disciplines. Deep learning is a key technology in many fields, among which general matrix multiplication (GEMM) is the most important and time-consuming operation in deep learning. For example, the convolution operation in convolutional neural networks can be converted into a general matrix multiplication (GEMM) using the Im2col method, and its execution time is about 60% of that of CNN.

[0005] The general matrix multiplication (GEMM) can be abstracted as an M×K×N load, where the shapes of the two matrices to be multiplied are (M, K) and (K, N), respectively. In many deep learning applications, the general matrix multiplication GEMM has a specific shape. For example, in the decoding phase of a large generative language model, the value of the M dimension is equal to the batch size, which is usually a relatively small value, while the sizes of the K and N dimensions can reach thousands or even tens of thousands. For example, in a linear layer in Llama2-7B, K = 4096 and N = 11008. In this case, the general matrix multiplication evolves into a dwarf matrix-matrix multiplication (Flat GEMM) with a smaller M dimension.

[0006] In order to reasonably allocate parallel computing resources, general matrix multiplication operator libraries usually use slicing to divide the dimensions of the matrix into smaller granularities and place them on specific parallel computing units. For example, a commonly used slicing granularity is B M =128, B K =128, B N =256; However, directly applying the slicing method in the existing general matrix multiplication operator library to the calculation of dwarf matrix-matrix multiplication faces challenges such as low computational parallelism and performance limited by memory access. The challenge of low computational parallelism is that since the minimum computational granularity of the tensor core is 8, the M-dimensional slice size B MThe size of the slice in the M dimension will exceed the original size of the matrix, so the number of slices in the M dimension is 1, which directly leads to a decrease in the total number of slices, making it difficult to fill the parallel computing unit and unable to fully utilize the computing parallelism of the hardware; the performance is limited by memory access. The challenge is that in the implementation of the existing general matrix multiplication operator library, B M The minimum is 64, which ensures a high compute-to-memory ratio, thereby translating the high computational throughput of Tensor Cores into actual computations. However, dwarf matrix-matrix multiplication requires reducing the slicing granularity in the M dimension, which directly leads to a decrease in its compute-to-memory ratio and its performance is limited by memory access. Therefore, an optimization method for general matrix multiplication is anticipated.

[0007] Summary of the Invention

[0008] In response to the problems existing in the prior art, the present invention provides a general matrix multiplication optimization method, system, device and medium to improve the parallelism of operations and reduce memory access limitations.

[0009] In a first aspect, embodiments of the present invention provide

[0010] A general matrix multiplication optimization method comprises the following steps:

[0011] Based on the impact of dimension slice size on parallelism and memory access ratio, the dimension slice strategy is obtained;

[0012] Based on the dimension slicing strategy, at least one target dimension in the general matrix multiplication is sliced ​​to obtain optimized dimension slices, thereby completing the optimization of the matrix multiplication.

[0013] Furthermore, the dimension slicing strategy is that the number of target dimension slices is positively correlated with the size of the target dimension.

[0014] Furthermore, the dimension slicing strategy is to use fewer target dimension slices when the target dimension is smaller, and to use more target dimension slices when the target dimension is larger.

[0015] Furthermore, the process of obtaining the dimension slicing strategy based on the impact of the dimension slice size on the parallelism and memory access ratio is as follows:

[0016] Based on multiple dimensions in general matrix multiplication, the total number of slices in all dimensions and the memory access amount of all slices are obtained;

[0017] The memory access ratio is obtained based on the total number of slices in all dimensions, the memory access amount of all slices, and the total number of multiplication and addition operations;

[0018] The degree of parallelism is obtained based on the total number of slices in all dimensions and the amount of memory accessed by all slices;

[0019] Based on the obtained memory access ratio and parallelism, we analyzed that the number of target dimension slices has opposite effects on the memory access ratio and parallelism.

[0020] Furthermore, the process of analyzing based on the obtained memory access ratio and parallelism is as follows:

[0021] Perform general matrix multiplication on target dimensions of different sizes and the corresponding number of slices of the target dimension, and obtain the visualization result.

[0022] Furthermore, the dimensions in the general matrix multiplication include M dimensions, K dimensions, and N dimensions;

[0023] The total number of slices across all dimensions is

[0024] The memory access amount of all slices is (M×B K +B K ×B N )×B+M×N;

[0025] Where B is the total number of slices in all dimensions; B K is the dimension slice granularity of K dimension; B N is the dimension slicing granularity of N dimensions; M is the value of M dimension, and both the dimension slicing granularity of M dimension and the value of M dimension are equal to the batch size; N is the value of N dimension.

[0026] Furthermore, the dimensions in the general matrix multiplication include M dimensions, K dimensions and N dimensions;

[0027] The memory access ratio is:

[0028] The total number of multiplication and addition operations is: (2×M×N×K);

[0029] The memory access ratio increases with B N increases as it increases, but has an upper bound of 2;

[0030] Among them, AI is the memory access ratio; B N is the dimension slice granularity of N dimension; M is the number of dimension slices of M dimension and the value of M dimension, both of which are equal to the batch size; K is the value of K dimension.

[0031] Furthermore, the dimensions in the general matrix multiplication include M dimensions, K dimensions and N dimensions;

[0032] The degree of parallelism is:

[0033] Among them, P is the degree of parallelism; B N B is the N-dimensional slice granularity; Kis the K-dimensional slicing granularity; K is the value of the K dimension; N is the value of the N dimension.

[0034] Furthermore, the dimension slicing strategy also includes a double buffering method;

[0035] The double buffering method allocates two separate buffers in shared memory and processes the loading and calculation of different slices in turn.

[0036] Furthermore, the two separate buffers are located in the same GPU block.

[0037] Furthermore, the double buffering method in the dimension slicing strategy is applied to larger target dimension processing scenarios.

[0038] In a second aspect, an embodiment of the present invention provides an optimization system for general matrix multiplication, including:

[0039] The preprocessing module is used to obtain the dimension slicing strategy based on the impact of the dimension slice size on the parallelism and memory access ratio;

[0040] The output module is used to slice at least one target dimension in the general matrix multiplication based on the dimension slicing strategy to obtain optimized dimension slices, thereby completing the optimization of the matrix multiplication.

[0041] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned method for optimizing general matrix multiplication when executing the computer program.

[0042] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned general matrix multiplication optimization method are implemented.

[0043] In the fifth aspect, an embodiment of the present invention also provides a computer program product, which includes a computing program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the steps of the optimization method of general matrix multiplication in the aforementioned first aspect or any implementation of the first aspect.

[0044] Other optional features and technical effects of the embodiments of the present invention are partially described below, and partially can be understood by reading this document.

[0045] Compared with the prior art, the present invention has the following beneficial technical effects:

[0046] The present invention provides a method, system, device and medium for optimizing general matrix multiplication, comprising the following steps: obtaining a dimension slicing strategy based on the influence of dimension slice size on parallelism and memory access ratio; slicing at least one target dimension in general matrix multiplication based on the dimension slicing strategy to obtain optimized dimension slices, thereby completing the optimization of matrix multiplication; the present application considers parallelism and memory access ratio as factors for the dimension slicing strategy, can comprehensively consider the size of the target dimension and the number of slices, and thus can fill parallel computing units, fully utilizing the computational parallelism of the hardware while improving the computational memory access ratio.

[0047] Furthermore, the double buffering method provided by the present invention can allocate two separate buffers in each GPU block shared memory when the target dimension is large, and use calculation to mask the delay of memory access. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. The elements shown are not limited to the scale shown in the drawings. The same or similar reference numerals in the drawings represent the same or similar elements, wherein:

[0049] FIG1 is a schematic flow chart of a general matrix multiplication optimization method according to an embodiment of the present invention;

[0050] FIG2 is a schematic diagram of a method flow for obtaining a dimension slicing strategy according to an embodiment of the present invention;

[0051] FIG3 is an embodiment of the present invention, wherein N dimension and B N A visual diagram showing the impact of changes on Flat GEMM computational performance.

[0052] FIG4 is a schematic diagram of a double buffer optimization mode according to an embodiment of the present invention;

[0053] FIG5 is a comparison diagram of the optimization effects of general matrix multiplication performed using the present application and the prior art in an embodiment of the present invention;

[0054] FIG6 is an image segmentation apparatus according to the present invention. DETAILED DESCRIPTION

[0055] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with specific embodiments and accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0056] As used herein, the term "including" and its variations represent open inclusion, i.e., "including but not limited to." Unless otherwise stated, the term "or" means "and / or." The term "based on" means "based at least in part on." The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment." The term "another embodiment" means "at least one additional embodiment." The terms "first," "second," etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0057] Figure 1 shows an optimization method 100 for general matrix multiplication according to this embodiment. As shown in Figure 1 , the method includes the following steps: at step S101, a dimension slicing strategy is obtained based on the influence of the dimension slice size on the parallelism and memory access ratio; then the process proceeds to step S102, at least one target dimension in the general matrix multiplication is sliced ​​based on the dimension slicing strategy to obtain an optimized dimension slice, thereby completing the optimization of the matrix multiplication.

[0058] It should be noted that, in this embodiment, the general matrix multiplication is dwarf matrix-matrix multiplication, wherein the general matrix multiplication is a common matrix multiplication, which is widely used in various operations in deep learning frameworks, such as convolution, fully connected layers, etc. In general matrix multiplication, the number of columns of the left matrix must be the same as the number of rows of the right matrix in order to perform multiplication operations; the dwarf matrix-matrix multiplication generally refers to a specific matrix multiplication method, which multiplies a smaller matrix (dwarf matrix) with another larger matrix. This multiplication method is generally used to convert matrix multiplication operations to better utilize computing resources and optimize performance; therefore, general matrix multiplication is a standard matrix multiplication operation, while dwarf matrix-matrix multiplication is generally a more specific matrix multiplication method used to optimize computing performance.

[0059] This application analyzes the impact of N-dimensional slice size on computational parallelism and computational memory access ratio, referred to as: N-dimensional slicing strategy for Flat GEMM, to achieve efficient slicing for Flat GEMM; it should be noted that the dimensions in the general matrix multiplication described in this embodiment include M dimensions, K dimensions, and N dimensions; since the number of slices in the M dimension is 1, in the calculation process of Flat GEMM, only K dimensions and N dimensions can be further sliced. Since K-dimensional slicing often brings atomic operations, introduces additional overhead, and is also prone to precision loss, it does not have high scalability. Therefore, the present invention focuses on N-dimensional slicing in the analysis.

[0060] Preferably, the dimension slicing strategy is that the number of target dimension slices is positively correlated with the size of the target dimension; specifically, the dimension slicing strategy is that when the target dimension is small, fewer target dimension slices are used, and when the target dimension is large, more target dimension slices are used; it should be noted that the larger and smaller described in this embodiment are relative concepts, and the object of comparison is the corresponding target dimension size and target dimension slice size when the slicing method in the general matrix multiplication operator library in the prior art is directly applied to the calculation of the dwarf matrix-matrix multiplication in this embodiment, which will result in low parallelism and performance being access-restricted. In different application scenarios, the larger and smaller, as well as fewer and more target dimension slices of the target dimension have different value ranges.

[0061] Preferably, Figure 2 shows the method 200 for obtaining the dimension slicing strategy based on the influence of the dimension slice size on the parallelism and the memory access ratio. As shown in Figure 2, at step S201, based on multiple dimensions in the general matrix multiplication, the total number of slices of all dimensions and the memory access of all slices are obtained; next, go to step S202, and obtain the memory access ratio based on the total number of slices of all dimensions and the memory access of all slices, as well as the total number of multiplication and addition operations; next, go to step S203, and obtain the parallelism based on the total number of slices of all dimensions and the memory access of all slices; next, go to step S204, and analyze based on the obtained memory access ratio and parallelism to obtain: the number of target dimension slices has opposite effects on the memory access ratio and the parallelism.

[0062] Furthermore, the dimensions in the general matrix multiplication include M dimensions, K dimensions and N dimensions;

[0063] The total number of slices across all dimensions is

[0064] The memory access amount of all slices is (M×B K +B K ×B N )×B+M×N;

[0065] Where B is the total number of slices in all dimensions; B K is the dimension slice granularity of K dimension; B N is the dimension slicing granularity of N dimensions; M is the dimension slicing granularity of M dimensions and the value of M dimensions, and the dimension slicing granularity of M dimensions and the value of M dimensions are both equal to the batch size; N is the value of N dimensions.

[0066] Furthermore, the dimensions in the general matrix multiplication include M dimensions, K dimensions and N dimensions;

[0067] The memory access ratio is:

[0068] The total number of multiplication and addition operations is: (2×M×N×K);

[0069] The memory access ratio increases with B N increases as it increases, but has an upper bound of 2;

[0070] Among them, AI is the memory access ratio; B N is the dimension slicing granularity of N dimensions; M is the dimension slicing granularity of M dimensions and the value of M dimensions, and the dimension slicing granularity of M dimensions and the value of M dimensions are both equal to the batch size; K is the value of K dimension.

[0071] Furthermore, the dimensions in the general matrix multiplication include M dimensions, K dimensions and N dimensions;

[0072] The degree of parallelism is:

[0073] Among them, P is the degree of parallelism; B N B is the N-dimensional slice granularity; K is the K-dimensional slicing granularity; K is the value of the K dimension; N is the value of the N dimension.

[0074] Preferably, the process of analyzing based on the obtained memory access ratio and parallelism is:

[0075] Perform a general matrix multiplication operation on target dimensions of different sizes and the number of slices corresponding to the target dimensions, and obtain the visualization operation result; It should be noted that in this embodiment, in order to further analyze this trade-off, the application combines the N dimension with the B dimension. N The impact of the change on the Flat GEMM calculation performance is visualized in Figure 3. It can be seen from the figure that B N The impact on Flat GEMM performance is determined by the size of N dimension: when N dimension is small, an excessively large B N This will result in fewer slices than the number of GPU stream processors, making it difficult to fully utilize the parallelism of the hardware. In this case, Flat GEMM is dominated by parallelism. When the N dimension is large, a too small B N This will lead to a decrease in the computational memory access ratio and an increase in the amount of memory access, making Flat GEMM the dominant memory access method and making it difficult to fully utilize the high computational throughput of Tensor Cores. Specifically, as shown in Figure 3, the values ​​of the K dimension in the general matrix multiplication are 4096 and 12288 respectively. As the value of the N dimension doubles, the number of dimension slices B increases. N The darker the color of the block, the stronger the parallelism dominance and memory access dominance. The darker the color of the overlapping area, the more balanced the parallelism dominance and memory access dominance are. This indicates the corresponding N-dimensional value and the number of dimension slices B. NIt can satisfy the needs of filling parallel computing units, fully utilize the computing parallelism of the hardware, and improve the computing memory access ratio; therefore, for the calculation of Flat GEMM, the N-dimensional slicing strategy is to take a smaller B when the N dimension is small. N To improve the parallelism of calculation; when the N dimension is large, take a slightly larger B N Improve the computation-to-memory access ratio. Since the computation-to-memory access ratio has an upper bound, when N is large, the computation of Flat GEMM is still dominated by memory access.

[0076] Preferably, the dimension slicing strategy also includes a double buffering method;

[0077] The double buffering method allocates two separate buffers in the shared memory and processes the loading and calculation of different slices in turn. The introduced double buffering method effectively supports Flat GEMM calculation and is referred to as the double buffering method for Flat GEMM.

[0078] Furthermore, the two separate buffers are located in the same GPU block.

[0079] Furthermore, the double buffering method in the dimension slicing strategy is applied to larger target dimension processing scenarios.

[0080] It should be noted that in this embodiment, in order to better solve the problem of memory access dominance when Flat GEMM is large in N dimension, a double buffering method is introduced. The specific method is to allocate two separate buffers in the shared memory to handle the loading and calculation of different slices in turn. When Flat GEMM calculation is performed in one buffer, the other buffer loads new slices for the calculation of the next Flat GEMM. Therefore, the calculation and memory access are overlapped.

[0081] Specifically, FIG4 shows a schematic diagram of the double-buffer optimization mode of Flat GEMM proposed in this application. For the convenience of the final accumulation, all slices of K dimensions are loaded on one GPU block for processing. The GPU block is Block, a programming unit for parallel computing of the GPU. One Block will be placed on one SM for calculation, such as A1, A2, A3..., and different slices of N dimensions all use different GPU blocks, such as C1, C2, C3...; this embodiment includes GPU Block1 and GPU Block2. Taking GPU Block1 as an example, the first slices A1 and B1 of each matrix in the K dimension are loaded into the buffer in the shared memory on the left, and Flat The GEMM calculation is performed between A1 and B1. At the same time, A2 and B2 are loaded into the buffer in the shared memory on the right, so that the calculation of A1 and B1 and the loading of A2 and B2 can overlap. In the next round, when A2 and B2 are calculated, A3 and B3 are again loaded into the buffer in the shared memory on the left to wait for calculation. Specifically, two separate buffers in the shared memory are used for loading and calculation respectively. When one of them loads the target, the other may be in the process of loading or calculating the target. When one buffer completes loading the target, it calculates the target, while the other buffer is also in the process of loading or calculating the target, thereby achieving overlapping and parallel loading and calculation. It should be further explained that similar processing is performed on different GPU blocks, except that since the number of slices in the M dimension is 1, different GPU blocks process different slices in the N dimension, that is, different GPU blocks will sequentially process the entire A matrix and different parts of the B matrix.

[0082] Figure 5 shows that the optimization scheme proposed in this patent is evaluated on the Flat GEMM loads of multiple generative large language models including Llama2-7B, OPT-6.7B and ChatGLM2-6B. The experimental results are shown in Figure 5. The vertical axis in the chart represents the acceleration ratio of the computing speed after the Flat GEMM loads of different generative large language models adopt the scheme of this application and the previous computing speed, and the horizontal axis represents the value of the M dimension. This embodiment uses two graphics processing units NVIDIA RTX 3090GPU and NVIDIA Tesla A100 GPU for processing and comparison respectively. It can be seen from the figure that compared with the general matrix multiplication operator library cuBLAS in the most prior art, the Flat GEMM calculation of this scheme is accelerated by 1.08 times to 1.42 times on the NVIDIA RTX 3090GPU and 1.04 times to 1.32 times on the NVIDIA Tesla A100 GPU.

[0083] The present invention provides an optimization system for general matrix multiplication, comprising:

[0084] A preprocessing module, configured to obtain a dimension slicing strategy based on an effect of the dimension slice size on the parallelism and the memory access ratio;

[0085] An output module is configured to slice at least one target dimension in general matrix multiplication based on a dimension slicing strategy to obtain optimized dimension slices, thereby completing the optimization of matrix multiplication.

[0086] FIG6 shows a schematic diagram of an electronic device 1000 that can implement a method or implement an embodiment of the present invention. In some embodiments, more or fewer electronic devices may be included than shown. In some embodiments, the method can be implemented using a single or multiple electronic devices. In some embodiments, the method can be implemented using cloud-based or distributed electronic devices.

[0087] As shown in Figure 6, electronic device 1000 includes a processor 1001, which can perform various appropriate operations and processes according to the programs and / or data stored in read-only memory (ROM) 1002 or the programs and / or data loaded from storage section 1008 into random access memory (RAM) 1003. Processor 1001 can be a multi-core processor or can include multiple processors. In some embodiments, processor 1001 can include a general-purpose main processor and one or more special coprocessors, such as a central processing unit (CPU), a graphics processing unit (GPU), a neural network processor (NPU), a digital signal processor (DSP), etc. In RAM 1003, various programs and data required for the operation of electronic device 1000 are also stored. Processor 1001, ROM 1002 and RAM 1003 are connected to each other via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0088] The processor and memory are used together to execute the program stored in the memory. When the program is executed by the computer, the methods, steps or functions described in the above embodiments can be implemented.

[0089] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, a touch screen, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. Removable media 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed. FIG. 6 schematically illustrates only some of the components, and does not mean that the computer system 1000 includes only the components shown in FIG.

[0090] The systems, devices, modules, or units described in the above embodiments may be implemented by a computer or its associated components. The computer may be, for example, a mobile terminal, a smartphone, a personal computer, a laptop computer, an in-vehicle human-computer interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server, or a combination thereof.

[0091] Although not shown, in an embodiment of the present invention, a storage medium is provided, wherein the storage medium stores a computer program, and the computer program is configured to execute any file difference-based compilation method according to any embodiment of the present invention when executed.

[0092] Storage media in embodiments of the present invention include permanent and non-permanent, removable and non-removable items that can be used to store information using any method or technology. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0093] The methods, programs, systems, and apparatuses of the embodiments of the present invention may be executed or implemented in a single or multiple networked computers, or may be practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks may be performed by remote processing devices connected via a communication network.

[0094] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, those skilled in the art will appreciate that the functional modules / units or controllers and related method steps described in the above embodiments may be implemented using software, hardware, or a combination of software / hardware.

[0095] Unless explicitly stated, the actions or steps of the methods, procedures, and methods described in accordance with the embodiments of the present invention do not have to be performed in a specific order and can still achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous.

[0096] In this document, multiple embodiments of the present invention are described, but for the sake of brevity, the description of each embodiment is not exhaustive, and the same or similar features or parts between the embodiments may be omitted. In this document, "one embodiment", "some embodiments", "example", "specific example", or "some examples" are intended to apply to at least one embodiment or example according to the present invention, but not all embodiments. The above terms do not necessarily mean to refer to the same embodiment or example. Those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are mutually contradictory.

[0097] While the exemplary systems and methods of the present invention have been specifically shown and described with reference to the foregoing embodiments, these are merely examples of the best modes for implementing the present systems and methods. Those skilled in the art will appreciate that various changes may be made to the embodiments of the systems and methods described herein when implementing the present systems and / or methods without departing from the spirit and scope of the present invention as defined in the appended claims.

[0098] Explanation of the accompanying drawings: 1000, electronic device; 1001, processor; 1002, read-only memory; 1003, memory; 1004, bus; 1005, I / O interface; 1006, input part; 1007, output part; 1008, storage part; 1009, communication part; 1010, drive; 1011, removable medium.

Claims

1. A general matrix multiplication optimization method, characterized in that: The following steps are involved: Based on the impact of dimension slice size on parallelism and memory access ratio, the dimension slice strategy is obtained; Based on the dimension slicing strategy, at least one target dimension in the general matrix multiplication is sliced ​​to obtain optimized dimension slices, thereby completing the optimization of the matrix multiplication.

2. The optimization method for general matrix multiplication according to claim 1, characterized in that: The dimension slicing strategy is that the number of target dimension slices is positively correlated with the size of the target dimension.

3. The optimization method for general matrix multiplication according to claim 1, characterized in that: The dimension slicing strategy is to use fewer target dimension slices when the target dimension is smaller, and to use more target dimension slices when the target dimension is larger.

4. The optimization method for general matrix multiplication according to claim 1, characterized in that: The process of obtaining the dimension slicing strategy based on the influence of the dimension slice size on the parallelism and memory access ratio is as follows: Based on multiple dimensions in general matrix multiplication, the total number of slices in all dimensions and the memory access amount of all slices are obtained; The memory access ratio is obtained based on the total number of slices in all dimensions, the memory access amount of all slices, and the total number of multiplication and addition operations; The degree of parallelism is obtained based on the total number of slices in all dimensions and the amount of memory accessed by all slices; Based on the obtained memory access ratio and parallelism, it is analyzed that the number of target dimension slices has opposite effects on the memory access ratio and parallelism.

5. The optimization method for general matrix multiplication according to claim 4, characterized in that: The process of analyzing based on the obtained memory access ratio and parallelism is as follows: Perform general matrix multiplication on target dimensions of different sizes and the corresponding number of slices of the target dimensions, and obtain the visualized operation results.

6. The optimization method for general matrix multiplication according to claim 4, characterized in that: The dimensions in the general matrix multiplication include M dimension, K dimension and N dimension; The total number of slices in all dimensions is The memory access amount of all slices is (M×B K +B K ×B N )×B+M×N; Where B is the total number of slices in all dimensions; B K B is the dimension slice granularity of K dimension; N is the dimension slice granularity of N dimension; M is the value of M dimension, the dimension slice granularity of M dimension and the value of M dimension are both equal to the batch size; N is the value of N dimension.

7. The optimization method for general matrix multiplication according to claim 4, characterized in that: The dimensions in the general matrix multiplication include M dimension, K dimension and N dimension; The memory access ratio is: The total number of multiplication and addition operations is: (2×M×N×K); The memory access ratio increases with B N increases as it increases, but has an upper bound of 2; Among them, AI is the memory access ratio; B N is the dimension slice granularity of N dimension; M is the value of M dimension, the dimension slice granularity of M dimension and the value of M dimension are both equal to the batch size; K is the value of K dimension.

8. The optimization method for general matrix multiplication according to claim 4, characterized in that: The dimensions in the general matrix multiplication include M dimension, K dimension and N dimension; The degree of parallelism is: Among them, P is the degree of parallelism; B N B is the dimension slice granularity of N dimensions; K is the dimension slice granularity of the K dimension; K is the value of the K dimension; N is the value of the N dimension.

9. The optimization method for general matrix multiplication according to claim 1, characterized in that: The dimension slicing strategy also includes a double buffering method; The double buffering method is to allocate two separate buffers in the shared memory and process the loading and calculation of different slices in turn.

10. The optimization method for general matrix multiplication according to claim 9, characterized in that: The two separate buffers are located in the same GPU block.

11. The optimization method for general matrix multiplication according to claim 9, characterized in that: The double buffering method in the dimension slicing strategy is applied to larger target dimension processing scenarios.

12. A general matrix multiplication optimization system, characterized in that: The optimization method of the general matrix multiplication according to any one of claims 1 to 11 comprises: A preprocessing module is used to obtain a dimension slicing strategy based on the impact of the dimension slice size on the parallelism and memory access ratio; The output module is used to slice at least one target dimension in the general matrix multiplication based on the dimension slicing strategy to obtain optimized dimension slices, thereby completing the optimization of the matrix multiplication.

13. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the processor implements the steps of the optimization method for general matrix multiplication according to any one of claims 1 to 11.

14. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the optimization method for general matrix multiplication according to any one of claims 1 to 11 are implemented.

15. A computer program product, comprising a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program comprises program instructions, which, when executed by a computer, cause the computer to perform the steps of the optimization method for general matrix multiplication as described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Single-precision matrix multiplication optimization method and system based on NVIDIA Kepler GPU assembly instruction

    CN106681694A

  • FT-2000+ based integer matrix multiplication kernel optimization method

    CN114090954A

  • General matrix multiplication optimization method, system, equipment and medium

    CN117407643A

  • Compiler-level general matrix multiplication configuration optimization

    US20210200521A1

Cited By

  • Parameter optimization method and device, electronic equipment, storage medium and computer program product

    CN120875051A

  • Memory architecture-oriented dual-precision general matrix multiplication optimization method and system

    CN121278223A