Data processing method and device, electronic equipment and storage medium

By splitting the matrix of the large model between multiple processors and solving the splitting solution based on the objective function, the problems of high hardware cost and large communication overhead in the multiplication operation of artificial intelligence large model matrix are solved, and the effect of efficiently processing large-scale models and reducing communication overhead is achieved.

CN120045338AActive Publication Date: 2025-05-27MOORE THREADS TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510521064.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

In the matrix multiplication operation of the artificial intelligence large model, a single processor cannot fully store the entire model, resulting in high hardware costs and the splitting method although lossless, it brings communication overhead.

Method used

By splitting the input matrix, the first parameter matrix, and the second parameter matrix between the multiple processors, and solving the splitting scheme based on the objective function, to minimize the total traffic. This objective function represents the relationship between the total traffic volume and the splitting scheme. The resulting splitting scheme can effectively reduce the communication overhead between processors.

Benefits of technology

A larger-scale model is implemented efficiently between multiple processors, reducing communication overhead between processors and improving system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045338A_ABST
    Figure CN120045338A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and provides a data processing method and device, electronic equipment and a storage medium, and the method comprises the steps: splitting an input matrix, a first parameter matrix and a second parameter matrix based on a splitting scheme through any one of a plurality of processors, and inputting a splitting result to the plurality of processors; the plurality of processors calculate the product of the input matrix, the first parameter matrix and the second parameter matrix based on the splitting result; the splitting scheme is obtained by solving a target function, the target function represents the relationship between the total communication traffic and the splitting scheme, and the total communication traffic is the total communication traffic among the plurality of processors in the process that the plurality of processors calculate the matrix product of the input matrix, the first parameter matrix and the second parameter matrix. According to the embodiment of the invention, through the splitting scheme obtained by solving the objective function, a larger-scale neural network model can be processed by using a plurality of processors, and the communication overhead among the plurality of processors is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a data processing method and apparatus, an electronic device, and a storage medium. Background Art

[0002] An artificial intelligence (AI) large model refers to a huge and complex neural network that needs to store more parameters to increase the depth and width of the model, thereby improving the performance of the model. The application scenarios of the artificial intelligence (AI) large model are very wide, covering fields such as medical care, finance, industry, education, and smart cities. For the inference of the large model, a special processor can be used for acceleration, such as a Graphics Processing Unit (GPU). The storage of such a processor is often expensive, resulting in a single processor being unable to store the entire model completely. To reduce the hardware cost, the large model can be processed by distillation, pruning, quantization, and splitting. Among them, compared with other processing methods, the splitting method is completely lossless, but splitting will bring additional communication overhead. Summary of the Invention

[0003] The present disclosure provides a technical solution for data processing.

[0004] According to one aspect of the present disclosure, there is provided a data processing method. The data processing method is applied to multiple processors, and the method includes: splitting an input matrix, a first parameter matrix, and a second parameter matrix based on a splitting scheme by any one of the multiple processors, and inputting the splitting result into the multiple processors, where the splitting scheme represents the number of splitting parts of each dimension of the input matrix, the first parameter matrix, and the second parameter matrix; calculating the product of the input matrix, the first parameter matrix, and the second parameter matrix by the multiple processors based on the splitting result; where the splitting scheme is obtained by solving an objective function, and the objective function represents the relationship between the total communication volume and the splitting scheme, and the total communication volume is the total communication volume between multiple processors during the process of calculating the matrix product of the input matrix, the first parameter matrix, and the second parameter matrix by the multiple processors.

[0005] In a possible implementation manner, the total communication volume is obtained based on the communication volume of the first-layer matrix product and the communication volume of the second-layer matrix product, where the first-layer matrix product is the product of the input matrix and the first parameter matrix, and the second-layer matrix product is the product of the product result of the input matrix and the first parameter matrix and the second parameter matrix.

[0006] In a possible implementation, the input matrix is a three-dimensional matrix, the first parameter matrix and the second parameter matrix are two-dimensional matrices, the number of elements in the third dimension of the input matrix, the first dimension of the first parameter matrix, and the second dimension of the second parameter matrix are the same, and the number of elements in the second dimension of the first parameter matrix and the first dimension of the second parameter matrix are the same; the objective function includes a first variable, a second variable, and a third variable, and the product of the first variable, the second variable, and the third variable is equal to the number of processors, and the splitting scheme is represented by the first variable, the second variable, and the third variable; the first variable is used to represent the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the first splitting direction, and the first splitting direction includes the direction of the first dimension of the input matrix; the second variable is used to represent the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the second splitting direction, and the second splitting direction includes the direction of the second dimension of the input matrix; the third variable is used to represent the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the third splitting direction, and the third splitting direction includes the direction of the third dimension of the input matrix.

[0007] In a possible implementation, splitting the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting scheme by any one of the multiple processors includes: in the first splitting direction, the second splitting direction, and the third splitting direction, splitting the input matrix, the first parameter matrix, and the second parameter matrix respectively according to the first variable, the second variable, and the third variable, so that each processor corresponds to different input sub-matrices, first parameter sub-matrices, and second parameter sub-matrices.

[0008] In a possible implementation, the communication volume of the first-layer matrix multiplication is obtained based on a first communication volume, a second communication volume, and a third communication volume, where: the first communication volume is the communication volume of performing a first collection operation on the input sub-matrices on each processor in the second splitting direction; the second communication volume is the communication volume of performing a second collection operation on the first parameter sub-matrices on each processor in the first splitting direction; the third communication volume is the communication volume of multiplying and summing the first collection result and the second collection result on each processor in the third splitting direction, and distributing the first multiplication and summation result of the first collection result and the second collection result to each processor, where the first collection result is the result obtained by performing the first collection operation, and the second collection result is the result obtained by performing the second collection operation.

[0009] In a possible implementation, the communication volume of the second-layer matrix multiplication is obtained based on a fourth communication volume and a fifth communication volume, where: the fourth communication volume is the communication volume of performing a third collection operation on the second parameter sub-matrices on each processor in the first splitting direction; the fifth communication volume is the communication volume of multiplying and summing the first multiplication and summation result and the third collection result on each processor in the second splitting direction, and splitting the second multiplication and summation result of the first multiplication and summation result and the third collection result to each processor, where the third collection result is the result obtained by performing the third collection operation.

[0010] In a possible implementation, the splitting scheme includes the first variable, the second variable, and the third variable that minimize the objective function.

[0011] In a possible implementation, the input matrix includes feature data in a deep learning task, and the feature data includes at least one of image feature data, speech feature data, and text feature data.

[0012] According to one aspect of the present disclosure, a data processing device is provided. The data processing device is applied to multiple processors, and the device includes: a splitting module, configured to split an input matrix, a first parameter matrix, and a second parameter matrix based on a splitting scheme through any one of the multiple processors, and input the splitting result to the multiple processors, where the splitting scheme represents the number of splitting parts of each dimension of the input matrix, the first parameter matrix, and the second parameter matrix; a calculation module, configured to calculate the product of the input matrix, the first parameter matrix, and the second parameter matrix by the multiple processors based on the splitting result; where the splitting scheme is obtained by solving an objective function, and the objective function represents the relationship between the total communication volume and the splitting scheme, and the total communication volume is the total communication volume between multiple processors during the process of the multiple processors calculating the matrix product of the input matrix, the first parameter matrix, and the second parameter matrix.

[0013] In a possible implementation, the total communication volume is obtained based on the communication volume of the first-layer matrix multiplication and the communication volume of the second-layer matrix multiplication, where the first-layer matrix multiplication is the product of the input matrix and the first parameter matrix, and the second-layer matrix multiplication is the product of the product result of the input matrix and the first parameter matrix and the second parameter matrix.

[0014] In a possible implementation, the input matrix is a three-dimensional matrix, the first parameter matrix and the second parameter matrix are two-dimensional matrices, the number of elements in the third dimension of the input matrix, the first dimension of the first parameter matrix, and the second dimension of the second parameter matrix is the same, and the number of elements in the second dimension of the first parameter matrix and the first dimension of the second parameter matrix is the same; the objective function includes a first variable, a second variable, and a third variable, and the product of the first variable, the second variable, and the third variable is equal to the number of processors, and the splitting scheme is represented by the first variable, the second variable, and the third variable; the first variable is used to represent the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the first splitting direction, and the first splitting direction includes the direction of the first dimension of the input matrix; the second variable is used to represent the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the second splitting direction, and the second splitting direction includes the direction of the second dimension of the input matrix; the third variable is used to represent the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the third splitting direction, and the third splitting direction includes the direction of the third dimension of the input matrix.

[0015] In a possible implementation, the splitting module is configured to: in the first splitting direction, the second splitting direction, and the third splitting direction, split the input matrix, the first parameter matrix, and the second parameter matrix respectively according to the first variable, the second variable, and the third variable, so that each processor corresponds to different input sub-matrices, first parameter sub-matrices, and second parameter sub-matrices.

[0016] In a possible implementation, the communication volume of the first-layer matrix multiplication is obtained based on a first communication volume, a second communication volume, and a third communication volume, where: the first communication volume is the communication volume of performing a first collection operation on the input sub-matrix on each processor in the second splitting direction; the second communication volume is the communication volume of performing a second collection operation on the first parameter sub-matrix on each processor in the first splitting direction; the third communication volume is the communication volume of multiplying and summing the first collection result and the second collection result on each processor in the third splitting direction, and distributing the first multiplication and summation result of the first collection result and the second collection result to each processor, where the first collection result is the result obtained by performing the first collection operation, and the second collection result is the result obtained by performing the second collection operation.

[0017] In a possible implementation, the communication volume of the second-layer matrix multiplication is obtained based on a fourth communication volume and a fifth communication volume, where: the fourth communication volume is the communication volume of performing a third collection operation on the second parameter sub-matrices on each processor in the first splitting direction; the fifth communication volume is the communication volume of multiplying and summing the first multiplication-sum result and the third collection result on each processor in the second splitting direction, and splitting the second multiplication-sum result of the first multiplication-sum result and the third collection result to each processor, where the third collection result is the result obtained by performing the third collection operation.

[0018] In a possible implementation, the splitting scheme includes the first variable, the second variable, and the third variable that minimize the objective function.

[0019] In a possible implementation, the input matrix includes feature data in a deep learning task, and the feature data includes at least one of image feature data, speech feature data, and text feature data.

[0020] According to one aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above method.

[0021] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.

[0022] The data processing method of the embodiments of the present disclosure can be applied to multiple processors. The method includes: through any one of the multiple processors, splitting an input matrix, a first parameter matrix, and a second parameter matrix based on a splitting scheme, and inputting the splitting results to the multiple processors, where the splitting scheme represents the number of splitting parts of each dimension of the input matrix, the first parameter matrix, and the second parameter matrix; the multiple processors calculate the products of the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting results; wherein, the splitting scheme is obtained by solving an objective function, and the objective function represents the relationship between the total communication volume and the splitting scheme, and the total communication volume is the total communication volume between the multiple processors during the process of the multiple processors calculating the matrix products of the input matrix, the first parameter matrix, and the second parameter matrix. The splitting scheme obtained by solving the objective function can utilize multiple processors to process larger-scale models (such as neural network models with a large number of parameters) and reduce the communication overhead between the multiple processors.

[0023] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Description of the Drawings

[0024] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings show embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0025] Figure 1 A flowchart showing a data processing method according to an embodiment of the present disclosure.

[0026] Figure 2 A schematic diagram showing a collection operation according to an embodiment of the present disclosure.

[0027] Figure 3 A schematic diagram showing a reduction-scatter sum operation according to an embodiment of the present disclosure.

[0028] Figure 4 A schematic diagram showing a full reduction sum operation according to an embodiment of the present disclosure.

[0029] Figure 5 A schematic diagram showing a full reduction multiply-accumulate operation according to an embodiment of the present disclosure.

[0030] Figure 6 A schematic diagram showing a reduction-scatter multiply-accumulate operation according to an embodiment of the present disclosure.

[0031] Figure 7 A schematic diagram showing the inference process of an objective function according to an embodiment of the present disclosure.

[0032] Figure 8 A schematic diagram comparing the total communication volume of a data processing method according to an embodiment of the present disclosure with related technologies.

[0033] Figure 9 A block diagram showing a data processing apparatus according to an embodiment of the present disclosure.

[0034] Figure 10 A block diagram showing an electronic device according to an embodiment of the present disclosure. Detailed Embodiments

[0035] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0036] As used herein, the term "exemplary" means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein need not be construed as superior or better than other embodiments.

[0037] As used herein, the term "and / or" is merely a description of an associated relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, as used herein, the term "at least one" means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.

[0038] In addition, to better illustrate the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0039] After splitting an artificial intelligence model (such as a neural network module) across different processors, the communication volume between the processors becomes an indicator affecting system performance. However, for a large artificial intelligence model, as the number of processors increases, the amount of computation on each processor decreases linearly, but the communication volume increases instead.

[0040] The splitting methods in the related art include pipeline parallelism by splitting according to model layers, model parallelism by directly splitting parameter matrices, etc. Among them, the model parallelism method can reduce the inference latency and can be divided into one-dimensional parallelism and two-dimensional parallelism. For example, based on ring-based splitting and collection, the communication volume can be made fixed through a one-dimensional matrix splitting algorithm. However, as the number of communication devices increases, the fixed handshake communication volume will gradually increase, and the fixed communication volume will still degrade the overall system performance. The method based on two-dimensional matrix splitting and matrix information collection can further reduce the communication volume, but this communication volume is still not optimal. If the number of processors n is considered as a variable, the minimum communication volume in the related art is inversely proportional to the square root of n, and this order of magnitude is clearly not optimal. In view of this, the embodiments of the present disclosure propose a data processing method that can effectively reduce the communication overhead between multiple processors.

[0041] Figure 1 A flowchart showing the data processing method according to an embodiment of the present disclosure is as Figure 1 shown, and the data processing method includes:

[0042] In step S11, any one of the multiple processors splits the input matrix, the first parameter matrix, and the second parameter matrix based on a splitting scheme, and inputs the splitting results into the multiple processors, where the splitting scheme represents the number of splitting parts for each dimension of the input matrix, the first parameter matrix, and the second parameter matrix.

[0043] In step S12, the multiple processors calculate the product of the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting results.

[0044] Among them, the splitting scheme is obtained by solving an objective function, where the objective function represents the relationship between the total communication volume and the splitting scheme, and the total communication volume is the total communication volume among the multiple processors during the process of the multiple processors calculating the matrix product of the input matrix, the first parameter matrix, and the second parameter matrix.

[0045] In the data processing method of the embodiments of the present disclosure, the splitting scheme obtained by solving the objective function can utilize multiple processors to process larger-scale models (such as neural network models with a large number of parameters) and reduce the communication overhead among the multiple processors.

[0046] In a possible implementation manner, the data processing method of the embodiments of the present disclosure can be applied to multiple processors. Each processor can be newly designed or obtained by improving an existing processor chip. The types of processor chips can include but are not limited to: Central Processing Unit (CPU), Graphic Processing Unit (GPU), General-Purpose Computing on Graphics Processing Units (GPGPU), Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA), or other programmable logic devices, and can also include a microprocessor or a processor of other conventional processors.

[0047] In a possible implementation manner, in large-scale calculations based on large models (such as neural network models with a large number of parameters), two-layer matrix multiplication occupies a large amount of computing resources. The two-layer matrix multiplication can be expressed as:

[0048] Y = act(XW 1 )W2 (1)

[0049] In formula (1), act represents an element-wise activation function, which can be a function running on the neurons of a neural network, responsible for mapping the input of a neuron to the output end and increasing the non-linearity of the neural network model. For example, it includes functions such as Gaussian Error Linear Units (gelu) function and Sigmoid function. X represents the input matrix, which can be a three-dimensional matrix with dimensions B×L×E (abbreviated as BLE), and W 1 represents the first parameter matrix, whose dimensions can be E×F (abbreviated as EF), and W 2 represents the second parameter matrix, whose dimensions can be F×E (abbreviated as FE), and Y represents the result of the product of two matrices, whose dimensions can be the same as those of the input matrix X, also B×L×E. Among them, the specific values of B, L, E, and F in the embodiments of the present disclosure are not limited and can be set according to the actual application scenario.

[0050] In the example, the input matrix X is a three-dimensional matrix, and the first parameter matrix W 1 , and the second parameter matrix W 2 are two-dimensional matrices (for example, the input matrix X has dimensions B×L×E, the first parameter matrix W 1 has dimensions E×F, and the second parameter matrix W 2 can have dimensions F×E). The number of elements in the third dimension of the input matrix X (for example, the dimension where the dimension E of the input matrix X is located), the first dimension of the first parameter matrix W 1 (for example, the dimension where the dimension E of the first parameter matrix W 1 is located), and the second dimension of the second parameter matrix W 2 (for example, the dimension where the dimension E of the second parameter matrix W 2 is located) are the same, and the number of elements in the second dimension of the first parameter matrix W 1 (for example, the dimension where the dimension F of the first parameter matrix W 1 is located) and the first dimension of the second parameter matrix W 2 (for example, the dimension where the dimension F of the second parameter matrix W 2 is located) are the same.

[0051] In the example, the input matrix includes feature data in a deep learning task, and the feature data includes at least one of image feature data, speech feature data, and text feature data. For example, in a scenario where a deep neural network is used for face recognition of a target object, the input matrix can be the image feature data of the target object (such as a face feature map); in a scenario where a deep neural network is used for speech recognition of a target object, the input matrix can be the speech feature data of the target object; in a scenario where a deep neural network is used for character recognition of a target document, the input matrix can be the text feature data of the target document; the embodiments of the present application do not limit the type of the input matrix.

[0052] For large-scale computing based on a large model, since a single processor cannot completely store the entire model, in step S11, any one of the multiple processors can read the objective function for determining the splitting scheme from the memory, and obtain the splitting scheme by solving the objective function. The splitting scheme can represent the input matrix X, the first parameter matrix W 1 , the second parameter matrix W 2 The number of splitting parts in each dimension. Then, any one of the multiple processors can split the input matrix X along the three dimensions of B, L, and E (for example, B corresponds to the horizontal direction, L corresponds to the vertical direction, and E corresponds to the depth direction) according to the splitting scheme, and split the first parameter matrix W 1 , the second parameter matrix W 2 along the two dimensions of E and F, and evenly and redundantly distribute the splitting results to multiple processors, so that each processor corresponds to a copy of the input sub-matrix, the first parameter sub-matrix, and the second parameter sub-matrix.

[0053] In a possible implementation manner, the objective function for determining the splitting scheme can represent the relationship between the total communication volume and the splitting scheme, and the total communication volume is the total communication volume between n (n>1) processors during the calculation of the matrix product of the input matrix X, the first parameter matrix W 1 , the second parameter matrix W 2 For example, the objective function can be expressed as:

[0054] T comm =BLE / xz + EF / yz + 2BLF / xy + EF / yz + BLE / xz (2)

[0055] In formula (2), T comm represents the total communication volume between n processors during the calculation of formula (1), BLE represents the size of the input matrix X, and EF represents the first parameter matrix W 1 (or the second parameter matrix W 2The size of (), the objective function includes a first variable x, a second variable y, and a third variable z. The product of the first variable x, the second variable y, and the third variable z is equal to the number of processors n, that is, xyz = n. The splitting scheme is represented by the first variable x, the second variable y, and the third variable z.

[0056] Among them, the first variable x is used to characterize the splitting fractions of the input matrix X, the first parameter matrix W 1 , and the second parameter matrix W 2 in the first splitting direction respectively. The first splitting direction includes the direction of the first dimension of the input matrix X, for example, the direction of the dimension where the size B of the input matrix X is located.

[0057] The second variable y is used to characterize the splitting fractions of the input matrix X, the first parameter matrix W 1 , and the second parameter matrix W 2 in the second splitting direction respectively. The second splitting direction includes the direction of the second dimension of the input matrix X, for example, the direction of the dimension where the size L of the input matrix X is located.

[0058] The third variable z is used to characterize the splitting fractions of the input matrix X, the first parameter matrix W 1 , and the second parameter matrix W 2 in the third splitting direction respectively. The third splitting direction includes the direction of the third dimension of the input matrix X, for example, the direction of the dimension where the size E of the input matrix X is located.

[0059] By solving the objective function based on the communication volume as shown in formula (2) under the condition of satisfying the constraint xyz = n, the splitting scheme can be obtained efficiently and quickly, which is beneficial to subsequent uniform splitting of the input matrix X, the first parameter matrix W 1 , and the second parameter matrix W 2 in three dimensions, making the splitting fractions xyz the same as the number of processors n, so as to make more full use of each processor and reduce the communication volume between processors.

[0060] In a possible implementation manner, step S11 may include: in the first splitting direction, the second splitting direction, and the third splitting direction, according to the first variable x, the second variable y, and the third variable z, respectively split the input matrix X, the first parameter matrix W 1 , and the second parameter matrix W 2 so that each processor corresponds to different input sub-matrices, first parameter sub-matrices, and second parameter sub-matrices respectively.

[0061] In the example, assume that processor i is any one of the n processors. Processor i can solve for the first variable x, the second variable y, and the third variable z through formula (2). Moreover, according to the input matrix X, the first splitting direction is the horizontal direction where the dimension B of the input matrix X is located, the second splitting direction is the vertical direction where the dimension L of the input matrix X is located, and the third splitting direction is the depth direction where the dimension E of the input matrix X is located.

[0062] Processor i can split the elements of the first dimension (the dimension where dimension B is located) of the input matrix X into x parts in the horizontal direction, split the elements of the second dimension (the dimension where dimension L is located) of the input matrix X into y parts in the vertical direction, and split the elements of the third dimension (the dimension where dimension E is located) of the input matrix X into z parts in the depth direction. In total, the input matrix X with the size of BLE is split into xyz = n parts, and the size of each part is B x L y E z . Processor i can input the n input sub-matrices obtained by splitting, each with the size of B x L y E z to the n processors respectively, so that each processor stores a different input sub-matrix.

[0063] Processor i can also split the elements of the first dimension (the dimension where dimension E is located) of the first parameter matrix W 1 into z parts in the depth direction, and split the elements of the second dimension (the dimension where dimension F is located) of the first parameter matrix W 1 into yx parts in the vertical and horizontal directions. In total, the first parameter matrix W with the size of EF 1 is split into xyz = n parts, and the size of each part is E z F yx . Processor i can input the n first parameter sub-matrices obtained by splitting, each with the size of E z F yx to the n processors respectively, so that each processor stores a different first parameter sub-matrix.

[0064] Processor i can similarly split the elements of the first dimension (the dimension where dimension F is located) of the second parameter matrix W 2 into y parts in the vertical direction, and split the elements of the second dimension (the dimension where dimension E is located) of the second parameter matrix W 2 into zx parts in the depth and horizontal directions. In total, the second parameter matrix W with the size of FE 2 is split into xyz = n parts, and the size of each part is F y E zx . Processor i can input the n second parameter sub-matrices obtained by splitting, each with the size of F y E zxThe second parameter sub-matrices are respectively input into n processors, so that each processor stores a different second parameter sub-matrix.

[0065] In this way, according to the first variable x, second variable y, and third variable z solved from the objective function, the input matrix X, the first parameter matrix W 1 , and the second parameter matrix W 2 are evenly split in three splitting dimensions, so that the number of splits xyz is the same as the number of processors n, which is beneficial to making full use of each processor and reducing the communication volume between processors.

[0066] The following introduces the reasoning process of the objective function, for example, including the process of three-dimensional splitting, collection, and calculation of the input matrix X, the first parameter matrix W 1 , and the second parameter matrix W 2 . To more clearly illustrate the reasoning process of the objective function, the following introduces several matrix communication operations and the overhead of each matrix communication operation.

[0067] Figure 2 The figure shows a schematic diagram of the collection operation according to an embodiment of the present disclosure. As Figure 2 shown, if there are n processors, such as processor 1 to processor n, each processor stores 1 / n of the matrix (see Figure 2 the upper part), and the collection operation collects each part of the matrix on each processor to obtain the complete matrix (see Figure 2 the lower part). If the size of each part of the matrix is w / n, then the overhead of the data volume (i.e., the communication volume) exchanged by n processors is w.

[0068] Figure 3 The figure shows a schematic diagram of the reduction-scatter summation operation according to an embodiment of the present disclosure. As Figure 3 shown, if there are n processors, such as processor 1 to processor n, each processor stores a different matrix (see Figure 3 the upper part), where the identifiers 11 to n1 are used to distinguish different parts of the matrix stored by processor 1, the identifiers 12 to n2 are used to distinguish different parts of the matrix stored by processor 2, and so on, and the identifiers 1n to nn are used to distinguish different parts of the matrix stored by processor n. The reduction-scatter summation operation calculates the sum of all matrices and evenly distributes the calculation results on n processors according to a certain dimension (see Figure 3 the lower part). If the size of each matrix is w / n, then the overhead of the data volume exchanged by n processors is approximately w / n.

[0069] Figure 4 The figure shows a schematic diagram of the all-reduce summation operation according to an embodiment of the present disclosure. As Figure 4As shown, if there are n processors, such as processor 1 to processor n, each processor stores a different matrix (see Figure 4 the upper part), the all-reduce sum operation calculates the sum of all matrices and distributes it to each of the n processors (see Figure 4 the lower part). If the size of each matrix is w / n, then the overhead of data interaction among the n processors is approximately 2w / n.

[0070] Figure 5 The figure shows a schematic diagram of the all-reduce multiply-add operation according to an embodiment of the present disclosure. As Figure 5 shown, if there are n processors, such as processor 1 to processor n, each processor stores two matrices to be multiplied (see Figure 5 the upper part), the all-reduce multiply-add operation calculates the product of the two matrices to be multiplied on each processor, then adds the products of the n processors, and distributes the multiplied and summed result to each of the n processors (see Figure 5 the lower part). This process is equivalent to each processor calculating the matrix multiplication operation separately, and then performing the all-reduce sum operation as Figure 4 shown. Therefore, if the size of each matrix is w / n, then the overhead of data interaction among the n processors is approximately 2w / n.

[0071] Figure 6 The figure shows a schematic diagram of the reduce-scatter multiply-add operation according to an embodiment of the present disclosure. As Figure 6 shown, if there are n processors, such as processor 1 to processor n, each processor stores two matrices to be multiplied (see Figure 6 the upper part), where labels 11 to n1 are used to distinguish different parts of one matrix stored by processor 1, and labels 11' to n1' are used to distinguish different parts of the other matrix stored by processor 1; labels 12 to n2 are used to distinguish different parts of one matrix stored by processor 2, and labels 12' to n2' are used to distinguish different parts of the other matrix stored by processor 2; and so on, labels 1n to nn are used to distinguish different parts of one matrix stored by processor n, and labels 1n' to nn' are used to distinguish different parts of the other matrix stored by processor n. The reduce-scatter multiply-add operation calculates the products of the two matrices to be multiplied on each processor, then adds the products of the n processors, and then distributes the multiplied and summed result to each of the n processors after splitting it according to a certain dimension (see Figure 6 the lower part). This process is equivalent to each processor calculating the matrix multiplication operation separately, and then performing the reduce-scatter sum operation as Figure 3 shown. Therefore, if the size of each matrix is w / n, then the overhead of data interaction among the n processors is approximately w / n.

[0072] Figure 7 A schematic diagram showing the reasoning process of the objective function according to an embodiment of the present disclosure, as Figure 7 shown, given that the size of the input matrix X is BLE, the size of the first parameter matrix W 1 is EF, and the size of the second parameter matrix W 2 is FE. It is possible to assume the first variable x, the second variable y, and the third variable z, which are used to represent the input matrix X, the first parameter matrix W 1 , and the second parameter matrix W 2 respectively split into x, y, and z parts in 3 splitting directions (for example, including the first splitting direction, the second splitting direction, and the third splitting direction), so that each of the n processors can correspond to a sub-input matrix of size B x L y E z , a first parameter sub-matrix of size E z F yx , and a second parameter sub-matrix of size F y E zx . Among them, xyz = n, so that the n processors can be logically mapped to a three-dimensional matrix, corresponding to x processors in the first splitting direction, y processors in the second splitting direction, and z processors in the third splitting direction.

[0073] In a possible implementation, as Figure 7 shown, the total communication volume is obtained based on the communication volume of the first-layer matrix multiplication and the communication volume of the second-layer matrix multiplication. The first-layer matrix multiplication is the product of the input matrix X and the first parameter matrix W 1 ; the second-layer matrix multiplication is the product of the product result XW 1 of the input matrix X and the first parameter matrix W 1 , and the second parameter matrix W 2 . In this way, by summing the communication volumes of the two-layer matrix multiplications, the total communication volume between multiple processors can be accurately determined.

[0074] Among them, as shown in formula (1), since an act operation (element-wise activation function operation) is performed after the matrix multiplication of the input matrix X and the first parameter matrix W 1 , and the act operation is performed separately on each processor without communication volume. Therefore, the total communication volume obtained based on the communication volume of the first-layer matrix multiplication and the communication volume of the second-layer matrix multiplication is also the total communication volume of multiple processors for processing formula (1).

[0075] In a possible implementation, the communication volume of the first-layer matrix multiplication is obtained based on the first communication volume, the second communication volume, and the third communication volume. The first communication volume is the communication volume of performing a first collection operation on the input sub-matrices on each processor in the second splitting direction; the second communication volume is the communication volume of performing a second collection operation on the first parameter sub-matrices on each processor in the first splitting direction; the third communication volume is the communication volume of multiplying and summing the first collection result and the second collection result on each processor in the third splitting direction, and distributing the first multiplication and summation result of the first collection result and the second collection result to each processor. In this way, it is beneficial to accurately determine the communication volume of the first-layer matrix multiplication.

[0076] For example, as Figure 7 shown, in n processors, the input sub-matrices of size B x L y E z can be subjected to the first collection operation as Figure 2 shown (see the collection operation shown in Figure 2 ) in the second splitting direction (for example, the splitting direction corresponding to the second variable y), so that each processor in the second splitting direction stores the first collection result of size B x LE z . The first communication volume of this step is BLE / xz.

[0077] Synchronously, the first parameter sub-matrices of size E z F yx can be subjected to the second collection operation (see the collection operation shown in Figure 2 ) in the first splitting direction (for example, the splitting direction corresponding to the first variable x), so that each processor in the first splitting direction stores the second collection result of size E z F y . The second communication volume of this step is EF / yz.

[0078] Each processor separately obtains the first collection result of size B x LE z and the second collection result of size E z F y . In the third splitting direction (for example, the splitting direction corresponding to the third variable z), the first collection result of size B x LE z and the second collection result of size E z F y can be multiplied and summed, and the first multiplication and summation result of size B x LF y is distributed to each processor (as Figure 5The fully-reduced multiply-accumulate operation shown), so that each processor stores a first multiply-sum result of size B x LF y The third traffic volume corresponding to this multiply-sum distribution operation is 2BLF / xy.

[0079] By summing the first traffic volume of BLE / xz, the second traffic volume of EF / yz, and the third traffic volume of 2BLF / xy, the traffic volume of the first-layer matrix product can be obtained.

[0080] In a possible implementation, the traffic volume of the second-layer matrix product is obtained based on a fourth traffic volume and a fifth traffic volume. The fourth traffic volume is the traffic volume of a third collection operation on the second parameter sub-matrix on each processor in the first splitting direction; the fifth traffic volume is the traffic volume of multiplying and summing the first multiply-sum result and the third collection result on each processor in the second splitting direction, and splitting the second multiply-sum result of the first multiply-sum result and the third collection result to each processor.

[0081] For example, as Figure 7 shown, a third collection operation (see the collection operation shown) can be performed on the second parameter sub-matrix of size F in the first splitting direction (for example, the splitting direction corresponding to the first variable x), so that each processor in the first splitting direction stores a third collection result of size F y E zx The fourth traffic volume for this step is EF / yz. Figure 2 shown, a third collection operation (see the collection operation shown) can be performed on the second parameter sub-matrix of size F in the first splitting direction (for example, the splitting direction corresponding to the first variable x), so that each processor in the first splitting direction stores a third collection result of size F y E z The fourth traffic volume for this step is EF / yz.

[0082] Then, in the second splitting direction (for example, the splitting direction corresponding to the second variable y), the first multiply-sum result of size B x LF y and the third collection result of size F y E z can be multiplied and summed, and the second multiply-sum result of size B x LE z can be split to each processor (such as the reduced-divergent multiply-accumulate operation shown), which is equivalent to evenly splitting the second multiply-sum result of size B Figure 6 shown, which is equivalent to evenly splitting the second multiply-sum result of size B x LE z after multiplication along the second splitting direction, so that each processor obtains the split target result B x L y E zIt is still three-dimensionally split, facilitating the next round of matrix multiplication. The fifth traffic volume corresponding to this multiplication and summation splitting operation is BLE / xz. Among them, the target results of size B among n processors x L y E z can be used to determine the product result of the two-layer matrix product.

[0083] By summing the fourth traffic volume EF / yz and the fifth traffic volume BLE / xz, the traffic volume of the second-layer matrix product can be obtained.

[0084] In this way, by adding these five traffic volumes corresponding to the two-layer matrix product, the objective function shown in formula (2) can be deduced.

[0085] In a possible implementation manner, the splitting scheme includes the first variable x, the second variable y, and the third variable z that minimize the objective function. By minimizing the objective function, it is beneficial to efficiently and quickly determine the first variable x, the second variable y, and the third variable z that minimize the total traffic volume T among multiple processors comm that minimize.

[0086] When actually determining the specific splitting scheme, since the first variable x, the second variable y, and the third variable z are integers and need to satisfy that the product of the first variable x, the second variable y, and the third variable z is equal to the number of processors n, that is: xyz = n, by performing a minimum search on the possible integer values of the first variable x, the second variable y, and the third variable z for the total traffic volume T comm minimum search, that is, by minimizing the objective function shown in formula (2), the first variable x, the second variable y, and the third variable z are solved.

[0087] Among them, in actual applications, the optimal integer values of the first variable x, the second variable y, and the third variable z are near a certain real value, and a search can be performed near it to accelerate the search speed.

[0088] In actual applications, the deduced objective function can be pre-written into the memory so that in the case of a two-layer matrix product operation based on big data, in step S11, through any one of the multiple processors, the splitting scheme is obtained by solving the objective function read from the memory, and the input matrix, the first parameter matrix, and the second parameter matrix are split based on the splitting scheme, and the splitting results are input to the multiple processors. In step S11, when the splitting results are input to the multiple processors, in step S12, the multiple processors calculate the products of the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting results. Among them, the splitting results include input sub-matrices, first parameter sub-matrices, and second parameter sub-matrices.

[0089] Take, for example, the two matrix multiplications of a feed-forward network operation in the prefill part during the inference operation of a large-scale neural network model (such as including the ChatGLM-6B language model). Assume that the input matrix X is language feature data, the size of the first dimension of the input matrix X is B = 8, the size of the second dimension is E = 4096, and the size of the third dimension is L = 2048. Among them, B can represent the batch, L can represent the length of the language feature, and E can represent the number of channels; and, the first parameter matrix W 1 has a size of E = 4096 in the first dimension and a size of F = 4096×4 in the second dimension, and the second parameter matrix W 2 has a size of F = 4096×4 in the first dimension and a size of E = 4096 in the second dimension.

[0090] Single-precision (FP32) inference can be performed on n = 256 GPU cards. Assume that the best splitting scheme obtained by searching through step S11 is: the first variable x = 8, the second variable y = 16, and the third splitting variable z = 2.

[0091] In this way, since the size of the input matrix X is 8×2048×4096, it is evenly split onto 8×16×2 = 256 GPU cards. After splitting, the size of the input sub-matrix on each GPU card is 1×128×2048. Similarly, since the first parameter matrix W 1 has a size of 4096×(4096×4), the size of the first parameter sub-matrix split onto each GPU card is 2048×128; since the size of the second parameter matrix W2 is (4096×4)×4096, the size of the second parameter sub-matrix split onto each GPU card is 1024×256.

[0092] In step S12, it can be according to the schematic diagram of the inference process of the objective function as shown in Figure 7 . Based on the input sub-matrix, the first parameter sub-matrix, and the second parameter sub-matrix, 256 GPU cards calculate the product of the input matrix X, the first parameter matrix W 1 , and the second parameter matrix W 2 .

[0093] A first collection operation can be performed on an input sub - matrix of size 1×128×2048 along a second splitting direction (the splitting direction corresponding to the second variable y), so that a first collection result of size 1×2048×2048 is obtained on each GPU card; and, a second collection is performed on a first parameter sub - matrix of size 2048×128 along a first splitting direction (the splitting direction corresponding to the first variable x), so that a second collection result of size 2048×1024 is obtained on each GPU card. Then, a full reduction multiply - add operation can be performed on the first collection result of size 1×2048×2048 and the second collection result of size 2048×1024 along a third splitting direction (the splitting direction corresponding to the third variable z), so that each GPU card obtains a first multiply - sum result of size 1×2048×1024.

[0094] A third collection operation is performed on a second parameter sub - matrix of size 1024×256 along the first splitting direction (the splitting direction corresponding to the first variable x), so that a third collection result of size 1024×2048 is obtained on each GPU card. Then, a reduction - divergence multiply - add operation can be performed on the first multiply - sum result of size 1×2048×1024 and the third collection result of size 1024×2048 along the second splitting direction (the splitting direction corresponding to the second variable y), so that each GPU card obtains a target result of size 1×128×2048. Based on the target result of each GPU card, the input matrix X and the first parameter matrix W 1 and the second parameter matrix W 2 of the two - layer matrix multiplication result.

[0095] Next, the communication complexity of the data processing method of the present disclosure embodiment is analyzed. Since the product of the first variable x, the second variable y, and the third variable z is equal to the number of processors n, xyz = n can be used to transform the objective function in formula (2), and the objective function shown in formula (3) is obtained:

[0096] T comm =BLE / xz + EF / yz + 2BLF / xy + EF / yz + BLE / xz = 2 / n(BLEy + EFx + BLFz) (3)

[0097] In formula (3), T comm represents the total communication volume between n processors, BLE represents the size of the input matrix X, and EF represents the size of the first parameter matrix W 1 (or the second parameter matrix W 2 ).

[0098] Since formula (3) needs to satisfy the constraint condition xyz = n, the constraint condition and the objective function can be rewritten using logarithmic transformation.

[0099] a=ln(EFx),b=ln(BLEy),c=ln(BLFz) (4)

[0100] In formula (4), a represents the total communication volume T in formula (3) comm The second term EFx takes the logarithm, b represents the total communication volume T in formula (3) comm The first term BLEy takes the logarithm, c represents the total communication volume T in formula (3) comm The third term BLFz of takes the logarithm. Substituting formula (4) into formula (3), the objective function becomes:

[0101] T comm =2 / n (e a +e b +e c ) (5)

[0102] Among them, a+b+c=ln(B 2 L 2 E 2 F 2 n), using the convex function Jessen inequality, we can know that the total communication volume T comm The minimum value is obtained when a=b=c, and its value is: min(T comm )=6(BLEF) 2 / 3 (1 / n 2 / 3 ), where min is the minimization function, B represents the size of the input matrix X in the first dimension, L represents the size of the input matrix X in the second dimension, and E represents the size of the input matrix X in the third dimension, i.e., the first parameter matrix W 1 In the first dimension, the second parameter matrix W 2 In the second dimension, F represents the first parameter matrix W 1 The size of the second dimension, that is, the second parameter matrix W 2 Size in the first dimension.

[0103] It can be seen that the minimum complexity of the total communication volume of data processing in the embodiment of the present disclosure for the processor cluster size n is O(1 / n 2 / 3 ), with a communication volume that is inversely proportional to 2 / 3 of the number of processors.

[0104] In the related technology, for example, the split and collect calculation solution from Google, whether it is two-dimensional splitting or one-dimensional splitting, plus its corresponding matrix collection method, its optimal communication volume is 8BLE×1 / n 1 / 2 The minimum complexity of the optimal communication volume for the processor cluster size n is O(1 / n 1 / 2 ).

[0105] The optimal complexity O(1 / n) of the traffic volume of the data processing method according to the embodiments of the present disclosure 2 / 3 ) is better than the optimal complexity O(1 / n) in the related art 1 / 2 ), and the overall traffic volume of the data processing method according to the embodiments of the present disclosure is one order of magnitude smaller than that of the related art.

[0106] Figure 8 The schematic diagram showing the comparison of the total traffic volume of the data processing method according to the embodiments of the present disclosure with the related art is as Figure 8 shown. When the first dimension size B of the input matrix X is 8, the second dimension size L of the input matrix X is 2048, and the third dimension size E of the input matrix X is 4096 (that is, the size of the first parameter matrix W 1 in the first dimension, and the size of the second parameter matrix W 2 in the second dimension), the first parameter matrix W 1 The second dimension size F is 4096×4 (that is, the size of the second parameter matrix W 2 in the first dimension), the total traffic volume curve of the data processing method according to the embodiments of the present disclosure, and the total traffic volume curve in the related art (such as the two-dimensional splitting method) change with the increase of the processor data n. As Figure 8 shown, when the number of processors exceeds 5, the data processing method according to the embodiments of the present disclosure can produce advantages. When the number of processors exceeds 256 and reaches a large scale, the advantages of the data processing method according to the embodiments of the present disclosure are more obvious.

[0107] In summary, in the data processing method according to the embodiments of the present disclosure, the splitting scheme obtained by solving the objective function can utilize multiple processors to process larger-scale models (such as neural network models with a large number of parameters) and reduce the communication overhead between multiple processors.

[0108] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not elaborate. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.

[0109] In addition, the present disclosure also provides a data processing device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any data processing method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method part and will not be elaborated here.

[0110] Figure 9 The block diagram showing the data processing device according to the embodiments of the present disclosure is as Figure 9As shown, the data processing device is applied to multiple processors, and the device includes:

[0111] A splitting module 91: configured to split an input matrix, a first parameter matrix, and a second parameter matrix based on a splitting scheme by any one of the multiple processors, and input the splitting results into the multiple processors, where the splitting scheme represents the number of splitting parts of each dimension of the input matrix, the first parameter matrix, and the second parameter matrix;

[0112] A calculation module 92, configured to calculate the product of the input matrix, the first parameter matrix, and the second parameter matrix by the multiple processors based on the splitting results;

[0113] Wherein, the splitting scheme is obtained by solving an objective function, the objective function represents the relationship between the total communication volume and the splitting scheme, and the total communication volume is the total communication volume between multiple processors during the process of calculating the matrix product of the input matrix, the first parameter matrix, and the second parameter matrix by the multiple processors.

[0114] In a possible implementation manner, the total communication volume is obtained based on the communication volume of the first-layer matrix product and the communication volume of the second-layer matrix product, where the first-layer matrix product is the product of the input matrix and the first parameter matrix, and the second-layer matrix product is the product of the product result of the input matrix and the first parameter matrix and the second parameter matrix.

[0115] In a possible implementation manner, the input matrix is a three-dimensional matrix, the first parameter matrix and the second parameter matrix are two-dimensional matrices, the number of elements in the third dimension of the input matrix, the first dimension of the first parameter matrix, and the second dimension of the second parameter matrix are the same, and the number of elements in the second dimension of the first parameter matrix and the first dimension of the second parameter matrix are the same; the objective function includes a first variable, a second variable, and a third variable, and the product of the first variable, the second variable, and the third variable is equal to the number of processors, and the splitting scheme is represented by the first variable, the second variable, and the third variable; the first variable is used to characterize the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the first splitting direction, and the first splitting direction includes the direction of the first dimension of the input matrix; the second variable is used to characterize the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the second splitting direction, and the second splitting direction includes the direction of the second dimension of the input matrix; the third variable is used to characterize the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the third splitting direction, and the third splitting direction includes the direction of the third dimension of the input matrix.

[0116] In a possible implementation, the splitting module 91 is configured to: split the input matrix, the first parameter matrix, and the second parameter matrix in the first splitting direction, the second splitting direction, and the third splitting direction according to the first variable, the second variable, and the third variable respectively, so that each processor corresponds to a different input sub-matrix, first parameter sub-matrix, and second parameter sub-matrix.

[0117] In a possible implementation, the communication volume of the first-layer matrix multiplication is obtained based on a first communication volume, a second communication volume, and a third communication volume, where: the first communication volume is the communication volume of performing a first collection operation on the input sub-matrices on each processor in the second splitting direction; the second communication volume is the communication volume of performing a second collection operation on the first parameter sub-matrices on each processor in the first splitting direction; the third communication volume is the communication volume of multiplying and summing the first collection result and the second collection result on each processor in the third splitting direction, and distributing the first multiplication and summation result of the first collection result and the second collection result to each processor, where the first collection result is the result obtained by performing the first collection operation, and the second collection result is the result obtained by performing the second collection operation.

[0118] In a possible implementation, the communication volume of the second-layer matrix multiplication is obtained based on a fourth communication volume and a fifth communication volume, where: the fourth communication volume is the communication volume of performing a third collection operation on the second parameter sub-matrices on each processor in the first splitting direction; the fifth communication volume is the communication volume of multiplying and summing the first multiplication and summation result and the third collection result on each processor in the second splitting direction, and splitting the second multiplication and summation result of the first multiplication and summation result and the third collection result to each processor, where the third collection result is the result obtained by performing the third collection operation.

[0119] In a possible implementation, the splitting scheme includes the first variable, the second variable, and the third variable that minimize the objective function.

[0120] In a possible implementation, the input matrix includes feature data in a deep learning task, and the feature data includes at least one of image feature data, voice feature data, and text feature data.

[0121] This method has a specific technical association with the internal structure of a computer system and can solve technical problems such as how to improve the operation efficiency or execution effect of hardware (including reducing the amount of data storage, reducing the amount of data transmission, and increasing the hardware processing speed), thereby obtaining a technical effect of improving the internal performance of the computer system in line with natural laws.

[0122] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0123] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0124] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the above methods.

[0125] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above methods.

[0126] The electronic device can be provided as a terminal, a server or other forms of devices. For example, it includes user equipment (UE), mobile devices, user terminals, terminals, cellular phones, cordless phones, personal digital assistants (PDAs), handheld devices, computing devices, in-vehicle devices, wearable devices, etc.

[0127] Figure 10 The block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to Figure 10 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to execute the above methods.

[0128] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows Server TM ), the graphical user interface-based operating system launched by Apple Inc. (Mac OS X TM ), the multi-user and multi-process computer operating system (Unix TM ), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM ), or the like.

[0129] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.

[0130] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0131] The computer-readable storage medium may be a tangible device that can retain and store instructions used by an instruction execution device. The computer-readable storage medium may be, for example, (but is not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0132] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0133] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0134] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0135] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0136] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0137] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0138] The computer program product may be implemented specifically by hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is embodied as a computer storage medium. In another alternative embodiment, the computer program product is embodied as a software product, such as a Software Development Kit (SDK), etc.

[0139] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or likenesses may be referred to each other. For the sake of brevity, they are not repeated herein.

[0140] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.

[0141] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0142] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A data processing method, characterized in that: The data processing method is applied to multiple processors, and the method comprises: Splitting the input matrix, the first parameter matrix, and the second parameter matrix based on a splitting scheme by any one of the multiple processors, and inputting the splitting results to the multiple processors, wherein the splitting scheme indicates the number of splits of each dimension of the input matrix, the first parameter matrix, and the second parameter matrix; The multiple processors calculate the product of the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting result; Among them, the splitting scheme is obtained by solving the objective function, and the objective function represents the relationship between the total communication volume and the splitting scheme. The total communication volume is the total communication volume between multiple processors during the process of multiple processors calculating the matrix product of the input matrix, the first parameter matrix, and the second parameter matrix.

2. The method according to claim 1, characterized in that The total communication volume is obtained based on the communication volume of the first layer matrix product and the communication volume of the second layer matrix product, wherein the first layer matrix product is the product of the input matrix and the first parameter matrix, and the second layer matrix product is the product of the product of the input matrix, the first parameter matrix and the second parameter matrix.

3. The method according to claim 2, characterized in that The input matrix is ​​a three-dimensional matrix, the first parameter matrix and the second parameter matrix are two-dimensional matrices, the third dimension of the input matrix, the first dimension of the first parameter matrix, and the second dimension of the second parameter matrix have the same number of elements, and the second dimension of the first parameter matrix and the first dimension of the second parameter matrix have the same number of elements; The objective function includes a first variable, a second variable, and a third variable, the product of the first variable, the second variable, and the third variable is equal to the number of processors, and the splitting scheme is represented by the first variable, the second variable, and the third variable; The first variable is used to represent the number of splits of the input matrix, the first parameter matrix, and the second parameter matrix in a first splitting direction, respectively, and the first splitting direction includes the direction of the first dimension of the input matrix; The second variable is used to represent the number of splits of the input matrix, the first parameter matrix, and the second parameter matrix in a second splitting direction, respectively, and the second splitting direction includes the direction of the second dimension of the input matrix; The third variable is used to characterize the number of splits of the input matrix, the first parameter matrix, and the second parameter matrix in a third splitting direction, respectively. The third splitting direction includes the direction of the third dimension of the input matrix.

4. The method according to claim 3, characterized in that The step of splitting the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting scheme by any one of the multiple processors includes: In the first splitting direction, the second splitting direction, and the third splitting direction, the input matrix, the first parameter matrix, and the second parameter matrix are split according to the first variable, the second variable, and the third variable, respectively, so that each processor corresponds to a different input sub-matrix, first parameter sub-matrix, and second parameter sub-matrix, respectively.

5. The method according to claim 4, characterized in that The communication volume of the first layer matrix product is obtained based on the first communication volume, the second communication volume, and the third communication volume, wherein: The first communication volume is the communication volume of performing a first collection operation on the input sub-matrix on each processor in the second splitting direction; The second communication volume is the communication volume of performing a second collection operation on the first parameter submatrix on each processor in the first splitting direction; The third communication volume is the communication volume that multiplies and sums the first collection result and the second collection result on each processor in the third split direction, and distributes the first multiplication and sum result of the first collection result and the second collection result to each processor, wherein the first collection result is the result obtained by executing the first collection operation, and the second collection result is the result obtained by executing the second collection operation.

6. The method according to claim 5, characterized in that The communication volume of the second layer matrix product is obtained based on the fourth communication volume and the fifth communication volume, wherein: The fourth communication volume is the communication volume of performing a third collection operation on the second parameter submatrix on each processor in the first splitting direction; The fifth communication volume is the communication volume of multiplying and summing the first multiplication and summation result and the third collection result on each processor in the second splitting direction, and dividing the second multiplication and summation result of the first multiplication and summation result and the third collection result to each processor, wherein the third collection result is the result obtained by executing the third collection operation.

7. The method according to any one of claims 3 to 6, characterized in that The splitting scheme includes the first variable, the second variable, and the third variable that minimize the objective function.

8. The method according to any one of claims 1 to 6, characterized in that The input matrix includes feature data in a deep learning task, and the feature data includes at least one of image feature data, speech feature data, and text feature data.

9. A data processing device, characterized in that: The data processing device is applied to multiple processors, and the device includes: A splitting module, configured to split the input matrix, the first parameter matrix, and the second parameter matrix based on a splitting scheme through any processor among the multiple processors, and input the splitting results to the multiple processors, wherein the splitting scheme indicates the number of splits of each dimension of the input matrix, the first parameter matrix, and the second parameter matrix; A calculation module, configured for the multiple processors to calculate the product of the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting result; Among them, the splitting scheme is obtained by solving the objective function, and the objective function represents the relationship between the total communication volume and the splitting scheme. The total communication volume is the total communication volume between multiple processors during the process of multiple processors calculating the matrix product of the input matrix, the first parameter matrix, and the second parameter matrix.

10. An electronic device, characterized in that: include: Multiple processors; a memory for storing a plurality of processor executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 8.

11. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Communication optimization method and system for distributed deep learning operator parallel training

    CN115996173A

  • Parallel processing method and device of model, first computing equipment and electronic equipment

    CN116820577A

  • Data processing method and device, electronic equipment, storage medium and product

    CN118069742A

  • Method, apparatus and computer program to carry out a training procedure in a convolutional neural network

    US20200125933A1