Data Processing Method, Apparatus, Electronic Device, and Storage Medium
By solving the objective function between multiple processors and uniformly splitting the input and parameter matrix, the problem of large communication overhead of large artificial intelligence models is solved, and more efficient system performance is achieved.
Patent Information
- Application Number
- CN202510521064.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-24
AI Technical Summary
When the prior art splits large artificial intelligence models, there is a problem of excessive communication overhead, which leads to a degradation of system performance, especially when increasing the number of processors, the increase in communication volume is unnecessary.
By solving the objective function, the input matrix, the first parameter matrix and the second parameter matrix are uniformly split between multiple processors, and the large-scale neural network model is calculated using multiple processors, and the communication overhead between processors is reduced.
The communication volume between multiple processors is effectively reduced and system performance is improved, especially when the number of processors increases, the communication overhead is reduced by an order of magnitude, which is better than the prior art.
Smart Images

Figure CN120045338B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a data processing method, an apparatus, an electronic device, and a storage medium. Background Art
[0002] An artificial intelligence (AI) large model refers to a huge and complex neural network that needs to store more parameters to increase the depth and width of the model, thereby improving the performance of the model. The application scenarios of artificial intelligence (AI) large models are very extensive and can cover fields such as healthcare, finance, industry, education, and smart cities. For the inference of large models, special processors can be used for acceleration, such as including a Graphics Processing Unit (GPU). Such processors often have expensive storage, resulting in a single processor being unable to store the entire model completely. To reduce hardware costs, the large model can be processed by distillation, pruning, quantization, and splitting. Among them, compared with other processing methods, the splitting method is completely lossless, but splitting brings additional communication overhead. Summary of the Invention
[0003] The present disclosure proposes a technical solution for data processing.
[0004] According to one aspect of the present disclosure, a data processing method is provided. The data processing method is applied to multiple processors. The method includes: splitting an input matrix, a first parameter matrix, and a second parameter matrix based on a splitting scheme by any one of the multiple processors, and inputting the splitting result into the multiple processors. The splitting scheme represents the number of splitting parts of each dimension of the input matrix, the first parameter matrix, and the second parameter matrix; the multiple processors calculate the product of the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting result; wherein, the splitting scheme is obtained by solving an objective function, and the objective function represents the relationship between the total communication volume and the splitting scheme. The total communication volume is the total communication volume between multiple processors during the process of calculating the matrix product of the input matrix, the first parameter matrix, and the second parameter matrix by the multiple processors.
[0005] In a possible implementation manner, the total communication volume is obtained based on the communication volume of the first-layer matrix product and the communication volume of the second-layer matrix product, where the first-layer matrix product is the product of the input matrix and the first parameter matrix, and the second-layer matrix product is the product of the product result of the input matrix and the first parameter matrix and the second parameter matrix.
[0006] In a possible implementation, the input matrix is a three-dimensional matrix, the first parameter matrix and the second parameter matrix are two-dimensional matrices, the number of elements in the third dimension of the input matrix, the first dimension of the first parameter matrix, and the second dimension of the second parameter matrix are the same, and the number of elements in the second dimension of the first parameter matrix and the first dimension of the second parameter matrix are the same; the objective function includes a first variable, a second variable, and a third variable, and the product of the first variable, the second variable, and the third variable is equal to the number of processors, and the splitting scheme is represented by the first variable, the second variable, and the third variable; the first variable is used to represent the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the first splitting direction, and the first splitting direction includes the direction of the first dimension of the input matrix; the second variable is used to represent the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the second splitting direction, and the second splitting direction includes the direction of the second dimension of the input matrix; the third variable is used to represent the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the third splitting direction, and the third splitting direction includes the direction of the third dimension of the input matrix.
[0007] In a possible implementation, splitting the input matrix, the first parameter matrix, and the second parameter matrix by any one of the multiple processors based on the splitting scheme includes: in the first splitting direction, the second splitting direction, and the third splitting direction, splitting the input matrix, the first parameter matrix, and the second parameter matrix according to the first variable, the second variable, and the third variable respectively, so that each processor corresponds to different input sub-matrices, first parameter sub-matrices, and second parameter sub-matrices.
[0008] In a possible implementation, the communication volume of the first-layer matrix multiplication is obtained based on a first communication volume, a second communication volume, and a third communication volume, where: the first communication volume is the communication volume of performing a first collection operation on the input sub-matrices on each processor in the second splitting direction; the second communication volume is the communication volume of performing a second collection operation on the first parameter sub-matrices on each processor in the first splitting direction; the third communication volume is the communication volume of multiplying and summing the first collection result and the second collection result on each processor in the third splitting direction, and distributing the first multiplication and summation result of the first collection result and the second collection result to each processor, where the first collection result is the result obtained by performing the first collection operation, and the second collection result is the result obtained by performing the second collection operation.
[0009] In a possible implementation, the communication volume of the second-layer matrix multiplication is obtained based on a fourth communication volume and a fifth communication volume, where: the fourth communication volume is the communication volume of performing a third collection operation on the second parameter sub-matrices on each processor in the first splitting direction; the fifth communication volume is the communication volume of multiplying and summing the first multiplication and summation result and the third collection result on each processor in the second splitting direction, and splitting the second multiplication and summation result of the first multiplication and summation result and the third collection result to each processor, where the third collection result is the result obtained by performing the third collection operation.
[0010] In a possible implementation, the splitting scheme includes the first variable, the second variable, and the third variable that minimize the objective function.
[0011] In a possible implementation, the input matrix includes feature data in a deep learning task, and the feature data includes at least one of image feature data, voice feature data, and text feature data.
[0012] According to one aspect of the present disclosure, a data processing device is provided. The data processing device is applied to multiple processors. The device includes: a splitting module, configured to split an input matrix, a first parameter matrix, and a second parameter matrix based on a splitting scheme through any one of the multiple processors, and input the splitting result to the multiple processors. The splitting scheme represents the splitting fractions of each dimension of the input matrix, the first parameter matrix, and the second parameter matrix; a calculation module, configured to calculate the product of the input matrix, the first parameter matrix, and the second parameter matrix by the multiple processors based on the splitting result; where the splitting scheme is obtained by solving an objective function, and the objective function represents the relationship between the total communication volume and the splitting scheme, and the total communication volume is the total communication volume between multiple processors during the process of calculating the matrix product of the input matrix, the first parameter matrix, and the second parameter matrix by the multiple processors.
[0013] In a possible implementation, the total communication volume is obtained based on the communication volume of the first-layer matrix multiplication and the communication volume of the second-layer matrix multiplication, where the first-layer matrix multiplication is the product of the input matrix and the first parameter matrix, and the second-layer matrix multiplication is the product of the product result of the input matrix and the first parameter matrix and the second parameter matrix.
[0014] In a possible implementation, the input matrix is a three-dimensional matrix, the first parameter matrix and the second parameter matrix are two-dimensional matrices, the number of elements in the third dimension of the input matrix, the first dimension of the first parameter matrix, and the second dimension of the second parameter matrix are the same, and the number of elements in the second dimension of the first parameter matrix and the first dimension of the second parameter matrix are the same; the objective function includes a first variable, a second variable, and a third variable, and the product of the first variable, the second variable, and the third variable is equal to the number of processors, and the splitting scheme is represented by the first variable, the second variable, and the third variable; the first variable is used to represent the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the first splitting direction, and the first splitting direction includes the direction of the first dimension of the input matrix; the second variable is used to represent the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the second splitting direction, and the second splitting direction includes the direction of the second dimension of the input matrix; the third variable is used to represent the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the third splitting direction, and the third splitting direction includes the direction of the third dimension of the input matrix.
[0015] In a possible implementation, the splitting module is configured to: in the first splitting direction, the second splitting direction, and the third splitting direction, split the input matrix, the first parameter matrix, and the second parameter matrix respectively according to the first variable, the second variable, and the third variable, so that each processor corresponds to different input sub-matrices, first parameter sub-matrices, and second parameter sub-matrices.
[0016] In a possible implementation, the communication volume of the first-layer matrix multiplication is obtained based on a first communication volume, a second communication volume, and a third communication volume, where: the first communication volume is the communication volume of performing a first collection operation on the input sub-matrix on each processor in the second splitting direction; the second communication volume is the communication volume of performing a second collection operation on the first parameter sub-matrix on each processor in the first splitting direction; the third communication volume is the communication volume of multiplying and summing the first collection result and the second collection result on each processor in the third splitting direction, and distributing the first multiplication and summation result of the first collection result and the second collection result to each processor, where the first collection result is the result obtained by performing the first collection operation, and the second collection result is the result obtained by performing the second collection operation.
[0017] In a possible implementation, the communication volume of the second-layer matrix multiplication is obtained based on a fourth communication volume and a fifth communication volume, where: the fourth communication volume is the communication volume of performing a third collection operation on the second parameter sub-matrices on each processor in the first splitting direction; the fifth communication volume is the communication volume of multiplying and summing the first multiplication and summation result and the third collection result on each processor in the second splitting direction, and splitting the second multiplication and summation result of the first multiplication and summation result and the third collection result to each processor, where the third collection result is the result obtained by performing the third collection operation.
[0018] In a possible implementation, the splitting scheme includes the first variable, the second variable, and the third variable that minimize the objective function.
[0019] In a possible implementation, the input matrix includes feature data in a deep learning task, and the feature data includes at least one of image feature data, speech feature data, and text feature data.
[0020] According to one aspect of the present disclosure, an electronic device is provided, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to call the instructions stored in the memory to execute the above method.
[0021] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.
[0022] The data processing method of the embodiments of the present disclosure can be applied to multiple processors. The method includes: through any one of the multiple processors, splitting an input matrix, a first parameter matrix, and a second parameter matrix based on a splitting scheme, and inputting the splitting result into the multiple processors, where the splitting scheme represents the number of splitting parts of each dimension of the input matrix, the first parameter matrix, and the second parameter matrix; the multiple processors calculate the product of the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting result; wherein, the splitting scheme is obtained by solving an objective function, and the objective function represents the relationship between the total communication volume and the splitting scheme, and the total communication volume is the total communication volume between the multiple processors during the process of the multiple processors calculating the matrix product of the input matrix, the first parameter matrix, and the second parameter matrix. The splitting scheme obtained by solving the objective function can utilize multiple processors to process larger-scale models (such as neural network models with a large number of parameters) and reduce the communication overhead between the multiple processors.
[0023] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Description of the Drawings
[0024] The accompanying drawings herein are incorporated into the specification and form a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0025] Figure 1 A flowchart showing a data processing method according to an embodiment of the present disclosure.
[0026] Figure 2 A schematic diagram showing a collection operation according to an embodiment of the present disclosure.
[0027] Figure 3 A schematic diagram showing a reduction-scatter summation operation according to an embodiment of the present disclosure.
[0028] Figure 4 A schematic diagram showing a full reduction summation operation according to an embodiment of the present disclosure.
[0029] Figure 5 A schematic diagram showing a full reduction multiply-accumulate operation according to an embodiment of the present disclosure.
[0030] Figure 6 A schematic diagram showing a reduction-scatter multiply-accumulate operation according to an embodiment of the present disclosure.
[0031] Figure 7 A schematic diagram showing the inference process of an objective function according to an embodiment of the present disclosure.
[0032] Figure 8 A schematic diagram comparing the total communication volume of a data processing method according to an embodiment of the present disclosure with related technologies.
[0033] Figure 9 A block diagram showing a data processing apparatus according to an embodiment of the present disclosure.
[0034] Figure 10 A block diagram showing an electronic device according to an embodiment of the present disclosure. Detailed Embodiments
[0035] The various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Like reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0036] As used herein, the term "exemplary" means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein need not be construed as superior or better than other embodiments.
[0037] As used herein, the term "and / or" is merely a description of an associated relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set consisting of A, B, and C.
[0038] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can still be implemented without some specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0039] After splitting an artificial intelligence model (such as a neural network module) into different processors, the communication volume between the processors becomes an indicator affecting system performance. However, for a large artificial intelligence model, as the number of processors increases, the amount of computation on each processor will decrease linearly, but the communication volume will instead increase.
[0040] The splitting methods in related technologies include pipeline parallelism by splitting according to model layers, model parallelism by directly splitting parameter matrices, etc. Among them, the model parallelism method can reduce the inference latency and can be divided into one-dimensional parallelism and two-dimensional parallelism. For example, based on ring-based splitting and collection, the communication volume can be made fixed through a one-dimensional matrix splitting algorithm. However, as the number of communication devices increases, the fixed handshake communication volume will gradually increase, and the fixed communication volume will still drag down the overall system performance. The method based on two-dimensional matrix splitting and matrix information collection can further reduce the communication volume, but this communication volume is still not optimal. If the number of processors n is considered as a variable, the minimum communication volume in related technologies is inversely proportional to the square root of n, and this magnitude is obviously not optimal. In view of this, the embodiments of the present disclosure propose a data processing method that can effectively reduce the communication overhead between multiple processors.
[0041] Figure 1 A flowchart showing the data processing method according to an embodiment of the present disclosure is as Figure 1 shown, and the data processing method includes:
[0042] In step S11, any one of the multiple processors splits the input matrix, the first parameter matrix, and the second parameter matrix based on a splitting scheme, and inputs the splitting results into the multiple processors. The splitting scheme represents the number of splitting parts for each dimension of the input matrix, the first parameter matrix, and the second parameter matrix.
[0043] In step S12, the multiple processors calculate the product of the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting results.
[0044] Among them, the splitting scheme is obtained by solving an objective function, which represents the relationship between the total communication volume and the splitting scheme. The total communication volume is the total communication volume between multiple processors during the process of the multiple processors calculating the matrix product of the input matrix, the first parameter matrix, and the second parameter matrix.
[0045] In the data processing method of the embodiments of the present disclosure, the splitting scheme obtained by solving the objective function can utilize multiple processors to process larger-scale models (such as neural network models with a large number of parameters) and reduce the communication overhead between multiple processors.
[0046] In a possible implementation manner, the data processing method of the embodiments of the present disclosure can be applied to multiple processors. Each processor can be a newly designed one or an improved one obtained from an existing processor chip. The types of processor chips can include but are not limited to: Central Processing Unit (CPU), Graphic Processing Unit (GPU), General-Purpose Computing on Graphics Processing Units (GPGPU), Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA), or other programmable logic devices, and can also include a microprocessor or a processor of other conventional processors.
[0047] In a possible implementation manner, in large-scale calculations based on large models (such as neural network models with a large number of parameters), the two-layer matrix product occupies a large amount of computing resources. The two-layer matrix product can be expressed as:
[0048] Y = act(XW1)W2 (1)
[0049] In formula (1), act represents an element-wise activation function, which can be a function running on the neurons of a neural network, responsible for mapping the input of a neuron to the output end and increasing the non-linearity of the neural network model. For example, it includes functions such as Gaussian Error Linear Units (gelu) function, Sigmoid function, etc. X represents the input matrix, which can be a three-dimensional matrix with dimensions of B×L×E (abbreviated as BLE). W1 represents the first parameter matrix, whose dimensions can be E×F (abbreviated as EF). W2 represents the second parameter matrix, whose dimensions can be F×E (abbreviated as FE). Y represents the result of the product of two matrices, and its dimensions can be the same as those of the input matrix X, also B×L×E. Among them, the specific values of B, L, E, and F in the embodiments of the present disclosure are not limited and can be set according to the actual application scenario.
[0050] In the example, the input matrix X is a three-dimensional matrix, the first parameter matrix W1 and the second parameter matrix W2 are two-dimensional matrices (for example, the input matrix X has dimensions of B×L×E, the first parameter matrix W1 has dimensions of E×F, and the second parameter matrix W2 has dimensions of F×E). The number of elements in the third dimension of the input matrix X (for example, the dimension where the dimension E of the input matrix X is located), the first dimension of the first parameter matrix W1 (for example, the dimension where the dimension E of the first parameter matrix W1 is located), and the second dimension of the second parameter matrix W2 (for example, the dimension where the dimension E of the second parameter matrix W2 is located) are the same. The number of elements in the second dimension of the first parameter matrix W1 (for example, the dimension where the dimension F of the first parameter matrix W1 is located) and the first dimension of the second parameter matrix W2 (for example, the dimension where the dimension F of the second parameter matrix W2 is located) are the same.
[0051] In the example, the input matrix includes feature data in a deep learning task, and the feature data includes at least one of image feature data, voice feature data, and text feature data. For example, in the scenario of using a deep neural network for face recognition of a target object, the input matrix can be the image feature data of the target object (such as a face feature map); in the scenario of using a deep neural network for speech recognition of a target object, the input matrix can be the voice feature data of the target object; in the scenario of using a deep neural network for character recognition of a target document, the input matrix can be the text feature data of the target document; the embodiments of the present application do not limit the type of the input matrix.
[0052] For large-scale computing based on large models, since a single processor cannot store the entire model completely, in step S11, any one of the multiple processors can read the objective function for determining the splitting scheme from the memory, and obtain the splitting scheme by solving the objective function. The splitting scheme can represent the number of splitting parts of the input matrix X, the first parameter matrix W1, and the second parameter matrix W2 in each dimension. Then, any one of the multiple processors can split the input matrix X along the three dimensions of BLE (for example, B corresponds to the horizontal direction, L corresponds to the vertical direction, and E corresponds to the depth direction) according to the splitting scheme, split the first parameter matrix W1 and the second parameter matrix W2 along the two dimensions of EF, and evenly and redundantly distribute the splitting results to multiple processors, so that each processor corresponds to an input sub-matrix, a first parameter sub-matrix, and a second parameter sub-matrix.
[0053] In a possible implementation, the objective function for determining the splitting scheme can represent the relationship between the total communication volume and the splitting scheme. The total communication volume is the total communication volume between n (n>1) processors during the calculation of the matrix product of the input matrix X, the first parameter matrix W1, and the second parameter matrix W2. For example, the objective function can be expressed as:
[0054] T comm =BLE / xz + EF / yz + 2BLF / xy + EF / yz + BLE / xz (2)
[0055] In formula (2), T comm represents the total communication volume between n processors during the calculation of formula (1). BLE represents the size of the input matrix X, EF represents the size of the first parameter matrix W1 (or the second parameter matrix W2). The objective function includes a first variable x, a second variable y, and a third variable z. The product of the first variable x, the second variable y, and the third variable z is equal to the number of processors n, that is, xyz = n. The splitting scheme is represented by the first variable x, the second variable y, and the third variable z;
[0056] Among them, the first variable x is used to characterize the number of splitting parts of the input matrix X, the first parameter matrix W1, and the second parameter matrix W2 in the first splitting direction. The first splitting direction includes the direction of the first dimension of the input matrix X, for example, the direction of the dimension where the size B of the input matrix X is located.
[0057] The second variable y is used to characterize the number of splitting parts of the input matrix X, the first parameter matrix W1, and the second parameter matrix W2 in the second splitting direction. The second splitting direction includes the direction of the second dimension of the input matrix X, for example, the direction of the dimension where the size L of the input matrix X is located.
[0058] The third variable z is used to represent the number of split parts of the input matrix X, the first parameter matrix W1, and the second parameter matrix W2 in the third splitting direction, and the third splitting direction includes the direction of the third dimension of the input matrix X, for example, the direction of the dimension where the size E of the input matrix X is located.
[0059] By solving the traffic-based objective function shown in formula (2) under the condition that xyz = n is satisfied, a splitting scheme can be obtained efficiently and quickly, which is beneficial to the subsequent uniform splitting of the input matrix X, the first parameter matrix W1, and the second parameter matrix W2 in three dimensions, making the number of split parts xyz the same as the number of processors n, so as to make more full use of each processor and reduce the traffic between processors.
[0060] In a possible implementation manner, step S11 may include: splitting the input matrix X, the first parameter matrix W1, and the second parameter matrix W2 in the first splitting direction, the second splitting direction, and the third splitting direction respectively according to the first variable x, the second variable y, and the third variable z, so that each processor corresponds to different input sub-matrices, first parameter sub-matrices, and second parameter sub-matrices.
[0061] In the example, assume that processor i is any one of the n processors. Processor i can solve the first variable x, the second variable y, and the third variable z through formula (2). And, according to the input matrix X, it can be known that the first splitting direction is the horizontal direction where the size B of the input matrix X is located, the second splitting direction is the vertical direction where the size L of the input matrix X is located, and the third splitting direction is the depth direction where the size E of the input matrix X is located.
[0062] Processor i can split the elements of the first dimension (the dimension where the size B is located) of the input matrix X into x parts in the horizontal direction, split the elements of the second dimension (the dimension where the size L is located) of the input matrix X into y parts in the vertical direction, and split the elements of the third dimension (the dimension where the size E is located) of the input matrix X into z parts in the depth direction. In total, the input matrix X with the size of BLE is split into xyz = n parts, and the size of each part is B x L y E z . Processor i can input the n input sub-matrices with the sizes of B x L y E z into the n processors respectively, so that each processor stores a different input sub-matrix.
[0063] Processor i can also split the elements of the first parameter matrix W1 in the depth direction into z parts along the first dimension (the dimension where the size E is located), and split the elements of the second dimension (the dimension where the size F is located) of the first parameter matrix W1 into yx parts in the vertical and horizontal directions, so that the first parameter matrix W1 with the size of EF is split into xyz = n parts in total, and the size of each part is E z F yx 。Processor i can input the n first parameter sub - matrices with the sizes of E z F yx into n processors respectively, so that each processor stores a different first parameter sub - matrix.
[0064] Similarly, processor i can split the elements of the second parameter matrix W2 in the vertical direction into y parts along the first dimension (the dimension where the size F is located), and split the elements of the second dimension (the dimension where the size E is located) of the second parameter matrix W2 into zx parts in the depth and horizontal directions, so that the second parameter matrix W2 with the size of FE is split into xyz = n parts in total, and the size of each part is F y E zx 。Processor i can input the n second parameter sub - matrices with the sizes of F y E zx into n processors respectively, so that each processor stores a different second parameter sub - matrix.
[0065] In this way, according to the first variable x, the second variable y, and the third variable z obtained by solving the objective function, the input matrix X, the first parameter matrix W1, and the second parameter matrix W2 can be evenly split in three splitting dimensions, making the number of splitting parts xyz the same as the number of processors n, which is beneficial to making full use of each processor and reducing the communication volume between processors.
[0066] The following introduces the reasoning process of the objective function, for example, including the process of three - dimensional splitting, collection, and calculation of the input matrix X, the first parameter matrix W1, and the second parameter matrix W2. To more clearly illustrate the reasoning process of the objective function, the following introduces several matrix communication operations and the overhead of each matrix communication operation.
[0067] Figure 2 FIG. shows a schematic diagram of the collection operation according to an embodiment of the present disclosure. As Figure 2 shown, if there are n processors, such as processor 1 to processor n, and each processor stores 1 / n of the matrix (see Figure 2 the upper part), the collection operation collects each part of the matrix on each processor to obtain the complete matrix (see Figure 2 the lower part). If the size of each part of the matrix is w / n, then the overhead of the data volume (i.e., the communication volume) for n processors to interact is w.
[0068] Figure 3 A schematic diagram showing a reduction-divergent summation operation according to an embodiment of the present disclosure. As Figure 3 shown, if there are n processors, such as processor 1 to processor n, each processor stores a different matrix (see Figure 3 the upper part), where identifiers 11 to 1n are used to distinguish different parts of the matrix stored by processor 1, identifiers 21 to 2n are used to distinguish different parts of the matrix stored by processor 2, and so on, identifiers n1 to nn are used to distinguish different parts of the matrix stored by processor n. The reduction-divergent summation operation calculates the sum of all matrices and evenly distributes the calculation results among the n processors according to a certain dimension (see Figure 3 the lower part). If the size of each matrix is w / n, then the overhead of data interaction among the n processors is approximately w / n.
[0069] Figure 4 A schematic diagram showing a full reduction summation operation according to an embodiment of the present disclosure. As Figure 4 shown, if there are n processors, such as processor 1 to processor n, each processor stores a different matrix (see Figure 4 the upper part), the full reduction summation operation calculates the sum of all matrices and distributes it to each of the n processors (see Figure 4 the lower part). If the size of each matrix is w / n, then the overhead of data interaction among the n processors is approximately 2w / n.
[0070] Figure 5 A schematic diagram showing a full reduction multiply-accumulate operation according to an embodiment of the present disclosure. As Figure 5 shown, if there are n processors, such as processor 1 to processor n, each processor stores two matrices to be multiplied (see Figure 5 the upper part), the full reduction multiply-accumulate operation calculates the product of the two matrices to be multiplied on each processor, then adds the products of the n processors, and distributes the multiply-sum result to each of the n processors (see Figure 5 the lower part). This process is equivalent to each processor separately calculating the matrix multiplication operation and then performing the full reduction summation operation as shown in Figure 4 shown. Therefore, if the size of each matrix is w / n, then the overhead of data interaction among the n processors is approximately 2w / n.
[0071] Figure 6 A schematic diagram showing a reduction-divergent multiply-accumulate operation according to an embodiment of the present disclosure. As Figure 6 shown, if there are n processors, such as processor 1 to processor n, each processor stores two matrices to be multiplied (see Figure 6Upper part), where identifiers 11 to n1 are used to distinguish different parts of a matrix stored in processor 1, and identifiers 11' to n1' are used to distinguish different parts of another matrix stored in processor 1; identifiers 12 to n2 are used to distinguish different parts of a matrix stored in processor 2, and identifiers 12' to n2' are used to distinguish different parts of another matrix stored in processor 2; and so on, identifiers 1n to nn are used to distinguish different parts of a matrix stored in processor n, and identifiers 1n' to nn' are used to distinguish different parts of another matrix stored in processor n. The reduction divergent multiply-add operation calculates the products of two products to be multiplied on each processor, then adds the products of the n processors, and then distributes the sum result of the product summation to each of the n processors after splitting according to a certain dimension (see Figure 6 Lower part). This process is equivalent to each processor calculating the matrix multiplication operation separately, and then performing the reduction divergent summation operation as shown in Figure 3 So if the size of each matrix is w / n, then the overhead of the data interaction volume of the n processors is approximately w / n.
[0072] Figure 7 A schematic diagram showing the reasoning process of the objective function according to an embodiment of the present disclosure, as shown in Figure 7 As shown, it is known that the size of the input matrix X is BLE, the size of the first parameter matrix W1 is EF, and the size of the second parameter matrix W2 is FE. It can be assumed that the first variable x, the second variable y, and the third variable z are used to represent that the input matrix X, the first parameter matrix W1, and the second parameter matrix W2 are split into x, y, and z parts respectively in 3 splitting directions (for example, including the first splitting direction, the second splitting direction, and the third splitting direction), so that each of the n processors can correspond to a sub-input matrix of size B x L y E z The input sub-matrix of, the size of the first parameter sub-matrix of E z F yx The first parameter sub-matrix of, the size of the second parameter sub-matrix of F y E zx Among them, xyz = n, so that the n processors can be logically mapped to a three-dimensional matrix, corresponding to x processors in the first splitting direction, y processors in the second splitting direction, and z processors in the third splitting direction.
[0073] In a possible implementation manner, as shown in Figure 7As shown, the total traffic is obtained based on the traffic of the first-layer matrix product and the traffic of the second-layer matrix product. The first-layer matrix product is the product of the input matrix X and the first parameter matrix W1; the second-layer matrix product is the product of the product result XW1 of the input matrix X and the first parameter matrix W1 and the second parameter matrix W2. In this way, by summing the traffic of the two-layer matrix product, the total traffic between multiple processors can be accurately determined.
[0074] Among them, as shown in formula (1), since an act operation (element-wise activation function operation) is performed after the matrix multiplication of the input matrix X and the first parameter matrix W1, and the act operation is performed separately on each processor without traffic. Therefore, the total traffic obtained based on the traffic of the first-layer matrix product and the traffic of the second-layer matrix product, that is, the total traffic for multiple processors to process formula (1).
[0075] In a possible implementation manner, the traffic of the first-layer matrix product is obtained based on the first traffic, the second traffic, and the third traffic. The first traffic is the traffic for performing a first collection operation on the input sub-matrices on each processor in the second splitting direction; the second traffic is the traffic for performing a second collection operation on the first parameter sub-matrices on each processor in the first splitting direction; the third traffic is the traffic for multiplying and summing the first collection result and the second collection result on each processor in the third splitting direction, and distributing the first multiplication and summation result of the first collection result and the second collection result to each processor. In this way, it is beneficial to accurately determine the traffic of the first-layer matrix product.
[0076] For example, as Figure 7 shown, in n processes, a first collection operation (see the collection operation shown in x L y E z can be performed on the input sub-matrices of size B Figure 2 in the second splitting direction (for example, the splitting direction corresponding to the second variable y), so that each processor in the second splitting direction stores the first collection result of size B Figure 2 LE x LE z . The first traffic for this step is BLE / xz.
[0077] Synchronously, a second collection operation (see the collection operation shown in z F yx ) can be performed on the first parameter sub-matrices of size E Figure 2 in the first splitting direction (for example, the splitting direction corresponding to the first variable x), so that each processor in the first splitting direction stores the size of Ez F y The second collection result of
[0078] Each processor separately obtains a first collection result with a size of B x LE z and a second collection result with a size of E z F y For the second collection result, the second traffic volume at this step is EF / yz. For the first collection result with a size of B and the second collection result with a size of E x LE z and the second collection result with a size of E z F y it is possible to perform a multiplication and summation on the first collection result with a size of B x LF y and the second collection result with a size of E Figure 5 in the third splitting direction (for example, the splitting direction corresponding to the third variable z), and distribute the first multiplication and summation result with a size of B x LF y to each processor (such as the all-reduce multiply-add operation shown in
[0079] so that each processor stores the first multiplication and summation result with a size of B
[0080] The traffic volume of the first layer matrix multiplication can be obtained by summing the first traffic volume BLE / xz, the second traffic volume EF / yz, and the third traffic volume 2BLF / xy. In a possible implementation, the traffic volume of the second layer matrix multiplication is obtained based on a fourth traffic volume and a fifth traffic volume. The fourth traffic volume is the traffic volume of performing a third collection operation on the second parameter sub-matrix on each processor in the first splitting direction; the fifth traffic volume is the traffic volume of performing a multiplication and summation on the first multiplication and summation result and the third collection result on each processor in the second splitting direction, and splitting the second multiplication and summation result of the first multiplication and summation result and the third collection result to each processor.
[0081] For example, as Figure 7 shown, it is possible to perform a third collection operation on the second parameter sub-matrix with a size of F y E zx in the first splitting direction (for example, the splitting direction corresponding to the first variable x) (see the collection operation shown in Figure 2 so that each processor in the first splitting direction stores the third collection result with a size of F y E z The fourth traffic volume at this step is EF / yz.
[0082] Then, in the second splitting direction (e.g., the splitting direction corresponding to the second variable y), for the first multiply-sum result of size B x LF y and the third gather result of size F y E z a multiply-sum operation is performed, and the second multiply-sum result of size B x LE z is sliced to each processor (such as the reduction-scatter multiply-add operation shown in Figure 6 ), which is equivalent to uniformly slicing the second multiply-sum result of size B x LE z along the second splitting direction after multiplication, so that the sliced target result B x L y E z obtained by each processor is still three-dimensionally split, facilitating the next matrix multiplication. The fifth communication volume corresponding to this multiply-sum splitting operation is BLE / xz. Among them, the target results of size B x L y E z from n processors can be used to determine the product result of the two-layer matrix product.
[0083] By summing the fourth communication volume EF / yz and the fifth communication volume BLE / xz, the communication volume of the second-layer matrix product can be obtained.
[0084] In this way, by adding these five communication volumes corresponding to the two-layer matrix product, the objective function shown in formula (2) can be deduced.
[0085] In a possible implementation, the splitting scheme includes the first variable x, the second variable y, and the third variable z that minimize the objective function. By minimizing the objective function, it is beneficial to efficiently and quickly determine the first variable x, the second variable y, and the third variable z that minimize the total communication volume T comm among multiple processors.
[0086] When actually determining the specific splitting scheme, since the first variable x, the second variable y, and the third variable z are integers and need to satisfy that the product of the first variable x, the second variable y, and the third variable z is equal to the number of processors n, that is: xyz = n, by performing a minimum search on the possible integer values of the first variable x, the second variable y, and the third variable z for the total communication volume T comm i.e., by minimizing the objective function shown in formula (2), the first variable x, the second variable y, and the third variable z are solved.
[0087] Among them, in actual applications, the optimal integer values of the first variable x, the second variable y, and the third variable z are near a certain real value, and a search can be performed near it to accelerate the search speed.
[0088] In actual applications, the derived objective function can be pre-written into the memory in advance, so that in the case of a two-layer matrix multiplication operation based on big data, in step S11, any one of the multiple processors can solve the objective function read from the memory to obtain a splitting scheme, and based on the splitting scheme, the input matrix, the first parameter matrix, and the second parameter matrix are split, and the splitting results are input to the multiple processors. In step S11, when the splitting results are input to the multiple processors, in step S12, the multiple processors calculate the products of the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting results. Among them, the splitting results include input sub-matrices, first parameter sub-matrices, and second parameter sub-matrices.
[0089] Take, for example, two matrix multiplications in a feed-forward network operation in the pre-fill part of the inference operation of a large-scale neural network model (such as including the ChatGLM-6B language model). Assume that the known input matrix X is language feature data, the size of the first dimension of the input matrix X is B = 8, the size of the second dimension is E = 4096, and the size of the third dimension is L = 2048. Among them, B can represent the batch, L can represent the length of the language feature, and E can represent the number of channels; and, the size of the first dimension of the first parameter matrix W1 is E = 4096, the size of the second dimension is F = 4096×4, and the size of the first dimension of the second parameter matrix W2 is F = 4096×4, and the size of the second dimension is E = 4096.
[0090] Single-precision (FP32) inference can be performed on n = 256 GPU cards. Assume that the best splitting scheme obtained by searching through step S11 is: the first variable x = 8, the second variable y = 16, and the third splitting variable z = 2.
[0091] In this way, since the size of the input matrix X is 8×2048×4096, it is evenly split onto 8×16×2 = 256 GPU cards. After splitting, the size of the input sub-matrix on each GPU card is 1×128×2048. Similarly, since the size of the first parameter matrix W1 is 4096×(4096×4), the size of the first parameter sub-matrix split onto each GPU card is 2048×128; since the size of the second parameter matrix W2 is (4096×4)×4096, the size of the second parameter sub-matrix split onto each GPU card is 1024×256.
[0092] In step S12, it can be in accordance with Figure 7Schematic diagram of the reasoning process of the objective function shown. 256 GPU cards calculate the product of the input matrix X, the first parameter matrix W1, and the second parameter matrix W2 based on the input sub-matrix, the first parameter sub-matrix, and the second parameter sub-matrix.
[0093] A first collection operation can be performed on the input sub-matrix with a size of 1×128×2048 along the second splitting direction (the splitting direction corresponding to the second variable y), so that each GPU card obtains a first collection result with a size of 1×2048×2048; and, a second collection is performed on the first parameter sub-matrix with a size of 2048×128 along the first splitting direction (the splitting direction corresponding to the first variable x), so that each GPU card obtains a second collection result with a size of 2048 ×1024. Then, a full reduction multiply-add operation can be performed on the first collection result with a size of 1×2048×2048 and the second collection result with a size of 2048 ×1024 along the third splitting direction (the splitting direction corresponding to the third variable z), so that each GPU card obtains a first multiply-sum result with a size of 1×2048×1024.
[0094] A third collection operation is performed on the second parameter sub-matrix with a size of 1024×256 along the first splitting direction (the splitting direction corresponding to the first variable x), so that each GPU card obtains a third collection result with a size of 1024×2048. Then, a reduced-divergent multiply-add operation can be performed on the first multiply-sum result with a size of 1×2048×1024 and the third collection result with a size of 1024×2048 along the second splitting direction (the splitting direction corresponding to the second variable y), so that each GPU card obtains a target result with a size of 1×128×2048. The target result of each GPU card can be used to obtain the two-layer matrix product result of the input matrix X, the first parameter matrix W1, and the second parameter matrix W2.
[0095] Next, analyze the communication complexity of the data processing method of the present disclosure embodiment. Since the product of the first variable x, the second variable y, and the third variable z is equal to the number of processors n, xyz=n can be used to transform the objective function in formula (2) to obtain the objective function shown in formula (3):
[0096] T comm =BLE / xz+EF / yz+2BLF / xy+EF / yz+BLE / xz=2 / n(BLEy+EFx+BLFz) (3)
[0097] In formula (3), T comm represents the total communication volume between n processors, BLE represents the size of the input matrix X, and EF represents the size of the first parameter matrix W1 (or the second parameter matrix W2).
[0098] Since the formula (3) needs to satisfy the constraint condition xyz = n, the constraint condition and the objective function can be rewritten by using logarithmic transformation.
[0099] a = ln(EFx), b = ln(BLEy), c = ln(BLFz) (4)
[0100] In formula (4), a represents taking the logarithm of the second term EFx of the total traffic T in formula (3), b represents taking the logarithm of the first term BLEy of the total traffic T in formula (3), and c represents taking the logarithm of the third term BLFz of the total traffic T in formula (3). Substituting formula (4) into formula (3), the objective function becomes: comm of the second term EFx of the total traffic T comm of the first term BLEy of the total traffic T comm of the third term BLFz of the total traffic T. Substituting formula (4) into formula (3), the objective function becomes:
[0101] T comm = 2 / n (e a + e b + e c ) (5)
[0102] where a + b + c = ln(B 2 L 2 E 2 F 2 n). Using the Jensen's inequality for convex functions, it can be known that the minimum value of the total traffic T comm is obtained when a = b = c, and its value is: min(T comm ) = 6(BLEF) 2 / 3 (1 / n 2 / 3 ), where min is the minimization function, B represents the size of the input matrix X in the first dimension, L represents the size of the input matrix X in the second dimension, E represents the size of the input matrix X in the third dimension, that is, the size of the first parameter matrix W1 in the first dimension, the size of the second parameter matrix W2 in the second dimension, F represents the size of the first parameter matrix W1 in the second dimension, that is, the size of the second parameter matrix W2 in the first dimension.
[0103] It can be seen that the minimum complexity of the total traffic of data processing in the embodiments of the present disclosure with respect to the processor cluster size n is O(1 / n 2 / 3 ), and it has a traffic inversely proportional to 2 / 3 of the number of processor chips.
[0104] In the related art, for example, in the split and collect computing scheme from Google, whether it is two-dimensional splitting or one-dimensional splitting, plus its corresponding matrix collection method, the optimal traffic is 8BLE × 1 / n 1 / 2 , and the minimum complexity of this optimal traffic with respect to the processor cluster size n is O(1 / n 1 / 2 ).
[0105] The optimal complexity O(1 / n) of the traffic volume of the data processing method according to the embodiments of the present disclosure 2 / 3 ) is better than the optimal complexity O(1 / n) in the related art 1 / 2 ), and the overall traffic volume of the data processing method according to the embodiments of the present disclosure is one order of magnitude smaller than that of the related art.
[0106] Figure 8 The schematic diagram showing the comparison of the total traffic volume of the data processing method according to the embodiments of the present disclosure with the related art is as Figure 8 shown. In the case where the first dimension size B of the input matrix X is 8, the second dimension size L of the input matrix X is 2048, and the third dimension size E of the input matrix X is 4096 (i.e., the size of the first parameter matrix W1 in the first dimension and the size of the second parameter matrix W2 in the second dimension), and the second dimension size F of the first parameter matrix W1 is 4096×4 (i.e., the size of the second parameter matrix W2 in the first dimension), the total traffic volume curve of the data processing method according to the embodiments of the present disclosure and the total traffic volume curve in the related art (such as the two-dimensional splitting method) change with the increase of the processor data n. As Figure 8 shown, when the number of processors exceeds 5, the data processing method according to the embodiments of the present disclosure can produce advantages. When the number of processors exceeds 256 and reaches a large scale level, the advantages of the data processing method according to the embodiments of the present disclosure are more obvious.
[0107] In summary, in the data processing method according to the embodiments of the present disclosure, the splitting scheme obtained by solving the objective function can utilize multiple processors to process a larger scale model (such as a neural network model with a large number of parameters) and reduce the communication overhead between multiple processors.
[0108] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.
[0109] In addition, the present disclosure also provides a data processing device, an electronic device, a computer-readable storage medium, and a program, all of which can be used to implement any data processing method provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding records in the method part and will not be elaborated further.
[0110] Figure 9 The block diagram showing the data processing device according to the embodiments of the present disclosure is as Figure 9 shown. The data processing device is applied to multiple processors, and the device includes:
[0111] Splitting module 91: It is used to split the input matrix, the first parameter matrix, and the second parameter matrix based on a splitting scheme by any one of the multiple processors, and input the splitting results into the multiple processors. The splitting scheme represents the number of splitting parts of each dimension of the input matrix, the first parameter matrix, and the second parameter matrix.
[0112] Calculation module 92, which is used for the multiple processors to calculate the product of the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting results.
[0113] Wherein, the splitting scheme is obtained by solving an objective function, the objective function represents the relationship between the total communication volume and the splitting scheme, and the total communication volume is the total communication volume between multiple processors during the process of the multiple processors calculating the matrix product of the input matrix, the first parameter matrix, and the second parameter matrix.
[0114] In a possible implementation manner, the total communication volume is obtained based on the communication volume of the first-layer matrix product and the communication volume of the second-layer matrix product. Wherein, the first-layer matrix product is the product of the input matrix and the first parameter matrix, and the second-layer matrix product is the product of the product result of the input matrix and the first parameter matrix and the second parameter matrix.
[0115] In a possible implementation manner, the input matrix is a three-dimensional matrix, the first parameter matrix and the second parameter matrix are two-dimensional matrices, the number of elements in the third dimension of the input matrix, the first dimension of the first parameter matrix, and the second dimension of the second parameter matrix are the same, and the number of elements in the second dimension of the first parameter matrix and the first dimension of the second parameter matrix are the same; the objective function includes a first variable, a second variable, and a third variable, and the product of the first variable, the second variable, and the third variable is equal to the number of processors. The splitting scheme is represented by the first variable, the second variable, and the third variable; the first variable is used to characterize the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the first splitting direction, and the first splitting direction includes the direction of the first dimension of the input matrix; the second variable is used to characterize the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the second splitting direction, and the second splitting direction includes the direction of the second dimension of the input matrix; the third variable is used to characterize the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the third splitting direction, and the third splitting direction includes the direction of the third dimension of the input matrix.
[0116] In a possible implementation, the splitting module 91 is configured to: split the input matrix, the first parameter matrix, and the second parameter matrix respectively in the first splitting direction, the second splitting direction, and the third splitting direction according to the first variable, the second variable, and the third variable, so that each processor corresponds to a different input sub-matrix, first parameter sub-matrix, and second parameter sub-matrix.
[0117] In a possible implementation, the communication volume of the first-layer matrix multiplication is obtained based on a first communication volume, a second communication volume, and a third communication volume, where: the first communication volume is the communication volume of performing a first collection operation on the input sub-matrices on each processor in the second splitting direction; the second communication volume is the communication volume of performing a second collection operation on the first parameter sub-matrices on each processor in the first splitting direction; the third communication volume is the communication volume of multiplying and summing the first collection result and the second collection result on each processor in the third splitting direction, and distributing the first multiplication and summation result of the first collection result and the second collection result to each processor, where the first collection result is the result obtained by performing the first collection operation, and the second collection result is the result obtained by performing the second collection operation.
[0118] In a possible implementation, the communication volume of the second-layer matrix multiplication is obtained based on a fourth communication volume and a fifth communication volume, where: the fourth communication volume is the communication volume of performing a third collection operation on the second parameter sub-matrices on each processor in the first splitting direction; the fifth communication volume is the communication volume of multiplying and summing the first multiplication and summation result and the third collection result on each processor in the second splitting direction, and splitting the second multiplication and summation result of the first multiplication and summation result and the third collection result to each processor, where the third collection result is the result obtained by performing the third collection operation.
[0119] In a possible implementation, the splitting scheme includes the first variable, the second variable, and the third variable that minimize the objective function.
[0120] In a possible implementation, the input matrix includes feature data in a deep learning task, and the feature data includes at least one of image feature data, voice feature data, and text feature data.
[0121] This method has a specific technical association with the internal structure of a computer system and can solve the technical problem of how to improve the hardware operation efficiency or execution effect (including reducing the data storage volume, reducing the data transmission volume, and increasing the hardware processing speed, etc.), so as to obtain the technical effect of improving the internal performance of the computer system that conforms to the natural law.
[0122] In some embodiments, the functions or modules included in the apparatus provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0123] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0124] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein the processor is configured to call the instructions stored in the memory to execute the above methods.
[0125] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in the processor of the electronic device, the processor in the electronic device executes the above methods.
[0126] The electronic device can be provided as a terminal, a server or other forms of devices. For example, it includes user equipment (UE), mobile devices, user terminals, terminals, cellular phones, cordless phones, personal digital assistants (PDAs), handheld devices, computing devices, in-vehicle devices, wearable devices, etc.
[0127] Figure 10 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to Figure 10 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to execute the above methods.
[0128] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958. The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows Server TM ), the graphical user interface-based operating system launched by Apple Inc. (Mac OS X TM ), the multi-user and multi-process computer operating system (Unix TM ), the free and open-source Unix-like operating system (Linux TM ), the open-source Unix-like operating system (FreeBSD TM ), or the like.
[0129] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.
[0130] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0131] The computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, (but is not limited to) an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0132] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0133] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0134] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0135] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when the instructions are executed by the processor of the computer or other programmable data processing apparatus, an apparatus is created that implements the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, so that the computer-readable medium storing the instructions comprises a manufacture, which includes instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0136] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0137] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur in a different order than noted in the figures. For example, two consecutive boxes may, in fact, be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0138] The computer program product can be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is embodied as a computer storage medium. In another alternative embodiment, the computer program product is embodied as a software product, such as a Software Development Kit (SDK), etc.
[0139] The above descriptions of the various embodiments tend to emphasize the differences between the various embodiments. Their similarities or resemblances can be referred to each other. For the sake of brevity, they will not be elaborated herein.
[0140] Those skilled in the art will appreciate that, in the above method of specific implementation, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of the steps should be determined by their functions and possible internal logic.
[0141] If the technical solution of this application involves personal information, the product using the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using the technical solution of this application has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to the collection of his or her personal information; or on the device that processes personal information, the personal information processing rules are notified by obvious signs / information, and the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; among them, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0142] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A data processing method, characterized in that, The described data processing method is applied to multiple processors, and the method includes: Any one of the multiple processors splits the input matrix, the first parameter matrix, and the second parameter matrix based on a splitting scheme, and inputs the splitting results into the multiple processors. The splitting scheme represents the number of splitting parts for each dimension of the input matrix, the first parameter matrix, and the second parameter matrix. The multiple processors calculate the product of the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting results. Among them, the splitting scheme is obtained by solving an objective function, and the objective function represents the relationship between the total communication volume and the splitting scheme. The total communication volume is the total communication volume among multiple processors during the process of multiple processors calculating the matrix product of the input matrix, the first parameter matrix, and the second parameter matrix. Among them, the objective function includes a first variable, a second variable, and a third variable, and the product of the first variable, the second variable, and the third variable is equal to the number of processors. The splitting scheme is represented by the first variable, the second variable, and the third variable. The first variable is used to characterize the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the first splitting direction, and the first splitting direction includes the direction of the first dimension of the input matrix. The second variable is used to characterize the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the second splitting direction, and the second splitting direction includes the direction of the second dimension of the input matrix. The third variable is used to characterize the number of splitting parts of the input matrix, the first parameter matrix, and the second parameter matrix in the third splitting direction, and the third splitting direction includes the direction of the third dimension of the input matrix.
2. The method according to claim 1, characterized in that The total communication volume is obtained based on the communication volume of the first-layer matrix product and the communication volume of the second-layer matrix product. Among them, the first-layer matrix product is the product of the input matrix and the first parameter matrix, and the second-layer matrix product is the product of the product result of the input matrix and the first parameter matrix and the second parameter matrix.
3. The method according to claim 2, wherein The input matrix is a three-dimensional matrix, the first parameter matrix and the second parameter matrix are two-dimensional matrices. The number of elements in the third dimension of the input matrix, the first dimension of the first parameter matrix, and the second dimension of the second parameter matrix are the same, and the number of elements in the second dimension of the first parameter matrix and the first dimension of the second parameter matrix are the same.
4. The method according to claim 3, characterized in that, The step of any one of the multiple processors splitting the input matrix, the first parameter matrix, and the second parameter matrix based on the splitting scheme includes: In the first splitting direction, the second splitting direction, and the third splitting direction, the input matrix, the first parameter matrix, and the second parameter matrix are respectively split according to the first variable, the second variable, and the third variable, so that each processor corresponds to different input sub-matrices, first parameter sub-matrices, and second parameter sub-matrices.
5. The method according to claim 4, characterized in that, The communication volume of the first-layer matrix multiplication is obtained based on the first communication volume, the second communication volume, and the third communication volume, where: The first communication volume is the communication volume of performing a first collection operation on the input sub-matrices on each processor in the second splitting direction; The second communication volume is the communication volume of performing a second collection operation on the first parameter sub-matrices on each processor in the first splitting direction; The third communication volume is the communication volume of multiplying and summing the first collection result and the second collection result on each processor in the third splitting direction, and distributing the first multiplication and summation result of the first collection result and the second collection result to each processor, where the first collection result is the result obtained by performing the first collection operation, and the second collection result is the result obtained by performing the second collection operation.
6. The method according to claim 5, characterized in that, The communication volume of the second-layer matrix multiplication is obtained based on the fourth communication volume and the fifth communication volume, where: The fourth communication volume is the communication volume of performing a third collection operation on the second parameter sub-matrices on each processor in the first splitting direction; The fifth communication volume is the communication volume of multiplying and summing the first multiplication and summation result and the third collection result on each processor in the second splitting direction, and splitting the second multiplication and summation result of the first multiplication and summation result and the third collection result to each processor, where the third collection result is the result obtained by performing the third collection operation.
7. The method according to any one of claims 3 to 6, characterized in that, The splitting scheme includes the first variable, the second variable, and the third variable that minimize the objective function.
8. The method according to any one of claims 1 to 6, characterized in that, The input matrix includes feature data in a deep learning task, and the feature data includes at least one of image feature data, voice feature data, and text feature data.
9. A data processing device, characterized in that, The data processing device is applied to multiple processors, and the device includes: A splitting module, configured to split an input matrix, a first parameter matrix, and a second parameter matrix based on a splitting scheme through any one of the multiple processors, and input the splitting result into the multiple processors, where the splitting scheme represents the number of splitting parts of each dimension of the input matrix, the first parameter matrix, and the second parameter matrix; A calculation module, configured to calculate the product of the input matrix, the first parameter matrix, and the second parameter matrix by the multiple processors based on the splitting result; Wherein, the splitting scheme is obtained by solving an objective function, the objective function represents the relationship between the total communication volume and the splitting scheme, and the total communication volume is the total communication volume between multiple processors during the process of calculating the matrix product of the input matrix, the first parameter matrix, and the second parameter matrix by the multiple processors; Wherein, the objective function includes a first variable, a second variable, and a third variable, the product of the first variable, the second variable, and the third variable is equal to the number of processors, and the splitting scheme is represented by the first variable, the second variable, and the third variable. The first variable is used to represent the number of split parts of the input matrix, the first parameter matrix, and the second parameter matrix in the first splitting direction, and the first splitting direction includes the direction of the first dimension of the input matrix; The second variable is used to represent the number of split parts of the input matrix, the first parameter matrix, and the second parameter matrix in the second splitting direction, and the second splitting direction includes the direction of the second dimension of the input matrix; The third variable is used to represent the number of split parts of the input matrix, the first parameter matrix, and the second parameter matrix in the third splitting direction, and the third splitting direction includes the direction of the third dimension of the input matrix.
10. An electronic device, characterized in that, Comprising: A plurality of processors; A memory for storing instructions executable by the plurality of processors; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 8.
11. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Parallel processing method and device of model, first computing equipment and electronic equipment
CN116820577A
Data processing method and device, electronic equipment, storage medium and product
CN118069742A