Parallel processing method and device of model, electronic equipment and readable storage medium

By processing the split data submatrix in parallel on multiple computing devices, the pressure on memory and training speed of hyperscale deep learning models is solved, and the training of larger models and more efficient computing performance is achieved.

CN119960970AActive Publication Date: 2025-05-09BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Application Number
CN202411896113.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-09
Estimated Expiration
2044-12-20

Smart Images

  • Figure CN119960970A_ABST
    Figure CN119960970A_ABST
Patent Text Reader

Abstract

The invention provides a model parallel processing method, and relates to the technical field of artificial intelligence such as deep learning, natural language processing, image processing and large language models. The method is applied to a first computing device in N computing devices, and comprises the following steps: acquiring a target first data sub-matrix and a target second data sub-matrix; starting a calculation process of matrix multiplication, processing the target first data sub-matrix and the target second data sub-matrix, and copying first candidate data sub-matrixes in other N-1 pieces of calculation equipment at the same time; in response to a first processing result between the target first data sub-matrix and the target second data sub-matrix, processing the copied first candidate data sub-matrix and the corresponding target data sub-matrix; and in response to a second processing result between the first candidate data sub-matrix and the corresponding target data sub-matrix, obtaining a target processing result of the first computing device according to the first processing result and the second processing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to the field of artificial intelligence technology such as deep learning, natural language processing, image processing, and large language models. A parallel processing method, device, electronic device, and readable storage medium for a model are provided. Background Art

[0002] With the successful application of deep learning models in various fields, people have begun to pay attention to how to expand deep learning models to a larger scale to improve the data processing capabilities, accuracy and performance of the models. Based on this, ultra-large-scale deep learning models came into being. Ultra-large-scale deep learning models face the pressure of memory and training speed. However, the memory of a single computing device is very limited. Therefore, how to use the limited memory of each computing device to train a larger model is a technical problem that needs to be solved urgently. Summary of the invention

[0003] According to a first aspect of the present disclosure, a parallel training method for a model is provided, which is applied to a first computing device among N computing devices, comprising: obtaining a target first data submatrix among N first data submatrices, and a target second data submatrix among N second data submatrices; wherein the N first data submatrices are obtained by dividing the first data matrix according to a first dividing method, and the N second data submatrices are obtained by dividing the second data matrix according to a second dividing method; wherein N is a positive integer greater than or equal to 2; starting a matrix multiplication calculation process, processing the target first data submatrix and the target second data submatrix, and performing matrix multiplication on the other N-1 computing devices at the same time; The first candidate data submatrix is ​​copied; in response to obtaining a first processing result between the target first data submatrix and the target second data submatrix, the copied first candidate data submatrix and its corresponding target data submatrix are processed; in response to obtaining a second processing result between the first candidate data submatrix and its corresponding target data submatrix, a target processing result of the first computing device is obtained according to the first processing result and the second processing result; wherein the target processing result of the first computing device is used to be spliced ​​with the N-1 target processing results of the other N-1 computing devices to obtain the target processing result between the first data matrix and the second data matrix.

[0004] According to a second aspect of the present disclosure, a parallel processing device of a model is provided, wherein a first computing device among N computing devices comprises: an acquisition unit, for acquiring a target first data submatrix among the N first data submatrices, and a target second data submatrix among the N second data submatrices; wherein the N first data submatrices are obtained by dividing the first data matrix according to a first dividing method, and the N second data submatrices are obtained by dividing the second data matrix according to a second dividing method; wherein N is a positive integer greater than or equal to 2; a first processing unit, for starting a matrix multiplication calculation process, and while processing the target first data submatrix and the target second data submatrix, processing the target first data submatrix and the target second data submatrix in the other N-1 computing devices; a first candidate data submatrix is ​​copied; a second processing unit is used for processing the copied first candidate data submatrix and its corresponding target data submatrix in response to obtaining a first processing result between the target first data submatrix and the target second data submatrix; a third processing unit is used for obtaining a target processing result of the first computing device according to the first processing result and the second processing result in response to obtaining a second processing result between the first candidate data submatrix and its corresponding target data submatrix; wherein the target processing result of the first computing device is used to be spliced ​​with the N-1 target processing results of the other N-1 computing devices to obtain the target processing result between the first data matrix and the second data matrix.

[0005] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0006] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method as described above.

[0007] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the method as described above when executed by a processor.

[0008] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0010] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;

[0011] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;

[0012] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;

[0013] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure;

[0014] Figure 5 is a schematic diagram according to a fifth embodiment of the present disclosure;

[0015] Figure 6 is a schematic diagram according to a sixth embodiment of the present disclosure;

[0016] Figure 7 is a schematic diagram according to a seventh embodiment of the present disclosure;

[0017] Figure 8 is a schematic diagram according to an eighth embodiment of the present disclosure;

[0018] Fig. 9 is a schematic diagram according to a ninth embodiment of the present disclosure;

[0019] Fig.10 is a schematic diagram according to a tenth embodiment of the present disclosure;

[0020] Fig.11 is a schematic diagram according to an eleventh embodiment of the present disclosure;

[0021] Fig.12 is a schematic diagram according to a twelfth embodiment of the present disclosure;

[0022] Fig.13 is a schematic diagram according to a thirteenth embodiment of the present disclosure;

[0023] Fig.14 is a schematic diagram according to a fourteenth embodiment of the present disclosure;

[0024] Fig.15 is a schematic diagram according to a fifteenth embodiment of the present disclosure;

[0025] Fig.16 is a schematic diagram according to a sixteenth embodiment of the present disclosure;

[0026] Fig.17 is a schematic diagram according to a seventeenth embodiment of the present disclosure;

[0027] Fig.18 is a schematic diagram according to an eighteenth embodiment of the present disclosure;

[0028] Fig.19 is a schematic diagram according to a nineteenth embodiment of the present disclosure;

[0029] Fig. 20 is a schematic diagram according to the twentieth embodiment of the present disclosure;

[0030] Fig.21 is a schematic diagram according to the twenty-first embodiment of the present disclosure;

[0031] Fig. 22 is a schematic diagram according to the twenty-second embodiment of the present disclosure;

[0032] Fig.23 is a schematic diagram according to the twenty-third embodiment of the present disclosure;

[0033] Fig.24 is a schematic diagram according to the twenty-fourth embodiment of the present disclosure;

[0034] Fig.25 is a schematic diagram according to the twenty-fifth embodiment of the present disclosure;

[0035] Fig.26 is a schematic diagram according to the twenty-sixth embodiment of the present disclosure;

[0036] Fig. 27 is a schematic diagram according to the twenty-seventh embodiment of the present disclosure;

[0037] Fig.28 It is a block diagram of an electronic device used to implement the parallel processing method of the model according to the embodiment of the present disclosure. DETAILED DESCRIPTION

[0038] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and mechanisms is omitted in the following description.

[0039] Figure 1 Schematic diagram of the first embodiment of the present disclosure. Figure 1 As shown, the parallel processing method of the model of this embodiment is applied to the first computing device among N computing devices, and specifically includes the following steps:

[0040] S101, obtaining a target first data submatrix among N first data submatrices, and a target second data submatrix among N second data submatrices; wherein the N first data submatrices are obtained by dividing the first data matrix according to a first dividing method, and the N second data submatrices are obtained by dividing the second data matrix according to a second dividing method; and N is a positive integer greater than or equal to 2;

[0041] S102, starting a matrix multiplication calculation process, while processing the target first data submatrix and the target second data submatrix, copying the first candidate data submatrices in other N-1 computing devices;

[0042] S103, in response to obtaining a first processing result between the target first data submatrix and the target second data submatrix, processing the copied first candidate data submatrix and its corresponding target data submatrix;

[0043] S104. In response to obtaining a second processing result between the first candidate data submatrix and its corresponding target data submatrix, a target processing result of the first computing device is obtained according to the first processing result and the second processing result; wherein the target processing result of the first computing device is used to be spliced ​​with the N-1 target processing results of the other N-1 computing devices to obtain the target processing result between the first data matrix and the second data matrix.

[0044] In this embodiment, the first data matrix can be a feature matrix corresponding to the input data, or a weight matrix corresponding to the target model; the second data matrix can be a feature matrix corresponding to the input data, or a weight matrix corresponding to the target model.

[0045] In this embodiment, the target model is a deep learning model, and the elements in the weight matrix corresponding to the target model are the parameters in the deep learning model; it can be understood that the weight matrix in this embodiment can be the weight matrix of some network layers in the target model.

[0046] In this embodiment, if the target model is an image processing model and the input data is an image, the feature matrix corresponding to the input data can be a matrix composed of pixel values ​​of each pixel in the image; if the target model is a natural language processing model and the input data is text, the feature matrix corresponding to the input data can be a matrix composed of word vectors of each word in the text.

[0047] In this embodiment, the computing device may be a device with parallel computing capabilities, such as a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), a GPU-like device, or an XPU, which is not limited in this embodiment.

[0048] In this embodiment, the first computing device is one of N computing devices, and the N computing devices correspond one-to-one to N first data sub-matrices and N second data sub-matrices, respectively, where N is a positive integer greater than or equal to 2; wherein the data sub-matrices corresponding to different computing devices can be stored in the memory of the computing device.

[0049] In this embodiment, any segmentation method can be used to segment the first data matrix and the second data matrix; wherein the first segmentation method corresponding to the first data matrix and the second segmentation method corresponding to the second data matrix can be the same or different.

[0050] That is to say, this embodiment does not limit the division method of the data sub-matrices obtained by the computing device, so that the computing device can perform parallel processing on the data sub-matrices obtained by any division method, thereby expanding the usage scenarios and achieving the purpose of performing true distributed parallel matrix multiplication by computing devices such as GPU, NPU, GPU-like or XPU.

[0051] In this embodiment, the first splitting method can be row splitting (i.e., splitting the first data matrix in the row direction) or column splitting (i.e., splitting the first data matrix in the column direction); the second splitting method can be row splitting (i.e., splitting the second data matrix in the row direction) or column splitting (i.e., splitting the second data matrix in the column direction).

[0052] Among them, row cutting refers to cutting a data matrix of M rows and K columns into N data sub-matrices with M / N rows and K columns; column cutting refers to cutting a data matrix of M rows and K columns into N data sub-matrices with M rows and K / N columns.

[0053] In this embodiment, each first data submatrix among the N first data submatrices and each second data submatrix among the N second data submatrices are distributed to different computing devices. When the first computing device in this embodiment executes S101, it uses the distributed first data submatrix as the target first data submatrix and the distributed second data submatrix as the target second data submatrix.

[0054] After executing S101 to receive the target first data submatrix and the target second data submatrix, the first computing device in this embodiment executes S102 to start the matrix multiplication calculation process, processes the received target first data submatrix and the target second data submatrix, and simultaneously copies the first candidate data submatrices in the other N-1 computing devices.

[0055] When executing S102 , the first computing device in this embodiment may start the matrix multiplication calculation process by calling a general matrix multiplication (GEMM) kernel.

[0056] After executing S102 to complete the start of the matrix multiplication calculation process, the first computing device in this embodiment can copy the first candidate data submatrices in other N-1 computing devices while performing matrix multiplication on the target first data submatrix and the target second data submatrix.

[0057] That is to say, the first computing device in this embodiment communicates with other N-1 computing devices while performing matrix multiplication on the existing data sub-matrices, thereby copying the first candidate data sub-matrix from the other N-1 computing devices, and achieving overlap between computing and communication when the first computing device such as a GPU, NPU, GPU-like or XPU performs parallel processing.

[0058] When the first computing device in this embodiment executes S102 to process the received target first data submatrix and target second data submatrix, the implementation method that can be adopted is: according to the first preset block size, the target first data submatrix is ​​divided into multiple target first matrix blocks, and the target second data submatrix is ​​divided into multiple target second matrix blocks; based on the obtained multiple target first matrix blocks and multiple target second matrix blocks, the processing result between the target first data submatrix and the target second data submatrix is ​​obtained, and the obtained processing result is the matrix multiplication result.

[0059] In this embodiment, the first preset block size may be a block size that matches a Warp size (Warp is a basic unit for scheduling and execution in a GPU).

[0060] That is to say, this embodiment achieves the purpose of performing Warp-level calculations within the called general matrix multiplication kernel by dividing the data sub-matrices and then performing matrix multiplication calculations between the data sub-matrices according to the matrix blocks obtained by the division. This can improve the computing efficiency of the first computing device such as the GPU, NPU, GPU-like or XPU when performing matrix multiplication.

[0061] In this embodiment, the first candidate data submatrix to be copied by the first computing device from the other N-1 computing devices may be all or part of the first data submatrix corresponding to the other N-1 computing devices, or may be all or part of the second data submatrix corresponding to the other N-1 computing devices.

[0062] The first computing device in this embodiment, when executing S102 to copy the first candidate data submatrix in the other N-1 computing devices, can be implemented as follows: construct a segmentation method set according to the first segmentation method, the second segmentation method, and the third segmentation method corresponding to the output data matrix; determine the first candidate data submatrix according to the constructed segmentation method set; and copy the first candidate data submatrix in the other N-1 computing devices.

[0063] The first computing device in this embodiment accesses the memory of other N-1 computing devices to copy the first candidate data sub-matrices of other N-1 computing devices.

[0064] In this embodiment, the output data matrix is ​​the target processing result between the first data matrix and the second data matrix; the third segmentation method can be row segmentation (i.e. segmenting the output data matrix in the row direction) or column segmentation (i.e. segmenting the output data matrix in the column direction).

[0065] Since the present embodiment supports arbitrary segmentation of the input data matrix and the output data matrix (usually the matrix is ​​segmented evenly), the segmentation method set constructed in the present embodiment has 8 cases according to the segmentation method of the data matrix, specifically (row segmentation, row segmentation, row segmentation), (row segmentation, row segmentation, column segmentation), (row segmentation, column segmentation, row segmentation), (row segmentation, column segmentation, column segmentation), (column segmentation, row segmentation, row segmentation), (column segmentation, row segmentation, column segmentation), (column segmentation, column segmentation, row segmentation) and (column segmentation, column segmentation, column segmentation); in the present embodiment, different segmentation method sets correspond to different types of first candidate data sub-matrices, and therefore, through the constructed segmentation method set, the first computing device determines what type of first candidate data sub-matrix to be copied from the other N-1 computing devices.

[0066] In this embodiment, the first computing device obtains the first candidate data sub-matrix required for matrix multiplication calculations from other N-1 computing devices by copying, thereby avoiding the steps of sending and receiving sub-matrices between computing devices, and can reduce the time required for the first computing device such as GPU, NPU, GPU-like or XPU to obtain the first candidate data sub-matrix from other N-1 computing devices, thereby improving the efficiency of subsequent matrix multiplication calculations based on the first candidate data sub-matrix.

[0067] The first computing device in this embodiment may store the copied first candidate data sub-matrices corresponding to different other computing devices into the memory of the first computing device so as to be acquired in subsequent processing.

[0068] The first computing device in this embodiment, when executing S102 to copy the first candidate data sub-matrix in the other N-1 computing devices, can copy the first candidate data sub-matrix in the other N-1 computing devices multiple times according to the second preset block size, that is, copy the matrix block corresponding to the second preset block size from the other N-1 computing devices each time; wherein the second preset block size in this embodiment can be the same as the first preset block size or different, and this embodiment can set the second preset block size according to actual needs.

[0069] That is to say, the first computing device in this embodiment can copy the first candidate data submatrix in the other N-1 computing devices multiple times according to smaller blocks, so that after completing the processing between the target first data submatrix and the target second data submatrix, the first computing device can perform faster matrix multiplication calculations based on the copied first candidate data submatrix (or the matrix block corresponding to the first candidate data submatrix), which can further improve the overlap efficiency between the computing communications of the first computing device such as GPU, NPU, GPU-like or XPU.

[0070] After executing S102 , the first computing device in this embodiment executes S103 to process the copied first candidate data submatrix and its corresponding target data submatrix in response to obtaining a first processing result between the target first data submatrix and the target second data submatrix.

[0071] In this embodiment, if the first candidate data submatrix is ​​the first data submatrix in the other N-1 computing devices, the target data submatrix corresponding to the first candidate data submatrix is ​​the target second data submatrix in the first computing device; if the first candidate data submatrix is ​​the second data submatrix in the other N-1 computing devices, the target data submatrix corresponding to the first candidate data submatrix is ​​the target first data submatrix in the first computing device.

[0072] After the first computing device in this embodiment executes S103 to determine that the matrix multiplication calculation between the target first data submatrix and the target second data submatrix has been completed, it can immediately perform matrix multiplication calculation on the first candidate data submatrix copied from the other N-1 computing devices and its corresponding target data submatrix.

[0073] It can be understood that, when executing S103, the first computing device in this embodiment may also include the following contents: obtaining the copy status of the first candidate data submatrix; in response to determining that the copy status is an incomplete copy, while processing the copied first candidate data submatrix and its corresponding target data submatrix, continue to copy the remaining first candidate data submatrices in other N-1 computing devices.

[0074] After executing S103, the first computing device in this embodiment executes S104 in response to obtaining the second processing result between the first candidate data submatrix and its corresponding target data submatrix, and obtains the target processing result of the first computing device according to the first processing result and the candidate processing result.

[0075] In this embodiment, the target processing result of the first computing device is spliced ​​with the N-1 target processing results of other N-1 computing devices to obtain an output data matrix, which is the target processing result between the first data matrix and the second data matrix.

[0076] When the first computing device in this embodiment executes S104, the second processing result obtained is the processing result between the complete first candidate data sub-matrix and its corresponding target data sub-matrix.

[0077] The first computing device in this embodiment, after executing S104 to obtain the second processing result between the first candidate data submatrix and its corresponding target data submatrix, can close the matrix multiplication calculation process, and then obtain the target processing result corresponding to the first computing device based on the first processing result and the second processing result.

[0078] In some special cases, for example, when the first candidate data submatrix copied by the first computing device is the first data submatrix in other N-1 computing devices, the first computing device in this embodiment obtains the target processing result corresponding to the first computing device based on the first calculation result and the second calculation result copied from the other N-1 computing devices when executing S104.

[0079] That is to say, the first computing device in this embodiment can obtain the target processing result based on the first calculation result and the second calculation result calculated by itself, or it can obtain the target processing result based on the first calculation result calculated by itself and the second calculation result calculated by other computing devices, which can further improve the accuracy of the obtained target processing result.

[0080] When the parallel processing method of the above model is used in this embodiment, the general matrix multiplication kernel only needs to be called once as a whole, thereby avoiding the problem of multiple calls to the general matrix multiplication kernel by the first computing device such as GPU, NPU, GPU-like or XPU, and improving the computational efficiency of the first computing device such as GPU, NPU, GPU-like or XPU when performing matrix multiplication; and the first computing device such as GPU, NPU, GPU-like or XPU in this embodiment can perform matrix multiplication calculation when determining that the data sub-matrix required for the matrix multiplication calculation is ready, which will not affect the copying process of the data sub-matrix, so that when the model is processed in parallel, the calculation and communication of the first computing device such as GPU, NPU, GPU-like or XPU overlap with each other, and will not affect the computational efficiency of matrix multiplication, which can greatly improve the efficiency of the calculation and communication overlap of the first computing device such as GPU, NPU, GPU-like or XPU, thereby more efficiently achieving the purpose of the first computing device such as GPU, NPU, GPU-like or XPU performing distributed parallel matrix calculation.

[0081] In this embodiment, the weight matrix may include parameters of some network layers in the target model. Accordingly, the processing result between the feature matrix and the weight matrix may be the processing result of some network layers in the target model, that is, the phased processing result.

[0082] Therefore, in practical applications, when the target processing results of each computing device are obtained, it can be determined whether to splice the target processing results of each computing device according to the structure of the model. For example, it can be chosen to continue to maintain the segmentation state to process the next network layer, or it can be chosen to splice the processing results of each computing device to obtain the target processing result between the feature matrix and the weight matrix.

[0083] Figure 2 is a schematic diagram according to the second embodiment of the present disclosure. Figure 2 As shown in , in this embodiment, when executing S104 "obtaining a target processing result of the first computing device according to the first processing result and the second processing result", the implementation method that can be adopted is:

[0084] S201, constructing a segmentation method set according to the first segmentation method, the second segmentation method and a third segmentation method corresponding to the output data matrix;

[0085] S202, determining a second candidate data sub-matrix according to the segmentation method set;

[0086] S203, copying the second candidate data sub-matrix in the other N-1 computing devices;

[0087] S204: Obtain a target processing result of the first computing device according to the first processing result, the second processing result, and the second candidate data sub-matrix.

[0088] That is to say, the first computing device in this embodiment can also determine the second candidate data submatrix to be copied from the other N-1 computing devices according to the segmentation method set constructed by the segmentation method, and then obtain the corresponding target processing result according to the first processing result obtained by itself, the second processing result and the copied second candidate data submatrix.

[0089] In this embodiment, different sets of segmentation methods correspond to different second candidate data sub-matrices; wherein the second candidate data sub-matrix is ​​a second calculation result obtained by matrix multiplication calculation by other N-1 computing devices.

[0090] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure. Figure 3 As shown in , in this embodiment, when executing S102 "copying the first candidate data submatrix in other N-1 computing devices", the implementation method that can be adopted is:

[0091] S301, obtaining a preset cycle order among the N computing devices;

[0092] S302 . According to the preset loop order, sequentially copy the first candidate data sub-matrices in the computing devices located after the first computing device.

[0093] That is to say, the first computing device in this embodiment copies the first candidate data sub-matrix in the computing device located after it in sequence according to the preset loop order between the N computing devices, thereby ensuring the orderly progress of the data copy process in the first computing device such as the GPU, NPU, GPU-like or XPU.

[0094] In this embodiment, the preset loop sequence refers to the sequence of loops formed between the devices. For example, if the N computing devices include computing device 0, computing device 1, computing device 2, and computing device 3, the preset loop sequence is the sequence of the loop formed by computing device 0-computing device 1-computing device 2-computing device 3-computing device 0.

[0095] When the first computing device in this embodiment executes S302, it copies the first candidate data submatrix in the computing device located after the first computing device in sequence, so that after completing the copying of the candidate data submatrix in the current computing device, it copies the candidate data submatrix in the computing device next to the current computing device, and continues in this way.

[0096] Figure 4 is a schematic diagram according to a fourth embodiment of the present disclosure. Figure 4 , a schematic diagram of the first method of legally splitting the first data matrix, the second data matrix and the output data matrix is ​​shown, wherein the first splitting method is row splitting, the second splitting method is row splitting, and the third splitting method is row splitting.

[0097] Figure 5 is a schematic diagram according to a fifth embodiment of the present disclosure. Figure 5 , a schematic diagram of a second method of legally splitting the first data matrix, the second data matrix and the output data matrix is ​​shown, wherein the first splitting method is row splitting, the second splitting method is row splitting, and the third splitting method is column splitting.

[0098] Figure 6 A schematic diagram of a sixth embodiment according to the present disclosure. Figure 6 , a schematic diagram of a third method of legally splitting the first data matrix, the second data matrix and the output data matrix is ​​shown, wherein the first splitting method is row splitting, the second splitting method is column splitting, and the third splitting method is row splitting.

[0099] Figure 7 is a schematic diagram according to a seventh embodiment of the present disclosure. Figure 7 , a schematic diagram of a fourth method of legally splitting the first data matrix, the second data matrix and the output data matrix is ​​shown, wherein the first splitting method is row splitting, the second splitting method is column splitting, and the third splitting method is column splitting.

[0100] Figure 8 is a schematic diagram according to an eighth embodiment of the present disclosure. Figure 8 , a fifth schematic diagram of legal matrix segmentation of the first data matrix, the second data matrix and the output data matrix is ​​shown, wherein the first segmentation method is column segmentation, the second segmentation method is row segmentation, and the third segmentation method is row segmentation.

[0101] Fig. 9 is a schematic diagram of a ninth embodiment of the present disclosure. Fig. 9 , a sixth schematic diagram of legal matrix segmentation of the first data matrix, the second data matrix and the output data matrix is ​​shown, wherein the first segmentation method is column segmentation, the second segmentation method is row segmentation, and the third segmentation method is column segmentation.

[0102] Fig.10 is a schematic diagram according to a tenth embodiment of the present disclosure. Fig.10 , a seventh schematic diagram of legal matrix segmentation of the first data matrix, the second data matrix and the output data matrix is ​​shown, wherein the first segmentation method is column segmentation, the second segmentation method is column segmentation, and the third segmentation method is row segmentation.

[0103] Fig.11 is a schematic diagram according to an eleventh embodiment of the present disclosure. Fig.11 , a schematic diagram of an eighth method of legally splitting the first data matrix, the second data matrix and the output data matrix is ​​shown, wherein the first splitting method is column splitting, the second splitting method is column splitting, and the third splitting method is column splitting.

[0104] Fig.12 is a schematic diagram of a twelfth embodiment of the present disclosure. Fig.12 A schematic diagram of the slicing communication module of this embodiment is shown in FIG. 1 ; in this embodiment, the slicing unit in the slicing communication module is used to use any legal slicing method (for example, any slicing method in the above-mentioned embodiments 4 to 11) to slice the first data matrix, the second data matrix and the output data matrix; the communication unit in the slicing communication module is used to select a target communication method from the candidate communication methods in the corresponding slicing method set according to the slicing method set used by the slicing unit.

[0105] Fig.13 is a schematic diagram of a thirteenth embodiment of the present disclosure. Fig.13 : A schematic diagram of selecting a target communication mode when the first method is used to legally split the matrix in this embodiment is shown in FIG: the communication unit selects a target communication mode from a plurality of first candidate communication modes corresponding to the first split mode set according to a first split mode set obtained by the first split mode being row splitting, the second split mode being row splitting, and the third split mode being row splitting.

[0106] Fig.14 is a schematic diagram according to a fourteenth embodiment of the present disclosure. Fig.14 : A schematic diagram of selecting a target communication mode when the second method of legally dividing a matrix is ​​used in this embodiment is shown in FIG: the communication unit selects a target communication mode from a plurality of second candidate communication modes corresponding to the second dividing mode set according to a second dividing mode set obtained by dividing the matrix by rows in the first dividing mode, by rows in the second dividing mode, and by columns in the third dividing mode.

[0107] Fig.15 is a schematic diagram according to a fifteenth embodiment of the present disclosure. Fig.15 : A schematic diagram of selecting a target communication mode when the third method is used to legally split the matrix in this embodiment is shown in FIG: the communication unit selects a target communication mode from a plurality of third candidate communication modes corresponding to the third split mode set according to a third split mode set obtained by using the first split mode as row splitting, the second split mode as column splitting, and the third split mode as row splitting.

[0108] Fig.16 is a schematic diagram according to a sixteenth embodiment of the present disclosure. Fig.16: A schematic diagram of selecting a target communication mode when the fourth method of legally dividing a matrix is ​​used in this embodiment is shown in FIG: the communication unit selects a target communication mode from a plurality of fourth candidate communication modes corresponding to the fourth dividing mode set according to a fourth dividing mode set obtained by dividing the matrix by rows, the second dividing mode by columns, and the third dividing mode by columns.

[0109] Fig.17 is a schematic diagram according to the seventeenth embodiment of the present disclosure. Fig.17 : A schematic diagram of selecting a target communication mode when the fifth method of legally splitting a matrix is ​​used in this embodiment is shown: the communication unit selects a target communication mode from a plurality of fifth candidate communication modes corresponding to the fifth split mode set according to a fifth split mode set obtained by the first split mode being column splitting, the second split mode being row splitting, and the third split mode being row splitting.

[0110] Fig.18 is a schematic diagram according to the eighteenth embodiment of the present disclosure. Fig.18 : A schematic diagram of selecting a target communication mode when the sixth method of legally dividing a matrix is ​​used in this embodiment is shown in FIG: the communication unit selects a target communication mode from a plurality of sixth candidate communication modes corresponding to the sixth dividing mode set according to a sixth dividing mode set obtained by dividing the matrix by columns as the first dividing mode, by rows as the second dividing mode, and by columns as the third dividing mode.

[0111] Fig.19 is a schematic diagram of a nineteenth embodiment of the present disclosure. Fig.19 : A schematic diagram of selecting a target communication mode when the seventh legal segmentation of a matrix is ​​performed in this embodiment is shown in FIG: the communication unit selects a target communication mode from a plurality of seventh candidate communication modes corresponding to the seventh segmentation mode set according to a seventh segmentation mode set obtained by using the first segmentation mode as column segmentation, the second segmentation mode as column segmentation, and the third segmentation mode as row segmentation.

[0112] Fig. 20 is a schematic diagram of the twentieth embodiment of the present disclosure. Fig. 20 : A schematic diagram of selecting a target communication mode when the eighth method is used to legally split a matrix in this embodiment is shown in the figure: the communication unit selects a target communication mode from a plurality of seventh candidate communication modes corresponding to the seventh split mode set according to a seventh split mode set obtained by the first split mode being column splitting, the second split mode being column splitting, and the third split mode being row splitting.

[0113] Fig.21 is a schematic diagram according to the twenty-first embodiment of the present disclosure. Fig.21The figure shows a processing flow chart of the first computing device performing parallel processing of the model based on a target communication method selected in the thirteenth embodiment; in this embodiment, the first segmentation method is row segmentation, the second segmentation method is row segmentation, and the third segmentation method is row segmentation.

[0114] like Fig.21 As shown in , this embodiment includes two computing devices, namely GPU1 and GPU2, wherein GPU1 is a first computing device, and GPU2 is other computing devices corresponding to GPU1; the target first data submatrix obtained by GPU1 is A1 composed of A11 and A12, and the target second data submatrix is ​​B1; the target first data submatrix obtained by GPU2 is A2 composed of A21 and A22, and the target second data submatrix is ​​B2.

[0115] In this embodiment, GPU1 performs matrix multiplication calculation based on the obtained A11 and B1, and simultaneously copies B2 (i.e., the first candidate data submatrix) from GPU2; after GPU1 determines that the matrix multiplication calculation between A11 and B1 is completed to obtain (A11*B1=C11), GPU1 continues to perform matrix multiplication calculation based on A12 and the copied B2 to obtain (A12*B2=C22); GPU1 obtains (C11+C12=C1) based on the obtained C11 (i.e., the first calculation result) and C12 (i.e., the second calculation result), and C1 is the target calculation result of the first computing device.

[0116] Fig. 22 is a schematic diagram of the twenty-second embodiment of the present disclosure. Fig. 22 A processing flow chart is shown in which the first computing device performs parallel processing of the model based on another target communication method selected in the thirteenth embodiment; in this embodiment, the first segmentation method is row segmentation, the second segmentation method is row segmentation, and the third segmentation method is row segmentation.

[0117] like Fig. 22 As shown in , this embodiment includes two computing devices, namely GPU1 and GPU2, wherein GPU1 is a first computing device, and GPU2 is other computing devices corresponding to GPU1; the target first data submatrix obtained by GPU1 is A1 composed of A11 and A12, and the target second data submatrix is ​​B1; the target first data submatrix obtained by GPU2 is A2 composed of A21 and A22, and the target second data submatrix is ​​B2.

[0118] In this embodiment, GPU1 performs matrix multiplication calculation based on the obtained A11 and B1, and simultaneously copies A21 (i.e., the first candidate data submatrix) from GPU2; after GPU1 determines that the matrix multiplication calculation between A11 and B1 is completed to obtain (A11*B1=C11), it performs calculation (A21*B1=C22); after GPU1 determines that the calculation of C22 is completed, it copies C12 from GPU2; GPU1 obtains (C11+C12=C1) based on the obtained C11 (i.e., the first calculation result) and C12, and C1 is the target calculation result of the first computing device.

[0119] Fig.23 is a schematic diagram according to the twenty-third embodiment of the present disclosure. Fig.23 The figure shows a processing flow chart of the first computing device performing parallel processing of the model based on a target communication method selected in the fourteenth embodiment; in this embodiment, the first segmentation method is row segmentation, the second segmentation method is row segmentation, and the third segmentation method is column segmentation.

[0120] like Fig.23 As shown in, this embodiment includes two computing devices, namely GPU1 and GPU2, wherein GPU1 is a first computing device, and GPU2 is other computing devices corresponding to GPU1; the target first data submatrix obtained by GPU1 is A1 composed of A11 and A12, and the target second data submatrix is ​​B1 composed of B11 and B12; the target first data submatrix obtained by GPU2 is A2 composed of A21 and A22, and the target second data submatrix is ​​B2 composed of B21 and B22.

[0121] In this embodiment, GPU1 performs matrix multiplication calculation according to the obtained A11 and B11, and A11 and B12, and simultaneously copies B21 and B22 (i.e., the first candidate data submatrix) from GPU2; after GPU1 determines that the matrix multiplication calculation between A11 and B11 is completed to obtain (A11*B11=C11_1), and the matrix multiplication result between A11 and B12 is obtained to obtain (A11*B12=C21_1), it continues to calculate the matrix multiplication result according to A12 and the copy. GPU1 performs matrix multiplication calculation on B21, A12 and the copied B22 to obtain (A12*B21=C11_2) and (A12*B22=C21_2) respectively; GPU1 obtains C11 according to C11_1 (i.e. the first calculation result) and C11_2 (i.e. the second calculation result), and obtains (C11+C12=C1) according to C12 copied from GPU2 (i.e. the second candidate data submatrix), and C1 is the target calculation result of the first computing device.

[0122] Fig.24 is a schematic diagram of the twenty-fourth embodiment of the present disclosure. Fig.24 A processing flow chart is shown in which the first computing device performs parallel processing of the model based on another target communication method selected in the fourteenth embodiment; in this embodiment, the first segmentation method is row segmentation, the second segmentation method is row segmentation, and the third segmentation method is column segmentation.

[0123] like Fig.24 As shown in , this embodiment includes two computing devices, namely GPU1 and GPU2, wherein GPU1 is a first computing device, and GPU2 is other computing devices corresponding to GPU1; the target first data submatrix obtained by GPU1 is A1 composed of A11 and A12, and the target second data submatrix is ​​B1; the target first data submatrix obtained by GPU2 is A2 composed of A21 and A22, and the target second data submatrix is ​​B2.

[0124] In this embodiment, GPU1 performs matrix multiplication calculation based on the obtained A11 and B11, and simultaneously copies B2 (i.e., the first candidate data submatrix) from GPU2; after GPU1 determines that the matrix multiplication calculation between A11 and B11 is completed to obtain (A11*B11=C11_C21_1), GPU1 continues to perform matrix multiplication calculation based on the copied B2 and A12 to obtain (A12*B2=C11_C21_2), and obtains C11_C21 based on C11_C21_2 and C11_C21_1; after GPU1 divides C11_C21 into C11 and C21, it obtains (C11+C12=C1) based on C12 copied from GPU2 and C11 obtained by the division, and C1 is the target calculation result of the first computing device.

[0125] Fig.25 It is a schematic diagram of the 25th embodiment of the present disclosure. Fig.25 : A framework diagram of the first computing device performing parallel processing of the model is shown in FIG. Fig.25 The first computing device in the N-1 computing devices performs operations on the first candidate data sub-matrices in the other N-1 computing devices while performing matrix multiplication calculations, thereby achieving overlap of calculation and communication, and the first computing device also continuously performs matrix multiplication calculations based on the copied first candidate data sub-matrix, thereby improving the calculation efficiency of matrix multiplication.

[0126] Fig.2626 is a schematic diagram according to the present disclosure. This embodiment shows the overlap between computing communications of computing devices based on the Wrap level. In this embodiment, GPU1 and GPU2 further divide the sub-matrix for matrix multiplication into matrix blocks corresponding to the Warp level, and then perform matrix multiplication calculations based on the matrix blocks, and perform IPC Copy during the calculation process to copy the corresponding data from other computing devices for subsequent matrix multiplication calculations.

[0127] Fig. 27 is a schematic diagram according to the twenty-seventh embodiment of the present disclosure. Fig. 27 As shown, the parallel processing device 2700 of the model of this embodiment is located in the first computing device among N computing devices, and includes:

[0128] An acquisition unit 2701 is used to acquire a target first data submatrix from among the N first data submatrices, and a target second data submatrix from among the N second data submatrices; wherein the N first data submatrices are obtained by dividing the first data matrix in a first dividing manner, and the N second data submatrices are obtained by dividing the second data matrix in a second dividing manner; and N is a positive integer greater than or equal to 2;

[0129] A first processing unit 2702 is used to start a matrix multiplication calculation process, and while processing the target first data submatrix and the target second data submatrix, copy the first candidate data submatrix in other N-1 computing devices;

[0130] A second processing unit 2703 is configured to process the copied first candidate data submatrix and its corresponding target data submatrix in response to obtaining a first processing result between the target first data submatrix and the target second data submatrix;

[0131] The third processing unit 2704 is used to obtain the target processing result of the first computing device according to the first processing result and the second processing result in response to obtaining the second processing result between the first candidate data submatrix and its corresponding target data submatrix; wherein the target processing result of the first computing device is used to be spliced ​​with the N-1 target processing results of the other N-1 computing devices to obtain the target processing result between the first data matrix and the second data matrix.

[0132] In this embodiment, each first data submatrix of the N first data submatrices and each second data submatrix of the N second data submatrices are distributed to different computing devices, and the acquisition unit 2701 acquires the distributed first data submatrix as the target first data submatrix and the distributed second data submatrix as the target second data submatrix.

[0133] In this embodiment, after the first computing device obtains the target first data submatrix and the target second data submatrix by the acquisition unit 2701, the first processing unit 2702 starts the matrix multiplication calculation process, processes the received target first data submatrix and the target second data submatrix, and simultaneously copies the first candidate data submatrices in the other N-1 computing devices.

[0134] The first processing unit 2702 may start the matrix multiplication calculation process by calling a general matrix multiplication (GEMM) kernel.

[0135] After completing the start of the matrix multiplication calculation process, the first processing unit 2702 can copy the first candidate data submatrices in other N-1 computing devices while performing matrix multiplication on the target first data submatrix and the target second data submatrix.

[0136] That is to say, the first processing unit 2702 communicates with other N-1 computing devices while performing matrix multiplication on the existing data sub-matrices, thereby copying the first candidate data sub-matrix from the other N-1 computing devices, and can achieve overlap between calculation and communication when the first computing device such as GPU, NPU, GPU-like or XPU performs parallel processing.

[0137] When the first processing unit 2702 processes the received target first data submatrix and target second data submatrix, the implementation method that can be adopted is: according to the first preset block size, the target first data submatrix is ​​divided into multiple target first matrix blocks, and the target second data submatrix is ​​divided into multiple target second matrix blocks; based on the obtained multiple target first matrix blocks and multiple target second matrix blocks, the processing result between the target first data submatrix and the target second data submatrix is ​​obtained, and the obtained processing result is the matrix multiplication result.

[0138] In this embodiment, the first preset block size may be a block size that matches a Warp size (Warp is a basic unit for scheduling and execution in a GPU).

[0139] That is to say, the first processing unit 2702 achieves the purpose of performing Warp-level calculations within the called general matrix multiplication core by dividing the data sub-matrices and then performing matrix multiplication calculations between the data sub-matrices according to the matrix blocks obtained by the division. This can improve the computing efficiency of the first computing device such as the GPU, NPU, GPU-like or XPU when performing matrix multiplication.

[0140] In this embodiment, the first candidate data submatrix to be copied by the first processing unit 2702 from the other N-1 computing devices may be all or part of the first data submatrix corresponding to the other N-1 computing devices, or may be all or part of the second data submatrix corresponding to the other N-1 computing devices.

[0141] When the first processing unit 2702 copies the first candidate data submatrix in the other N-1 computing devices, the implementation method that can be adopted is: construct a segmentation method set according to the first segmentation method, the second segmentation method and the third segmentation method corresponding to the output data matrix; determine the first candidate data submatrix according to the constructed segmentation method set; and copy the first candidate data submatrix in the other N-1 computing devices.

[0142] The first processing unit 2702 accesses the memory of other N-1 computing devices to copy the first candidate data sub-matrices of other N-1 computing devices.

[0143] In this embodiment, the output data matrix is ​​the target processing result between the first data matrix and the second data matrix; the third segmentation method can be row segmentation (i.e. segmenting the output data matrix in the row direction) or column segmentation (i.e. segmenting the output data matrix in the column direction).

[0144] In this embodiment, different sets of segmentation methods correspond to different types of first candidate data submatrices. Therefore, through the constructed set of segmentation methods, the first computing device determines what type of first candidate data submatrix to copy from the other N-1 computing devices.

[0145] In this embodiment, the first processing unit 2702 obtains the first candidate data sub-matrix required for matrix multiplication calculation from the other N-1 computing devices by copying, thereby avoiding the steps of sending and receiving sub-matrices between computing devices, and can reduce the time required for the first computing device such as GPU, NPU, GPU-like or XPU to obtain the first candidate data sub-matrix from the other N-1 computing devices, thereby improving the efficiency of subsequent matrix multiplication calculations based on the first candidate data sub-matrix.

[0146] The first processing unit 2702 may store the copied first candidate data sub-matrices corresponding to different other computing devices into the memory of the first computing device so as to be obtained during subsequent processing.

[0147] When the first processing unit 2702 copies the first candidate data sub-matrix in the other N-1 computing devices, it can copy the first candidate data sub-matrix in the other N-1 computing devices multiple times according to the second preset block size, that is, copy the matrix block corresponding to the second preset block size from the other N-1 computing devices each time; wherein, the second preset block size in this embodiment may be the same as or different from the first preset block size, and this embodiment may set the second preset block size according to actual needs.

[0148] That is to say, the first processing unit 2702 can copy the first candidate data submatrix in the other N-1 computing devices multiple times according to smaller blocks, so that after completing the processing between the target first data submatrix and the target second data submatrix, the first computing device can perform faster matrix multiplication calculations based on the copied first candidate data submatrix (or the matrix block corresponding to the first candidate data submatrix), which can further improve the overlap efficiency between the computing communications of the first computing device such as GPU, NPU, GPU-like or XPU.

[0149] When the first processing unit 2702 copies the first candidate data sub-matrices in the other N-1 computing devices, the implementation method that can be adopted is: obtaining a preset loop order between the N computing devices; and according to the preset loop order, sequentially copying the first candidate data sub-matrices in the computing devices located after the first computing device.

[0150] That is, the first processing unit 2702 copies the first candidate data sub-matrix in the computing device located behind it in sequence according to the preset loop order among the N computing devices, so as to ensure the orderly progress of the data copying process.

[0151] The first processing unit 2702 sequentially copies the first candidate data submatrix in the computing devices located after the first computing device, and after completing the copying of the candidate data submatrix in the current computing device, copies the candidate data submatrix in the computing device next to the current computing device, and continues in this manner.

[0152] After the first computing device in this embodiment completes the execution of the first processing unit 2702, the second processing unit 2703 processes the copied first candidate data submatrix and its corresponding target data submatrix in response to obtaining the first processing result between the target first data submatrix and the target second data submatrix.

[0153] In this embodiment, if the first candidate data submatrix is ​​the first data submatrix in the other N-1 computing devices, the target data submatrix corresponding to the first candidate data submatrix is ​​the target second data submatrix in the first computing device; if the first candidate data submatrix is ​​the second data submatrix in the other N-1 computing devices, the target data submatrix corresponding to the first candidate data submatrix is ​​the target first data submatrix in the first computing device.

[0154] After determining that the matrix multiplication calculation between the target first data submatrix and the target second data submatrix has been completed, the second processing unit 2703 can immediately perform matrix multiplication calculation on the first candidate data submatrix copied from the other N-1 computing devices and its corresponding target data submatrix.

[0155] It is understandable that the second processing unit 2703 can also perform the following: obtain the copy status of the first candidate data submatrix; in response to determining that the copy status is an incomplete copy, while processing the copied first candidate data submatrix and its corresponding target data submatrix, continue to copy the remaining first candidate data submatrices in other N-1 computing devices.

[0156] After the first computing device in this embodiment completes the execution of the second processing unit 2703, the third processing unit 2704 responds to the second processing result between the first candidate data submatrix and its corresponding target data submatrix, and obtains the target processing result of the first computing device based on the first processing result and the candidate processing result.

[0157] In this embodiment, the target processing result of the first computing device is spliced ​​with the N-1 target processing results of other N-1 computing devices to obtain an output data matrix, which is the target processing result between the first data matrix and the second data matrix.

[0158] The second processing result obtained by the third processing unit 2704 is the processing result between the complete first candidate data sub-matrix and its corresponding target data sub-matrix.

[0159] After obtaining the second processing result between the first candidate data submatrix and its corresponding target data submatrix, the third processing unit 2704 can close the matrix multiplication calculation process, and then obtain the target processing result corresponding to the first computing device based on the first processing result and the second processing result.

[0160] In some special cases, for example, when the first candidate data submatrix copied by the first computing device is the first data submatrix in other N-1 computing devices, the third processing unit 2704 also obtains the target processing result corresponding to the first computing device based on the first calculation result and the second calculation result copied from the other N-1 computing devices.

[0161] That is to say, the third processing unit 2704 can obtain the target processing result based on the first calculation result and the second calculation result calculated by itself, or it can obtain the target processing result based on the first calculation result calculated by itself and the second calculation result calculated by other computing devices, which can further improve the accuracy of the obtained target processing result.

[0162] When the third processing unit 2704 obtains the target processing result of the first computing device based on the first processing result and the second processing result, the implementation method that can be adopted is: constructing a segmentation method set according to the first segmentation method, the second segmentation method and the third segmentation method corresponding to the output data matrix; determining the second candidate data sub-matrix according to the segmentation method set; copying the second candidate data sub-matrix in other N-1 computing devices; and obtaining the target processing result of the first computing device according to the first processing result, the second processing result and the second candidate data sub-matrix.

[0163] That is to say, the third processing unit 2704 can also determine the second candidate data submatrix to be copied from the other N-1 computing devices according to the segmentation method set constructed by the segmentation method, and then obtain the corresponding target processing result according to the first processing result obtained by itself, the second processing result and the copied second candidate data submatrix.

[0164] In this embodiment, different sets of segmentation methods correspond to different second candidate data sub-matrices; wherein the second candidate data sub-matrix is ​​a second calculation result obtained by matrix multiplication calculation by other N-1 computing devices.

[0165] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0166] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0167] like Fig.28, is a block diagram of an electronic device according to a parallel processing method of a model of an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0168] like Fig.28 As shown, the device 2800 includes a computing unit 2801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 2802 or a computer program loaded from a storage unit 2808 into a random access memory (RAM) 2803. In RAM 2803, various programs and data required for the operation of the device 2800 can also be stored. The computing unit 2801, ROM 2802, and RAM 2803 are connected to each other via a bus 2804. An input / output (I / O) interface 2805 is also connected to the bus 2804.

[0169] A number of components in the device 2800 are connected to the I / O interface 2805, including: an input unit 2806, such as a keyboard, a mouse, etc.; an output unit 2807, such as various types of displays, speakers, etc.; a storage unit 2808, such as a disk, an optical disk, etc.; and a communication unit 2809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 2809 allows the device 2800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0170] The computing unit 2801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 2801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 2801 performs the various methods and processes described above, such as the parallel processing method of the model. For example, in some embodiments, the parallel processing method of the model may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 2808.

[0171] In some embodiments, part or all of the computer program may be loaded and / or installed on the device 2800 via the ROM 2802 and / or the communication unit 2809. When the computer program is loaded into the RAM 2803 and executed by the computing unit 2801, one or more steps of the parallel processing method of the model described above may be performed. Alternatively, in other embodiments, the computing unit 2801 may be configured to execute the parallel processing method of the model in any other appropriate manner (e.g., by means of firmware).

[0172] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0173] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a parallel processing device of a general-purpose computer, a special-purpose computer, or other programmable model, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.

[0174] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0175] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0176] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0177] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server for a distributed system, or a server combined with a blockchain.

[0178] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0179] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A parallel processing method of a model, applied to a first computing device among N computing devices, comprising: Obtain a target first data submatrix from among the N first data submatrices, and a target second data submatrix from among the N second data submatrices; wherein the N first data submatrices are obtained by dividing the first data matrix in a first dividing manner, and the N second data submatrices are obtained by dividing the second data matrix in a second dividing manner; the first dividing manner and the second dividing manner are legal dividing manners that support matrix multiplication calculation of the divided data submatrices; and N is a positive integer greater than or equal to 2; Initiate a matrix multiplication calculation process, and while processing the target first data submatrix and the target second data submatrix, copy the first candidate data submatrix in other N-1 computing devices; In response to obtaining a first processing result between the target first data submatrix and the target second data submatrix, processing the copied first candidate data submatrix and its corresponding target data submatrix; In response to obtaining a second processing result between the first candidate data submatrix and its corresponding target data submatrix, a target processing result of the first computing device is obtained based on the first processing result and the second processing result; wherein the target processing result of the first computing device is used to be spliced ​​with the N-1 target processing results of the other N-1 computing devices to obtain the target processing result between the first data matrix and the second data matrix.

2. The method according to claim 1, wherein: The processing of the target first data sub-matrix and the target second data sub-matrix includes: According to a first preset block size, the target first data submatrix is ​​divided into a plurality of target first matrix blocks, and the target second data submatrix is ​​divided into a plurality of target second matrix blocks; A processing result between the target first data sub-matrix and the target second data sub-matrix is ​​obtained according to the plurality of target first matrix blocks and the plurality of target second matrix blocks.

3. The method according to claim 1, wherein: The copying of the first candidate data submatrix in the other N-1 computing devices includes: Constructing a segmentation method set according to the first segmentation method, the second segmentation method, and a third segmentation method corresponding to the output data matrix; Determining the first candidate data sub-matrix according to the segmentation mode set; The first candidate data sub-matrix in the other N-1 computing devices is copied.

4. The method according to claim 3, wherein: The copying of the first candidate data submatrix in the other N-1 computing devices includes: According to a second preset block size, the first candidate data sub-matrix in the other N-1 computing devices is copied multiple times.

5. The method according to claim 4, wherein: The processing of the copied first candidate data submatrix and its corresponding target data submatrix includes: Obtaining a copy state of the first candidate data submatrix; In response to determining that the copy status is an incomplete copy, while processing the copied first candidate data submatrix and its corresponding target data submatrix, continue to copy the remaining first candidate data submatrices in the other N-1 computing devices.

6. The method according to claim 1, wherein: Obtaining a target processing result of the first computing device according to the first processing result and the second processing result includes: Constructing a segmentation method set according to the first segmentation method, the second segmentation method, and a third segmentation method corresponding to the output data matrix; Determining a second candidate data submatrix according to the segmentation method set; Copying the second candidate data sub-matrices in the other N-1 computing devices; A target processing result of the first computing device is obtained according to the first processing result, the second processing result and the second candidate data sub-matrix.

7. The method according to claim 1, wherein: The copying of the first candidate data submatrix in the other N-1 computing devices includes: Obtaining a preset cycle order among the N computing devices; According to the preset loop order, the first candidate data sub-matrices in the computing devices located after the first computing device are copied in sequence.

8. A parallel processing device for a model, a first computing device located in N computing devices, comprising: An acquisition unit, used for acquiring a target first data submatrix among N first data submatrices, and a target second data submatrix among N second data submatrices; wherein the N first data submatrices are obtained by dividing the first data matrix according to a first dividing method, and the N second data submatrices are obtained by dividing the second data matrix according to a second dividing method; the first dividing method and the second dividing method are legal dividing methods that support matrix multiplication calculation of the divided data submatrices; and N is a positive integer greater than or equal to 2; A first processing unit is used to start a matrix multiplication calculation process, and while processing the target first data submatrix and the target second data submatrix, copy the first candidate data submatrix in other N-1 computing devices; a second processing unit, configured to process the copied first candidate data submatrix and its corresponding target data submatrix in response to obtaining a first processing result between the target first data submatrix and the target second data submatrix; A third processing unit is used to obtain a target processing result of the first computing device in response to obtaining a second processing result between the first candidate data submatrix and its corresponding target data submatrix, based on the first processing result and the second processing result; wherein the target processing result of the first computing device is used to be spliced ​​with the N-1 target processing results of the other N-1 computing devices to obtain the target processing result between the first data matrix and the second data matrix.

9. The device according to claim 8, wherein: When processing the target first data sub-matrix and the target second data sub-matrix, the first processing unit specifically performs: According to a first preset block size, the target first data submatrix is ​​divided into a plurality of target first matrix blocks, and the target second data submatrix is ​​divided into a plurality of target second matrix blocks; A processing result between the target first data sub-matrix and the target second data sub-matrix is ​​obtained according to the plurality of target first matrix blocks and the plurality of target second matrix blocks.

10. The device according to claim 8, wherein: When copying the first candidate data submatrix in the other N-1 computing devices, the first processing unit specifically performs: Constructing a segmentation method set according to the first segmentation method, the second segmentation method, and a third segmentation method corresponding to the output data matrix; Determining the first candidate data sub-matrix according to the segmentation mode set; The first candidate data sub-matrix in the other N-1 computing devices is copied.

11. The device according to claim 10, wherein: When copying the first candidate data submatrix in the other N-1 computing devices, the first processing unit specifically performs: According to a second preset block size, the first candidate data sub-matrix in the other N-1 computing devices is copied multiple times.

12. The device according to claim 11, wherein When processing the copied first candidate data submatrix and its corresponding target data submatrix, the second processing unit specifically performs: Obtaining a copy state of the first candidate data submatrix; In response to determining that the copy status is an incomplete copy, while processing the copied first candidate data submatrix and its corresponding target data submatrix, continue to copy the remaining first candidate data submatrices in the other N-1 computing devices.

13. The device according to claim 8, wherein: When the third processing unit obtains the target processing result of the first computing device according to the first processing result and the second processing result, the third processing unit specifically performs: Constructing a segmentation method set according to the first segmentation method, the second segmentation method, and a third segmentation method corresponding to the output data matrix; Determining a second candidate data submatrix according to the segmentation method set; Copying the second candidate data sub-matrices in the other N-1 computing devices; A target processing result of the first computing device is obtained according to the first processing result, the second processing result and the second candidate data sub-matrix.

14. The device according to claim 8, wherein: When copying the first candidate data submatrix in the other N-1 computing devices, the first processing unit specifically performs: Obtaining a preset cycle order among the N computing devices; According to the preset loop order, the first candidate data sub-matrices in the computing devices located after the first computing device are copied in sequence.

15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

17. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Parallel processing method and device of model, first computing equipment and electronic equipment

    CN116820577A

  • Operation resource processing method and related equipment

    CN118113972A

  • Matrix processing method, processor, system on chip, electronic equipment and storage medium

    CN118656575A

  • Model calculation method and related device

    CN118798275A

  • Operation resource processing method and related device

    WO2024114304A1

Cited By

  • Data processing method and device, electronic equipment, storage medium and program product

    CN121900974A

  • Data processing method and device, electronic equipment, storage medium and program product

    CN121900974B