Data processing method and device based on many-core processor, equipment and medium
By cutting large matrices into submatrices and passing elements between the cores of the multi-core processor, multiple loading problems caused by insufficient storage space are solved, and the efficiency of matrix multiplication calculation is improved.
Patent Information
- Application Number
- CN202411984321.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-06-06
AI Technical Summary
In large matrix multiplication calculation, due to the limited local storage space of the core, sub-matrixes need to be loaded in multiple times, which increases the time to read data and reduces the calculation efficiency.
The product calculation of the sub-matrix is completed by cutting the first matrix and the second matrix to be performed according to the arrangement dimensions of the cores of the multi-core processor and passing elements between the cores.
Reduce the number of data loading times, improve data calculation efficiency, and avoid multiple loading problems caused by insufficient storage space.
Smart Images

Figure CN120104185A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to a data processing method, device, equipment and medium based on a many-core processor. Background Art
[0002] At present, in some technologies, when performing multiplication calculations on a large matrix, the result matrix is divided into multiple sub-result matrices according to the number of cores of the processor, and each core is responsible for the calculation of one or more sub-result matrices. During the calculation of the sub-result matrices, the core needs to load one or more sub-matrices obtained by splitting from the large matrix. In the case where the local storage space of the core is limited, it is necessary to load multiple times. This technical solution of repeatedly loading sub-matrices in multiple times increases the time consumption of data reading and greatly reduces the efficiency of data calculation. Summary of the invention
[0003] In view of this, the present disclosure provides a data processing method based on a many-core processor, a data processing device based on a many-core processor, an electronic device and a computer-readable storage medium, which can reduce the number of data loading times and improve data computing efficiency.
[0004] In a first aspect, the present disclosure provides a data processing method based on a many-core processor, wherein the many-core processor includes multiple cores; the method includes:
[0005] Obtain a first matrix and a second matrix to be multiplied;
[0006] According to the arrangement dimensions of the plurality of cores, the first matrix is cut into at least one first sub-matrix, and the second matrix is cut into at least one second sub-matrix;
[0007] For any of the first submatrix and any of the second submatrix, input elements of the first submatrix and the second submatrix into the multiple cores, and determine a product result of the first submatrix and the second submatrix by transferring elements between the cores;
[0008] Based on the product result of the first submatrix and the second submatrix, a product result of the first matrix and the second matrix is obtained.
[0009] In a second aspect, the present disclosure provides a data processing device based on a many-core processor, wherein the many-core processor includes multiple cores; the device includes:
[0010] A matrix acquisition module, used to acquire a first matrix and a second matrix to be multiplied;
[0011] A matrix cutting module, used for cutting the first matrix into at least one first sub-matrix and cutting the second matrix into at least one second sub-matrix according to the arrangement dimensions of the plurality of cores;
[0012] a submatrix multiplication module, configured to input elements of the first submatrix and the second submatrix into the multiple cores for any of the first submatrix and any of the second submatrix, and determine a product result of the first submatrix and the second submatrix by transferring elements between the cores;
[0013] A matrix multiplication module is used to obtain a product result of the first matrix and the second matrix based on the product result of the first sub-matrix and the second sub-matrix.
[0014] In a third aspect, the present disclosure provides an electronic device, comprising: a memory and a processor, wherein the memory and the processor are communicatively connected to each other, computer instructions are stored in the memory, and the processor executes the above method by executing the computer instructions.
[0015] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the above method.
[0016] In the technical solutions of some embodiments of the present application, the first matrix and the second matrix to be multiplied are divided into the first sub-matrix and the second sub-matrix according to the arrangement dimension of the core, so that the elements in the first sub-matrix and the second sub-matrix can be guaranteed to be in a one-to-one correspondence with the core, that is, each core can save an element in the first sub-matrix and an element in the second sub-matrix respectively, so that in each core, the element product calculation between the first sub-matrix and the second sub-matrix can be performed respectively. Furthermore, by transferring elements between cores, the elements in each core can be updated to complete all element product calculations between the first sub-matrix and the second sub-matrix, and obtain the product result of the first sub-matrix and the second sub-matrix. Since the method disclosed in the present invention relies on multiple cores to jointly complete the element product calculation of a group of sub-matrices, and in the calculation process, the element update is performed by data transfer, therefore, when calculating one of the sub-result matrices, it is only necessary to load the sub-matrix once, without multiple data loading, thereby reducing the number of data loading times and improving data calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related technologies, the drawings required for use in the specific embodiments or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 It is a schematic diagram of matrix multiplication;
[0019] Figure 2 is a schematic diagram of data transfer between cores provided by some embodiments of the present disclosure;
[0020] Figure 3 is a flowchart of a data processing method provided by some embodiments of the present disclosure;
[0021] Figure 4 is a cutting schematic diagram of a first matrix and a second matrix provided by some embodiments of the present disclosure;
[0022] Figure 5 is a schematic diagram of distributing elements in a first sub-matrix and a second sub-matrix in a core provided by an embodiment of the present disclosure;
[0023] Figure 6 is a schematic diagram of a sub-result matrix provided by an embodiment of the present disclosure;
[0024] Figure 7 is a schematic diagram of the distribution of elements in the core after element transfer provided by an embodiment of the present disclosure;
[0025] Figure 8 is a schematic diagram of matrix merging provided by some embodiments of the present disclosure;
[0026] Fig. 9 is a schematic diagram of components of a many-core processor provided by some embodiments of the present disclosure;
[0027] Fig.10 is a schematic diagram of component scheduling of a many-core processor provided by some embodiments of the present disclosure;
[0028] Fig.11 is a module schematic diagram of a data processing device provided by some embodiments of the present disclosure;
[0029] Fig.12 It is a schematic diagram of the structure of an electronic device provided by some embodiments of the present disclosure. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0031] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0032] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "some embodiments" or "the embodiment" should be understood as "at least some embodiments". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.
[0033] Herein, unless explicitly stated, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.
[0034] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0035] It is understandable that before using the technical solutions disclosed in the various embodiments of the present disclosure, the types, scopes of use, usage scenarios, etc. of the information involved in the present disclosure should be informed to relevant users and their authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations. The relevant users may include any type of right holders, such as individuals, enterprises, and groups.
[0036] For example, in response to receiving an active request from a user, a prompt message is sent to the relevant user to clearly prompt the relevant user that the operation requested to be performed will require obtaining and using the information of the relevant user, so that the relevant user can independently choose whether to provide information to software or hardware such as an electronic device, application, server or storage medium that executes the operation of the technical solution of the present disclosure based on the prompt message.
[0037] As an optional but non-limiting implementation, in response to receiving an active request from a relevant user, a prompt message is sent to the relevant user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide information to the electronic device.
[0038] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0039] Combined with reference Figure 1 , which is a schematic diagram of matrix multiplication. Figure 1 In the figure, when matrix A and matrix B are multiplied, the result matrix C shown in the direction of the arrow can be obtained. According to the principle of matrix multiplication, the calculation formula for each element in the result matrix C is as follows:
[0040] D0=A00*B00+A01*B10+A02*B20+……+A07*B70
[0041] D1=A00*B01+A01*B11+A02*B21+……+A07*B71
[0042] …
[0043] D8=A10*B00+A11*B10+A12*B20+……+A17*B70
[0044] D9=A10*B01+A11*B11+A12*B21+……+A17*B70
[0045] …
[0046] And so on. The dimension of the result matrix is related to the dimension of matrix A and matrix B. For example, if the dimension of matrix A is M*K and the dimension of matrix B is K*N, then the dimension of the result matrix is M*N.
[0047] In some current technologies, the result matrix is divided into multiple sub-result matrices according to the number of cores in the processor. Each core is responsible for calculating one or more sub-result matrices. For example, if the processor has 3 cores, then Figure 1The result matrix C is divided into 6 sub-result matrices, and each color area represents a sub-result matrix. Among them, the first core can calculate the sub-result matrix marked with green and blue, the second core can calculate the sub-result matrix marked with yellow and red, and the third core can calculate the sub-result matrix marked with purple and gray. Finally, the sub-result matrices calculated by each core are merged to obtain the result matrix C. This method of calculating the result matrix in parallel with multiple cores can improve the calculation efficiency.
[0048] However, this technology requires the core to have a large local storage space to store the sub-matrices of matrix A and matrix B. For example, when calculating the sub-result matrices marked with green and blue, the core needs to load the 0th, 1st, 2nd, 6th, and 7th rows of matrix A and the 0th, 1st, 2nd, and 3rd columns of matrix B into the local storage space. In the case of insufficient local storage space, the required sub-matrices need to be loaded in multiple times. For example, the first time, the 0th row of matrix A and the 0th and 1st columns of matrix B are loaded, and the second time, the 1st row of matrix A and the 2nd and 3rd columns of matrix B are loaded. This technology of loading in multiple times increases the time consumption of data reading and reduces the efficiency of data calculation, and needs to be improved.
[0049] In view of this, the present disclosure provides a data processing method based on a many-core processor, which can reduce the number of data loading times and improve data calculation efficiency. Among them, the data processing method can be applied to electronic devices. Electronic devices include but are not limited to tablets, laptops, desktop computers, servers, etc. The many-core processor may include at least one core group, and each core group may include multiple cores. The arrangement dimension of the multiple cores in each core group can be M rows and K columns, where the values of M and K are integers greater than 0. For example, the many-core processor may include 3 core groups, and each core group may include 32 cores. The 32 cores in each core group can be arranged in 4 rows and 8 columns. Data can be transferred or exchanged between the cores in each row or column through channels. For example, in combination with reference to Figure 2 , a schematic diagram of data transfer between cores provided for some embodiments of the present disclosure. Figure 2 In the example of the first row, based on the channel, core 7 can pass data to core 6, core 6 can pass data to core 5, ..., core 0 can pass data to core 7. Let's take the first column as an example again. Based on the channel, core 24 can pass data to core 16, core 16 can pass data to core 8, ..., core 0 can pass data to core 24.
[0050] For each core group, the method of the present disclosure may be performed separately. For ease of description, in the subsequent description and some schematic diagrams of the present application, each core group includes 32 cores, and the 32 cores are arranged in 4 rows and 8 columns as an example for illustration.
[0051] Combined with reference Figure 3 , which is a flow chart of a data processing method provided in some embodiments of the present disclosure. Figure 3 In the data processing method, the data processing method comprises the following steps:
[0052] Step S301, obtaining a first matrix and a second matrix to be multiplied.
[0053] Specifically, the first matrix and the second matrix may be matrices generated during the operation of the target electronic device executing the method of the present disclosure, or matrices sent to the target electronic device by other devices, or matrices input by a user through a human-computer interaction interface of the target electronic device. The present disclosure does not limit the way to obtain the first matrix and the second matrix.
[0054] Step S302 : cutting the first matrix into at least one first sub-matrix and cutting the second matrix into at least one second sub-matrix according to the arrangement dimensions of the plurality of cores.
[0055] Specifically, in this embodiment, when the arrangement dimension of multiple cores is M rows and K columns, the first matrix can be cut into at least one first sub-matrix of M rows and K columns, and the second matrix can be cut into at least one second sub-matrix of K rows and M columns.
[0056] For easier understanding, please refer to Figure 4 , which is a cutting schematic diagram of the first matrix and the second matrix provided in some embodiments of the present disclosure. Figure 4 In the example, the arrangement dimension of the cores in the core group is 4 rows and 8 columns, the first matrix can be cut into multiple first sub-matrices with 4 rows and 8 columns, and the second matrix can be cut into multiple second sub-matrices with 4 columns and 8 rows. For example, the gray area in the first matrix can be regarded as one of the first sub-matrices obtained by cutting, and the gray area in the second matrix can be regarded as one of the second sub-matrices obtained by cutting.
[0057] Step S303: for any first sub-matrix and any second sub-matrix, input elements of the first sub-matrix and the second sub-matrix into multiple cores, and determine the product result of the first sub-matrix and the second sub-matrix by transferring elements between the cores.
[0058] In this embodiment, when the elements of the first submatrix and the second submatrix are input into multiple cores, each row element of the first submatrix can be input into the corresponding core row according to the one-to-one correspondence between the matrix rows and the core rows, and each column element of the second submatrix can be input into the corresponding core column according to the one-to-one correspondence between the matrix columns and the core rows.
[0059] Specifically, in practical applications, the m1th row of the first submatrix corresponds to the m1th row of cores, and the n1th column of the second submatrix corresponds to the n1th row of cores. That is, the 1st row of the first submatrix corresponds to the 1st row of cores, the 2nd row of the first submatrix corresponds to the 2nd row of cores, and so on. Alternatively, the 1st column of the second submatrix corresponds to the 1st row of cores, the 2nd column of the second submatrix corresponds to the 2nd row of cores, and so on.
[0060] Of course, it is also possible that the m1th row of the first submatrix corresponds to the m2th row of the core (m1 is not equal to m2), and the n1th column of the second submatrix corresponds to the n2th row of the core (n1 is not equal to n2). For example, the 1st row of the first submatrix corresponds to the 2nd row of the core, and the 2nd row of the first submatrix corresponds to the 3rd row of the core. Alternatively, the 1st column of the second submatrix corresponds to the 2nd row of the core, and the 2nd column of the second submatrix corresponds to the 3rd row of the core.
[0061] As long as the matrix rows of the first submatrix and the core rows satisfy a one-to-one correspondence, and the matrix columns of the second submatrix and the core rows satisfy a one-to-one correspondence, the present disclosure does not limit the specific correspondence.
[0062] For any row in the first submatrix, the elements in the row can be sequentially placed in the corresponding core row. For example, when the first row of the first submatrix corresponds to the first row of cores, the first element in the first row can be placed in the first core of the first row, the second element in the first row can be placed in the second core of the first row, and so on.
[0063] Similarly, for any column in the second submatrix, the elements in the column can be sequentially placed in the corresponding core row. For example, when the first column of the second submatrix corresponds to the first row of the core, the first element in the first column can be placed in the first core of the first row, the second element in the first column can be placed in the second core of the first row, and so on.
[0064] Thus, each core may include one element of the first submatrix and one element of the second submatrix. For ease of understanding, assume Figure 4 The gray area of the first matrix in is the first submatrix, and Figure 4The gray area of the second matrix in is the second submatrix. According to the corresponding relationship, after the first submatrix and the second submatrix are placed in the core, the distribution of the elements in the first submatrix and the second submatrix in the core can be as follows: Figure 5 shown.
[0065] Furthermore, according to the matrix multiplication principle, Figure 4 After multiplying the first submatrix and the second submatrix in the gray area, the product should be as follows Figure 6 The sub-result matrix shown. Figure 6 The calculation formulas for each element are as follows:
[0066] D0=A00*B00+A01*B10+A02*B20+……+A06*B60+A07*B70
[0067] D1=A00*B01+A01*B11+A02*B21+……+A06*B61+A07*B71
[0068] D2=A00*B02+A01*B12+A02*B22+……+A06*B62+A07*B72
[0069] D3=A00*B03+A01*B13+A02*B23+……+A06*B63+A07*B73
[0070] D4=A10*B00+A11*B10+A12*B20+……+A16*B60+A17*B70
[0071] D6=A10*B02+A10*B12+A10*B22+……+A10*B62+A10*B72
[0072] And so on.
[0073] Combining the above calculation formula and Figure 5 It can be seen that after putting the elements in the first sub-matrix and the second sub-matrix into the core according to the corresponding relationship, multiplying the elements in each core and adding the product results of the elements corresponding to each core in each row, one of the elements in the sub-result matrix can be obtained.
[0074] For example, Figure 5 In the first row, the element product result corresponding to the 0th core is A00*B00, the element product result corresponding to the 1st core is A01*B10, and so on. Adding the element product results corresponding to each core in the first row, we get Figure 6 D0 in the sub-result matrix shown.
[0075] Similarly, Figure 5 Add the product of the elements corresponding to each core in the second row to get Figure 6 The resulting matrix is shown in D5.
[0076] Similarly, Figure 5 Add the product of the elements corresponding to each core in the third row to get Figure 6 D15 in the resulting matrix is shown.
[0077] Furthermore, assuming that Figure 5 The elements of the first submatrix in each core remain unchanged, and the elements of the second submatrix are transferred between the cores in the same column, so the elements of the second submatrix in each core can be updated. Figure 5 For example. In the column direction, the elements of each second submatrix in the next row of cores are transferred to the previous row of cores, and the elements of each second submatrix in the first row of cores are transferred to the last row of cores. Figure 5 The element distribution changes shown are Figure 7 shown.
[0078] Based on the updated second sub-matrix elements in each core, the elements in each core can be multiplied again, and the product results of the elements corresponding to each core in each row can be added together to obtain the remaining elements in the sub-result matrix.
[0079] For example, Figure 7 In the first row, the element product result corresponding to the 0th core is A00*B01, the element product result corresponding to the 1st core is A01*B11, and so on. Adding the element product results corresponding to each core in the first row, we get Figure 6 D1 in the sub-result matrix shown.
[0080] Similarly, Figure 7 Add the product of the elements corresponding to each core in the second row to get Figure 6 The resulting matrix is shown as D6.
[0081] Similarly, Figure 7 Add the product of the elements corresponding to each core in the third row to get Figure 6 The result matrix is shown as D11.
[0082] By analogy, continue to transfer the elements of the second submatrix in the column direction, and you can get Figure 6 The other elements in the result matrix are shown.
[0083] For any combination of the first submatrix and the second submatrix, the above step S303 can be performed respectively to obtain the corresponding sub-result matrix. That is, in the case where there are multiple first submatrices or multiple second submatrices, the number of sub-result matrices can be multiple. For example, assuming that the first submatrices A1 and A2 are obtained by cutting from the first matrix, and the second submatrices B1 and B2 are obtained by cutting from the second matrix, then the first submatrix A1 and the second submatrix B1 are multiplied to obtain the corresponding sub-result matrix A1B1, the first submatrix A1 and the second submatrix B2 are multiplied to obtain the corresponding sub-result matrix A1B2, the first submatrix A2 and the second submatrix B1 are multiplied to obtain the corresponding sub-result matrix A2B1, and the first submatrix A2 and the second submatrix B2 are multiplied to obtain the corresponding sub-result matrix A2B2.
[0084] It should be noted that the above is explained by taking the elements of the first submatrix in the core as an example, while the elements of the second submatrix are transferred between the cores in the same column. In practical applications, the elements of the second submatrix in the core may be kept unchanged, while the elements of the first submatrix are transferred between the cores in the same column, to obtain the corresponding sub-result matrix.
[0085] Based on the above description, the element transfer between cores in step S303 may include:
[0086] Keep the elements of the first submatrix in each core unchanged and transfer the elements of the second submatrix between cores in the same column; or
[0087] The elements of the second submatrix in each core are kept unchanged, and the elements of the first submatrix are transferred between cores in the same column.
[0088] And, determining the product result of the first sub-matrix and the second sub-matrix by transferring elements between the cores may include:
[0089] When each element transfer is completed, for any core, the element that remains unchanged in the core is multiplied with the element transferred to the core to obtain the element product result corresponding to the core;
[0090] Based on the element product results corresponding to each row of the core, the product result of the first sub-matrix and the second sub-matrix is determined.
[0091] And, determining the product result of the first submatrix and the second submatrix based on the product result of the elements corresponding to each row of the core may include:
[0092] The product results of the elements in the same row of the core are accumulated, and the accumulated result is used as one of the elements in the sub-result matrix.
[0093] Step S304: obtaining a product result of the first matrix and the second matrix based on the product result of the first sub-matrix and the second sub-matrix.
[0094] Specifically, the multiple sub-result matrices in step S303 may be merged to obtain a result matrix, and the result matrix is used as the product result of the first matrix and the second matrix, wherein different sub-result matrices are located in different areas of the result matrix.
[0095] For example, refer to Figure 8 , which is a schematic diagram of matrix merging provided in some embodiments of the present disclosure. Figure 8 In , after multiplying the first submatrix in the gray area of the first matrix with the second submatrix in the gray area of the second matrix, a sub-result matrix located in the yellow area of the result matrix can be obtained; after multiplying the first submatrix in the gray area of the first matrix with the second submatrix in the green area of the second matrix, a sub-result matrix located in the green area of the result matrix can be obtained; after multiplying the first submatrix in the green area of the first matrix with the second submatrix in the gray area of the second matrix, a sub-result matrix located in the blue area of the result matrix can be obtained; after multiplying the first submatrix in the green area of the first matrix with the second submatrix in the green area of the second matrix, a sub-result matrix located in the gray area of the result matrix can be obtained.
[0096] In summary, in the technical solutions of some embodiments of the present application, the first matrix and the second matrix to be multiplied are divided into the first sub-matrix and the second sub-matrix according to the arrangement dimension of the core, which can ensure that the elements in the first sub-matrix and the second sub-matrix are in a one-to-one correspondence with the core, that is, each core can respectively save an element in the first sub-matrix and an element in the second sub-matrix, so that in each core, the element product calculation between the first sub-matrix and the second sub-matrix can be performed respectively. Furthermore, by transferring elements between cores, the elements in each core can be updated to complete all element product calculations between the first sub-matrix and the second sub-matrix, and obtain the product result of the first sub-matrix and the second sub-matrix. Since the method disclosed in the present invention relies on multiple cores to jointly complete the element product calculation of a group of sub-matrices, and in the calculation process, the element update is performed by data transfer, therefore, when calculating one of the sub-result matrices, it is only necessary to load the sub-matrix once, without multiple data loading, thereby reducing the number of data loading times and improving data calculation efficiency.
[0097] In general, the biggest differences between the present disclosure and some technologies are as follows:
[0098] 1) Different matrix partitioning methods. Some technologies first partition the result matrix based on the number of cores in the processor, while the method disclosed herein partitions the first matrix and the second matrix to be multiplied according to the arrangement dimension of the cores in the processor.
[0099] 2) The product calculation method of the sub-matrix is different. In some technologies, each core is used to calculate one or more sub-result matrices. In other words, the cores of some technologies are independent of each other, and each core loads the data it needs into the local storage space, and calculates the sub-result matrix based on the data in its local storage space; while in the method disclosed in the present invention, multiple cores share data by means of data transmission, and when calculating a sub-result matrix, it is completed by the cooperation of multiple cores, that is, each core completes the calculation of some elements in the sub-result matrix. For ease of understanding, the following is explained by way of example. For example, assume that the processor includes 3 cores. In some technologies, the sub-result matrix A1 is calculated by core 1, the sub-result matrix A2 is calculated by core 2, and the sub-result matrices A3 and A4 are calculated by core 3. These cores each load the data they need to calculate the sub-result matrix. However, in the method disclosed herein, the cores 1, 2, and 3 may first calculate the sub-result matrix A1 through mutual cooperation, and then calculate the sub-result matrix A2 through mutual cooperation, and then calculate the sub-result matrix A3 through mutual cooperation, and finally calculate the sub-result matrix A4 through mutual cooperation.
[0100] In the method disclosed in the present invention, since the data of the two sub-matrices for matrix multiplication calculation are stored in different cores in a dispersed manner, the requirement for the local storage space of the core can be greatly reduced. That is, the calculation of the sub-result matrix can be completed when the local storage space of the core is relatively efficient, and there is no need to load data multiple times, which is more practical.
[0101] Combined with reference Fig. 9 , which is a module schematic diagram of a many-core processor provided in some embodiments of the present disclosure. Fig. 9 In the multi-core processor, DMA, ACE, RMA, and SIMD components are included. ACE is responsible for performing the multiplication operation of the first sub-matrix and the second sub-matrix, DMA is responsible for loading the first sub-matrix and the second sub-matrix from the device-side memory (such as HBM) to the local storage space of the core, RMA is responsible for transferring the elements of the second sub-matrix between the local storage spaces of the cores in the same column, and SIMD is responsible for the accumulation operation of the intermediate results. Based on these components, the following can be implemented: Fig.10 The maximum pipeline parallel scheduling shown can greatly improve data processing efficiency.
[0102] Corresponding to the method, the present disclosure also provides a data processing device based on a multi-core processor. Fig.11 , which is a module schematic diagram of a data processing device provided in some embodiments of the present disclosure. Fig.11In the data processing device, the data processing device comprises:
[0103] A matrix acquisition module 111, used to acquire a first matrix and a second matrix to be multiplied;
[0104] A matrix cutting module 112, configured to cut the first matrix into at least one first sub-matrix and the second matrix into at least one second sub-matrix according to the arrangement dimensions of the plurality of cores;
[0105] The sub-matrix multiplication module 113 is used for inputting the elements of the first sub-matrix and the second sub-matrix into a plurality of cores for any first sub-matrix and any second sub-matrix, and determining the product result of the first sub-matrix and the second sub-matrix by transferring the elements between the cores;
[0106] The matrix multiplication module 114 is used to obtain a product result of the first matrix and the second matrix based on the product result of the first sub-matrix and the second sub-matrix.
[0107] In some embodiments, the arrangement dimension of the multiple cores is M rows and K columns; the sub-matrix multiplication module 113 is specifically used for:
[0108] Cut the first matrix into at least one first sub-matrix with M rows and K columns;
[0109] The second matrix is cut into at least one second sub-matrix having K rows and M columns.
[0110] In some embodiments, the sub-matrix multiplication module 113 is specifically used for:
[0111] According to the one-to-one correspondence between matrix rows and core rows, each row element of the first submatrix is input into the corresponding core row;
[0112] According to the one-to-one correspondence between matrix columns and core rows, each column element of the second submatrix is input into the corresponding core column.
[0113] In some embodiments, each core includes one element of the first sub-matrix and one element of the second sub-matrix; the sub-matrix multiplication module 113 is specifically used for:
[0114] Keep the elements of the first submatrix in each core unchanged and transfer the elements of the second submatrix between cores in the same column; or
[0115] The elements of the second submatrix in each core are kept unchanged, and the elements of the first submatrix are transferred between cores in the same column.
[0116] In some embodiments, the sub-matrix multiplication module 113 is specifically used for:
[0117] When each element transfer is completed, for any core, the element that remains unchanged in the core is multiplied with the element transferred to the core to obtain the element product result corresponding to the core;
[0118] Based on the element product results corresponding to each row of the core, the product result of the first sub-matrix and the second sub-matrix is determined.
[0119] In some embodiments, the product of the first sub-matrix and the second sub-matrix is a sub-result matrix; the sub-matrix multiplication module 113 is specifically used for:
[0120] The product results of the elements in the same row of the core are accumulated, and the accumulated result is used as one of the elements in the sub-result matrix.
[0121] In some embodiments, when there are multiple first sub-matrices or multiple second sub-matrices, the number of sub-result matrices is multiple; the matrix multiplication module 114 is specifically used for:
[0122] A plurality of sub-result matrices are merged to obtain a result matrix, and the result matrix is used as a product result of the first matrix and the second matrix, wherein different sub-result matrices are located in different areas of the result matrix.
[0123] The present disclosure also provides an electronic device having the above Figure 7 The data processing device shown is based on a multi-core processor.
[0124] Combined with reference Fig.12 , is a schematic diagram of the structure of an electronic device provided by some embodiments of the present disclosure. Fig.12 As shown, the electronic device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Fig.12 A processor 10 is taken as an example.
[0125] The processor 10 may be a first PCIe device, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0126] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0127] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0128] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0129] The electronic device further comprises a communication interface 30 for the electronic device to communicate with other devices or a communication network.
[0130] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium and downloaded through a network, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0131] A part of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the existence of computer program instructions in computer-readable media includes, but is not limited to, source files, executable files, installation package files, etc., and accordingly, the way in which computer program instructions are executed by a computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0132] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A data processing method based on a many-core processor, characterized in that: The many-core processor includes a plurality of cores; the method includes: Obtain a first matrix and a second matrix to be multiplied; According to the arrangement dimensions of the plurality of cores, the first matrix is cut into at least one first sub-matrix, and the second matrix is cut into at least one second sub-matrix; For any of the first submatrix and any of the second submatrix, input elements of the first submatrix and the second submatrix into the multiple cores, and determine a product result of the first submatrix and the second submatrix by transferring elements between the cores; Based on the product result of the first submatrix and the second submatrix, a product result of the first matrix and the second matrix is obtained.
2. The method according to claim 1, characterized in that: The arrangement dimension of the multiple cores is M rows and K columns; The step of cutting the first matrix into at least one first sub-matrix and cutting the second matrix into at least one second sub-matrix according to the arrangement dimensions of the plurality of cores comprises: Cutting the first matrix into at least one first sub-matrix with M rows and K columns; Cutting the second matrix into at least one second sub-matrix having K rows and M columns; The values of M and K are integers greater than 0.
3. The method according to claim 2, characterized in that The step of inputting the elements of the first sub-matrix and the second sub-matrix into the plurality of cores comprises: According to the one-to-one correspondence between matrix rows and core rows, input each row element of the first submatrix into the corresponding core row; According to the one-to-one correspondence between matrix columns and core rows, each column element of the second submatrix is input into the corresponding core column respectively.
4. The method according to claim 3, characterized in that Each of the cores includes one element of the first sub-matrix and one element of the second sub-matrix; The element transfer between cores includes: Keeping the elements of the first submatrix in each of the cores unchanged, transferring the elements of the second submatrix between the cores in the same column; or The elements of the second submatrix in each of the cores are kept unchanged, and the elements of the first submatrix are transferred between the cores in the same column.
5. The method according to claim 4, characterized in that The determining the product result of the first sub-matrix and the second sub-matrix by transferring elements between cores includes: When each element transfer is completed, for any of the cores, the elements that remain unchanged in the core and the elements transferred to the core are multiplied to obtain the element product result corresponding to the core; Based on the element product results corresponding to each row of the core, a product result of the first sub-matrix and the second sub-matrix is determined.
6. The method according to claim 5, characterized in that The product result of the first sub-matrix and the second sub-matrix is a sub-result matrix; The determining the product result of the first submatrix and the second submatrix based on the product result of the elements corresponding to each row of the core includes: The product results of the elements in the same row of the core are accumulated, and the accumulated result is used as one of the elements in the sub-result matrix.
7. The method according to claim 1, characterized in that A product result of the first submatrix and the second submatrix is a sub-result matrix, and when there are multiple first submatrices or multiple second submatrices, the number of the sub-result matrices is multiple; The obtaining the product result of the first matrix and the second matrix based on the product result of the first submatrix and the second submatrix includes: A plurality of the sub-result matrices are merged to obtain a result matrix, and the result matrix is used as a product result of the first matrix and the second matrix, wherein different sub-result matrices are located in different areas of the result matrix.
8. A data processing device based on a multi-core processor, characterized in that: The many-core processor includes a plurality of cores; the device includes: A matrix acquisition module, used to acquire a first matrix and a second matrix to be multiplied; A matrix cutting module, used for cutting the first matrix into at least one first sub-matrix and cutting the second matrix into at least one second sub-matrix according to the arrangement dimensions of the plurality of cores; a submatrix multiplication module, configured to input elements of the first submatrix and the second submatrix into the multiple cores for any of the first submatrix and any of the second submatrix, and determine a product result of the first submatrix and the second submatrix by transferring elements between the cores; A matrix multiplication module is used to obtain a product result of the first matrix and the second matrix based on the product result of the first sub-matrix and the second sub-matrix.
9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the data processing method based on a multi-core processor according to any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the data processing method based on a many-core processor according to any one of claims 1 to 7.