Parallel computing program processing method and computing equipment
By adding cutting subroutines and adjusting parallel parameters during the migration process, the high computational complexity problem during CUDA to CANN migration is solved, and a more efficient parallel computing program migration is achieved.
Patent Information
- Application Number
- CN202510309055.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-29
AI Technical Summary
During the migration process, when the prior art migrates the computing program of the computing device deploying CUDA to the computing device deploying CANN, there is a problem of high migration complexity of parallel computing programs. In particular, because CUDA supports multi-dimensional parallel parameters, CANN usually uses a single parallel parameter, resulting in complex computing logic and data processing.
By adding a cutting subroutine during the migration process, the original calculation parameters have different dimensions and generate target calculation parameters with the same dimensions, and adjust the parallel parameters and index logic to adapt to the CANN's computing architecture and reduce the complexity of index logic and parallel computing logic.
It reduces the migration complexity of parallel computing programs when migrating from CUDA to CANN, simplifies computing logic and index processing, and improves computing efficiency.
Smart Images

Figure CN120386555A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of servers, and in particular, to a method for processing parallel computing programs and a computing device. Background Art
[0002] Parallel computing program migration refers to the process of migrating a parallel computing program from one computing device to another. At present, many application scenarios may involve parallel computing program migration. For example, in an enterprise large-scale deployment scenario, deploying computing device A has lower costs than deploying computing device B. Therefore, the computing program of computing device A can be migrated to computing device B to reduce the enterprise's deployment costs.
[0003] During the migration process, the migration of computing logic is involved, that is, the computing logic of one computing device is migrated to another computing device. Computing logic is used to implement parallel computing. For example, for a computing device deployed with a compute unified device architecture (CUDA), parallel computing is implemented by the CUDA kernel function, and for a computing device deployed with a compute architecture for neural networks (CANN), parallel computing is implemented by the CANN kernel function. Since CUDA supports multi-dimensional parallel parameters, while CANN usually uses a single parallel parameter, and the CUDA kernel function can directly perform parallel computing, while the CANN kernel function also requires additional data preprocessing operations such as space allocation, data alignment, and data transfer, etc. Therefore, migrating the computing program of a computing device deployed with CUDA to a computing device deployed with CANN has a relatively high migration complexity for the parallel computing program. Summary of the Invention
[0004] Embodiments of this application provide a method for processing parallel computing programs and a computing device, which are used to reduce the migration complexity of parallel computing programs.
[0005] In a first aspect, embodiments of this application provide a method for processing parallel computing programs, which is applied to a computing device. The computing device includes a first accelerator, and the first accelerator is used to execute parallel computing.
[0006] Specifically, the computing device obtains an original computing program, where the original computing program is a parallel computing program written based on a second accelerator. The original computing program is used to perform parallel computing on original computing parameters based on the second accelerator; the original computing program is adjusted to obtain a target computing program; the target computing program is a parallel computing program written based on a first accelerator. The target computing program includes a cutting subroutine, and the cutting subroutine is used to cut the original computing parameters to obtain target computing parameters with the same dimensions; the target computing program is used to perform parallel computing on the target computing parameters based on the first accelerator.
[0007] Thus, when the parallel computing program migrates from the second accelerator to the first accelerator, a cutting subroutine is added. By means of cutting, the original computing parameters with different original dimensions are made to generate target computing parameters with the same dimensions. For each accelerator, if the dimensions of the computing parameters are the same, for example, all are two-dimensional data structures or all are one-dimensional data structures, then each computing parameter can use the same indexing logic to determine the index, without the need to design additional offsets or other indexing logics for parameters with different dimensions. And for computing parameters with the same dimensions, only the computation needs to be performed at the corresponding index. Therefore, making the dimensions of the computing parameters for parallel computing the same through cutting can reduce the complexity of the indexing logic and the parallel computing logic. That is to say, it directly reduces the parallel computing complexity, thereby reducing the migration complexity of the parallel computing program from the second accelerator to the first accelerator.
[0008] In a specific implementation, the original computing program includes a first subroutine and a second subroutine, where the first subroutine is used to set the parallel parameter of the parallel computing to a first parallel parameter; the second subroutine is used to perform parallel computing on the original computing parameters using the second accelerator according to the first parallel parameter. The target computing program further includes a first sub-target program and a second sub-target program; the first sub-target program is used to set the parallel parameter to a second parallel parameter; the second sub-target program is used to perform parallel computing on the target computing parameters using the first accelerator according to the second parallel parameter.
[0009] The computing device adjusts the parallel parameters set in the first subroutine from the first parallel parameters to the second parallel parameters, and adjusts the first parallel computation in the second subroutine to the second parallel computation; the first parallel computation indicates performing parallel computation on the original computation parameters using the second accelerator according to the first parallel parameters; the second parallel computation indicates performing parallel computation on the target computation parameters using the first accelerator according to the second parallel parameters; converting the data formats of the adjusted first subroutine and second subroutine into the target data format to obtain the first sub-target program and the second sub-target program; the target data format is the data format supported by the first accelerator. That is, the computing device migrates the parallel strategy and the computing logic of the parallel computation in the original computation program to the first accelerator by modifying the parallel parameters to meet the parallel parameters of the first accelerator and converting the data format to the data format supported by the first accelerator.
[0010] In another specific implementation, if the first parallel parameter is a multi-dimensional parallel parameter and the second parallel parameter is a single parallel parameter; adjusting the first index logic in the first parallel computation to the second index logic, and adjusting the first computation logic in the first parallel computation to the second computation logic; the second parallel computation includes the second index logic and the second computation logic; the first index logic is to determine the index based on the first parallel parameter, and the second index logic is to determine the index based on the second parallel parameter and by adding a target offset; the first computation logic indicates performing parallel computation on the original computation parameters, and the second computation logic indicates performing parallel computation on the target computation parameters. Thus, the computing device migrates the computing logic and the index logic in the original computation program to the first accelerator by adjusting the first index logic to the second index logic and adjusting the first computation logic to the second computation logic, and the first accelerator performs parallel computation based on the migrated computing logic and index logic.
[0011] In yet another specific implementation, the target computation program further includes a splicing subroutine; the splicing subroutine is used to embed the parallel computation result into the area corresponding to the original computation parameters. That is, by adding a splicing subroutine in the target computation program, the correct output of the parallel computation result is ensured by adding the splicing subroutine.
[0012] In still another specific implementation, the target computation program further includes a third sub-target program; the third sub-target program is used to instruct the processor to copy the target computation parameters to the first accelerator.
[0013] In still another specific implementation, the target computation program further includes a fourth sub-target program; the fourth sub-target program is used to instruct the first accelerator to copy the parallel computation result to the processor.
[0014] In yet another specific implementation, if the dimensions of the original calculation parameters are different, the computing device cuts the original calculation parameters to obtain target calculation parameters. That is, the computing device cuts the calculation parameters with different dimensions to ensure that the target calculation parameters obtained after cutting have the same dimensions, which helps to reduce the difficulty of parallel computing migration.
[0015] In yet another specific implementation, if the dimensions of the original calculation parameters are different, determine the target elements in the original calculation parameters; the target elements indicate the elements that need to perform parallel computing; cut out the target elements from the original calculation parameters to obtain the target calculation parameters.
[0016] In yet another specific implementation, the first accelerator is a neural network processor NPU, and the NPU deploys a neural network computing architecture CANN for performing parallel computing through a target computing program;
[0017] The second accelerator is a graphics processing unit GPU, and the GPU deploys a compute unified device architecture CUDA for performing parallel computing through a target computing program.
[0018] In a second aspect, an embodiment of the present application provides a computing device, including:
[0019] A memory for storing a program;
[0020] A processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of the first aspect.
[0021] In a third aspect, the present application provides a computer storage medium for storing a computer program. When the computer program is executed, it is used to implement the method provided by any one of the implementation manners in the first aspect of the present application.
[0022] In a fourth aspect, the present application provides a computer program product containing instructions. When it runs on at least one computing device, it causes at least one computing device to implement the method provided by any one of the implementation manners in the first aspect of the present application.
[0023] Any of the above provided parallel computing program processing methods, corresponding computing devices, computer-readable storage media, computer program products, etc. are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0025] Figure 2 It is a system architecture diagram of a computing device provided by an embodiment of the present application;
[0026] Figure 3 A parallel computing program for implementing parallel computing provided by an embodiment of the present application;
[0027] Figure 4 An interaction diagram of a method for processing a parallel computing program provided by an embodiment of the present application;
[0028] Figure 5A A schematic diagram of calculation parameters provided by an embodiment of the present application;
[0029] Figure 5B A comparison schematic diagram of executing parallel computing provided by an embodiment of the present application;
[0030] Figure 5C Another comparison schematic diagram of executing parallel computing provided by an embodiment of the present application;
[0031] Figure 6A An interaction diagram of a method for processing a parallel computing program provided by an embodiment of the present application;
[0032] Figure 6B A logical schematic diagram of a target computing program provided by an embodiment of the present application;
[0033] Figure 7 A structural schematic diagram of a device for processing a parallel computing program provided by an embodiment of the present application. Detailed implementation manners
[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0035] First, introduce the professional terms involved in the embodiments of the present application.
[0036] Parallel computing: Parallel computing is a computing method that splits a computing task into multiple subtasks, and the multiple subtasks are simultaneously executed by different processors. Parallel computing can significantly improve the computing speed and efficiency of computing devices, and is increasingly widely used in fields with high computing requirements such as artificial intelligence, big data processing, and machine learning.
[0037] Kernel function: A function that indicates parallel execution on a Graphics Processing Unit (GPU) or a Neural Processing Unit (NPU). The kernel function can be executed in parallel by multiple threads, with each thread processing different parts of the data. The kernel function improves the computational speed of intensive computational tasks by decomposing the computational tasks into multiple blocks and performing parallel processing on the GPU or NPU (denoted as GPU / NPU).
[0038] GPU: A processor used for image processing and image rendering. In addition, the GPU can process general-purpose parallel computing data by deploying CUDA. Specifically, CUDA enables developers to write parallel computing code on the GPU, thereby leveraging the parallel processing capabilities of the GPU to improve computational performance.
[0039] NPU: A processor that executes artificial intelligence and neural network computing tasks. The NPU can accelerate application programs of the artificial intelligence processor by deploying CANN. Among them, CANN is a heterogeneous computing architecture that can execute various applications and services including parallel computing.
[0040] Parallel computing program: Refers to a series of instructions, codes, or rules for parallel processing of data in an application program or algorithm, including but not limited to the following: algorithms, control structures, functions and procedures, data operations, parallel and synchronization mechanisms, etc. In the embodiments of this application, the parallel computing program is a parallel computing program written by developers for implementing parallel computing through a computing platform. Through the parallel computing program, parallel computing tasks can be achieved.
[0041] The embodiments of this application provide a method for processing a parallel computing program. When migrating from a parallel computing program deployed on a second accelerator to a first accelerator, a cutting subroutine is added. By cutting, the original computing parameters with different dimensions are made to generate target computing parameters with the same dimensions. Since the indexing logics of the first computing device and the second computing device are the same when the computing parameter sizes are the same, the computing logic only needs to add a common offset to each computing parameter. Therefore, the parallel computing complexities of the first computing device and the second computing device are low, and the complexity of parallel computing is low, thereby reducing the migration complexity of the parallel computing program from the second computing device to the first computing device.
[0042] The following introduces the application scenarios of the embodiments of this application with reference to the accompanying drawings.
[0043] Exemplarily, the appendix Figure 1 is a schematic diagram of an application scenario provided by the embodiments of this application. Specifically, it relates to a computing device 10, and the computing device 10 is used to perform parallel computing.
[0044] In actual use, parallel computing relies on an accelerator for parallel computing. For example, the accelerator can be a GPU or an NPU, which is not specifically limited in the embodiments of the present application. However, the accelerator cannot run independently and needs to rely on a processor (Central Processing Unit, CPU) to complete the full parallel computing.
[0045] As Figure 1 shown, the computing device 10 includes a Host side and a Device side. Among them, the Host side dominates the system operation, is responsible for task scheduling and resource management, etc., and the Device side executes specific tasks and is controlled by the Host side, usually a dedicated hardware or subsystem.
[0046] In the embodiments of the present application, the Host side includes a CPU 101 and a Dynamic Random-Access Memory (DRAM) 102, and the Device side includes an accelerator 103 and a DRAM 104.
[0047] Among them, the DRAM 102 on the Host side is used to store the host application program, enabling the CPU to execute related operations based on the application program. The DRAM 104 on the Device side is used to store the device-side application program, enabling the accelerator to perform parallel computing based on the application program.
[0048] Specifically, in parallel computing, the CPU 101 on the Host side performs the following operations: first data transfer, setting parallel policies, and second data transfer, etc.
[0049] Among them, the first data transfer refers to transferring the computing data from the Host side to the accelerator 103.
[0050] Setting parallel policies refers to setting the parallel parameters of parallel computing. The purpose is to control the distribution of parallel computing tasks according to the parallel parameters. The parallel parameters determine the data range of each thread or thread block (and also determine the index of the thread or thread block), that is, the position or dimension of each thread or thread block in the grid. The grid determines the total parallelism, that is, how many parallel operations of thread blocks can be executed simultaneously.
[0051] In one example, the CPU 101 also calls a kernel function while setting the parallel policy to control the accelerator to perform parallel computing operations based on the kernel function.
[0052] The second data transfer refers to controlling the accelerator 103 to copy the parallel computing result from the Device side to the Host side, so that the CPU 101 can process the parallel computing result.
[0053] The accelerator 103 on the device side is used to perform parallel computing. The accelerator can be a GPU or an NPU, or other devices capable of performing parallel computing, and the embodiments of the present application do not specifically limit this.
[0054] In the embodiments of the present application, the accelerator 103 is used to perform the following operations: parallel computing implementation (subsequent explanations will be given taking the kernel function implementation as an example). Parallel computing implementation refers to performing parallel computing on computing parameters based on the set parallel strategy. Specifically, based on the set parallel strategy, the kernel function is used to perform parallel computing on the computing parameters.
[0055] Specifically, in parallel computing, the CPU on the host side performs the following operations: memory allocation, memory release, data transfer, and kernel function call, etc. Memory allocation refers to allocating memory for the parallel computing task. Kernel function call refers to calling the kernel function and setting the parallel strategy, so that the accelerator on the device side can perform parallel computing based on the kernel function and the parallel strategy. Memory release refers to releasing the previously allocated memory.
[0056] The accelerator on the device side is used to perform: performing parallel computing based on the kernel function and the parallel strategy.
[0057] In one example, the DRAM 104 on the device side may include global memory and local memory, and the accelerator 103 includes an Artificial Intelligence Core (AIcore). For specific reference, see Figure 2 as shown. Figure 2 It is a system architecture diagram of a computing device provided by the embodiments of the present application.
[0058] Global memory refers to the memory area that can be accessed by all computing units in the computing device. The storage space of global memory is relatively large, but the access speed of global memory is relatively slow. Local memory refers to the private memory area of a specific computing unit. In the embodiments of the present application, local memory refers to the private memory that the accelerator can access. The storage space of local memory is relatively small, but the access speed is relatively fast. AI Core refers to a hardware unit specifically used to perform artificial intelligence computing, such as matrix computing and convolution computing.
[0059] When performing parallel computing, the CPU 101 first transfers the parallel computing data from the DRAM 102 on the host side to the global memory on the device side for storage. Then, the accelerator 103 extracts the parallel computing data from the global memory and sends the extracted parallel computing data to the local memory. Next, all the AI Cores in the accelerator 103 that perform parallel computing retrieve the parallel computing data from the local memory for calculation and return the calculation results to the global memory. Among them, the AI Core performs parallel computing on the parallel computing data based on the kernel function and parallel parameters. Finally, the accelerator 103 returns the parallel computing results in the global memory to the host side. Thus, the CPU 101 and the accelerator 103 work together to complete the parallel computing.
[0060] Exemplarily, attached Figure 3 is a parallel computing program for implementing parallel computing provided by an embodiment of the present application. This parallel computing program is applied to Figure 2 or Figure 1 the computing device shown, and this parallel computing program includes the following content:
[0061] S310. The CPU 101 performs memory allocation.
[0062] In the embodiment of the present application, the CPU 101 allocates memory for the host side and the device side to ensure that there is sufficient memory space to execute the parallel computing task.
[0063] Specifically, memory is allocated in the DRAM 102 on the host side to store the data and programs of the parallel computing task, and memory is allocated in the DRAM 104 on the device side, including global memory and local memory, to store the data required for the accelerator 103 to perform parallel computing.
[0064] S320. The CPU 10 creates a thread pool and initializes the thread pool.
[0065] Specifically, the CPU 101 reads the parallel computing data from the DRAM 102 on the host side and copies the read data to the global memory on the device side so that the accelerator 103 can access and process the parallel computing data.
[0066] S330. The CPU 101 calls the kernel function and sets the parallel strategy.
[0067] Among them, setting the parallel strategy means setting the parameters of parallel computing.
[0068] S340. The accelerator 103 performs parallel computing.
[0069] Exemplarily, the AI core in the accelerator 103 retrieves the parallel computing data from the local memory and implements parallel computing based on the kernel function and the parallel strategy.
[0070] The S350 synchronizes the parallel computing results with the CPU 101.
[0071] The CPU 101 synchronizes the parallel computing results so that the computing results of all AI cores have been written back to the global memory.
[0072] The S360 controls the CPU 101 to transfer the parallel computing results of the accelerator 103 from the device side to the host side.
[0073] The CPU 101 reads the parallel computing results from the global memory on the device side and transfers the parallel computing results back to the DRAM 102 on the host side.
[0074] The S370, the CPU 101 performs memory release.
[0075] The CPU 101 releases the memory previously allocated for the host side and the device side to avoid memory leakage.
[0076] Thus, parallel computing can be achieved through the above parallel computing program.
[0077] However, in actual use, many application scenarios involve the migration of parallel computing programs. For example, migrating the parallel computing program executed by one accelerator to another accelerator for execution. For the convenience of description, the accelerator before migration is called the second accelerator, and the accelerator after migration is called the first accelerator. Among them, the first accelerator and the second accelerator are different accelerators, and they adopt different parallel parameters.
[0078] Exemplarily, if the first accelerator is an NPU, CANN is deployed on the NPU, and the first accelerator performs parallel computing through CANN. The parallel parameter executed by CANN is single. The second accelerator is a GPU, CUDA is deployed on the GPU, and the second accelerator performs parallel computing through CUDA. The parallel parameters executed by CUDA are multiple, for example, 2, which are (x, y) respectively.
[0079] Due to different parallel parameters, there will be great differences in the computing logics of the two during actual parallel computing. Therefore, when migrating the parallel computing program, the computing logic needs to be modified.
[0080] Specifically, the computing device 10 obtains the original computing program. The original computing program is the parallel computing code written through the second accelerator. The original computing program is used to calculate the original computing parameters based on the second accelerator. Specifically, if the second accelerator is a GPU and CUDA is deployed on the GPU, the original computing program is a parallel computing program based on CUDA and is used to perform parallel computing.
[0081] The computing device 10 adjusts the original computing program to obtain a target computing program. The target computing program is used to implement parallel computing through a first accelerator. In the embodiments of the present application, the target computing program includes a cutting subroutine. The cutting subroutine is used to cut the original computing parameters to obtain target computing parameters with the same dimensional size. The target program is used to calculate the target parameters based on the first accelerator.
[0082] In one example, the original computing program includes at least two subroutines, namely a first subroutine and a second subroutine. The first subroutine is used to set a parallel strategy, that is, to set parallel parameters for parallel computing when running on a second accelerator (the parallel parameters set by the first subroutine are hereinafter referred to as first parallel parameters). The second subroutine is used to calculate the original computing parameters using the second accelerator according to the first parallel parameters.
[0083] The target computing program includes a first sub-target program, a cutting subroutine, and a second sub-target program.
[0084] The first sub-target program is used to set parallel parameters (hereinafter referred to as second parallel parameters) when running on the first accelerator. In the embodiments of the present application, the computing device 10 can obtain the first sub-target program by adjusting the first subroutine. In one example, the computing device 10 can adjust the parallel parameters set in the first subroutine from the first parallel parameters to the second parallel parameters to achieve this.
[0085] The second sub-target program is used to perform parallel computing on the target computing parameters using the first accelerator according to the second parallel parameters. In the embodiments of the present application, the second sub-target program can be implemented by adjusting the second subroutine. Specifically, the computing device 10 can adjust the ports of the second subroutine so that the adjusted second subroutine can be executed on the second accelerator. Then, the following operations are performed on the adjusted second subroutine: adjusting the first parallel parameters to the second parallel parameters, and adjusting the original computing parameters to the target computing parameters, to obtain the second sub-target program.
[0086] Therefore, when the parallel computing program migrates from the first accelerator to the second accelerator, a cutting subroutine is added. By means of cutting, the original computing parameters with different dimensions are made to generate target computing parameters with the same dimensions. For each accelerator, if the dimensions of the computing parameters are the same, for example, all are two-dimensional data structures or all are one-dimensional data structures, then each computing parameter can use the same indexing logic to determine the index, without the need to design additional offsets or other indexing logics for parameters with different dimensions. And for computing parameters with the same dimensions, only the calculation needs to be performed at the corresponding index. Therefore, making the dimensions of the computing parameters for parallel computing the same through cutting can reduce the complexity of the indexing logic and the parallel computing logic. That is to say, it directly reduces the parallel computing complexity, thereby reducing the migration complexity of the parallel computing program from the second accelerator to the first accelerator.
[0087] It should be noted that in actual applications, the accelerator in the computing device can be replaced from the second accelerator to the first accelerator, or it can also be applied to migrating the parallel computing program of the second computing device to the first computing device for execution. Among them, the second computing device includes a second processor and a second accelerator, and the first computing device includes a first accelerator and a first processor, etc. In addition, the embodiments of the present application can also be applied to other application programs, and the embodiments of the present application do not specifically limit.
[0088] It should be noted that the computing device 10 provided in the embodiments of the present application can be a server, specifically an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), as well as large databases and artificial intelligence platforms. When the above server is a server cluster or distributed system composed of multiple physical servers, the multiple physical servers can form a blockchain, and each physical server is a node on the blockchain. The physical type of the service area can be a rack server, a high-density server, a GPU server, a tower server, or a blade server, a whole cabinet server, etc., and the present application does not specifically limit.
[0089] The following will describe the parallel computing program processing method provided in the embodiments of the present application with reference to the drawings. The following will take the example of migrating the parallel computing program running on the second accelerator to the first accelerator in the same computing device to perform parallel computing for illustrative purposes.
[0090] Appendix Figure 4 FIG. is a schematic diagram for implementing a parallel computing program processing method provided in an embodiment of the present application. The method includes:
[0091] S410. The computing device migrates the original computing program to the first accelerator.
[0092] The original computing program refers to a parallel computing program written by a developer through a computing platform deployed on a second accelerator for parallel computing of original computing parameters. For example, if the second accelerator is a GPU and CUDA is deployed on the second accelerator, the original computing program is a CUDA-based parallel computing program.
[0093] In one example, the original computing program may include a first subroutine and a second subroutine. The first subroutine is used to set the parallel parameter to a first parallel parameter. The dimension of the first parallel parameter is the dimension supported by the second accelerator. For example, if the second accelerator is a GPU and the executed parallel parameter dimension is multi-dimensional, the dimension of the first parallel parameter can be 2 or more dimensions. Exemplarily, the first parallel parameter can be 2-dimensional, namely (x, y).
[0094] The parallel parameter is used to determine the index of each thread or thread block in the grid. Exemplarily, Figure 5A FIG. Figure 5A is a schematic diagram of a kind of computing parameter provided by an embodiment of the present application. As
[0095] shown, the computing parameter includes array a, array b, and array c. For a, the grid is 3×3. For example, if the first parallel parameter is (x, y), the index of each thread or thread block in the grid can be determined through formula (1).
[0096] idx_a = y * col + x (1)
[0097] The second subroutine is used to perform parallel computing on the original computing parameters by using the second accelerator according to the first parallel parameter. In one example, the second subroutine includes a first index logic and a first computing logic.
[0098] The first index logic indicates a computing index based on the first parallel parameter. To better illustrate the way of computing the index based on the first parallel parameter, the following takes the first parallel parameter being 2-dimensional (x, y) and Figure 5A as an example for schematic illustration.
[0099] Continuing as Figure 5AAs shown, b is a 4×4 array and c is a 5×5 array. For array b, based on the first parallel parameter and formula (2), calculate the index idx_b of the thread or thread block in array b. For array c, based on the first parallel parameter and formula (3), calculate the index idx_c of the thread or thread block in array c.
[0100] idx_b = (y + 1) * col + x (2)
[0101] idx_c = (y + 1) * col + (x + 1) (3)
[0102] The first calculation logic is used to perform parallel calculations on the original calculation parameters. By way of example, continuing as Figure 5A shown, the first calculation logic is used to perform parallel calculations on the original calculation parameters (array a, array b, and array c). Specifically, the thread (x, y) adds the number at the corresponding index in array a to the number at the corresponding index in array b, and places the addition result at the corresponding index position in array c. The specific parallel calculation method is shown in formula (4).
[0103] c[idx_c] = a[idx_a] + b[idx_b] (4)
[0104] For example, when x = 1 and y = 1, referring to formulas (1) to (3), we get idx_a = 4, idx_b = 9, and idx_c = 20. The thread (1, 1) executes to add the number at index 4 in array a to the number at index 9 in array b, and places the addition result at index 20 in array c.
[0105] In the embodiments of the present application, the original parallel calculation program further includes but is not limited to the first data transfer (abbreviated as the third subroutine) that copies the data on the host side to the second accelerator for parallel calculation, synchronizes the parallel calculation results, and the second data transfer (abbreviated as the fourth subroutine) that copies the parallel calculation results from the second accelerator to the processor, and so on. Exemplarily, the original parallel calculation program is the parallel calculation program as Figure 3 shown.
[0106] S420. The computing device adjusts the original computing program to obtain a target computing program.
[0107] The target computing program is used to implement parallel calculation through the first accelerator. For example, if the second accelerator is a GPU deployed with CUDA and the first accelerator is an NPU deployed with CANN, then the original computing program is a parallel computing program based on CUDA, and the target computing program is a parallel computing program based on CANN.
[0108] In an embodiment of the present application, the target calculation program includes a cutting subroutine. The cutting subroutine is used to cut the original calculation parameters to obtain target calculation parameters with the same dimension. The original calculation parameters refer to the calculation parameters passed to the kernel function in the original calculation program and are used to perform parallel calculations.
[0109] Exemplarily, if the original calculation parameters are array a, array b, and array c. Among them, array a is a 3×3 dimensional array, array b is a 4×4 dimensional array, and array c is a 5×5 dimensional array. The cutting subroutine cuts the original calculation parameters to obtain array a1, array b1, and array c1. The dimensions of array a1, array b1, and array c1 are the same, for example, all are 3×3 dimensional arrays. The target calculation program can use the same index logic during parallel calculation without designing additional offsets or other index logics for parameters with different dimensions. And for calculation parameters with the same dimension, only calculations need to be performed at the corresponding indexes.
[0110] Exemplarily, if the original calculation program includes a first subroutine and a second subroutine, the target calculation program may further include a first sub-target program and a second sub-target program.
[0111] Among them, the first sub-target program is used to set the parallel parameter to the second parallel parameter. The second parallel parameter is the parallel parameter supported by the first accelerator. For example, if the first accelerator is an NPU, the second parallel parameter is single.
[0112] The second sub-target program is used to perform parallel calculation on the target calculation parameters by using the first accelerator according to the second parallel parameter. Among them, the dimensions of the target calculation parameters are the same.
[0113] In an embodiment of the present application, if the dimensions of the original calculation parameters are different, each calculation parameter needs to use a different index logic to determine the index. For example Figure 5A As shown, array a uses formula (1) to determine the index, array b uses formula (2) to determine the index, and array c uses formula (3) for indexing, that is, additional offsets need to be designed for parameters with different dimensions, which brings complexity to index calculation and index logic. When the computing device migrates the original calculation program to a calculation program that can be executed on the first accelerator, it will also increase the operation complexity of the migration.
[0114] Exemplarily, append Figure 5BA comparison schematic diagram for performing parallel computing provided by an embodiment of this application. Optionally, when the array a is migrated to meet the requirements of the computing framework, it is adjusted from 3×3 to 3×16, the array b is adjusted to 4×16, and the array c is adjusted to 5×16. After the switch, the index logic of the array a switches to determine the index according to formula (5), the index logic of the array b switches to determine the index according to formula (6), and the index logic of the array c switches to determine the index according to formula (7).
[0115] bidx_a = GetBlockKIdx() (5)
[0116] bidx_b = GetBlockKIdx() + 1 (6)
[0117] bidx_c = GetBlockKIdx() + 2 (7)
[0118] Among them, GetBlockKIdx() indicates obtaining the x - dimension index of the current thread block in the grid. For example, GetBlockKIdx(0) means that the x - dimension index of the current thread block in the grid is 0.
[0119] That is, for computing parameters with different dimensions and when the second parallel parameter is single, the computing program needs to calculate an offset for each computing parameter once. For example, the row offset of b is 1 and the column offset of c is 2. In terms of the computing logic, the second accelerator directly calculates as shown in formula (4). For the computing logic of the first accelerator, an additional offset will be calculated separately for each original computing parameter to adapt to its dimension, as specifically shown in formula (8). The more parameters there are, the greater the difference in the complexity of parallel computing.
[0120] c[idx_c][2:5] = a[idx_a][0:3] + b[idx_b][0:3] (8)
[0121] Among them, [0:3] in formula (8) means selecting consecutive elements starting from index 0 to index 3 (excluding index 3), and [2:5] means selecting consecutive elements starting from index 2 to index 5 (excluding index 5). The greater the complexity of parallel computing, the greater the resulting migration difficulty.
[0122] In view of this, the computing device cuts the original computing parameters to obtain target computing parameters with the same dimension. Exemplarily, Figure 5C Another comparison schematic diagram for performing parallel computing provided by an embodiment of this application. As Figure 5CAs shown, the arrays a, b, and c are sliced to obtain target calculation parameters, which are arrays a1, b1, and b2. All three arrays are 3×3, so the same indexing logic can be used to calculate the indices. Figure 5C As shown, idx_a, idx_b, and idx_c all determine the indices using the indexing logic "y*col + x" and do not require designing additional offsets.
[0123] Among them, for the target calculation parameters a1, b1, and c1 with the same dimensions and when the second parallel parameter is single, the indices can be determined based on the index GetBlock() determined by the second parallel parameter and the target offset (referred to as the second indexing logic), rather than directly globally calculating the indices of all elements. This method reduces the index complexity. And the parallel calculation formula at this time is as shown in formula (9), and the parallel calculation can be directly performed at the corresponding indices without calculating the offset, so the parallel calculation complexity is also reduced at this time.
[0124] c[idx_c][0:3] = a[idx_a][0:3] + b[idx_b][0:3] (9)
[0125] Furthermore, the slicing subroutine can be further refined as follows: Determine whether the dimensions of the original calculation parameters are the same. If the dimensions of the original calculation parameters are different, slice the original calculation parameters to obtain the target calculation parameters. If the dimensions of the original calculation parameters are the same, use the original calculation parameters as the target calculation parameters.
[0126] Furthermore, the computing device can slice the original calculation parameters in the following way: Determine the target elements in the original calculation parameters. Among them, the target elements indicate the elements that need to be parallelly calculated. Specifically, the computing device identifies which elements have dependency relationships and uses the elements with dependency relationships as the target elements for slicing.
[0127] It should be noted that the embodiments of the present application can adjust the original calculation program to the target calculation program in the following way: The computing device adjusts the parallel parameter in the original program from the first parallel parameter to the second parallel parameter. Adjust the first parallel calculation in the second subroutine to the second parallel calculation, where the first parallel calculation indicates parallelly calculating the original calculation parameters using the second accelerator according to the first parallel parameter. The second parallel calculation indicates parallelly calculating the target calculation parameters using the first accelerator according to the second parallel parameter. The computing device converts the data formats of the adjusted first subroutine and second subroutine into the target data format, where the target data format is the data format supported by the first accelerator.
[0128] It should be noted that the target calculation program may further include a splicing subroutine, which is used to embed the parallel calculation results into the area corresponding to the original calculation parameters. By adding the splicing subroutine, the correct output of the parallel calculation results is ensured.
[0129] The target calculation program may further include a third sub-target program and / or a fourth sub-target program. Among them, the third sub-target program is used to instruct the processor to copy the target calculation parameters to the first accelerator and perform parallel calculations. The fourth sub-target program is used to instruct the first accelerator to copy the parallel calculation results to the processor for output.
[0130] In summary, through slicing, it is ensured that the dimensions of the calculation parameters for parallel calculation are the same, and the data structures of the calculation parameters with the same dimensions are consistent. For example, they are all three-dimensional arrays or all two-dimensional arrays. This enables the use of the same indexing logic during parallel calculation without the need to design additional indexing logic for parameters with different dimensions. And during parallel calculation, the computing device usually needs to calculate the data position processed by each thread according to the parallel parameters. Therefore, for calculation parameters with the same dimensions, parallel calculation can be achieved only by performing parallel calculation at the corresponding index, which directly reduces the complexity of parallel calculation and thus reduces the migration complexity.
[0131] To enable those skilled in the art to better understand the parallel computing program processing method provided in the embodiments of the present application, the following takes the second accelerator as a GPU and deploys CUDA as an example, and takes the first accelerator as an NPU and deploys CANN as an example for illustrative description. The original calculation program is as Figure 3 shown. Among them, the first subroutine takes the kernel function call as an example, and the second subroutine takes the GPU to implement the kernel function as an example.
[0132] Appendix Figure 6A is a schematic diagram of the implementation of a parallel computing program processing method provided in the embodiments of the present application. The method includes:
[0133] S610. The computing device obtains the original calculation program.
[0134] Exemplarily, the original calculation program is as Figure 3 shown, including two parts: the calculation program on the host side and the calculation program on the device side. Among them, the calculation program on the host side is executed by the processor and includes: memory allocation, the first data transfer (abbreviated as the third subroutine), kernel function call (setting the parallel strategy), synchronizing the parallel calculation results, the second data transfer (abbreviated as the fourth subroutine), and memory release, etc.
[0135] The calculation program on the device side is executed by the GPU and includes: kernel function implementation to perform parallel calculation operations.
[0136] S620. The computing device adjusts the data format of the original calculation program to the target data format.
[0137] The first computing device deploys CANN, and the second computing device deploys CUDA. Among them, the parallel computing program written through CANN uses the C++ language and requires the use of the API dedicated to the CANN framework. The parallel computing program written by CUDN uses the C language. Therefore, the computing device can first adjust the data format of the original computing program to the target data format supported by CANN.
[0138] It should be noted that in the embodiments of the present application, the target data format of the original computing program can be adjusted first, or after all computing programs are adjusted, the data format of the adjusted computing program can be adjusted to the target data format. The embodiments of the present application do not specifically limit this.
[0139] S630. The computing device performs port mapping on the computing program on the host side in the original computing program to obtain a preliminarily adjusted computing program.
[0140] The computing program on the host side has its corresponding ports in CUDA and CANN respectively. Therefore, when the computing device migrates the computing program on the host side from CUDA to CANN, it first performs port mapping on the computing program on the host side so that the port-mapped computing program is compatible with the NPU.
[0141] Specifically, the computing device can adjust the first data transmission to a third sub-target program and the second data transmission to a fourth sub-target program based on the port mapping relationship. Among them, the third sub-target program instructs the processor to copy the target computing parameters to the first accelerator (specifically, the NPU); among them, the cutting subprogram is added before the third sub-target program. The fourth sub-target program instructs the first accelerator (specifically, the NPU) to copy the parallel computing result to the processor.
[0142] Exemplarily, as shown in Table 1, it is a port mapping relationship table provided by the embodiments of the present application.
[0143] Table 1
[0144]
[0145] It should be noted that the port mapping in Table 1 is only for illustrative purposes and can be adjusted according to needs in actual use.
[0146] S640. The computing device adjusts the preliminarily adjusted computing program to obtain a target computing program.
[0147] The computing device adjusts the preliminarily adjusted computing program, and the specific adjustment methods include:
[0148] Add a cutting subroutine and a splicing subroutine on the host side. Among them, the cutting subroutine is located before the first data transfer, and the splicing subroutine is located after the second data transfer. In one example, the cutting subroutine is also called the cutting data operation, and the splicing subroutine is also called the operation of splicing output data.
[0149] The splicing subroutine instructs to embed the obtained parallel computing result into the area corresponding to the original parallel computing parameter. Exemplarily, as shown in the appendix Figure 5C The obtained parallel computing result is assigned to the area specified by c1. Since c1 is the cut array, the processor also needs to embed the parallel computing result in c1 into the corresponding area in the original parallel computing parameter c0 to achieve splicing output.
[0150] The target computing program obtained after the above adjustment is as Figure 6B shown.
[0151] Appendix Figure 6B is the logical schematic diagram of the target computing program provided by the embodiment of the present application, including:
[0152] The host side executes:
[0153] S310. Memory allocation.
[0154] S6A: Cut the input data
[0155] S320. The first data transfer.
[0156] S330. Call the kernel function and set the parallel strategy.
[0157] S350. Synchronize the parallel computing result.
[0158] S360. Control the parallel computing result of the NPU to be transferred from the device side to the host side.
[0159] S6B: Splice the output data.
[0160] S370. Execute memory release.
[0161] After S330, the device side executes S340.
[0162] S340. The NPU executes parallel computing.
[0163] In the embodiment of the present application, during the process of performing parallel computing based on CANN, the index logic (cutting input, splicing output) that is not easy to process on the NPU is moved to the CPU side for execution, further simplifying the index link in the CANN parallel strategy. Although the development and computing workload on the CPU side will increase, the overall difficulty of parallel computing migration is greatly reduced.
[0164] In addition, an embodiment of the present application further provides a parallel computing program processing device.
[0165] Appendix Figure 7 It is a schematic structural diagram of a parallel computing program processing device provided by an embodiment of the present application. The device 700 is applied to a computing device. The computing device includes a first accelerator for performing parallel computing. The device 700 includes:
[0166] An acquisition unit 701, configured to acquire an original computing program; the original computing program is a parallel computing program written based on a second accelerator; the original computing program is used to perform parallel computing on original computing parameters based on the second accelerator;
[0167] An adjustment unit 702, configured to adjust the original computing program to obtain a target computing program; the target computing program is a parallel computing program written based on the first accelerator; the target computing program includes a cutting subroutine; the cutting subroutine is used to cut the original computing parameters to obtain target computing parameters with the same dimensions; the target computing program is used to perform parallel computing on the target computing parameters based on the first accelerator.
[0168] Optionally, if the original computing program includes a first subroutine and a second subroutine, the first subroutine is used to set the parallel parameter of the parallel computing to a first parallel parameter; the second subroutine is used to perform parallel computing on the original computing parameters using the second accelerator according to the first parallel parameter;
[0169] The target computing program further includes a first sub-target program and a second sub-target program; the first sub-target program is used to set the parallel parameter to a second parallel parameter; the second sub-target program is used to perform parallel computing on the target computing parameters using the first accelerator according to the second parallel parameter;
[0170] The adjustment unit 702 is specifically configured to: adjust the parallel parameter set in the first subroutine from the first parallel parameter to the second parallel parameter, and adjust the first parallel computing in the second subroutine to the second parallel computing; the first parallel computing indicates performing parallel computing on the original computing parameters using the second accelerator according to the first parallel parameter; the second parallel computing indicates performing parallel computing on the target computing parameters using the first accelerator according to the second parallel parameter; convert the data formats of the adjusted first subroutine and second subroutine into a target data format to obtain the first sub-target program and the second sub-target program; the target data format is the data format supported by the first accelerator.
[0171] Optionally, if the first parallel parameter is a multi-dimensional parallel parameter and the second parallel parameter is a single parallel parameter;
[0172] The adjustment unit 702 is further configured to: adjust the first index logic in the first parallel computation to a second index logic, and adjust the first computation logic in the first parallel computation to a second computation logic; the second parallel computation includes the second index logic and the second computation logic; the first index logic determines an index based on a first parallel parameter, and the second index logic determines an index based on a second parallel parameter and by adding a target offset; the first computation logic indicates performing parallel computation on original computation parameters, and the second computation logic indicates performing parallel computation on target computation parameters.
[0173] Optionally, the target computation program further includes a splicing subprogram; the splicing subprogram is configured to embed the parallel computation result into a region corresponding to the original computation parameter.
[0174] Optionally, the target computation program further includes a third sub-target program; the third sub-target program is configured to instruct the processor to copy the target computation parameter to the first accelerator.
[0175] Optionally, the target computation program further includes a fourth sub-target program; the fourth sub-target program is configured to instruct the first accelerator to copy the parallel computation result to the processor.
[0176] Optionally, the adjustment unit 702 is further configured to: if the dimensions of the original computation parameters are different, cut the original computation parameters to obtain the target computation parameters. Specifically, if the dimensions of the original computation parameters are different, determine target elements in the original computation parameters; the target elements indicate the elements that need to perform parallel computation; cut out the target elements from the original computation parameters to obtain the target computation parameters.
[0177] When the parallel computation program provided by the embodiment of the present application migrates from the second accelerator to the first accelerator, the device 700 adds a cutting subprogram. By means of cutting, the original computation parameters with different original dimensions generate target computation parameters with the same dimension. For each accelerator, if the dimensions of the computation parameters are the same, for example, all are two-dimensional data structures or all are one-dimensional data structures, then each computation parameter can use the same index logic to determine the index, and there is no need to design additional offsets or other index logics for parameters with different dimensions. And for computation parameters with the same dimension, only the computation needs to be performed at the corresponding index. Therefore, by cutting to make the dimensions of the computation parameters for parallel computation the same, the complexity of the index logic and the parallel computation logic can be reduced. That is to say, the complexity of the parallel computation is directly reduced, thereby reducing the migration complexity of the parallel computation program from the second accelerator to the first accelerator.
[0178] An embodiment of the present application also provides a computer program product containing instructions. The computer program product may be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on a computing device, it causes the computing device to execute the above-mentioned parallel computing program processing method. An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the above-mentioned parallel computing program processing method.
[0179] The descriptions of the processes or structures corresponding to the above respective drawings each have their own focuses. For parts not detailed in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.
[0180] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for processing a parallel computing program, characterized in that, Applied to a computing device, the computing device includes a first accelerator for performing parallel computing, and the method includes: Obtain an original computing program; the original computing program is a parallel computing program written based on a second accelerator; the original computing program is used to perform parallel computing on original computing parameters based on the second accelerator; Adjust the original computing program to obtain a target computing program; the target computing program is a parallel computing program written based on the first accelerator; the target computing program includes a cutting subroutine; the cutting subroutine is used to cut the original computing parameters to obtain target computing parameters with the same dimension; the target computing program is used to perform parallel computing on the target computing parameters based on the first accelerator.
2. The processing method according to claim 1, wherein The original computing program includes a first subroutine and a second subroutine. The first subroutine is used to set the parallel parameter of parallel computing to a first parallel parameter; the second subroutine is used to perform parallel computing on the original computing parameters by using the second accelerator according to the first parallel parameter; The target computing program further includes a first sub-target program and a second sub-target program; the first sub-target program is used to set the parallel parameter to a second parallel parameter; the second sub-target program is used to perform parallel computing on the target computing parameters by using the first accelerator according to the second parallel parameter; The adjusting the original computing program to obtain a target computing program includes: Adjust the parallel parameter set in the first subroutine from the first parallel parameter to the second parallel parameter, and adjust the first parallel computing in the second subroutine to a second parallel computing; the first parallel computing indicates performing parallel computing on the original computing parameters by using the second accelerator according to the first parallel parameter; the second parallel computing indicates performing parallel computing on the target computing parameters by using the first accelerator according to the second parallel parameter; Convert the data formats of the adjusted first subroutine and second subroutine into a target data format to obtain the first sub-target program and the second sub-target program; the target data format is a data format supported by the first accelerator.
3. The processing method according to claim 2, characterized in that, If the first parallel parameter is a multi-dimensional parallel parameter and the second parallel parameter is a single parallel parameter; The adjusting the first parallel computing in the second subroutine to a second parallel computing includes: Adjust the first index logic in the first parallel computing to a second index logic, and adjust the first computing logic in the first parallel computing to a second computing logic; the second parallel computing includes the second index logic and the second computing logic; The first index logic is to determine an index based on the first parallel parameter, and the second index logic is to determine an index based on the second parallel parameter and by adding a target offset; the first computing logic indicates performing parallel computing on the original computing parameters, and the second computing logic indicates performing parallel computing on the target computing parameters.
4. The processing method according to any one of claims 1-3, characterized in that, The target calculation program further includes a splicing subroutine; the splicing subroutine is used to embed the parallel calculation result into the area corresponding to the original calculation parameter.
5. The processing method according to any one of claims 1-4, characterized in that, The target calculation program further includes a third sub-target program; the third sub-target program is used to instruct the processor to copy the target calculation parameter to the first accelerator.
6. The processing method according to any one of claims 1-5, characterized in that, The target calculation program further includes a fourth sub-target program; the fourth sub-target program is used to instruct the first accelerator to copy the parallel calculation result to the processor.
7. The processing method according to any one of claims 1-6, characterized in that, The cutting subroutine is used to cut the original calculation parameter to obtain a target calculation parameter with the same dimension, including: If the dimensions of the original calculation parameters are different, the original calculation parameters are cut to obtain the target calculation parameters.
8. The processing method according to claim 7, wherein if the dimensions of the original calculation parameters are different, the original calculation parameters are cut to obtain the target calculation parameters, including: If the dimensions of the original calculation parameters are different, determine the target elements in the original calculation parameters; The target elements indicate the elements that need to perform parallel calculations; Cut out the target elements from the original calculation parameters to obtain the target calculation parameters.
9. The processing method according to any one of claims 1-8, wherein the first accelerator is a neural network processor NPU, and the NPU deploys a neural network computing architecture CANN for performing parallel calculations through the target calculation program; The second accelerator is a graphics processor GPU, and the GPU deploys a computing unified device architecture CUDA for performing parallel calculations through the target calculation program.
10. A computing device, characterized in that, It includes a processor and a memory, and the processor is coupled to the memory; The memory is used to store programs; The processor is used to execute the programs stored in the memory. When the programs are executed, the processor is used to implement the method according to any one of claims 1-9.