Modeling source code distribution method for model management platform and storage medium
Through an automatic allocation method based on an algebraic framework and estimated execution time, the problem of users' difficulty in efficiently utilizing CPU and GPU devices is solved, and the code execution efficiency and real-time improvement is achieved.
Patent Information
- Application Number
- CN202411970022.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
AI Technical Summary
When users write source code for the model management platform, it is difficult for users to efficiently utilize CPU and GPU devices, resulting in performance bottlenecks and real-time requirements being difficult to meet.
The automatic allocation method based on the source code algebra framework and estimated execution time is adopted. By detecting the data access mode of the loop body, implementing data access mode conversion, evaluating the execution time, and automatically allocating the model source code to the CPU or GPU according to the execution time.
Make full use of the advantages of CPU and GPU devices to improve code execution efficiency and real-time performance, and shorten the overall execution time of the program.
Smart Images

Figure CN119987778A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing and code optimization, and in particular to an automatic distribution method of modeling source code for a model management platform. Background Art
[0002] In recent years, artificial intelligence modeling and management platforms (referred to as "model management platforms") have been increasingly widely used in enterprises and institutions, such as Tencent's TI-ONE platform and Alibaba Cloud's PAI platform. The hardware resources that support model management platforms usually include computing devices such as CPUs and GPUs. When users write business model source code, it is difficult to decide which fragments of the source code are more efficient to execute on GPU devices, so all codes are usually placed on GPU devices for execution.
[0003] GPU devices are good at processing codes with high data parallelism, strong computational intensity and high bandwidth utilization. Therefore, allocating code snippets with the above characteristics in the source code to the GPU for execution can better meet the real-time requirements of model training.
[0004] However, data transmission between CPU and GPU devices can easily become a performance bottleneck. Even if the source code snippet meets the above characteristics, it is necessary to comprehensively consider whether it is suitable for allocation to the GPU device for execution. Providing users with automatic code snippet allocation tools at the modeling source code level and making full use of the respective advantages of CPU and GPU devices is of great significance to meeting users' real-time requirements for modeling. The existing technology has no relevant allocation control strategy that can meet users' requirements. Summary of the invention
[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a method for automatically allocating modeling source code for a model management platform. In view of the fact that users cannot efficiently utilize heterogeneous computing devices such as CPUs and GPUs when writing modeling codes using a model management platform, a method for automatically allocating source code based on a source code algebraic framework and estimated execution time is designed, including: detecting data access patterns of loop bodies based on algebraic descriptions, implementing data access pattern conversion on loop bodies, evaluating the execution time of converted loop bodies, and automatically allocating model source code based on the execution time.
[0006] In order to achieve the above object, the technical solution adopted by the present invention is: a method for automatically allocating modeling source code for a model management platform, comprising the following steps:
[0007] S1. Obtain a code snippet of a model to be assigned, and construct an algebraic description of a loop body in the model source code snippet;
[0008] S2, detect the data access pattern of the loop body based on algebraic description;
[0009] S3, implement data access mode conversion on loop body LB;
[0010] S4, evaluate the execution time of the converted loop body;
[0011] S5. Allocate the model source code to the CPU and GPU for execution based on the evaluation time.
[0012] Step S5 includes: the execution time of the loop body LB on the CPU is TCPU LB , if TCPU LB ≤T change +TGPU LB , the loop body LB is assigned to the CPU for execution; otherwise, it is assigned to the GPU device for execution; the total execution time of LB on the GPU device is TGPU LB , T change For conversion time
[0013] The step S1 comprises:
[0014] (1) Constructing the loop body data access matrix
[0015] Traverse the source code and extract the loop body (LB1, LB2, ..., LB N ) and the arrays contained in the loop body (AR1,AR2,...,AR M ), construct the following LB-AR matrix, where MP ij Representing LB i About AR j Data access mode:
[0016]
[0017] (2) Construction of MP ij The algebraic equation
[0018] Assume that the loop body LB is a W-layer nested loop and the iteration vector of LB is r LB =(r1,r2,…,r W ) T , r1 represents the LB first-layer loop index variable, r2 represents the LB second-layer loop index variable, r W represents the loop index variable of the Wth layer of LB. The data access vector MP of the Z-dimensional array AR = {mp1, mp2, ..., mp Z} T , mp1 represents the AR first dimension access variable, mp2 represents the AR second dimension access variable, and mpZ represents the AR Zth dimension access variable. Since the AR data access vector is an affine function of the LB iteration vector, MP can be expressed as:
[0019] MP=FAR r LB +sf
[0020] The matrix F of size Z×W AR is the access matrix of array AR, matrix F AR Each row represents the data access mode of each dimension of array AR; sf is an offset of size Z×1, representing the starting position of LB accessing AR.
[0021] Step S2 includes: (1) detecting an array AR that can be converted into perfect continuous data access: detecting matrix F AR For the main diagonal elements, if the first element is 1, check whether the elements other than the last element are 0; if this condition is not met, it cannot be converted to perfect continuous access. Detection matrix F AR For the secondary diagonal elements, if the first element is 1, check whether the last element is 1 and the other elements are 0; if this condition is not met, it cannot be converted to perfect continuous access;
[0022] (2) Detect whether the array AR has self-dependence and eliminate the self-dependence by introducing an intermediate array
[0023] The self-dependence of array AR means updating AR itself with its own value. If array AR has self-dependence, an intermediate array needs to be introduced to eliminate the self-dependence.
[0024] Step 3 includes: if the data access mode of the loop body LB for the array AR can be converted into perfect continuous access, then performing access mode conversion on it:
[0025] The access matrix of Z-dimensional data AR is F AR After the perfect continuous access mode conversion (C, a) is implemented, its access matrix becomes F' AR , needs to satisfy F' AR =CF AR +a, where C is a transformation matrix of size Z×W and a is a calibration matrix of size Z×W.
[0026] The conversion matrix C and calibration matrix a are:
[0027] (1) If the data access mode of the array AR is perfect continuous access, no conversion is required. In this case, C is the identity matrix and a is the zero matrix.
[0028] (2) If the data access mode to the array AR is monotonically decreasing access, then C is the inverse of the element corresponding to the monotonically decreasing access in the identity matrix, and a is a zero matrix.
[0029] (3) If the data access mode to the array AR is offset access with an offset of s, then C is the identity matrix and a is the last element of the main diagonal of the zero matrix changed to -s.
[0030] (4) If the data access mode to the array AR is a slanted access with a slant of D times, then C is the element in the identity matrix corresponding to the slanted access expanded by D times and then inverted, and a is a zero matrix.
[0031] (5) If the data access mode of the array AR is a stride access with a stride of S, then C is the reciprocal of the element in the identity matrix corresponding to the stride access expanded by S times, and a is a zero matrix.
[0032] (6) If the data access mode of array AR is column-based access, then C is the transposed matrix of the identity matrix and a is the zero matrix.
[0033] Step S4 includes: (1) scanning the LB-AR matrix of the source code to detect whether the mp vector in each MP element in the matrix is a perfect continuous access; if it is not a perfect continuous access, judging whether it can be converted to a perfect continuous access according to the conditions in step S2; if it can be converted, implementing step S3 and recording the conversion time T change ;
[0034] (2) Calculate the execution time T of each loop body LB in the source code LB , if the access mode of a LB to AR is perfect continuous access, then T LB Excluding conversion time T change Otherwise, you need to add T change The total execution time of the LB;
[0035] (3) Calculate the execution time of the loop body after LB parallelization, and evaluate the execution time allocated to the GPU device after LB parallelization.
[0036] The execution time of the calculation loop body LB after parallelization includes:
[0037] (1) Evaluate data transmission time
[0038] Assuming that the amount of read data of the loop body LB to be assigned to the GPU device for execution is RS (in MB), the amount of write data is WS (in MB), and the actual transmission speed of the PCI-E bus is TS (in MB / s), then the data transmission overhead of assigning LB to the GPU device for execution is
[0039] (2) Evaluate instruction execution time
[0040] Assume that the number of stream processing units on the GPU device is N SM , the number of stream processors is NSP (A stream processor contains multiple stream processing units), then the instruction execution time allocated by LB to the GPU device et i is the execution time of the ith instruction on the GPU device, NUM i is the number of the i-th instruction, F is the number of instruction types, and WARP is the warp size in the CUDA model, which is a fixed value of 32.
[0041] (3) Evaluate memory access overhead
[0042] Since the data access of the loop body LB has been converted to perfect continuous access in step 3, when the amount of data processed by the warp does not exceed the amount of data that can be returned by one video memory read, the video memory access request can be completed in one transmission; if it exceeds, multiple transmissions are required. Therefore, the maximum value of the video memory access time can be approximated as:
[0043]
[0044] Where D WARP is the amount of data processed by a warp, d MEM is the amount of data that can be read in one video memory transfer, t MEM The time required for one video memory transfer.
[0045] (4) Evaluate whether computation and memory access can overlap
[0046] Each stream processor in the GPU device can execute multiple warps in a round-robin manner. When one or more warps are suspended due to accessing the video memory, the stream processor can execute other warps, avoiding the idleness of the GPU device computing unit. It can be seen that the video memory access time of a warp can be hidden by the computing time of other warps. When the computing time of a warp is greater than the video memory access time, the video memory access time can be completely hidden. At this time, the video memory access time does not affect the overall execution time of LB on the GPU device, T MEM_MAX In this state, it is 0. When the computation time of a warp is less than the video memory access time, the actual video memory access time T MEM for:
[0047] T MEM =T MEM_MAX -T compute
[0048] (5) Estimate the number of warps and the total execution time of LB on the GPU device
[0049] Since the iteration vector of LB is r LB =(r1,r2,…,rW ) T , so the total number of LB cycles is r1×r2×…×r w In theory, a GPU device thread processes one loop, so a total of r1×r2×…×r w GPU device threads. A GPU warp contains 32 GPU device threads, so a total of (r1×r2×…×r w )÷32 warps, denoted as N warp .
[0050] When N warp When the number of stream processors is less than or equal to that of the GPU device, it means that all warps can run simultaneously, that is, there is no warp rotation. At this time, the video memory access of a warp cannot be hidden by the calculation of other warps. The total execution time of LB on the GPU device is:
[0051] TGPU LB =T transfer +T compute +T MEM_MAX
[0052] When N warp When it is larger than the stream processor of the GPU device, it means that all warps cannot run at the same time, that is, there is warp rotation. At this time, the video memory access of a warp may be hidden by the calculation of other warps. The total execution time of LB on the GPU device is:
[0053]
[0054] A computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1 to 8.
[0055] The advantages of the present invention are: making full use of the respective advantages of CPU and GPU devices, allocating according to the respective characteristics of CPU and GPU, improving the execution efficiency of code, improving the real-time performance of code execution, and having important significance for meeting the real-time performance requirements of users for modeling. Automatic code allocation does not require developers to understand the GPU hardware architecture, and has the advantages of high efficiency and speed. The technical effect is to shorten the overall execution time of the program. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The following is a brief description of the contents expressed in the drawings of the present invention and the symbols in the drawings:
[0057] Figure 1 The figure is a flow chart of the method for automatically allocating modeling source code for a model management platform according to the present invention. DETAILED DESCRIPTION
[0058] The specific implementation of the present invention will be further explained in detail below by describing the optimal embodiment with reference to the accompanying drawings.
[0059] Aiming at the problem that users cannot efficiently utilize heterogeneous computing devices such as CPU and GPU when writing modeling codes using a model management platform, the present invention designs a source code automatic allocation method based on a source code algebra framework and estimated execution time. Figure 1 As shown, the allocation method includes: detecting the data access mode of the loop body based on the algebraic description, implementing data access mode conversion on the loop body, evaluating the execution time of the converted loop body, and automatically allocating the model source code based on the execution time. The specific scheme is as follows: Step 1 Construct an algebraic description of the loop body in the model source code
[0060] Sub-step 1 constructs a loop body data access matrix. Loop body is a computer professional term. The four elements of the loop body include: loop variable initialization, loop condition, loop body and loop variable update.
[0061] Traverse the source code and extract the loop body (LB1, LB2, ..., LB N ) and the arrays contained in the loop body (AR1,AR2,...,AR M ), construct the following LB-AR matrix, where MP ij Representing LB i About AR j Data access mode:
[0062]
[0063] Sub-step 2: Build MP ij The algebraic equation
[0064] Assume that the loop body LB is a W-layer nested loop and the iteration vector of LB is r LB =(r1,r2,…,r W ) T , r1 represents the LB first-layer loop index variable, r2 represents the LB second-layer loop index variable, r W Represents the loop index variable of the Wth layer of LB. The data access vector MP of the Z-dimensional array AR (Z-dimensional array means that the array contained in the loop body is Z-dimensional, and Z-dimensional array means that the array includes Z dimensions.) = {mp1, mp2, ..., mp Z} T , mp1 represents the AR first dimension access variable, mp2 represents the AR second dimension access variable, and mpZ represents the AR Zth dimension access variable. Since the AR data access vector is an affine function of the LB iteration vector, MP can be expressed as:
[0065] MP=F AR r LB +sf
[0066] The matrix F of size Z×W AR is the access matrix of array AR, matrix F AR Each row represents the data access mode of each dimension of array AR; sf is an offset of size Z×1, representing the starting position of LB accessing AR.
[0067] Step 2: Detect the data access pattern of the loop body based on the algebraic description
[0068] To achieve high bandwidth, the GPU device memory returns data in the form of continuous blocks. Therefore, in order to give full play to the performance of the GPU device, the loop body LB's access to the array AR should meet perfect continuity, that is, LB starts accessing the first element of the first dimension of AR, and subsequent accesses monotonically increase the offset 1 to access the continuously arranged adjacent data of AR. In order to achieve perfect continuous access, it is necessary to meet the following requirements:
[0069] (1) Matrix F AR The first and last elements of the main diagonal are 1, and the rest are 0.
[0070] (2) Matrix F AR All the sub-diagonal elements are 0
[0071] (3) Vector sf = {0, 0, ..., 0} T
[0072] Sub-step 1 detects array AR that can be converted to perfect sequential data access
[0073] Detection matrix F AR For the main diagonal elements, if the first element is 1, check whether the elements other than the last element are 0; if this condition is not met, it cannot be converted to perfect continuous access.
[0074] Detection matrix F AR For the off-diagonal elements, if the first element is 1, check whether the last element is 1 and the other elements are 0; if this condition is not met, it cannot be converted to perfect continuous access.
[0075] Sub-step 2 detects whether array AR has self-dependence and eliminates the self-dependence by introducing an intermediate array
[0076] The self-dependence of an array AR is to update AR itself with its own value, for example:
[0077] AR1[i][j][k]=AR1[i][j][k]+AR2[i][j][k]
[0078] If the array AR is self-dependent, an intermediate array needs to be introduced to remove the self-dependence:
[0079] AR3[i][j][k]=AR1[i][j][k]
[0080] AR1[i][j][k]=AR3[i][j][k]+AR2[i][j][k]
[0081] Step 3: Implement data access mode conversion on loop body LB
[0082] If the data access mode of the loop body LB for the array AR can be converted to perfect continuous access, the access mode conversion is performed on it:
[0083] The access matrix of Z-dimensional data AR is F AR After the perfect continuous access mode conversion (C, a) is implemented, its access matrix becomes F' AR , needs to satisfy F' AR =CF AR +a, where C is a transformation matrix of size Z × W and a is a calibration matrix of size Z × W. For common data access patterns, the transformation matrix C and calibration matrix a are:
[0084] (1) If the data access mode of the array AR is perfect continuous access, no conversion is required. In this case, C is the identity matrix and a is the zero matrix.
[0085] (2) If the data access mode to the array AR is monotonically decreasing access, then C is the inverse of the element corresponding to the monotonically decreasing access in the identity matrix, and a is a zero matrix.
[0086] (3) If the data access mode to the array AR is offset access with an offset of s, then C is the identity matrix and a is the last element of the main diagonal of the zero matrix changed to -s.
[0087] (4) If the data access mode to the array AR is a slanted access with a slant of D times, then C is the element in the identity matrix corresponding to the slanted access expanded by D times and then inverted, and a is a zero matrix.
[0088] (5) If the data access mode of the array AR is a stride access with a stride of S, then C is the reciprocal of the element in the identity matrix corresponding to the stride access expanded by S times, and a is a zero matrix.
[0089] (6) If the data access mode of array AR is column-based access, then C is the transposed matrix of the identity matrix and a is the zero matrix.
[0090] Step 4: Evaluate the execution time of the converted loop body
[0091] Sub-step 1 scans the LB-AR matrix of the source code and detects whether the mp vector in each MP element in the matrix is a perfect continuous access. If it is not a perfect continuous access, determine whether it can be converted to a perfect continuous access according to the conditions in step 2. If it can be converted, implement step 3 and record the conversion time T change .
[0092] Sub-step 2 calculates the execution time T of each loop body LB in the source code LB , if the access mode of a LB to AR is perfect continuous access, then T LB Excluding conversion time T change Otherwise, you need to add T change The total execution time of this LB.
[0093] Sub-step 3 calculates the execution time of the loop body after LB parallelization
[0094] For each loop body LB in the source code, it can be assigned to the CPU or the GPU for execution. If it is assigned to the GPU for parallel execution, the execution time may be significantly shortened, but the data transmission overhead between the CPU and GPU may reduce the effect of parallel computing. Therefore, it is necessary to evaluate the execution time after LB is parallelized (i.e. assigned to the GPU). The execution of the loop body LB on the GPU device needs to go through the following process: (1) The CPU transfers the data accessed by LB to the GPU device; (2) The GPU device starts multi-threaded execution of the calculation; (3) After the calculation is completed, the result is transferred back to the CPU.
[0095] In this step, the instruction execution overhead is the measured average execution time of a certain instruction (such as addition, subtraction, multiplication and division, etc.) executed 1000 times on the CPU or GPU. The data transmission bandwidth is the measured average bandwidth of transmitting 10MB, 50MB, 100MB, 200MB and 500MB data 100 times respectively. The video memory access bandwidth is the measured average bandwidth of accessing 1MB, 5MB, 10MB, 20MB and 30MB data 100 times respectively.
[0096] (1) Evaluate data transmission time
[0097] Assuming that the amount of read data of the loop body LB to be assigned to the GPU device for execution is RS (in MB), the amount of write data is WS (in MB), and the actual transmission speed of the PCI-E bus is TS (in MB / s), then the data transmission overhead of assigning LB to the GPU device for execution is
[0098] (2) Evaluate instruction execution time
[0099] Assume that the number of stream processing units on the GPU device is N SM , the number of stream processors is N SP (A stream processor contains multiple stream processing units), then the instruction execution time allocated by LB to the GPU device et i is the execution time of the ith instruction on the GPU device, NUM i is the number of the i-th instruction, F is the number of instruction types, and WARP is the warp size in the CUDA model, which is a fixed value of 32.
[0100] (3) Evaluate memory access overhead
[0101] Since the data access of the loop body LB has been converted to perfect continuous access in step 3, when the amount of data processed by the warp does not exceed the amount of data that can be returned by one video memory read, the video memory access request can be completed in one transmission; if it exceeds, multiple transmissions are required. Therefore, the maximum value of the video memory access time can be approximated as:
[0102]
[0103] Where D WARP is the amount of data processed by a warp, d MEM is the amount of data that can be read in one video memory transfer, t MEM The time required for one video memory transfer.
[0104] (4) Evaluate whether computation and memory access can overlap
[0105] Each stream processor in the GPU device can execute multiple warps in a round-robin manner. When one or more warps are suspended due to accessing the video memory, the stream processor can execute other warps, avoiding the idleness of the GPU device computing unit. It can be seen that the video memory access time of a warp can be hidden by the computing time of other warps. When the computing time of a warp is greater than the video memory access time, the video memory access time can be completely hidden. At this time, the video memory access time does not affect the overall execution time of LB on the GPU device, T MEM_MAX In this state, it is 0. When the computation time of a warp is less than the video memory access time, the actual video memory access time T MEM for:
[0106] T MEM =T MEM_MAX -T compute
[0107] (5) Estimate the number of warps and the total execution time of LB on the GPU device
[0108] Since the iteration vector of LB is r LB =(r1,r2,…,r W ) T , so the total number of LB cycles is r1×r2×…×r w In theory, a GPU device thread processes one loop, so a total of r1×r2×…×r w GPU device threads. A GPU warp contains 32 GPU device threads, so a total of (r1×r2×…×r w )÷32 warps, denoted as N warp .
[0109] When N warp When it is less than or equal to the number of stream processors of the GPU device, it means that all warps can run at the same time, that is, there is no warp rotation. At this time, the video memory access of a warp cannot be hidden by the calculation of other warps. The total execution time of LB on the GPU device is:
[0110] TGPU LB =T transfer +T compute +T MEM_MAX
[0111] When N warp When it is larger than the stream processor of the GPU device, it means that all warps cannot run at the same time, that is, there is warp rotation. At this time, the video memory access of a warp may be hidden by the calculation of other warps. The total execution time of LB on the GPU device is:
[0112]
[0113] Step 5: Automatically distribute model source code based on execution time
[0114] Assume that the execution time of the loop body LB on the CPU is TCPU LB , if TCPU LB ≤T change +TGPU LB , the loop body LB is assigned to the CPU for execution. Otherwise, it is assigned to the GPU device for execution. In fact, there are a lot of model codes. The model code is divided into different loop bodies, and then divided into CPU and GPU for execution based on the loop body. The code fragment is composed of the loop body, so dividing the loop body here can complete the division of the code fragment.
[0115] Obviously, the specific implementation of the present invention is not limited to the above-mentioned methods. As long as various non-substantial improvements are made using the method concept and technical solution of the present invention, they are all within the protection scope of the present invention.
Claims
1. A modeling source code distribution method for a model management platform, characterized by: The steps include: S1. Obtain a code snippet of a model to be assigned, and construct an algebraic description of a loop body in the model source code snippet; S2, detect the data access pattern of the loop body based on algebraic description; S3, implement data access mode conversion on loop body LB; S4, evaluate the execution time of the converted loop body; S5. Allocate the model source code to the CPU and GPU for execution based on the evaluation time.
2. A modeling source code distribution method for a model management platform as claimed in claim 1, characterized in that: Step S5 includes: the execution time of the loop body LB on the CPU is TCPU LB , if TCPU LB ≤T change +TGPU LB , the loop body LB is assigned to the CPU for execution; otherwise, it is assigned to the GPU device for execution; the total execution time of LB on the GPU device is TGPU LB , T change For conversion time.
3. The modeling source code distribution method for a model management platform according to claim 1, characterized in that: The step S1 comprises: (1) Constructing the loop data access matrix Traverse the source code and extract the loop body (LB1, LB2, ..., LB N ) and the arrays contained in the loop body (AR1,AR2,...,AR M ), construct the following LB-AR matrix, where MP ij Representing LB i About AR j Data access mode: (2) Construction of MP ij The algebraic equation Assume that the loop body LB is a W-layer nested loop and the iteration vector of LB is r LB =(r1,r2,…,r W ) T , r1 represents the LB first-layer loop index variable, r2 represents the LB second-layer loop index variable, r W represents the loop index variable of the Wth layer of LB; the data access vector MP of the Z-dimensional array AR = {mp1, mp2, ..., mp Z } T , mp1 represents the AR first dimension access variable, mp2 represents the AR second dimension access variable, and mpZ represents the AR Zth dimension access variable; since the AR data access vector is an affine function of the LB iteration vector, MP can be expressed as: MP=F AR r LB +sf The matrix F of size Z×W AR is the access matrix of array AR, matrix F AR Each row represents the data access mode of each dimension of array AR; sf is an offset of size Z×1, representing the starting position of LB accessing AR.
4. The modeling source code distribution method for a model management platform according to claim 1, characterized in that: Step S2 includes: (1) detecting an array AR that can be converted into perfect continuous data access: detecting matrix F AR The main diagonal elements, if the first element is 1, detect whether the other elements except the last element are 0; if this condition is not met, it cannot be converted to perfect continuity access; the detection matrix F AR For the secondary diagonal elements, if the first element is 1, check whether the last element is 1 and the other elements are 0; if this condition is not met, it cannot be converted to perfect continuous access; (2) Detect whether the array AR has self-dependence and eliminate the self-dependence by introducing an intermediate array The self-dependence of array AR means updating AR itself with its own value. If array AR has self-dependence, an intermediate array needs to be introduced to eliminate the self-dependence.
5. The modeling source code distribution method for a model management platform according to claim 1, characterized in that: Step 3 includes: if the data access mode of the loop body LB for the array AR can be converted into perfect continuous access, then performing access mode conversion on it: The access matrix of Z-dimensional data AR is F AR After the perfect continuous access mode conversion (C, a) is implemented, its access matrix becomes F' AR , needs to satisfy F' AR =CF AR +a, where C is a transformation matrix of size Z×W and a is a calibration matrix of size Z×W.
6. The modeling source code distribution method for a model management platform according to claim 5, characterized in that: The conversion matrix C and calibration matrix a are: (1) If the data access mode of the array AR is perfect continuous access, no conversion is required. In this case, C is the identity matrix and a is the zero matrix. (2) If the data access mode to the array AR is monotonically decreasing access, then C is the inverse of the element corresponding to the monotonically decreasing access in the identity matrix, and a is a zero matrix; (3) If the data access mode to the array AR is offset access with an offset of s, then C is the identity matrix, and a is the last element of the main diagonal of the zero matrix changed to -s; (4) If the data access mode to the array AR is a tilted access with a tilt of D times, then C is the element in the identity matrix corresponding to the tilted access expanded D times and inverted, and a is a zero matrix; (5) If the data access mode of the array AR is a stride access with a stride of S, then C is the reciprocal of the element in the identity matrix corresponding to the stride access expanded by S times, and a is a zero matrix; (6) If the data access mode of array AR is column-based access, then C is the transposed matrix of the identity matrix and a is the zero matrix.
7. The modeling source code distribution method for a model management platform according to claim 1, characterized in that: Step S4 includes: (1) scanning the LB-AR matrix of the source code to detect whether the mp vector in each MP element in the matrix is a perfect continuous access; if it is not a perfect continuous access, judging whether it can be converted to a perfect continuous access according to the conditions in step S2; if it can be converted, implementing step S3 and recording the conversion time T change ; (2) Calculate the execution time T of each loop body LB in the source code LB , if the access mode of a LB to AR is perfect continuous access, then T LB Excluding conversion time T change Otherwise, you need to add T change The total execution time of the LB; (3) Calculate the execution time of the loop body after LB parallelization, and evaluate the execution time allocated to the GPU device after LB parallelization.
8. The modeling source code distribution method for a model management platform according to claim 1, characterized in that: The execution time of the calculation loop body LB after parallelization includes: (1) Evaluate data transmission time Assuming that the amount of read data of the loop body LB to be assigned to the GPU device for execution is RS (in MB), the amount of write data is WS (in MB), and the actual transmission speed of the PCI-E bus is TS (in MB / s), then the data transmission overhead of assigning LB to the GPU device for execution is (2) Evaluate instruction execution time Assume that the number of stream processing units on the GPU device is N SM , the number of stream processors is N SP (A stream processor contains multiple stream processing units), then the instruction execution time allocated by LB to the GPU device et i is the execution time of the ith instruction on the GPU device, NUM i is the number of the i-th instruction, F is the number of instruction types, WARP is the warp size in the CUDA model, which is a fixed value of 32; (3) Evaluate memory access overhead Since the data access of the loop body LB has been converted into perfect continuous access in step 3, when the amount of data processed by the warp does not exceed the amount of data that can be returned by one video memory read, the video memory access request can be completed in one transmission; if it exceeds, multiple transmissions are required; therefore, the maximum value of the video memory access time can be approximated as: Where D WARP is the amount of data processed by a warp, d MEM is the amount of data that can be read in one video memory transfer, t MEM The time required for one video memory transfer; (4) Evaluate whether computation and memory access can overlap Each stream processor in the GPU device can execute multiple warps in a round-robin manner. When one or more warps are suspended due to accessing the video memory, the stream processor can execute other warps, avoiding the idleness of the computing unit of the GPU device. It can be seen that the video memory access time of a warp can be hidden by the computing time of other warps. When the computing time of a warp is greater than the video memory access time, the video memory access time can be completely hidden. At this time, the video memory access time does not affect the overall execution time of LB on the GPU device. MEM_MAX In this state, it is 0; when the calculation time of the warp is less than the video memory access time, the actual video memory access time T MEM for: T MEM =T MEM_MAX -T compute (5) Estimate the number of warps and the total execution time of LB on the GPU device Since the iteration vector of LB is r LB =(r1,r2,…,r W ) T , so the total number of LB cycles is r1×r2×…×r w times; one GPU device thread processes one loop, so a total of r1×r2×…×r needs to be generated w GPU device threads; a GPU warp contains 32 GPU device threads, so a total of (r1×r2×…×r w )÷32 warps, denoted as N warp ; When N warp When the number of stream processors is less than or equal to that of the GPU device, it means that all warps can run simultaneously, that is, there is no warp rotation. At this time, the video memory access of a warp cannot be hidden by the calculation of other warps. The total execution time of LB on the GPU device is: TGPU LB =T transfer +T compute +T MEM_MAX When N warp When it is larger than the stream processor of the GPU device, it means that all warps cannot run at the same time, that is, there is warp rotation. At this time, the video memory access of a warp may be hidden by the calculation of other warps. The total execution time of LB on the GPU device is:
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.