Task processing method and apparatus
By analyzing and utilizing the subtask dependencies in grid computing tasks, and allocating computing resources to perform subtasks, the problems of low resolution efficiency of sparse triangle equation systems and large access stocks in high-performance computing are solved, and the effect of quickly processing grid computing tasks and improving task processing efficiency is achieved.
Patent Information
- Application Number
- PCT/CN2024/129354
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-07
- Filing Date
- 2024-11-01
- Publication Date
- 2025-06-12
AI Technical Summary
In the field of high-performance computing, there are problems such as low efficiency and large access stocks for solving sparse triangular equations during task processing.
By analyzing the dependencies between subtasks in grid computing tasks, determining dependency rules, and allocating computing resources to perform subtasks based on these rules, thereby improving the efficiency of task processing.
This method can quickly process grid computing tasks, improve task processing efficiency, reduce access stocks, and improve computing performance.
Smart Images

Figure CN2024129354_12062025_PF_FP_ABST
Abstract
Description
Task processing method and device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 7, 2023, with application number 202311680958.5 and application name “A Task Processing Method and Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a task processing method and device. Background Art
[0003] In the field of high performance computing (HPC), computing tasks in most applications are converted into solving partial differential equations after mathematical and physical modeling.
[0004] The idea of solving partial differential equations is to numerically discretize the partial differential equations and transform the problem into solving a large-scale sparse linear system of equations. Among them, one method of solving a sparse linear system of equations is the iterative method, which is used to solve the sparse linear system of equations. The solution problem of sparse linear system of equations includes the solution problem of sparse triangular equations (sparse triangular solve, SpTRSV).
[0005] Currently, for the computational task of solving sparse triangular equations, the amount of memory access during task processing is large and the task processing efficiency is low.
[0006] Summary of the Invention
[0007] The present application provides a task processing method and device that can quickly process grid computing tasks and improve the efficiency of task processing.
[0008] This application adopts the following technical solutions:
[0009] In a first aspect, the present application provides a task processing method, comprising: obtaining a grid computing task in an application, the grid computing task including multiple subtasks, and the dependency relationship between the multiple subtasks conforming to the dependency rules; analyzing the dependency relationship between the first subtask and other subtasks in the multiple subtasks, and determining the dependency rules; and analyzing the dependency relationship between other subtasks in the multiple subtasks according to the dependency rules; and allocating computing resources to the multiple subtasks according to the dependency relationship between the multiple subtasks, and executing the multiple subtasks.
[0010] In this application, since the dependency relationship between multiple subtasks of a grid computing task conforms to the dependency rules, by analyzing the dependency relationship between one subtask and other subtasks among the multiple subtasks and obtaining the dependency rules, the dependency relationship between other subtasks can be quickly analyzed based on the dependency rules, and then computing resources are allocated to the multiple subtasks according to the dependency relationship between the multiple subtasks, and multiple subtasks are executed. In this way, the grid computing task can be processed quickly and the efficiency of task processing can be improved.
[0011] In one possible implementation, the grid computing task is to solve triangular equations using matrix multiplication.
[0012] In one possible implementation, the task processing method provided in this application further includes: for each of the multiple subtasks, deleting duplicate dependencies between the subtask and other subtasks, thereby pruning the dependencies between the multiple subtasks. Pruning redundant dependencies can simplify the dependencies and facilitate subsequent resource allocation.
[0013] In one possible implementation, the method provided in the present application also includes: for a first subtask and a second subtask that are executed in parallel and have no dependency relationship, adjusting the storage location of the data of the first subtask or the data of the second subtask so that the storage locations of the data of the first subtask and the data of the second subtask are continuous, thereby saving memory bandwidth.
[0014] In one possible implementation, the grid of the above-mentioned grid computing task is a structural grid, and multiple vertices of the structural grid correspond one-to-one to multiple subtasks. The dependency relationship between the first subtask and other subtasks in the multiple subtasks is analyzed, and the dependency rules are determined, including: analyzing at least one other vertex on which the first vertex corresponding to the first subtask depends; and determining the dependency rules based on the positional relationship between the first vertex and at least one other vertex.
[0015] In this application, since the points in the structural grid of the grid computing task have a one-to-one correspondence with the subtasks, determining the dependency rules of the subtasks by analyzing the positional relationship of the points with dependency relationships in the structural grid helps to quickly analyze the dependency relationships between multiple subtasks.
[0016] In a possible implementation, the structured grid is a three-dimensional structured grid, and the relationship between the index number p of the subtask of the grid computing task and the coordinates (i, j, k) of the vertices of the structured grid satisfies:
[0017] p=i+j×n x +k×n x ×n y
[0018] Among them, n x Indicates the number of grid cells in the x-axis direction of the structured grid, n y Indicates the number of grid cells in the structured grid in the y-axis direction.
[0019] In a second aspect, the present application provides a computing device comprising various modules for implementing the method described in the first aspect and one of its possible implementations, such as an acquisition module, an analysis module, a processing module, a cropping module, an adjustment module, etc.
[0020] The computing device has the functionality to implement the behaviors in the method examples of any one of the first aspect and its possible implementations. The functionality can be implemented by hardware or by hardware executing corresponding software implementations. The hardware or software includes one or more modules corresponding to the functionality.
[0021] In a third aspect, the present application provides a computing device comprising a memory and at least one processor connected to the memory, the memory being used to store computer program code, the computer program code comprising computer instructions, which, when executed by at least one processor, enable the computing device to execute the method of the first aspect and any one of its possible implementations.
[0022] In a fourth aspect, the present application provides a computer-readable storage medium storing computer instructions, which, when executed on a computer, execute the method of the first aspect and any one of its possible implementations.
[0023] In a fifth aspect, the present application provides a computer program product, which includes computer instructions. When the computer instructions are run on a computer, the method of the first aspect and any one of its possible implementations is executed.
[0024] In a sixth aspect, the present application provides a chip system, comprising: a processor for calling and running a computer program from a memory, so that a computing device equipped with the chip system executes the method of the first aspect and any one of its possible implementations.
[0025] It should be understood that the beneficial effects achieved by the technical solutions of the second to sixth aspects of this application and the corresponding possible implementation methods can be referred to the technical effects of the first aspect and its corresponding possible implementation methods mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] FIG1 is a schematic diagram of a structural grid according to an embodiment of the present application;
[0027] FIG2 is a schematic diagram showing the positions of non-zero elements of a sparse triangular matrix provided in an embodiment of the present application;
[0028] FIG3 is a schematic diagram of the positions of non-zero elements of a sparse triangular matrix and task dependencies provided by an embodiment of the present application;
[0029] FIG4 is a schematic diagram showing the relationship between a CSR format and the positions of non-zero elements of a coefficient matrix provided in an embodiment of the present application;
[0030] FIG5 is a schematic diagram of the hardware structure of a computing device provided in an embodiment of the present application;
[0031] FIG6 is a flowchart of a task processing method according to an embodiment of the present application;
[0032] FIG7 is a second schematic diagram of a structural grid provided in an embodiment of the present application;
[0033] FIG8 is a schematic diagram of coordinates of points in a structured grid provided in an embodiment of the present application;
[0034] FIG9 is a second flowchart of the task processing method provided in an embodiment of the present application;
[0035] FIG10 is a schematic diagram showing the relationship between the coordinates of a point in a structured grid and the index of a subtask provided in an embodiment of the present application;
[0036] FIG11 is a third flowchart of the task processing method provided in an embodiment of the present application;
[0037] FIG12 is a schematic diagram of a task dependency clipping result provided by an embodiment of the present application;
[0038] FIG13 is a schematic diagram of one of the analysis results of a dependency relationship provided in an embodiment of the present application;
[0039] FIG14 is a second schematic diagram of a dependency analysis result provided in an embodiment of the present application;
[0040] FIG15 is a fourth flowchart of the task processing method provided in an embodiment of the present application;
[0041] FIG16 is a schematic diagram of a row structure of a structured matrix provided in an embodiment of the present application;
[0042] FIG17 is a schematic diagram of a structured matrix provided in an embodiment of the present application;
[0043] FIG18 is a schematic diagram of a solution process of a sparse triangular system of equations provided in an embodiment of the present application;
[0044] FIG19 is a schematic diagram of a structure of a computing device according to an embodiment of the present application;
[0045] FIG20 is a second schematic diagram of the structure of the computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0046] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0047] In the description and claims of the embodiments of this application, the terms "first" and "second" are used to distinguish different objects, rather than to describe a specific order of objects. For example, the terms "first vertex" and "second vertex" are used to distinguish different vertices, rather than to describe a specific order of vertices.
[0048] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0049] In the description of the embodiments of the present application, unless otherwise specified, “multiple” means two or more, and “multiple” can also be described as “at least two”.
[0050] The task processing method provided in the embodiment of the present application can be applied in the field of high-performance computing (HPC). For example, for computing tasks in applications (such as meteorological applications, etc.), a large number of computing problems can be efficiently solved by this method. In the field of HPC, most of the computing problems in applications can be converted into problems of solving partial differential equations. One solution idea for partial differential equations is to numerically discretize them and convert the problem into a problem of solving a sparse linear system of equations. Further, many methods for solving a sparse linear system of equations include computing tasks of solving triangular equations (i.e., solving a sparse triangular system of equations) by matrix multiplication, and efficiently processing such computing tasks has an important impact on the computing performance of HPC.
[0051] First, some technical terms involved in a task processing method and device provided in an embodiment of the present application are explained.
[0052] 1. Grid and grid computing tasks
[0053] First, let's introduce the concept of meshing. The idea behind meshing stems from discretization. This involves dividing a continuous computational domain into a number of finite subregions, solving for the physical variables in each subregion separately, and ultimately obtaining the physical variables for the entire computational domain. Dividing the computational domain into multiple subregions can be understood as the process of meshing, with each subregion being equivalent to the smallest grid cell.
[0054] In the embodiments of the present application, tasks that complete calculations based on the concept of a grid are defined as grid computing tasks. For example, for solving a system of equations, the concept of a grid can be used to discretize the calculation area into a number of finite sub-areas, that is, divided into multiple grid cells, to complete the solution of the system of equations. Therefore, the task of solving the system of equations is a grid computing task. In the embodiments of the present application, the calculation area of the computing task can be called the grid of the computing task, and the grid of the computing task includes multiple grid cells.
[0055] Optionally, the grid may be two-dimensional, three-dimensional, or multi-dimensional. A two-dimensional grid may be, for example, a quadrilateral grid, and a three-dimensional grid may be, for example, a cuboid or cube grid.
[0056] It can be understood that the grid may include a structured grid and an unstructured grid. The embodiment of the present application relates to a structured grid, and only the structured grid is introduced below, while the unstructured grid is not introduced.
[0057] For a structured grid, the grid area is divided into multiple grid cells, and each point in the grid area (referring to the vertex of the grid cell in the structured grid) has one or more neighboring grid cells, and the neighboring grid cells are grid cells adjacent to the point in all directions. A structured grid means that all internal points in the grid area have the same number of neighboring grid cells of the same type. That is to say, for any two points in the grid area, for example, point A and point B, the number of neighboring grid cells of point A is the same as the number of neighboring grid cells of point B, and the type of neighboring grid cells of point A is the same as the type of neighboring grid cells of point B. It should be noted that for a point, for example, point A, the types of multiple neighboring grid cells of point A are also the same.
[0058] In the embodiments of this application, the structural grid of the 3D19 template is used as an example for illustration. Referring to Figure 1, the multiple grid cells in the structural grid of the 3D19 template are rectangular parallelepipeds, each vertex has 8 neighboring grid cells, and each grid cell is a cube. The number of points in a point association group is 19.
[0059] 2. Sparse Linear Equations
[0060] First, we briefly introduce the concept of linear equations. Linear equations can be expressed in the following matrix equation form:
[0061] Ax=b
[0062] Where A represents the coefficient matrix of the linear equation system, A∈n×n; x is the solution vector to be solved in the linear equation system, x∈n×1; b is the right-hand side vector of the linear equation system, b∈n×1.
[0063] It can be understood that the process of solving a system of linear equations refers to the process of obtaining the solution vector x through the coefficient matrix A and the right-hand side vector b.
[0064] For the coefficient matrix A of the above linear equations, when the values of most elements in the coefficient matrix A are 0, that is, when the n 2 When the proportion of non-zero elements in the elements is low (for example, less than 0.1%), the coefficient matrix A is considered to be a sparse matrix. Correspondingly, when the coefficient matrix A in the above linear equation system is a sparse matrix, the linear equation system is a sparse linear equation system.
[0065] 3. Sparse triangular equations
[0066] Based on the above introduction to the concept of sparse linear equations, sparse triangular equations refer to linear equations whose coefficient matrices are sparse triangular matrices. Sparse triangular equations can be expressed as the following matrix equation:
[0067] Wx=b
[0068] Where W represents the coefficient matrix of the sparse triangular system of equations, W is a triangular matrix, and W is a sparse triangular matrix. W∈n×n; x is the solution vector to be solved in the linear system of equations, x∈n×1; b is the right-hand side vector of the linear system of equations, b∈n×1
[0069] Optionally, in the sparse triangular equation system, the coefficient matrix W can be a sparse lower triangular matrix or a sparse upper triangular matrix, which is not limited in the embodiments of the present application.
[0070] Typically, there are two methods for solving the aforementioned sparse linear equations: direct methods and iterative methods. Iterative methods gradually approach the solution through iterative corrections. In iterative methods, the preconditioner is a crucial component, accelerating convergence and addressing ill-conditioned matrices.
[0071] Optionally, iterative methods may include, but are not limited to, conjugate gradient methods, generalized minimum residual methods, generalized conjugate residual methods, and Chebyshev methods. Preconditioners may include, but are not limited to, additive Schwarz, block Jacobian, Jacobian, successive over-relaxation (SOR), incomplete lower-upper factorization (ILU), and multigrid algorithms. Among them, the ILU and SOR algorithms are the most basic preconditioners, with simple algorithmic processes and good convergence. The SOR and ILU algorithms can also be nested with other preconditioners.
[0072] The above-mentioned SOR algorithms may include Gauss-Seidel, symmetric Gauss-Seidel, symmetric successive over-relaxation, etc. ILU algorithms include ILU(0), ILU(1), ILU(k), etc.
[0073] The following briefly introduces the core points of the SOR algorithm and ILU algorithm.
[0074] The core of the SOR algorithm is: First, split the coefficient matrix A of the sparse linear system into three matrices: a lower triangular matrix L, an upper triangular matrix U, and a diagonal matrix D, that is, A = L + U + D. This converts the original sparse linear system into (L + U + D) x = b. It can be seen that solving a sparse linear system involves solving multiple sparse triangular systems (where the diagonal matrix D can be combined with the lower triangular matrix L or the upper triangular matrix U into a triangular matrix, and the diagonal matrix D can also be regarded as a triangular matrix). Secondly, according to different SOR algorithms, the sparse triangular system of equations is solved.
[0075] Taking the Gauss-Seidel algorithm as an example, the iteration formula of the Gauss-Seidel algorithm is (L+D)×x(k+1)=bU×x(k) , x(k) is the solution vector obtained at the kth iteration, x(k+1) is the solution vector obtained at the k+1th iteration, and L+D is the lower triangular part of the coefficient matrix A.
[0076] The core of the ILU algorithm is: first, decompose the coefficient matrix A of the sparse linear equation system into two matrices, a lower triangular matrix L and an upper triangular matrix U, that is, A = L × U, so that the original sparse linear equation system is transformed into LUx = b; then, assuming Ux = y, the solution process of LUx = b includes solving the two sparse triangular equation systems Ly = b and Lx = y.
[0077] 4. Compressed sparse row (CSR) format
[0078] The CSR format is a sparse storage format for sparse matrices. For example, the coefficient matrix A in the above-mentioned sparse linear equations can be stored in the CSR format. Using the sparse storage format to store matrices can save memory space.
[0079] It is understandable that when the proportion of non-zero elements (or non-zero elements) in the coefficient matrix A is very low, if the traditional two-dimensional array (a[n][n]) is used to store the coefficient matrix A, a lot of memory space will be wasted. Therefore, for sparse matrices, a sparse storage format is generally used. The CSR format is a commonly used sparse storage format.
[0080] The CSR format uses three one-dimensional arrays to represent the coefficient matrix A. The three one-dimensional arrays are a, d, and c. Array a includes nnz elements, array d also includes nnz elements, and array c includes n+1 elements. nnz represents the total number of non-zero elements in the coefficient matrix A, and n represents the number of rows in the coefficient matrix A. Array a is used to store the values of all non-zero elements in the coefficient matrix A, array d is used to store the index of the column in the coefficient matrix A where each non-zero element is located, and array c is used to store the index of the position of the first non-zero element in each row of the coefficient matrix A in the above array a, as well as the number of non-zero elements in the coefficient matrix A.
[0081] a[k] is the k-th element in array a, and represents the value of the k-th non-zero element of coefficient matrix A. d[k] is the k-th element in array d, and represents the column index of the k-th non-zero element of coefficient matrix A, k = 0, 1, ..., nzz-1. c[t] is the t-th element in array c, and represents the position index of the first non-zero element in the t-th row of coefficient matrix A in array a, t = 0, 1, ..., n-1. c[n] is the n-th element (i.e., the last element) in array c, and represents the total number of non-zero elements in coefficient matrix A.
[0082] It should be noted that in the embodiments of the present application, for arrays, matrices and vectors, the position indexes of the elements therein all start numbering from 0. For example, if a vector includes n elements, the position indexes of the elements of the vector are 0 to n-1, and the kth element of the vector refers to the element with index number k in the vector. The following embodiments will not explain them one by one.
[0083] For example, for the following coefficient matrix A:
[0084] The three one-dimensional arrays corresponding to the coefficient matrix A are:
[0085] a=[1, 2, 3, 6, 7, 9]; where a[2]=3, indicating that the value of the second non-zero element of the coefficient matrix A is 3.
[0086] d=[0, 2, 1, 2, 1, 3]; where d[5]=3, indicating that the column index of the fifth non-zero element of the coefficient matrix A is 3.
[0087] c=[0, 2, 3, 4, 6]; where c[2]=3, indicating that the index of the first non-zero element in the second row of the coefficient matrix A (i.e., non-zero element 6) in array a is 3.
[0088] Based on the above description of the CSR format, we can see that the storage space of a traditional two-dimensional array requires O(n 2 ), while the CSR format only takes O(n+2nnz). When 2nnz is much smaller than n 2 When using CSR format to store the coefficient matrix A, the memory usage can be significantly reduced.
[0089] 5. Dependencies in Sparse Triangular Equations
[0090] As can be seen from the description of the above embodiments, in the HPC field, most computing problems, after transformation, will include the problem of solving sparse triangular equations. In the embodiments of the present application, the task of solving the sparse triangular equations involved in the application is referred to as a computing task of the application, and the process of solving a component in the solution vector of the sparse triangular equations is defined as a subtask of solving the sparse triangular equations. In other words, a computing task includes multiple subtasks. In the process of solving each component, the solution of the component may depend on other components.
[0091] The following is an example of a simple sparse triangular equation system. For the following example sparse triangular equation system:
[0092] The above sparse triangular equations can also be expressed as:
[0093] Combined with the CSR format of the coefficient matrix of the sparse triangular equations, the non-zero element array a = [1, 2, 1, 3, 1, 4, 1], the above sparse triangular equations can also be expressed as:
[0094] For the solution vector x in the sparse triangular system of equations, the four components of the solution vector are x0, x1, x2, and x3, corresponding to the four subtasks: subtask 0, subtask 1, subtask 2, and subtask 3. Referring to Figure 2, the positions of the non-zero elements of the coefficient matrix (sparse lower triangular matrix) are shown. Combining Figure 2 with the above sparse triangular system of equations, we can directly draw the following conclusions:
[0095] The dependency relationship between subtask 0 (corresponding to solving component x0) and other subtasks is: subtask 0 does not depend on other subtasks.
[0096] The dependency relationship between subtask 1 (corresponding to solving component x1) and other subtasks is: subtask 1 depends on subtask 0.
[0097] The dependency relationship between subtask 2 (corresponding to solving component x2) and other subtasks is: subtask 2 depends on subtask 1.
[0098] The dependency relationship between subtask 3 (corresponding to solving component x3) and other subtasks is: subtask 3 depends on subtask 2.
[0099] In summary, we can draw the following conclusion: for subtask i, the index number of the subtask on which subtask i depends is equal to the column number of the non-zero element in the i-th row in the coefficient matrix.
[0100] For example, referring to the position of the non-zero elements of the coefficient matrix in Figure 2, for subtask 0, the column number of the non-zero elements in the 0th row of the coefficient matrix is 0, which determines that subtask 0 does not depend on other subtasks; for subtask 1, the column numbers of the non-zero elements in the 1st row of the coefficient matrix are 0 and 1, which determines that subtask 1 depends on subtask 0; for subtask 2, the column numbers of the non-zero elements in the 2nd row of the coefficient matrix are 1 and 2, which determines that subtask 2 depends on subtask 1; for subtask 3, the column numbers of the non-zero elements in the 3rd row of the coefficient matrix are 2 and 3, which determines that subtask 3 depends on subtask 2.
[0101] For example, referring to FIG3 , FIG3 (a) illustrates the non-zero elements of the coefficient matrix (lower triangular matrix) of a sparse triangular equation system. In the figure, the shaded positions are all positions of non-zero elements (black shades indicate diagonal positions). According to FIG3 , the sparse triangular equation system includes 16 subtasks, and the coefficient matrix includes 16 rows. According to FIG3 (a), the dependency relationship between the 16 subtasks can be analyzed. The schematic diagram of the dependency relationship between the 16 subtasks is FIG3 (b). The dependency relationship includes four layers in sequence, and the dependency relationship between the rows of subtasks is indicated by arrows. For example, subtask 0 and subtask 2 point to subtask 3 through arrows, indicating that subtask 3 depends on subtask 0 and subtask 2. As shown in (b) of Figure 3, subtasks 0, 1, and 2 do not depend on other subtasks. That is, subtasks 0, 1, and 2 are independent of each other and are located in the first layer of the dependency graph. Subtasks 3, 4, 5, 6, 7, and 8 are independent of each other and are located in the second layer of the dependency graph. Subtasks 9, 10, 11, 12, and 13 are independent of each other and are located in the third layer of the dependency graph. Subtasks 14 and 15 are independent of each other and are located in the fourth layer of the dependency graph. Each subtask in the latter layer depends on some subtasks in the former layer (at least one former layer).
[0102] In summary, by analyzing the coefficient matrix of the sparse triangular system, we can derive the dependencies between the subtasks of the sparse triangular system. Some of the subtasks are independent, while others are interdependent. Based on these dependencies, we can solve each subtask in parallel using multiple threads, achieving high-performance solutions to the sparse triangular system.
[0103] Currently, the coefficient matrix of a sparse triangular system of equations is typically stored in the CSR format, rather than directly storing the coefficient matrix. During the solution of a sparse triangular system of equations, the dependencies between subtasks are analyzed based on the three arrays stored in the CSR format, and threads are then assigned to the subtasks based on these dependencies. The three one-dimensional digits in the CSR format are a, d, and c. Array a stores the values of all nonzero elements in the coefficient matrix, array d stores the index of the column in which each nonzero element resides in the coefficient matrix, and array c stores the index of the position of the first nonzero element in each row of the coefficient matrix in array a, as well as the number of nonzero elements in the coefficient matrix.
[0104] For example, the three arrays of the coefficient matrix of a sparse triangular equation stored in CSR format are: a = {1, 2, 1, 3, 1, 4, 1}, d = {0, 0, 1, 1, 2, 2, 3}, and c = {0, 1, 3, 5, 7}. According to a, the dimension of the coefficient matrix is 4×4, and the total number of non-zero elements is 7. Therefore, when analyzing the dependencies between the subtasks of the sparse triangular equation system, it is necessary to analyze each subtask one by one.
[0105] Referring to (a) in FIG4 , according to array c, the first non-zero element of the 0th row of the coefficient matrix has a position index of 0 in array a, and the first non-zero element of the 1st row has a position index of 1 in array a. It can be seen that the 0th row of the coefficient matrix includes 1 non-zero element, and the column index of the non-zero element read from array d is 0; further, since the first non-zero element of the 1st row of the coefficient matrix has a position index of 1 in array a, and the first non-zero element of the 2nd row has a position index of 3 in array a, it can be seen that the 0th row of the coefficient matrix includes 1 non-zero element. Row 1 includes 2 non-zero elements, and the column indices of the non-zero elements read from array d are 0 and 1; further, since the position index of the first non-zero element of the 2nd row of the coefficient matrix in array a is 3, and the position index of the first non-zero element of the 3rd row in array a is 5, it can be seen that the 2nd row of the coefficient matrix includes 2 non-zero elements, and the column indices of the non-zero elements read from array d are 1 and 2; finally, the 3rd row (the last row) of the coefficient matrix contains 2 non-zero elements, and the column indices of the non-zero elements read from array d are 2 and 3.
[0106] Refer to (b) in Figure 4. According to the above analysis results, the position of the non-zero elements in each row of the coefficient matrix can be obtained, and thus the dependency relationship between the subtasks of the sparse triangular equation system can be obtained, such as subtask 0 does not depend on other subtasks, subtask 1 depends on subtask 0, subtask 2 depends on subtask 1, and subtask 3 depends on subtask 2.
[0107] In the process of analyzing the dependencies between the subtasks using the array of coefficient matrices of a sparse triangular matrix stored in CSR format, for each subtask of the coefficient matrix, when analyzing the subtasks on which the subtask depends, it is necessary to read the relevant elements from arrays c and d. When the dimension of the coefficient matrix is relatively high, the dependencies are analyzed subtask by subtask, which makes the efficiency of solving the sparse triangular equation system low. In addition, the amount of memory access in the process of analyzing the dependencies is relatively large, and the computational memory access ratio is low (the computational memory access ratio is used to reflect the computational intensity of a program relative to memory access. The computational memory access ratio can be the ratio of the number of computational operations to the number of memory access operations). It should be understood that the problem of solving the sparse triangular equation system is a problem with limited memory access. When the computational memory access ratio is high, the performance of the solution method is better, otherwise, the performance is poor.
[0108] In response to the problem of low efficiency in computing task processing in the prior art, an embodiment of the present application provides a task processing method for processing grid computing tasks, wherein a task processing device obtains a grid computing task in an application, wherein the grid computing task includes multiple subtasks, and the dependency relationship between the multiple subtasks complies with the dependency rule; then, the dependency relationship between the first subtask and the other subtasks in the multiple subtasks is analyzed, and the dependency rule is determined; then, the dependency relationship between the other subtasks in the multiple subtasks is analyzed according to the dependency rule; then, computing resources are allocated to the multiple subtasks according to the dependency relationship between the multiple subtasks, and the multiple subtasks are executed. In this method, since the dependency relationship between the multiple subtasks of the grid computing task complies with the dependency rule, the dependency relationship between one subtask and the other subtasks in the multiple subtasks is analyzed and the dependency rule is obtained, and then, based on the dependency rule, the dependency relationship between the other subtasks can be quickly analyzed and obtained, and then computing resources are allocated to the multiple subtasks according to the dependency relationship between the multiple subtasks, and the multiple subtasks are executed. In this way, the grid computing task can be processed quickly and the efficiency of task processing can be improved.
[0109] Optionally, the hardware device used to execute the task processing method provided in the embodiments of the present application can be a device with processing and storage functions, such as a computing device, such as a server, a desktop computer, etc. For example, the method can be integrated into various applications that need to process grid computing tasks, such as a third-party mathematical library, which is applied in some solver frameworks in the form of a mathematical library and then indirectly called by upper-level applications, or called by the solver framework or upper-level applications in the form of code.
[0110] Taking the hardware device for executing the task processing method as a computing device as an example, Figure 5 is a schematic diagram of the hardware structure of a computing device provided in an embodiment of the present application. The various components shown in Figure 5 can be implemented in hardware, software, or a combination of hardware and software including one or more signal processing and / or dedicated integrated circuits.
[0111] As shown in Figure 5, the computing device may include: a processor 501, a memory 502, and a communication interface 503. The processor 501, the memory 502, and the communication interface 503 may be connected to each other via a bus 504, or in other ways.
[0112] The processor 501 is the control center of the computing device. The processor 501 may be a general-purpose central processing unit (CPU) or other general-purpose processors. The general-purpose processor may be a microprocessor or any conventional processor.
[0113] The controller in processor 501 is the nerve center and command center of the computing device. Based on instruction opcodes and timing signals, the controller generates operational control signals to control instruction fetching and execution. Optionally, processor 501 may also include a memory for storing instructions and data. Exemplarily, processor 501 may include one or more CPUs, such as CPU 0 and CPU 1 shown in FIG5 .
[0114] The memory 502 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer. In the embodiment of the present application, the memory 502 can store information such as computer instructions.
[0115] In one possible implementation, the memory 502 may exist independently of the processor 501. The memory 502 may be connected to the processor 501 via a bus 504 and used to store data (e.g., data in a CSR format), instructions, or program codes. When the processor 501 calls and executes the instructions or program codes stored in the memory 502, the relevant steps of the method provided in the embodiment of the present application can be implemented.
[0116] In another possible implementation, the memory 502 may also be integrated with the processor 501 .
[0117] The communication interface 503 can be a transceiver module for communicating with other devices or communication networks (for example, communicating with the log management platform or node information processing platform shown in Figure 5), such as Ethernet, RAN, wireless local area networks (WLAN), etc. The communication interface 503 can receive instructions, messages or data. The transceiver module can be a device such as a transceiver or a transceiver. Optionally, the communication interface 503 can also be a transceiver circuit located in the processor 501, for realizing signal input and signal output of the processor. The communication interface 503 can be a wired interface (port), such as a fiber distributed data interface (FDDI), a gigabit Ethernet (GE) interface, or the communication interface 503 can also be a wireless interface.
[0118] Bus 504 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. This bus can be classified as an address bus, a data bus, a control bus, etc. Buses can also be classified as serial buses and parallel buses. For ease of illustration, FIG5 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0119] Optionally, the computing device in the embodiment of the present application may further include an input / output interface 505, which is used to connect to an input device and receive information input by a user through the input device (e.g., coefficient matrix, right-end vector, template information of the structure grid, etc.). Input devices include but are not limited to keyboards, touch screens, microphones, etc. The input / output interface 505 is also used to connect to an output device to output processing results of the processor.
[0120] It should be noted that the computing device shown in FIG5 is merely an example of a computing device, and the computing device may have more or fewer components than those shown in FIG5 , may combine two or more components, or may have a different component configuration.
[0121] In combination with the above, the task processing method provided by the embodiment of the present application is described in detail below. As shown in FIG6 , the method includes S601 - S604 .
[0122] S601: Obtain a grid computing task in an application.
[0123] A grid computing task includes multiple subtasks, and the dependencies between the subtasks conform to the dependency rules.
[0124] In the embodiment of the present application, the grid of the grid computing task is a structural grid, and the multiple vertices of the structural grid correspond one-to-one to the multiple subtasks of the grid computing task. It should be understood that based on the structural characteristics of the structural grid described in the above embodiment, for any point in the structural grid of the grid computing task (the vertex of the grid unit), hereinafter referred to as the target point, the target point has one or more associated points, and the associated points of the target point are points that have an associated relationship with the target point. In the embodiment of the present application, the target point and the associated points of the target point can be formed into a point association group.
[0125] It should be noted that, except for the points on the boundary surface of the structured grid, other points of the structured grid, when used as target points, all have the same number of associated points, and the positional relationships between the associated points and the target points are also the same.
[0126] Since the multiple vertices of the structural grid correspond one-to-one to the multiple subtasks of the grid computing task, there is also a dependency relationship between the multiple subtasks. The dependency relationship between the subtasks can be obtained through the positional relationship between the points. For example, for a subtask corresponding to a target point, the subtask that the subtask depends on is the subtask corresponding to the associated point of the target point, and any of the multiple subtasks can determine the subtask it depends on according to the same positional relationship. It can be seen that for the multiple subtasks of the grid computing task, the dependency relationship between the multiple subtasks conforms to the dependency rule. The dependency rule can be understood as the positional relationship between the target point and the associated point mentioned above. The dependency rule can also be called a dependency template.
[0127] If the grid computing task is a task of solving equations (i.e., solving a system of equations) using matrix multiplication, then the total number of points in a point association group in the structural grid of the grid computing task is equal to the total number of non-zero elements in the target row (the target row is the row corresponding to the target point) of the coefficient matrix of the system of equations.
[0128] Optionally, the structural grid includes different templates, and the number and position of associated points of the target point in the structural grids of different templates are different, corresponding to the coefficient matrix of the equation group, the number of non-zero elements in each row of the coefficient matrix is different.
[0129] For example, if the grid computing task is to solve the sparse linear equation group Ax=b, if the size of the coefficient matrix A is 27×27, the size of the solution vector x is 27×1, and the size of the right-hand vector b is 27×1. Taking the structural grid of the sparse linear equation as the structural grid of the 3D19 template as an example, the size of the structural grid is 3×3×3, that is, the structural grid is a cube, and contains 2 grid cells in three directions (i.e., x, y, z directions), and each grid cell is also a cube. Referring to Figure 7, a structural grid of a sparse linear equation group is shown, which includes 8 grid cells and 27 vertices. A target point (a point on a non-boundary surface) in the structural grid has 18 associated points. Therefore, corresponding to the sparse linear equation group, the corresponding row in the coefficient matrix of the sparse linear equation group has 19 non-zero elements.
[0130] For the convenience of description, as shown in Figure 7, the 27 vertices in the structured grid are numbered 0-26. The 27 points can be recorded as point 0, point 1, ..., point 26 respectively. Taking point 13 in the figure as an example, the point association group with point 13 as the target point includes 19 points, and the point numbers are {0, 1, 3, 4, 5, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 21, 22, 23, 25}.
[0131] For the sparse linear system of equations, after splitting or decomposing the coefficient matrix into a lower triangular matrix and an upper triangular matrix, solving the sparse linear system of equations includes solving the sparse lower triangular system of equations and the sparse upper triangular system of equations. For the 13th row in the coefficient matrix of the sparse linear system of equations, among the 18 points in the point association group other than point 13, some points correspond to the associated points of point 13 in the sparse lower triangular system of equations (or the points on which point 13 depends), and another part of the 18 points correspond to the associated points of point 13 in the sparse upper triangular system of equations (or the points on which point 13 depends). Continuing to refer to Figure 1, for the sparse lower triangular equation system corresponding to the sparse linear equation system, point 13 in the structural grid depends on 9 points, and the 9 points are numbered {1, 3, 4, 5, 7, 9, 10, 11, 12}; for the sparse upper triangular equation system corresponding to the sparse linear equation system, point 13 in the structural grid depends on 9 points, and the 9 points are numbered {14, 15, 16, 17, 19, 21, 22, 23, 25}.
[0132] In one implementation, the grid computing task in the embodiments of the present application is to solve triangular equations using matrix multiplication, i.e., to solve a triangular system of equations. The triangular system of equations may be a sparse triangular system of equations, and a component of the solution vector for solving the triangular system of equations is referred to as a subtask of the triangular system of equations. For ease of description, the following embodiments illustrate the grid computing task of solving a sparse triangular system of equations as an example.
[0133] If the grid computing task is solving a sparse triangular equation system, the above-mentioned acquisition application of one grid computing task includes acquiring structural grid information of the sparse triangular equation system and a coefficient matrix of the sparse triangular equation system.
[0134] In the embodiment of the present application, the sparse triangular system of equations can be expressed as Wx = b, where W is the coefficient matrix, and W is a triangular matrix, W∈n×n; x is the solution vector, x∈n×1; and b is the right-hand side vector, b∈n×1. Solving a component of the solution vector x is considered a subtask of solving the sparse triangular system of equations. Therefore, the sparse triangular system of equations includes n subtasks, and solving the sparse triangular system of equations is equivalent to solving n subtasks.
[0135] Optionally, the coefficient matrix W can be an upper triangular matrix or a lower triangular matrix, which is not limited in the embodiment of the present application.
[0136] Optionally, the structured grid information may include template information of the structured grid, size information of the structured grid, and the like.
[0137] The structural grid information is used to indicate the distribution of multiple subtasks of a grid computing task within the structural grid of the grid computing task. For the structural grid of a sparse triangular system of equations, each vertex in the structural grid corresponds to a subtask. Specifically, if a vertex in the structural grid corresponds to a subtask, the index of a vertex in the structural grid is the same as the row index of a corresponding component in the solution vector. For example, vertex 0 in the structural grid corresponds to component 0 of the solution vector (i.e., subtask 0).
[0138] In the embodiment of the present application, in the coordinate system of the structured grid of the sparse triangular equation system, the relationship between the coordinates of each vertex and the index of the vertex is defined, such as the component x of the solution vector x of the sparse triangular equation system. p The corresponding subtask has an index number of p, which corresponds to the coordinates (i, j, k) of the vertices of the structure grid, where p = 0, 2, ..., n-1. Thus, based on the correspondence between the subtask index and the coordinates of the vertices of the structure grid, it is possible to associate points in the structure grid with subtasks, and then, by analyzing the dependencies between the points, quickly analyze the dependencies between multiple subtasks of the grid computing task.
[0139] If the structural grid is a three-dimensional structural grid, as shown in Figure 8, the above structural grid is a cube (including 6 faces, front, back, left, right, top, and bottom). For the structural grid of the cube, the coordinate system of the structural grid is defined, where the horizontal right direction is defined as the positive direction of the x-axis, the vertical upward direction is defined as the positive direction of the z-axis, and the direction perpendicular to the x-axis and z-axis pointing to the back of the cube is defined as the positive direction of the y-axis.
[0140] In one implementation, the size information of the structured grid includes: the number n of grid cells in the x-axis direction of the structured grid x , the number of grid cells in the structured grid in the y-axis direction n y And the number of grid cells n in the structure grid in the z-axis direction z When the coefficient matrix of the sparse triangular equation system is a lower triangular matrix, the relationship between the index number p of the subtask and the coordinates (i, j, k) of the vertex of the structure grid satisfies:
[0141] p=i+j×n x +k×n x ×n y
[0142] For example, for the sparse triangular equation system Wx=b, if the size of the coefficient matrix W is 27×27, the size of the solution vector x is 27×1, and the dimension of the right-hand vector b is 27×1, still taking the structural grid of the 3D19 template as an example, the structural grid of the sparse triangular equation system is the structural grid shown in Figure 7 above. The relationship between the coordinates of the points in the structural grid and the index of the subtask is shown in Figure 10 (Figure 10 illustrates the coordinates of some points). According to the above calculation formula, the index of the subtask corresponding to each point can be calculated. For example, for the vertex (1, 1, 1) in the structural grid, the index of the subtask corresponding to the vertex is calculated as 13 according to the calculation formula.
[0143] S602: Analyze the dependency relationship between the first subtask and other subtasks in the multiple subtasks, and determine the dependency rule.
[0144] In an embodiment of the present application, analyzing the dependency relationship between the first subtask and other subtasks among multiple subtasks is to analyze the distribution of the multiple subtasks indicated by the structural grid information in the structural grid, and determine which subtasks among the other subtasks are the tasks that the first subtask depends on.
[0145] In one implementation, as shown in FIG9 in combination with FIG6 , the above S602 is implemented through S6021 - S6022 .
[0146] S6021: Analyze at least one other vertex on which the first vertex corresponding to the first subtask in the structural grid of the grid computing task depends.
[0147] In an embodiment of the present application, for any two vertices in the structural grid, such as the first vertex and the second vertex, in the structural grid, the positional relationship between at least one vertex on which the first vertex depends and the first vertex is a first positional relationship, and the positional relationship between at least one vertex on which the second vertex depends and the second vertex is a second positional relationship. According to the known characteristics of the structural grid, it can be seen that the first positional relationship is the same as the second positional relationship.
[0148] It can be understood that when the grid computing task is determined, the relationship between the points in the structural grid is also determined. For example, when the coefficient matrix of the sparse triangular equation group is determined, the relationship between the various points in the structural grid of the sparse triangular equation group is known.
[0149] Continuing to refer to FIG10 above, assuming that the first vertex corresponding to the first task is point 13, point 13 depends on 9 points, namely point 1, point 3, point 4, point 5, point 7, point 9, point 10, point 11, and point 12. Thus, it can be seen that subtask 13 depends on subtask 1, subtask 3, subtask 4, subtask 5, subtask 7, subtask 9, subtask 10, subtask 11, and subtask 12. Based on the conclusion of the dependency relationship between multiple subtasks of the sparse triangular equation system introduced in the above embodiment: for subtask i, the index number of the subtask on which subtask i depends is equal to the column number of the non-zero element in the i-th row of the coefficient matrix, then the column index of the non-zero element in the 13th row of the coefficient matrix is 1, 3, 4, 5, 7, 9, 10, 11, 12, and 13.
[0150] For ease of description, the positive x-axis of the structural grid's coordinate system is referred to as the right, the negative x-axis as the left, the positive y-axis as the back, the negative y-axis as the front, the positive z-axis as the top, and the negative z-axis as the bottom. In the 3D19 template's structural grid, points on the right depend on points on the left, points above depend on points below, and points in the back depend on points in the front.
[0151] Of the nine points that point 13 depends on, point 1 is located in front of and below point 13. Starting from point 13, first move forward one unit (the length of one side of a cube grid cell) and then move downward one unit. Similarly, we can analyze the positional relationships between point 13 and points 3, 4, 5, 7, 9, 10, 11, and 12, as shown in Table 1.
[0152] Table 1
[0153] S6022: Determine a dependency rule based on a positional relationship between the first vertex and at least one other vertex.
[0154] According to the positional relationship between the first vertex and at least one other vertex shown in Table 1 above, the dependency relationship between the first subtask and other subtasks can be known. The dependency relationship between the subtasks is determined by the positional relationship between the above points. It can be understood that the above positional relationship can be considered as the dependency rule that the subtasks comply with.
[0155] S603: Analyze the dependency relationships among the other subtasks in the plurality of subtasks according to the dependency rules.
[0156] Through the description of the above embodiment, according to the positional relationship shown in Table 1, the points on which each point in the structured grid depends can be determined, that is, the subtasks on which each subtask of the sparse triangular equation system depends can be determined. It should be noted that for the points on the boundary surface of the structured grid, the number of points on which they depend is less than 9. For the above-mentioned sparse triangular equation system, Table 2 shows the subtasks on which each subtask depends. The subtasks in Table 2 are the vertices in the structured grid. According to the point i corresponding to subtask i and the positional relationship analyzed above, the one or more points on which the point i depends can be determined, and the one or more subtasks on which the subtask i depends can be determined.
[0157] Table 2
[0158] After analyzing the dependencies between subtasks, we can know which subtasks are independent of each other, which subtasks are dependent on each other, and the number of layers of dependencies.
[0159] In an embodiment of the present application, the dependency rules that the subtasks comply with are analyzed based on the structural characteristics of the structural grid, that is, the distribution of multiple subtasks, and then the dependency relationship between the multiple subtasks is obtained. For the grid computing task of solving a sparse triangular equation group, there is no need to analyze the information of the coefficient matrix of the sparse triangular equation group one by one to obtain the dependency relationship. It can be seen that the method of analyzing subtask dependency based on the structural grid in the present application can more quickly analyze and obtain the dependency relationship between multiple subtasks.
[0160] Optionally, in combination with FIG9 , as shown in FIG11 , after S603 , the method provided in the embodiment of the present application further includes S605 .
[0161] S605 : For each of the multiple subtasks, delete duplicate dependencies between the subtask and other subtasks, so as to trim the dependencies between the multiple subtasks.
[0162] In an embodiment of the present application, there may be redundancy in the dependencies between subtasks, such as in at least one subtask on which a subtask depends, there are one or more subtasks that depend on other subtasks in at least one subtask. For example, for the above-mentioned subtask 13, this subtask 13 depends on subtask 1, subtask 3, subtask 4, subtask 5, subtask 7, subtask 9, subtask 10, subtask 11, and subtask 12, among which subtask 5 also depends on subtask 1 and subtask 4. Therefore, there is a repeated dependency, that is, a redundant dependency. Trimming redundant dependencies can simplify the dependencies and facilitate subsequent resource allocation.
[0163] Optionally, deleting duplicate dependencies between a subtask and other subtasks specifically involves, for each of the multiple subtasks, deleting any duplicate subtasks from at least one of the subtasks on which each subtask depends. For example, among the subtasks on which subtask 13 depends, at least subtask 1 and subtask 4 are duplicate dependencies.
[0164] In one implementation, for each of the multiple subtasks, the method of deleting the subtasks that are repeatedly dependent on at least one subtask on which each subtask depends further includes step 1 and step 2.
[0165] Step 1: In the structure grid, for each vertex corresponding to a subtask, determine the positional relationship between the vertex and the repeatedly dependent vertex among at least one vertex on which the vertex depends.
[0166] Continuing with the above example, as shown in Figure 12, according to the dependency rules of points in the structural grid of the 3D19 template, for subtask 13, that is, point 13 in the structural grid, among the points on which point 13 depends, points 1, 3, 4, 5, 9, and 10 are repeatedly dependent. Therefore, it is necessary to delete the corresponding subtasks in the subtasks on which subtask 13 depends, so that the dependency of subtask 13 is trimmed to point 13 depending on points 7, 11, and 12. Combined with Table 2 and Figure 12 above, the positional relationship between the repeatedly dependent points among the points on which point 13 depends and point 13 is: located below and in front of point 13, below and to the left, directly below, below and to the right, in front of and to the left, and directly in front of point 13.
[0167] Step 2: According to the above positional relationship, determine the subtasks that are repeatedly dependent on at least one subtask that each subtask depends on, and delete the subtasks that are repeatedly dependent on.
[0168] In an embodiment of the present application, in a structural grid, for all vertices in the structural grid, the positional relationship between the vertex and the at least one vertex on which the vertex depends is identical. Therefore, according to the positional relationship determined in step 1, the at least one subtask on which each subtask depends can be determined to be repeatedly dependent, and then the repeatedly dependent subtasks can be deleted to obtain the at least one subtask on which each subtask ultimately depends. Table 3 illustrates the dependency relationships between the subtasks in Table 2 above after being pruned.
[0169] Table 3
[0170] According to the pruned dependency relationships shown in Table 3, we can analyze that the dependency relationships between subtasks include 13 layers, which are recorded as The subtasks contained in each layer are shown in Figure 13.
[0171] In the above process of clipping the dependencies between subtasks, based on the structural feature that all vertices in the structural grid have the same positional relationship between the repeatedly dependent vertices and at least one of the vertices on which the vertex depends. After knowing the positional relationship, the repeatedly dependent vertices on which each vertex depends can be quickly determined according to the positional relationship, thereby realizing dependency clipping more quickly.
[0172] S604: Allocate computing resources to the multiple subtasks according to the dependency relationship between the multiple subtasks, and execute the multiple subtasks.
[0173] The computing resources allocated to the multiple subtasks mentioned above refer to allocating the multiple subtasks to different threads for execution, and the different threads process the subtasks in parallel.
[0174] During the execution of subtasks, based on the relationship between subtasks obtained from the above analysis, the execution order of multiple subtasks is determined as follows: subtasks on the same layer are independent of each other and can run in parallel, while subtasks on different layers have dependencies and need to be run layer by layer.
[0175] To facilitate understanding, Figure 14 intuitively illustrates the dependency relationship between multiple subtasks. For example, for the 27 subtasks shown in Figure 14, subtasks 0, 1, 2, 4, 5, 6, 7, 11, 8, 13, 14, 15, 16, 20, 17, 23, 25, and 26 can be assigned to thread 1, and subtasks 3, 9, 10, 12, 18, 19, 21, 22, and 24 can be assigned to thread 2. In this way, threads 1 and 2 can process subtasks in parallel, while subtasks can be processed serially within a single thread.
[0176] When the grid computing task is to solve a sparse triangular equation system, the execution of multiple subtasks specifically takes the coefficient matrix of the sparse triangular equation as input, and multiple threads execute the multiple subtasks of the sparse triangular equation in the execution order of the multiple subtasks to obtain the solution vector of the sparse triangular equation system.
[0177] For a sparse triangular system of equations Wx=b, the coefficient matrix W and the right-hand side vector b are known input information. The process of executing multiple subtasks of the sparse triangular system of equations in the order in which they are executed, and obtaining the solution vector for the sparse triangular system of equations, specifically includes: processing the subtasks in parallel by each thread, and processing the subtasks serially within the same thread. The coefficient matrix W can be obtained based on three one-dimensional arrays of the coefficient matrix in CSR format.
[0178] It should be understood that when a thread executes a corresponding subtask, it solves the corresponding components of the solution vector of the triangular equations based on the other subtasks that the subtask depends on, the non-zero elements of the row corresponding to the subtask in the coefficient matrix, and the corresponding components in the right-hand vector. For example, referring to the dependency diagram shown in FIG14 above, when a thread executes subtask 13, the data required to execute subtask 13 are the 1st, 3rd, 4th, 5th, 7th, 9th, 10th, 11th, and 12th components in the solution vector, as well as the b in the right-hand vector. 13 .
[0179] The following describes the processing of a subtask by taking a thread as an example. As shown in FIG15 , the process of a thread (thread 1) executing a subtask includes the following S1501 - S1507 .
[0180] S1501: Determine whether the current subtask depends on other subtasks.
[0181] When thread 1 executes the current subtask, it first determines whether the current subtask depends on other subtasks based on the dependency relationship between the subtasks obtained by analysis. If the current subtask depends on other subtasks, it executes S1502; if the current subtask does not depend on other subtasks, it executes S1505 and subsequent steps.
[0182] For example, if the current subtask is subtask 0 in FIG. 14 , and the dependency relationships determine that subtask 0 does not depend on other subtasks, then S1505 is executed to execute the subtask to obtain the corresponding component x0 in the solution vector. For another example, if the current subtask is subtask 4 in FIG. 14 , and it is known that subtask 4 depends on subtasks 2 and 3, then S1502 is continued.
[0183] S1502: Determine whether there are any unfinished subtasks among the subtasks that the current subtask depends on.
[0184] Since the current subtask depends on other subtasks, the execution of this subtask requires the data obtained after the other subtasks have finished executing. Therefore, it is necessary to determine whether the subtasks on which the current subtask depends have been completed. If there are no unfinished subtasks among the subtasks on which the current subtask depends (i.e., all the subtasks on which the current subtask depends have been completed), then S1505 is executed to execute the current subtask; if there are unfinished subtasks among the subtasks on which the current subtask depends, then S1503-S1504 are executed.
[0185] Optionally, a method for determining whether the subtask on which the current subtask depends is completed is that thread 1 continuously reads the subtask completion flag of the dependent subtask, where the subtask completion flag indicates whether the subtask is completed.
[0186] S1503: Read the completion flag of the subtask that the current subtask depends on.
[0187] Exemplarily, the completion flag of a subtask may be 0 or 1, where 1 indicates that the subtask has been completed and 0 indicates that the subtask has not been completed.
[0188] S1504: Determine whether the dependent subtask is completed based on the read completion flag of the subtask.
[0189] If the completion flag of the subtask on which the current subtask depends indicates that the subtask is not completed, thread 1 returns to S1503, continues to wait and reads the completion flag of the subtask; if the completion flag of the subtask on which the current subtask depends indicates that the subtask is completed, it is determined that the dependent subtask has been executed.
[0190] It should be noted that thread 1 needs to determine whether each subtask among the subtasks on which the current subtask depends has been completed, until there are no unfinished subtasks among all the subtasks on which the current subtask depends (that is, all have been completed), and then execute S1505 to solve the components of the solution vector corresponding to the subtask.
[0191] S1505: Execute the current subtask to obtain the components corresponding to the solution vector.
[0192] S1506: Modify the completion flag of the current subtask.
[0193] After thread 1 completes a subtask, it needs to modify the completion flag of the subtask so that when other subtasks are executed subsequently, if other subtasks depend on the subtask, thread 1 can determine whether the subtask is completed based on the subtask completion flag (that is, the process described in S1502 above).
[0194] After thread 1 completes the current subtask, thread 1 continues to execute S1507 to determine whether thread 1 has ended.
[0195] S1507: Determine whether the layer where the current subtask is located is the last layer of the thread.
[0196] It will be appreciated that in this embodiment of the present application, the dependencies between subtasks are hierarchical, and threads execute subtasks in order of layers. If the layer where the current subtask is located is the last layer in the dependency relationship, thread 1 terminates. If the layer where the current subtask is located is not the last layer in the thread, thread 1 continues to loop through steps S1501-S1507 to execute the subtasks in the next layer.
[0197] Each thread executes the corresponding subtask according to the above processing process, and the solution vector of the sparse triangular equation group can be obtained after all threads have completed the execution.
[0198] Typically, when storing a coefficient matrix in memory, the non-zero elements of each row are stored sequentially according to the order of the coefficient matrix rows. When a thread runs a subtask, it needs to read the non-zero elements of the row corresponding to the subtask in the coefficient matrix from memory (for example, to execute subtask 13, it needs to read the non-zero elements of the 13th row of the coefficient matrix). However, for a thread, the sequence numbers of the subtasks it executes may be discontinuous. Accordingly, the thread needs to jump from memory to read the non-zero elements of the row corresponding to the coefficient matrix. This will result in the thread being unable to access memory continuously during the execution of the subtask (i.e., discontinuous memory access), resulting in a waste of memory bandwidth.
[0199] Continuing with the example shown in Figure 14 above, thread 1 executes each subtask in the order of subtask sequence numbers 0, 1, 2, 4, 5, 6, 7, 11, 8, 13, 14, 15, 16, 20, 17, 23, 25, and 26. It can be seen that when thread 1 executes subtask 0, it reads the non-zero elements in row 0 of the coefficient matrix. After executing subtask 0, thread 1 executes subtask 1 and reads the non-zero elements in row 1 of the coefficient matrix. After executing subtask 1, thread 1 executes subtask 2 and reads the non-zero elements in row 2 of the coefficient matrix. It can be seen that the subtask sequence numbers are continuous during the execution of subtasks 0, 1, and 2, so there is no problem of discontinuous memory access. However, after executing subtask 2, thread 1 needs to execute subtask 4 and read the non-zero elements in row 4 of the coefficient matrix. In this case, thread 1 needs to skip the non-zero elements in row 3 of the coefficient matrix in memory and read the non-zero elements in row 4, resulting in discontinuous memory access.
[0200] Optionally, in response to discontinuous memory access, in an embodiment of the present application, after analyzing the dependencies between multiple subtasks, the task processing method provided in an embodiment of the present application further includes: adjusting the storage location of the task data for subtasks that have no dependencies between them, so that the storage location of the data of tasks that can be executed in parallel is continuous. For example, for a first subtask and a second subtask that are executed in parallel and have no dependencies between them, the storage location of the data of the first subtask or the data of the second subtask is adjusted so that the storage location of the data of the first subtask and the data of the second subtask are continuous.
[0201] For example, taking subtask 7 and subtask 11 executed by thread 1 in Figure 14 as an example, there is no dependency between subtask 7 and subtask 11, and thread 1 can execute subtask 7 and subtask 11 in parallel, then the position of the data in the 7th row and / or 11th row, and the 7th column and / or 11th column in the sparse matrix is adjusted so that the storage positions of the data of subtask 7 and subtask 11 are continuous.
[0202] In one implementation, for solving a sparse triangular system of equations, the data of the subtask is the data in the coefficient matrix of the sparse triangular system of equations. The method for adjusting the storage location of the task data may include: rearranging the coefficient matrix of the sparse triangular system of equations based on the execution order of multiple subtasks and the computing resources allocated to the multiple subtasks, so that memory access is continuous during the execution of the subtasks by the threads, fully utilizing the memory bandwidth. Specifically, the method for rearranging the coefficient matrix includes: for a thread, rearranging both the rows and columns of the coefficient matrix according to the subtask sequence number, and different threads rearranging the corresponding rows and columns in the coefficient matrix according to the thread sequence number.
[0203] The following example illustrates the rearrangement process of the coefficient matrix. For the example shown in Figure 14 above, subtasks 0, 1, 2, 4, 5, 6, 7, 11, 8, 13, 14, 15, 16, 20, 17, 23, 25, 26 correspond to thread 1, and subtasks 3, 9, 10, 12, 18, 19, 21, 22, 24 correspond to thread 2. Therefore, the rows of the coefficient matrix are arranged in the order of row index 0, 1, 2, 4, 5, 6, 7, 11, 8, 13, 14, 15, 16, 20, 17, 23, 25, 26, 3, 9, 10, 12, 18, 19, 21, 22, 24.
[0204] The columns of the coefficient matrix also need to be rearranged in the order of the rows.
[0205] For example, the original 4th row in the coefficient matrix is rearranged to the 3rd row, and accordingly, the 4th column in the coefficient matrix is rearranged to the 3rd column; the original 5th row in the coefficient matrix is rearranged to the 4th row, and accordingly, the 5th column in the coefficient matrix is rearranged to the 4th column; the original 6th row in the coefficient matrix is rearranged to the 5th row, and accordingly, the 6th column in the coefficient matrix is rearranged to the 5th column; the columns of the coefficient matrix are rearranged according to this rule.
[0206] It should be noted that in the embodiment of the present application, the rearrangement of the coefficient matrix is to rearrange the rows and columns of the coefficient matrix accordingly, which is only for optimizing the continuity of memory access and will not change the dependency relationship between subtasks and the execution order of subtasks, so it will not affect the solution result.
[0207] It is understandable that during the execution of the subtask, it is necessary to access the components of the corresponding positions of the solution vector according to the column index of the non-zero element of the coefficient matrix. For example, for the component x 13 , component x 13 Depends on the components x1, x3, x4, x5, x7, x9, x 10 , x 11 , x 12 , solve for the component x 13 When you need to read the components it depends on, the positions of these components are based on the component x 13 The column indices (1, 3, 4, 5, 7, 9, 10, 11, 12) of the non-zero entries corresponding to the 13th row in the coefficient matrix are determined.
[0208] In the embodiment of the present application, taking the structure grid of the 3D19 template as an example, the structured storage method of the coefficient matrix is A[nx, ny, nz, 0:9], which is a storage method for storing a structured coefficient matrix (which can be called a structured matrix). x , n y , n z is the size information of the structural grid, n x =3,n y =3,n z =3, the size of the structured matrix is 27×10, and 0:9 means that the column index of the structured matrix is from 0 to 9, with 10 elements.
[0209] The above structured matrix is related to the structured grid of the grid computing task, so it can be called a structured matrix. Corresponding to the structured grid, the number of rows of the structured matrix 27 corresponds to the 27 points in the structured grid, and the number of columns of the structured matrix 10 corresponds to the number of points in a point association group in the structured grid.
[0210] Referring to Figure 16, taking the pth row of the structured matrix as an example, the pth row corresponds to the pth point in the structured grid (the value of p is related to the coordinates of the pth point in the structured grid, see the above embodiment). According to the characteristics of the structured grid, each point on a non-critical surface depends on 9 other points. Therefore, of the 10 elements in the pth row, the element with column index 9 corresponds to the pth point, and the elements with column indexes 0-8 correspond to the other points on the structured grid that the pth point depends on. Assuming that the coordinates of the pth point are (i, j, k), the coordinates of the points in the structured grid corresponding to the elements with column indexes 0-8 are: (i, j-1, k-1), (i-1, j, k-1), (i, j, k-1), (i+1, j, k-1), (i, j+1, k-1), (i-1, j-1, k), (i, j-1, k), (i+1, j-1, k), (i-1, j, k). The relationship between the coordinates reflects the positional relationship between the pth point and the 9 points on which it depends, that is, the positional relationship obtained in the above embodiment. The 9 points are located in the front and lower part, lower left, directly below, lower right, lower back, left front, directly in front, right front, and left front of point p, respectively.
[0211] Since the coordinates of a point and the index of the point (i.e. the index of the row) have a mathematical relationship (p = i + j × n x +k×n x ×n y ), based on the coordinates of the points corresponding to the p-th row, the column indices of the non-zero elements in the p-th row can be calculated. Taking the 13th row as an example, the coordinates (i, j, k) of point 13 in the structured grid are (1, 1, 1). Based on the coordinates of the points in the 13th row, the positions of the non-zero elements in the 13th row of the coefficient matrix (i.e., the column indices of the non-zero elements) can be calculated, which are 1, 3, 4, 5, 7, 9, 10, 11, 12, and 13 respectively.
[0212] Figure 17 provides a simple schematic diagram of the structure of the structured matrix. The elements in the shaded position in the matrix indicate that non-zero elements are stored. The column number of the non-zero element is calculated through the coordinates of the point. The numerical value indicated by the shaded position in the figure represents the column number of the non-zero element. The value of the non-zero element is stored at the position of the non-zero element in the structured matrix.
[0213] As described above, the column coordinates of the non-zero elements in the coefficient matrix can be directly calculated based on the point coordinates in the structural grid corresponding to the elements in the coefficient matrix, and the relevant components are read based on the column coordinates. Compared with the CSR format, there is no need for secondary indexing in the process of accessing the solution vector (that is, first read the column index from the array d, and then read the component based on the column index), which can reduce the amount of memory access, speed up the memory access speed, and improve the calculation memory access ratio.
[0214] To summarize the above, referring to Figure 18, the solution process of the sparse triangular equations can include a dependency analysis stage, a data preparation stage, and a calculation stage, wherein the dependency analysis stage refers to analyzing the dependency between multiple subtasks based on the structural grid of the sparse triangular equations to obtain the dependency of multiple subtasks, the data preparation stage refers to the process of adjusting the position of the data in the coefficient matrix (or structured coefficient matrix) according to the analyzed dependency, and the calculation stage refers to the process of solving the sparse triangular equations based on the dependency between the subtasks and the coefficient matrix after position adjustment and other input information to obtain the solution vector.
[0215] Optionally, the idea of solving the sparse triangular equations can also be used to solve the matrix decomposition problem in the ILU algorithm. According to the description of the above embodiment, in the ILU algorithm, a matrix is decomposed into an upper triangular matrix and a lower triangular matrix. For example, the sparse matrix A is decomposed, LU=A. It can be seen that U or L can be regarded as the solution vector, and A can be regarded as the right-hand vector. The process of solving U or L is actually the process of solving the sparse triangular equations. Therefore, the method provided in the embodiment of the present application can also be used for ILU decomposition.
[0216] In the task processing method provided in the embodiment of the present application, since a vertex in the structural grid of the grid computing task corresponds to a subtask of the grid computing task, and the positional relationship between all vertices in the structural grid and the vertices on which they depend is the same, after analyzing the subtasks on which a subtask depends based on the structural grid, the other subtasks on which multiple subtasks depend can be quickly analyzed according to the same positional relationship, that is, the dependency relationship between multiple subtasks can be quickly analyzed through the structural grid of the grid computing task. In this way, the grid computing task can be processed quickly and the efficiency of business processing can be improved.
[0217] Furthermore, in the task processing method provided in the embodiment of the present application, there is no need to analyze the arrays stored in the CSR format one by one, but the dependencies between subtasks are analyzed based on the structural grid, which can significantly reduce the amount of memory access and improve the computational memory access ratio.
[0218] It is understandable that the above method is performed by a computing device, which includes a hardware structure and / or software module for performing each function in order to implement the above functions. Those skilled in the art should easily appreciate that, in combination with the method steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is performed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0219] The embodiment of the present application can divide the functional modules of the above-mentioned computing device according to the above-mentioned method example. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation.
[0220] In the case of dividing each functional module according to each function, FIG19 shows a possible structural diagram of the computing device involved in the above embodiment. The computing device includes an acquisition module 1901 , an analysis module 1902 , and a processing module 1903 .
[0221] Among them, the acquisition module 1901 is used to execute S601 in the above method embodiment; the analysis module 1902 is used to execute S602 (including S6021-S6022), S603, and S1501 in the above method embodiment; and the processing module 1903 is used to execute S604, S1505, and S1506 in the above method embodiment.
[0222] Optionally, the computing device provided in the embodiment of the present application further includes a clipping module 1904 and an adjustment module 1905. The clipping module 1904 is configured to execute S605 in the above method embodiment; the adjustment module 1905 is configured to adjust the storage location of the data of the first subtask or the data of the second subtask, for a first subtask and a second subtask that have no dependency relationship.
[0223] The various modules of the above-mentioned computing device can also be used to perform other actions in the above-mentioned method embodiment. All relevant contents of each step involved in the above-mentioned method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.
[0224] In the case of adopting an integrated unit, Figure 20 shows another possible structural diagram of the computing device involved in the above embodiment. As shown in Figure 20, the computing device provided in the embodiment of the present application may include: a processing module 2001 and a communication module 2002. The processing module 2001 can be used to control and manage the actions of the computing device. For example, the processing module 2001 can be used to support the acquisition module 1901, analysis module 1902, processing module 1903, cropping module 1904 and adjustment module 1905 in Figure 19 above to perform corresponding steps, and / or other processes for the technology described herein. The communication module 2002 can be used to support communication between the communication device and other network entities. As shown in Figure 20, the communication device may also include a storage module 2003 for storing computer instructions and data (such as coefficient matrices).
[0225] Among them, the processing module 2001 can be a processor or a controller (for example, the processing module 2001 can be the processor 501 in Figure 5). The above-mentioned processing module 2001 can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc. The communication module 2002 can be a communication interface (for example, the communication module 2002 can be the communication interface 503 in Figure 5). The storage module 2003 can be a memory (for example, the storage module 2003 can be the memory 502 in Figure 5). When the processing module 2001 is a processor, the communication module 2002 is a communication interface, and the storage module 2003 is a memory, the processor, transceiver, and memory can be connected via a bus.
[0226] For more details on how the modules included in the computing device implement the above functions, please refer to the descriptions in the previous method embodiments, which will not be repeated here. The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments.
[0227] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions in accordance with the embodiments of the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a magnetic disk, a magnetic tape), an optical medium (eg, a digital video disc (DVD)), or a semiconductor medium (eg, a solid state drive (SSD)).
[0228] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0229] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0230] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0231] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0232] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.
[0233] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A task processing method, characterized in that: include: Obtain a grid computing task in an application, wherein the grid computing task includes multiple subtasks, and the dependency relationship between the multiple subtasks complies with the dependency rule; Analyzing the dependency relationship between the first subtask and other subtasks in the plurality of subtasks, and determining the dependency rule; Analyze the dependency relationship between other subtasks in the multiple subtasks according to the dependency rule; Allocate computing resources to the multiple subtasks according to the dependencies between the multiple subtasks, and execute the multiple subtasks.
2. The method according to claim 1, characterized in that The grid computing task is to solve triangular equations using matrix multiplication.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: For each of the multiple subtasks, duplicate dependencies among dependencies between the subtask and other subtasks are deleted to trim the dependencies between the multiple subtasks.
4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: For a first subtask and a second subtask that are executed in parallel and have no dependency relationship, the storage location of the data of the first subtask or the data of the second subtask is adjusted to make the storage locations of the data of the first subtask and the data of the second subtask continuous.
5. The method according to any one of claims 1 to 4, characterized in that: The grid of the grid computing task is a structural grid, and the multiple vertices of the structural grid correspond one-to-one to the multiple subtasks; The analyzing the dependency relationship between the first subtask and other subtasks in the plurality of subtasks and determining the dependency rule includes: Analyzing at least one other vertex on which the first vertex corresponding to the first subtask depends; The dependency rule is determined according to a positional relationship between the first vertex and the at least one other vertex.
6. The method according to claim 5, characterized in that The structured grid is a three-dimensional structured grid, and the relationship between the index number p of the subtask of the grid computing task and the coordinates (i, j, k) of the vertices of the structured grid satisfies: p=i+j×n x +k×n x ×n y Among them, n x Represents the number of grid cells in the x-axis direction of the structural grid, n y Represents the number of grid cells in the structure grid in the y-axis direction.
7. A task processing device, characterized in that: include: Acquisition module, analysis module and processing module; The acquisition module is used to acquire a grid computing task in an application, wherein the grid computing task includes multiple subtasks, and the dependency relationship between the multiple subtasks complies with the dependency rule; The analysis module is used to analyze the dependency relationship between the first subtask and other subtasks in the multiple subtasks, and determine the dependency rule; The analysis module is further used to analyze the dependency relationship between other subtasks in the multiple subtasks according to the dependency rule; The processing module is used to allocate computing resources to the multiple subtasks according to the dependency relationships between the multiple subtasks, and execute the multiple subtasks.
8. The task processing device according to claim 7, characterized in that: The grid computing task is to solve triangular equations using matrix multiplication.
9. The task processing device according to claim 7 or 8, characterized in that: Also includes a cropping module; The trimming module is used for, for each of the multiple subtasks, deleting duplicate dependencies between the subtask and other subtasks, so as to trim the dependencies between the multiple subtasks.
10. The task processing device according to any one of claims 7 to 9, characterized in that: Also includes adjustment modules; The adjustment module is used to adjust the storage location of the data of the first subtask or the data of the second subtask, for a first subtask and a second subtask that are executed in parallel and have no dependency relationship, so that the storage locations of the data of the first subtask and the data of the second subtask are continuous.
11. The task processing device according to any one of claims 7 to 9, characterized in that: The grid of the grid computing task is a structural grid. The analysis module is specifically configured to analyze at least one other vertex on which the first vertex corresponding to the first subtask depends; and determine the dependency rule according to a positional relationship between the first vertex and the at least one other vertex.
12. The task processing device according to claim 11, characterized in that: The structured grid is a three-dimensional structured grid, and the relationship between the index number p of the subtask of the grid computing task and the coordinates (i, j, k) of the vertices of the structured grid satisfies: p=i+j×n x +k×n x ×n y Among them, n x Represents the number of grid cells in the x-axis direction of the structural grid, n y Represents the number of grid cells in the structure grid in the y-axis direction.
13. A computing device, characterized in that: The method comprises a memory and at least one processor connected to the memory, wherein the memory is used to store computer program code, and the computer program code comprises computer instructions. When the computer instructions are executed by the at least one processor, the processor executes the method according to any one of claims 1 to 6.
14. A computer-readable storage medium, characterized in that: Computer instructions are stored, and when the computer instructions are run on a computer, the method according to any one of claims 1 to 6 is executed.
Citation Information
Patent Citations
Task processing method and device
CN120123620A
Prediction cost optimization method for static analysis performance of scientific calculation program
CN105224452A
Lower trigonometric equation parallel solving method for structural grid sparse matrix
CN111079078A
High-throughput calculation method and system based on container technology
CN111897622A
Parallel computing method for directly solving structured triangular sparse linear equation set
CN114385972A
Cited By
Aircraft Cartesian grid data parallel processing method, device and equipment based on reordering and self-scheduling and storage medium
CN121300959A