Task processing method and device

By analyzing and utilizing the subtask dependencies in grid computing tasks and allocating computing resources, the problem of low task processing efficiency in solving sparse triangle equations in high-performance computing is solved, and more efficient task processing is achieved.

CN120123620APending Publication Date: 2025-06-10HUAWEI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311680958.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the field of high-performance computing, in the prior art, task processing efficiency is low in the process of solving sparse triangular equation systems, mainly due to the large amount of access and low calculation efficiency.

Method used

By analyzing the dependencies between subtasks in grid computing tasks, determining dependency rules, and allocating computing resources based on these rules to execute multiple subtasks, thereby improving the efficiency of task processing.

Benefits of technology

This method can quickly process grid computing tasks, improve task processing efficiency, reduce the stock of access, and improve the calculation and storage ratio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123620A_ABST
    Figure CN120123620A_ABST
Patent Text Reader

Abstract

The invention provides a task processing method and device, relates to the technical field of computers, and aims to quickly process a grid computing task and improve the task processing efficiency. The method comprises the steps that a grid computing task in an application is obtained, the grid computing task comprises a plurality of subtasks, and the dependency relationship among the subtasks conforms to a dependency rule; analyzing a dependency relationship between a first sub-task in the plurality of sub-tasks and other sub-tasks, and determining a dependency rule; the dependency relationship among other sub-tasks in the plurality of sub-tasks is analyzed according to a dependency rule; and allocating computing resources to the plurality of sub-tasks according to the dependency relationship among the plurality of sub-tasks, and executing the plurality of sub-tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a task processing method and apparatus. Background Art

[0002] In the field of high performance computing (HPC), the computing tasks in most applications are transformed into the solution of partial differential equations after mathematical and physical modeling.

[0003] The solution idea of partial differential equations is as follows: numerically discretize the partial differential equations to transform the problem into solving a large-scale sparse linear equation system. Among them, one method for solving the sparse linear equation system is the iterative method. The iterative method is used to solve the sparse linear equation system, and the solution problem of the sparse linear equation system includes the solution problem of sparse triangular equations (sparse triangular solve, SpTRSV).

[0004] Currently, for this computing task of solving sparse triangular equations, the memory access volume during the task processing is large, and the task processing efficiency is low. Summary of the Invention

[0005] This application provides a task processing method and apparatus, which can quickly process grid computing tasks and improve the task processing efficiency.

[0006] This application adopts the following technical solutions:

[0007] In a first aspect, this application provides a task processing method, including: obtaining a grid computing task in an application, where the grid computing task includes multiple subtasks, and the dependency relationship between the multiple subtasks conforms to a dependency rule; analyzing the dependency relationship between a first subtask and other subtasks among the multiple subtasks, and determining the dependency rule; and analyzing the dependency relationship between other subtasks among the multiple subtasks according to the dependency rule; and allocating computing resources to the multiple subtasks according to the dependency relationship between the multiple subtasks, and executing the multiple subtasks.

[0008] In this application, since the dependency relationship between the multiple subtasks of the grid computing task conforms to the dependency rule, by analyzing the dependency relationship between one subtask and other subtasks among the multiple subtasks and obtaining the dependency rule, and then based on the dependency rule, the dependency relationship between other subtasks can be quickly analyzed, and then according to the dependency relationship between the multiple subtasks, computing resources are allocated to the multiple subtasks and the multiple subtasks are executed. In this way, the grid computing task can be quickly processed and the task processing efficiency can be improved.

[0009] In a possible implementation manner, the above grid computing task is to solve triangular equations using matrix multiplication.

[0010] In one possible implementation, the method for task processing provided by this application further includes: for each of the multiple subtasks, deleting the duplicate dependencies in the dependencies between the subtasks to trim the dependencies between the multiple subtasks. Trimming the redundant dependencies can simplify the dependencies and contribute to subsequent resource allocation.

[0011] In one possible implementation, the method provided by this application further includes: for a first subtask and a second subtask that are executed in parallel and have no dependencies between them, adjusting the storage location of the data of the first subtask or the data of the second subtask so that the storage locations of the data of the first subtask and the data of the second subtask are consecutive, thereby saving memory bandwidth.

[0012] In one possible implementation, the grid of the above grid computing task is a structured grid, and multiple vertices of the structured grid correspond to multiple subtasks one by one. Then, analyze the dependencies between a first subtask among the multiple subtasks and other subtasks, and determine the dependency rules, including: analyzing at least one other vertex on which the first vertex corresponding to the first subtask depends; and determining the dependency rules according to the positional relationship between the first vertex and the at least one other vertex.

[0013] In this application, since the points in the structured grid of the grid computing task have a one-to-one correspondence with the subtasks, therefore, by analyzing the positional relationship of the points with dependencies in the structured grid to determine the dependency rules of the subtasks, it helps to quickly analyze the dependencies between multiple subtasks.

[0014] In one possible implementation, the above structured grid is a three-dimensional structured grid, and the relationship between the index number p of the subtask of the grid computing task and the coordinates (i, j, k) of the vertex of the structured grid satisfies:

[0015] p = i + j × n x + k × n x × n y

[0016] where n x represents the number of grid cells of the structured grid in the x-axis direction, and n y represents the number of grid cells of the structured grid in the y-axis direction.

[0017] In a second aspect, this application provides a computing device, which includes various modules for implementing the method described in the first aspect and one of its possible implementations, such as an acquisition module, an analysis module, a processing module, a trimming module, an adjustment module, etc.

[0018] The computing device has a function to implement the behavior in the method example of any one of the above-mentioned first aspect and its possible implementation manners. The function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above-mentioned function.

[0019] In a third aspect, the present application provides a computing device, including a memory and at least one processor connected to the memory. The memory is used to store computer program code, and the computer program code includes computer instructions. When the computer instructions are executed by the at least one processor, the computing device is caused to execute the method of any one of the first aspect and its possible implementation manners.

[0020] In a fourth aspect, the present application provides a computer-readable storage medium storing computer instructions, and when the computer instructions run on a computer, they execute the method of any one of the first aspect and its possible implementation manners.

[0021] In a fifth aspect, the present application provides a computer program product, which includes computer instructions, and when the computer instructions run on a computer, they execute the method of any one of the first aspect and its possible implementation manners.

[0022] In a sixth aspect, the present application provides a chip system, including: a processor, which is used to call and run a computer program from a memory, so that a computing device installed with the chip system executes the method of any one of the first aspect and its possible implementation manners.

[0023] It should be understood that for the beneficial effects obtained by the technical solutions of the second to sixth aspects of the present application and the corresponding possible implementation manners, reference can be made to the technical effects of the first aspect and its corresponding possible implementation manners described above, and details are not elaborated here. Description of the Drawings

[0024] Figure 1 One of the schematic diagrams of a structural grid provided by an embodiment of the present application;

[0025] Figure 2 A schematic diagram of the positions of non-zero elements of a sparse triangular matrix provided by an embodiment of the present application;

[0026] Figure 3 A schematic diagram of the positions of non-zero elements of a sparse triangular matrix and task dependency relationships provided by an embodiment of the present application;

[0027] Figure 4 A schematic diagram of the relationship between the CSR format and the positions of non-zero elements of a coefficient matrix provided by an embodiment of the present application;

[0028] Figure 5Schematic diagram of the hardware structure of a computing device provided by an embodiment of the present application;

[0029] Figure 6 One of the schematic flowcharts of the task processing method provided by an embodiment of the present application;

[0030] Figure 7 Another schematic diagram of a structured grid provided by an embodiment of the present application;

[0031] Figure 8 Schematic diagram of the coordinates of points in a structured grid provided by an embodiment of the present application;

[0032] Figure 9 Another schematic flowchart of the task processing method provided by an embodiment of the present application;

[0033] Figure 10 Schematic diagram of the relationship between the coordinates of points in a structured grid and the indexes of subtasks provided by an embodiment of the present application;

[0034] Figure 11 Another schematic flowchart of the task processing method provided by an embodiment of the present application;

[0035] Figure 12 Schematic diagram of the pruning result of the dependency relationship of a task provided by an embodiment of the present application;

[0036] Figure 13 One of the schematic diagrams of the analysis result of a dependency relationship provided by an embodiment of the present application;

[0037] Figure 14 Another schematic diagram of the analysis result of a dependency relationship provided by an embodiment of the present application;

[0038] Figure 15 Another schematic flowchart of the task processing method provided by an embodiment of the present application;

[0039] Figure 16 Schematic diagram of the row structure of a structured matrix provided by an embodiment of the present application;

[0040] Figure 17 Schematic diagram of a structured matrix provided by an embodiment of the present application;

[0041] Figure 18 Schematic diagram of the solution process of a sparse triangular system of equations provided by an embodiment of the present application;

[0042] Figure 19 One of the schematic diagrams of the structure of a computing device provided by an embodiment of the present application;

[0043] Figure 20 Another schematic diagram of the structure of a computing device provided by an embodiment of the present application. Detailed implementation manners

[0044] In this text, the term "and / or" is merely a correlative relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, these three situations.

[0045] In the description of the embodiments of this application, the terms "first", "second", etc. in the specification and claims are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first vertex and the second vertex, etc. are used to distinguish different vertices, rather than to describe a specific order of the vertices.

[0046] In the embodiments of this application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way.

[0047] In the description of the embodiments of this application, unless otherwise specified, the meaning of "a plurality" refers to two or more, and "a plurality" can also be described as "at least two".

[0048] The task processing method provided by the embodiments of this application can be applied in the field of high-performance computing (HPC). For example, for computing tasks in applications (such as meteorological applications, etc.), through this method, a large number of computing problems can be efficiently solved. In the HPC field, most of the computing problems in applications can be transformed into the problem of solving partial differential equations. One idea for solving partial differential equations is to discretize them numerically, transforming the problem into the problem of solving a sparse linear equation system. Further, many methods for solving sparse linear equation systems involve computational tasks of using matrix multiplication to solve triangular equations (i.e., solving sparse triangular equation systems), and efficiently processing such computational tasks has an important impact on the computing performance of HPC.

[0049] First, some technical terms involved in a task processing method and device provided by the embodiments of this application are explained.

[0050] 1. Grid and grid computing tasks

[0051] First, the concept of a grid is introduced. The idea of a grid stems from discretized solution. The continuous computational domain is discretized into several finite sub-regions, and the physical variables of each sub-region are solved respectively, and then the physical variables over the entire computational domain are obtained. Dividing the computational domain into multiple sub-regions can be understood as the process of grid division, and a sub-region is equivalent to a smallest grid cell.

[0052] In the embodiments of the present application, a task of performing calculations based on the idea of a grid is defined as a grid computing task. For example, for solving a system of equations, the idea of a grid can be adopted to discretize the calculation region into several finite sub-regions, that is, divided into multiple grid cells to complete the solution of the system of equations. Therefore, the task of solving the system of equations is a grid computing task. In the embodiments of the present application, the calculation region of a calculation task can be referred to as the grid of the calculation task, and the grid of the calculation task includes multiple grid cells.

[0053] Optionally, the grid can be two-dimensional, three-dimensional or multi-dimensional. A two-dimensional grid can be, for example, a quadrilateral grid, and a three-dimensional grid can be, for example, a cuboid or cube grid.

[0054] It can be understood that the grid can include structured grids and unstructured grids. The embodiments of the present application relate to structured grids, and only structured grids will be introduced below, and unstructured grids will not be introduced.

[0055] For a structured grid, the grid region is divided into multiple grid cells, and each point (referring to the vertex of the grid cell in the structured grid) within the grid region has one or more neighboring grid cells, and the neighboring grid cells are the grid cells adjacent to this point in all directions. A structured grid means that all internal points within the grid region have the same number and the same type of neighboring grid cells. That is to say, for any two points within the grid region, for example, for point A and point B, the number of neighboring grid cells of point A is the same as the number of neighboring grid cells of point B, and the types of neighboring grid cells of point A are the same as the types of neighboring grid cells of point B. It should be noted that for a point, for example, for point A, the types of the multiple neighboring grid cells of point A are also the same.

[0056] In the embodiments of the present application, the structured grid of the 3D19 template is taken as an example for illustration. Refer to Figure 1 , the multiple grid cells in the structured grid of the 3D19 template are cuboids, each vertex has 8 neighboring grid cells, and each grid cell is a cube, and the number of points in a point association group is 19.

[0057] 2. Sparse linear equations

[0058] First, the concept of a linear equation system is briefly introduced. A linear equation system can be represented in the following matrix equation form:

[0059] Ax = b

[0060] Where A represents the coefficient matrix of the linear equation system, A ∈ n×n; x is the solution vector to be solved in the linear equation system, x ∈ n×1; b is the right-hand vector of the linear equation system, b ∈ n×1.

[0061] Understandably, the process of solving a linear equation system refers to the process of obtaining the solution vector x by solving the coefficient matrix A and the right - hand vector b.

[0062] For the coefficient matrix A of the above - mentioned linear equation system, when the values of the vast majority of the elements in the coefficient matrix A are 0, that is to say, when the proportion of non - zero elements among the n elements of the coefficient matrix A is relatively low (for example, less than 0.1%), the coefficient matrix A is considered a sparse matrix. Correspondingly, when the coefficient matrix A in the above - mentioned linear equation system is a sparse matrix, this linear equation system is a sparse linear equation system. 2 3. Sparse triangular equation system

[0063] Based on the above introduction to the concept of a sparse linear equation system, a sparse triangular equation system refers to a linear equation system whose coefficient matrix is a sparse triangular matrix. The sparse triangular equation system can be expressed in the following matrix equation form:

[0064] Wx = b

[0065] where W represents the coefficient matrix of the sparse triangular equation system. W is a triangular matrix and W is a sparse triangular matrix. W ∈ n×n; x is the solution vector to be solved in the linear equation system, x ∈ n×1; b is the right - hand vector of the linear equation system, b ∈ n×1

[0066] Optionally, in the sparse triangular equation system, the coefficient matrix W can be a sparse lower triangular matrix or a sparse upper triangular matrix, which is not limited in the embodiments of the present application.

[0067] Generally, the above - mentioned methods for solving sparse linear equation systems include direct methods and iterative methods. The iterative method gradually approaches the solution of the equation by means of iterative correction. In the iterative method, the preconditioner is a very important module. The preconditioner can improve the convergence speed and handle ill - conditioned matrices.

[0068] Optionally, the iterative method can include but is not limited to the conjugate gradient method, the generalized minimum residual method, the generalized conjugate residual method, the Chebyshev method, etc. The preconditioner can include but is not limited to algorithms such as additive Schwarz, block Jacobi, Jacobi, successive over - relaxation (SOR), incomplete lower - upper factorization (ILU), and multigrid. Among them, the ILU algorithm and the SOR algorithm are the most basic preconditioners, with simple algorithm processes and good convergence. The SOR algorithm and the ILU algorithm can also be nested and used by other preconditioners.

[0069] ​

[0070] The above SOR algorithms may include algorithms such as the Gauss - Seidel algorithm, the symmetric Gauss - Seidel algorithm, and the symmetric successive over - relaxation algorithm. ILU - type algorithms include algorithms such as ILU(0), ILU(1), and ILU(k).

[0071] The following briefly introduces the core points of the SOR algorithm and the ILU algorithm.

[0072] The core of the SOR algorithm is: First, split the coefficient matrix A of the sparse linear equations into three matrices, namely a lower triangular matrix L, an upper triangular matrix U, and a diagonal matrix D, that is, A = L + U + D. Thus, the original sparse linear equations are transformed into (L + U + D)x = b. It can be seen that solving the sparse linear equations involves solving multiple sparse triangular equations (where the diagonal matrix D can be combined with the lower triangular matrix L or the upper triangular matrix U into a triangular matrix, and the diagonal matrix D can also be regarded as a triangular matrix). Second, according to different SOR algorithms, solve the sparse triangular equations.

[0073] Taking the Gauss - Seidel algorithm as an example, the iterative formula of the Gauss - Seidel algorithm is (L + D)×x(k + 1)=b - U×x(k), where x(k) is the solution vector obtained in the k - th iteration, x(k + 1) is the solution vector obtained in the (k + 1)-th iteration, and L + D is the lower triangular part of the coefficient matrix A.

[0074] The core of the ILU algorithm is: First, decompose the coefficient matrix A of the sparse linear equations into two matrices, namely a lower triangular matrix L and an upper triangular matrix U, that is, A = L×U. Thus, the original sparse linear equations are transformed into LUx = b; then, assume Ux = y, then the solution process of LUx = b involves solving two sparse triangular equations, Ly = b and Lx = y.

[0075] 4. Compressed Sparse Row (CSR) Format

[0076] The CSR format is a sparse storage format for sparse matrices. For example, for the coefficient matrix A in the above - mentioned sparse linear equations, it can be stored in the CSR format. Storing the matrix in a sparse storage format can save memory space.

[0077] It is understandable that when the proportion of non - zero elements (or non - zero entries) in the coefficient matrix A is very low, if the traditional two - dimensional array (a[n][n]) form is used to store the coefficient matrix A, a large amount of memory space will be wasted. Therefore, for sparse matrices, a sparse storage format is generally adopted, and the CSR format is a commonly used sparse storage format.

[0078] The CSR format uses three one - dimensional arrays to represent the coefficient matrix A. The three one - dimensional arrays are a, d, and c respectively. The array a contains nnz elements, the array d also contains nnz elements, and the array c contains n + 1 elements. Here, nnz represents the total number of non - zero elements in the coefficient matrix A, and n represents the number of rows of the coefficient matrix A. Among them, the array a is used to store the values of all non - zero elements in the coefficient matrix A, the array d is used to store the column index of each non - zero element in the coefficient matrix A, and the array c is used to store the position index of the first non - zero element in each row of the coefficient matrix A in the above - mentioned array a and the number of non - zero elements in the coefficient matrix A.

[0079] a[k] is the k - th element in the array a, and a[k] represents the value of the k - th non - zero element of the coefficient matrix A; d[k] is the k - th element in the array d, and d[k] represents the column index of the k - th non - zero element of the coefficient matrix A, where k = 0, 1, ……, nnz - 1. c[t] is the t - th element in the array c, and c[t] represents the position index of the first non - zero element in the t - th row of the coefficient matrix A in the array a, where t = 0, 1, ……, n - 1, and c[n] is the n - th element (i.e., the last element) in the array c, and c[n] represents the total number of non - zero elements in the coefficient matrix A.

[0080] It should be noted that in the embodiments of the present application, for arrays, matrices, and vectors, the numbering of the position indices of their elements starts from 0. For example, if a vector contains n elements, the position indices of the elements of the vector are 0 to n - 1, and the k - th element of the vector refers to the element with the index number k in the vector. This will not be elaborated one by one in the following embodiments.

[0081] Exemplarily, for the following coefficient matrix A:

[0082]

[0083] The three one - dimensional arrays corresponding to the coefficient matrix A are respectively:

[0084] a = [1, 2, 3, 6, 7, 9]; among them, a[2] = 3, indicating that the value of the 2 - nd non - zero element of the coefficient matrix A is 3.

[0085] d = [0, 2, 1, 2, 1, 3]; among them, d[5] = 3, indicating that the column index of the 5 - th non - zero element of the coefficient matrix A is 3.

[0086] c = [0, 2, 3, 4, 6]; where c[2] = 3, indicating that the index of the first non-zero element (i.e., non-zero element 6) in the second row of the coefficient matrix A in the array a is 3.

[0087] Based on the above description of the CSR format, it can be seen that the storage space of a traditional two-dimensional array requires O(n 2 ), while the CSR format only requires O(n + 2nnz). When 2nnz is much smaller than n 2 , storing the coefficient matrix A in the CSR format can significantly reduce the memory occupancy.

[0088] 5. Dependencies in Sparse Triangular Systems of Equations

[0089] According to the description of the above embodiments, in the field of HPC, most computational problems will involve the solution of sparse triangular systems of equations after transformation. In the embodiments of the present application, the task of solving the sparse triangular systems of equations involved in the application is called a computational task of the application, and the process of solving one component in the solution vector of the sparse triangular system of equations is defined as a subtask of solving the sparse triangular system of equations. That is to say, a computational task includes multiple subtasks. In the process of solving each component, the solution of this component may depend on other components.

[0090] The following uses a simple sparse triangular system of equations as an example to illustrate. For the following example of a sparse triangular system of equations:

[0091]

[0092] The above sparse triangular system of equations can also be expressed as:

[0093]

[0094] Combined with the CSR format of the coefficient matrix of this sparse triangular system of equations, the non-zero element array a = [1, 2, 1, 3, 1, 4, 1]. Thus, the above sparse triangular system of equations can also be expressed as:

[0095]

[0096] For the solution vector x in this sparse triangular system of equations, the 4 components of this solution vector are x 0 , x 1 , x 2 , x 3 , corresponding to 4 subtasks respectively, namely subtask 0, subtask 1, subtask 2, and subtask 3. Referring to Figure 2 , it shows the positions of the non-zero elements of the above coefficient matrix (sparse lower triangular matrix). Combining Figure 2 and the above sparse triangular system of equations, the following conclusions can be directly obtained:

[0097] Subtask 0 (corresponding to the solution component x 0 ) The dependency relationship between it and other subtasks is: Subtask 0 does not depend on other subtasks.

[0098] Subtask 1 (corresponding to the solution component x 1 ) The dependency relationship between it and other subtasks is: Subtask 1 depends on Subtask 0.

[0099] Subtask 2 (corresponding to the solution component x 2 ) The dependency relationship between it and other subtasks is: Subtask 2 depends on Subtask 1.

[0100] Subtask 3 (corresponding to the solution component x 3 ) The dependency relationship between it and other subtasks is: Subtask 3 depends on Subtask 2.

[0101] In summary, the following conclusion can be drawn: For subtask i, the index number of the subtask it depends on is equal to the column number of the non-zero element in the i-th row of the coefficient matrix.

[0102] For example, referring to Figure 2 the positions of the non-zero elements in the coefficient matrix, for Subtask 0, the column number of the non-zero element in the 0-th row of the coefficient matrix is 0, so it is determined that Subtask 0 does not depend on other subtasks; for Subtask 1, the column numbers of the non-zero elements in the 1-st row of the coefficient matrix are 0 and 1, so it is determined that Subtask 1 depends on Subtask 0; for Subtask 2, the column numbers of the non-zero elements in the 2-nd row of the coefficient matrix are 1 and 2, so it is determined that Subtask 2 depends on Subtask 1; for Subtask 3, the column numbers of the non-zero elements in the 3-rd row of the coefficient matrix are 2 and 3, so it is determined that Subtask 3 depends on Subtask 2.

[0103] Exemplarily, referring to Figure 3 , Figure 3 in which (a) shows the non-zero elements of the coefficient matrix (lower triangular matrix) of a sparse triangular system of equations. In the figure, the shaded positions are all the positions of the non-zero elements (the black shaded ones indicate the diagonal positions). According to Figure 3 it can be known that the sparse triangular system of equations includes 16 subtasks and the coefficient matrix includes 16 rows. According to Figure 3 in (a), the dependency relationship between the 16 subtasks can be analyzed. The schematic diagram of the dependency relationship between the 16 subtasks is Figure 3 in (b). The dependency relationship includes four layers in sequence, and the dependency relationship between the subtasks is indicated by arrows. For example, the arrows from Subtask 0 and Subtask 2 to Subtask 3 indicate that Subtask 3 depends on Subtask 0 and Subtask 2. As Figure 3As shown in (b) of [the figure], subtask 0, subtask 1, and subtask 2 do not depend on other subtasks, that is, subtask 0, subtask 1, and subtask 2 are independent of each other and are located in the first layer of the dependency graph; subtask 3, subtask 4, subtask 5, subtask 6, subtask 7, and subtask 8 are independent of each other and are located in the second layer of the dependency graph; subtask 9, subtask 10, subtask 11, subtask 12, and subtask 13 are independent of each other and are located in the third layer of the dependency graph; subtask 14 and subtask 15 are independent of each other and are located in the fourth layer of the dependency graph. Each subtask in the subsequent layer depends on some subtasks in the previous layer (at least one previous layer).

[0104] In summary, by analyzing the coefficient matrix of the sparse triangular equations, the dependency relationships between the subtasks of the sparse triangular equations can be obtained. Some of the multiple subtasks are independent of each other, and some subtasks are dependent on each other. According to the dependency relationships between the subtasks, each subtask can be solved in parallel with multiple threads, and the solution of the sparse triangular equations can be completed with high performance.

[0105] Currently, the coefficient matrix of the sparse triangular equations is usually stored in CSR format. Instead of directly storing the coefficient matrix, during the solution process of the sparse triangular equations, the dependency relationships between the subtasks are analyzed based on the three arrays stored in CSR format, and then the threads are allocated to the subtasks according to the dependency relationships. The three one-dimensional arrays in CSR format are a, d, and c. Array a is used to store the values of all non-zero elements of the coefficient matrix, array d is used to store the column indices where each non-zero element is located in the coefficient matrix, and array c is used to store the positions of the first non-zero element in each row of the coefficient matrix in array a and the number of non-zero elements in the coefficient matrix.

[0106] Exemplarily, the three arrays of the coefficient matrix of a sparse triangular equation stored in CSR format are: a = {1, 2, 1, 3, 1, 4, 1}, d = {0, 0, 1, 1, 2, 2, 3}, c = {0, 1, 3, 5, 7}. According to a, the dimension of the coefficient matrix is 4×4, and the total number of non-zero elements is 7. When analyzing the dependency relationships between the subtasks of the sparse triangular equations, each subtask needs to be analyzed one by one.

[0107] Reference Figure 4In (a) thereof, according to the array c, the position index of the first non-zero element in the 0th row of the coefficient matrix in the array a is 0, and the position index of the first non-zero element in the 1st row of the coefficient matrix in the array a is 1. It can be seen that the 0th row of the coefficient matrix includes 1 non-zero element, and the column index of reading this non-zero element from the array d is 0; further, since the position index of the first non-zero element in the 1st row of the coefficient matrix in the array a is 1, and the position index of the first non-zero element in the 2nd row of the coefficient matrix in the array a is 3, it can be seen that the 1st row of the coefficient matrix includes 2 non-zero elements, and the column indices of reading the non-zero elements from the array d are 0 and 1; further, since the position index of the first non-zero element in the 2nd row of the coefficient matrix in the array a is 3, and the position index of the first non-zero element in the 3rd row of the coefficient matrix in the array a is 5, it can be seen that the 2nd row of the coefficient matrix includes 2 non-zero elements, and the column indices of reading the non-zero elements from the array d are 1 and 2; finally, the 3rd row (the last row) of the coefficient matrix contains 2 non-zero elements, and the column indices of reading the non-zero elements from the array d are 2 and 3.

[0108] Reference Figure 4 In (b) thereof, according to the above analysis results, the positions of the non-zero elements in each row of the coefficient matrix can be obtained, so that the dependency relationships between the sub-tasks of the sparse triangular equations can be obtained. For example, sub-task 0 does not depend on other sub-tasks, sub-task 1 depends on sub-task 0, sub-task 2 depends on sub-task 1, and sub-task 3 depends on sub-task 2.

[0109] In the process of analyzing the dependency relationships between the sub-tasks through the array of the coefficient matrix of a sparse triangular matrix stored in CSR format above, for each sub-task of the coefficient matrix, when analyzing the sub-task it depends on, relevant elements need to be read from the arrays c and d. When the dimension of the coefficient matrix is relatively high, analyzing the dependency relationships sub-task by sub-task makes the solution efficiency of the sparse triangular equations relatively low. In addition, the memory access volume in the process of analyzing the dependency relationships is relatively large, and the computation-to-memory ratio is low (the computation-to-memory ratio is used to reflect the computational intensity of a program relative to memory access, and the computation-to-memory ratio can be the ratio of the number of computation operations to the number of memory access operations). It should be understood that the problem of solving the sparse triangular equations is a memory-access-limited problem. A higher computation-to-memory ratio indicates better performance of the solution method, otherwise, a lower computation-to-memory ratio indicates poorer performance.

[0110] In view of the problem of low efficiency in processing computing tasks in the prior art, an embodiment of the present application provides a task processing method for processing grid computing tasks. The task processing device obtains a grid computing task in an application. The grid computing task includes multiple subtasks, and the dependency relationships among the multiple subtasks conform to dependency rules. Then, it analyzes the dependency relationships between the first subtask and other subtasks among the multiple subtasks and determines the dependency rules. Furthermore, it analyzes the dependency relationships between other subtasks among the multiple subtasks according to the dependency rules. Then, it allocates computing resources to the multiple subtasks according to the dependency relationships among the multiple subtasks and executes the multiple subtasks. In this method, since the dependency relationships among the multiple subtasks of the grid computing task conform to the dependency rules, by analyzing the dependency relationships between one subtask and other subtasks among the multiple subtasks and obtaining the dependency rules, it is possible to quickly analyze the dependency relationships between other subtasks based on the dependency rules. Then, it allocates computing resources to the multiple subtasks according to the dependency relationships among the multiple subtasks and executes the multiple subtasks. In this way, it is possible to quickly process the grid computing tasks and improve the efficiency of task processing.

[0111] Optionally, the hardware device for executing the task processing method provided by the embodiment of the present application may be a device with processing and storage functions, such as a computing device, such as a server, a desktop computer, etc. For example, this method can be integrated into various application programs that need to process grid computing tasks, such as a third-party math library, and applied in some solver frameworks in the form of a math library, and then indirectly called by upper-layer applications. Another example is that it is called by the solver framework or upper-layer applications in the form of code.

[0112] Taking the hardware device for executing the task processing method as a computing device as an example, Figure 5 FIG. is a schematic diagram of the hardware structure of a computing device provided by an embodiment of the present application. Figure 5 The various components shown therein can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0113] As Figure 5 shown, the computing device may include: a processor 501, a memory 502, and a communication interface 503. Among them, the processor 501, the memory 502, and the communication interface 503 may be connected through a bus 504, or connected to each other in other ways.

[0114] Among them, the processor 501 is the control center of the computing device. The processor 501 may be a general-purpose central processing unit (CPU), or other general-purpose processors. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.

[0115] The controller in the processor 501 is the nerve center and command center of the computing device. The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of instruction fetching and execution. Optionally, a memory can also be provided in the processor 501 for storing instructions and data. Exemplarily, the processor 501 can include one or more CPUs, such as Figure 5 the CPU 0 and CPU 1 shown in

[0116] The memory 502 includes, but is not limited to, a random access memory (RAM), a read only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, or an optical memory, a magnetic disk storage medium, or any other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer. In the embodiments of the present application, the memory 502 can store information such as computer instructions.

[0117] In a possible implementation, the memory 502 can exist independently of the processor 501. The memory 502 can be connected to the processor 501 through the bus 504 for storing data (such as data in CSR format, etc.), instructions, or program code. When the processor 501 calls and executes the instructions or program code stored in the memory 502, the relevant steps in the method provided by the embodiments of the present application can be implemented.

[0118] In another possible implementation, the memory 502 can also be integrated with the processor 501.

[0119] The communication interface 503 can be a transceiver module for communicating with other devices or communication networks (such as communicating with Figure 5 the log management platform or node information processing platform shown in), such as communication through Ethernet, RAN, wireless local area networks (WLAN), etc. The communication interface 503 can receive instructions, messages, or data, etc. The transceiver module can be a device such as a transceiver or a transceiver. Optionally, the communication interface 503 can also be a transceiver circuit located in the processor 501 to implement the signal input and signal output of the processor. The communication interface 503 can be a wired interface (port), such as a fiber distributed data interface (FDDI), a gigabit Ethernet (GE) interface, or, the communication interface 503 can also be a wireless interface.

[0120] The bus 504 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus can also be divided into a serial bus and a parallel bus. For the sake of convenience of representation, Figure 5 it is only represented by a thick line in the figure, but it does not mean that there is only one bus or one type of bus.

[0121] Optionally, the computing device in the embodiments of the present application may further include an input / output interface 505. The input / output interface 505 is used to connect to an input device and receive information input by the user through the input device (such as a coefficient matrix, a right-hand vector, template information of a structured grid, etc.). The input device includes but is not limited to a keyboard, a touch screen, a microphone, etc. The input / output interface 505 is also used to connect to an output device to output the processing result of the processor.

[0122] It should be noted that Figure 5 the shown computing device is only an example of a computing device, and this computing device may have more or fewer components than Figure 5 those shown in the figure, two or more components may be combined, or different component configurations may be had.

[0123] Combined with the above content, the task processing method provided by the embodiments of the present application will be described in detail below. As Figure 6 shown, this method includes S601 - S604.

[0124] S601. Obtain a grid computing task in an application.

[0125] The grid computing task includes multiple subtasks, and the dependency relationships between the multiple subtasks conform to the dependency rules.

[0126] In the embodiments of the present application, the grid of the grid computing task is a structured grid, and multiple vertices of the structured grid correspond one-to-one to multiple subtasks of the grid computing task. It should be understood that based on the structural characteristics of the structured grid described in the above embodiments, for any point (the vertex of a grid cell) in the structured grid of the grid computing task, hereinafter referred to as the target point, the target point has one or more associated points, and the associated points of the target point are the points having an associated relationship with the target point. In the embodiments of the present application, a point association group can be formed by the target point and its associated points.

[0127] It should be noted that, except for the points on the boundary surface of the structured grid, when other points of the structured grid are used as target points, they all have the same number of associated points, and the positional relationship between the associated points and the target points is also the same.

[0128] Since multiple vertices of the structured grid correspond one-to-one to multiple subtasks of the grid computing task, there is also a dependency relationship between the multiple subtasks. The dependency relationship between the subtasks can be obtained through the positional relationship between the points. For example, for a subtask corresponding to a target point, the subtasks it depends on are the subtasks corresponding to the associated points of the target point, and for any one of the multiple subtasks, the subtasks it depends on can be determined according to the same positional relationship. It can be seen that for the multiple subtasks of the grid computing task, the dependency relationship between the multiple subtasks conforms to the dependency rule, and this dependency rule can be understood as the above-mentioned positional relationship between the target point and the associated point. The dependency rule can also be called a dependency relationship template.

[0129] If the grid computing task is a task of solving equations using matrix multiplication (i.e., solving a system of equations), then the total number of points in a point association group in the structured grid of this grid computing task is equal to the total number of non-zero elements in the target row (the target row is the row corresponding to the target point) of the coefficient matrix of the system of equations.

[0130] Optionally, the structured grid includes different templates. The number and positions of the associated points of the target points in the structured grids of different templates are different. Corresponding to the coefficient matrix of the system of equations, the number of non-zero elements in each row of the coefficient matrix is different.

[0131] Exemplarily, if the grid computing task is a task of solving the sparse linear equation system Ax = b, if the size of the coefficient matrix A is 27×27, the size of the solution vector x is 27×1, and the size of the right-hand vector b is 27×1. Taking the structured grid of this sparse linear equation as the structured grid of the 3D19 template as an example, the size of this structured grid is 3×3×3, that is, the structured grid is a cube, and it includes 2 grid cells in each of the three directions (i.e., the x, y, and z directions). Each grid cell is also a cube. Refer to Figure 7 , which shows the structured grid of a sparse linear equation system. This structured grid includes 8 grid cells and 27 vertices. A target point (a point that is not on the boundary surface) in the structured grid has 18 associated points. Therefore, corresponding to the sparse linear equation system, the corresponding row in the coefficient matrix of this sparse linear equation system has 19 non-zero elements.

[0132] For the sake of convenience of description, as Figure 7As shown, the 27 vertices in the structural grid are numbered from 0 to 26. The 27 points can be respectively denoted as point 0, point 1, ……, point 26. Taking point 13 in the figure as an example, the point association group with point 13 as the target point includes 19 points, and the numbers of the points are {0, 1, 3, 4, 5, 9, 10, 11, 12, 13, 14, 15, 16, 17, 19, 21, 22, 23, 25}.

[0133] For this sparse linear equation system, after splitting or decomposing the coefficient matrix into a lower triangular matrix and an upper triangular matrix, the solution of this sparse linear equation system includes the solution of a sparse lower triangular equation system and a sparse upper triangular equation system. For the 13th row in the coefficient matrix of the sparse linear equation system, among the 18 points other than point 13 in the point association group, some points correspond to the associated points of point 13 (or the points that point 13 depends on) in the sparse lower triangular equation system, and another part of the 18 points correspond to the associated points of point 13 (or the points that point 13 depends on) in the sparse upper triangular equation system. Continuing to refer to Figure 1 , for the sparse lower triangular equation system corresponding to the sparse linear equation system, point 13 in the structural grid depends on 9 points, and the numbers of the 9 points are {1, 3, 4, 5, 7, 9, 10, 11, 12}; for the sparse upper triangular equation system corresponding to the sparse linear equation system, point 13 in the structural grid depends on 9 points, and the numbers of the 9 points are {14, 15, 16, 17, 19, 21, 22, 23, 25}.

[0134] In one implementation, the grid computing task in the embodiments of the present application is to solve a triangular equation using matrix multiplication, that is, the task of solving a triangular equation system, and the triangular equation system can be a sparse triangular equation system. Solving one component of the solution vector of the triangular equation system is called a subtask of the triangular equation system. For the convenience of description, in the following embodiments, the solution of the sparse triangular equation system is taken as an example for illustration.

[0135] If the grid computing task is to solve a sparse triangular equation system, the above-mentioned obtaining of a grid computing task of an application includes obtaining the structural grid information of the sparse triangular equation system and the coefficient matrix of the sparse triangular equation system.

[0136] In the embodiments of the present application, the sparse triangular equation system can be expressed as Wx = b, where W is the coefficient matrix and W is a triangular matrix, W ∈ n×n; x is the solution vector, x ∈ n×1; b is the right-end vector, b ∈ n×1. Taking the solving process of one component of the solution vector x as a subtask of solving the sparse triangular equation system, it can be known that the sparse triangular equation system includes n subtasks, and the solving process of the sparse triangular equation system is to solve n subtasks.

[0137] Optionally, the coefficient matrix W can be an upper triangular matrix or a lower triangular matrix, which is not limited in the embodiments of the present application.

[0138] Optionally, the structured grid information may include template information of the structured grid, size information of the structured grid, and the like.

[0139] The structured grid information is used to indicate the distribution of multiple subtasks of a grid computing task in the structured grid of the grid computing task. For the structured grid of a sparse triangular equation set, one vertex in the structured grid corresponds to one subtask. Specifically, if one vertex in the structured grid corresponds to one subtask, then the index of one vertex in the structured grid is the same as the row index of the corresponding component in the solution vector. For example, vertex 0 in the structured grid corresponds to component 0 (i.e., subtask 0) of the solution vector.

[0140] In the embodiments of the present application, in the coordinate system of the structured grid of the sparse triangular equation set, the relationship between the coordinates of each vertex and the index of the vertex is defined. For example, for the component x of the solution vector x of the sparse triangular equation set p The corresponding subtask index number is p. There is a corresponding relationship between the subtask index number p and the coordinates (i, j, k) of the vertex of the structured grid, where p = 0, 2, ……, n - 1. Thus, according to the corresponding relationship between the index of the subtask and the coordinates of the vertex of the structured grid, the points in the structured grid can be corresponding to the subtasks, and then the dependency relationship between multiple subtasks of the grid computing task can be quickly analyzed by analyzing the dependency relationship between the points.

[0141] If the structured grid is a three-dimensional structured grid, such as Figure 8 shown, the above structured grid is a cube (including 6 faces, the front face, the back face, the left face, the right face, the upper face, and the lower face). For the structured grid of the cube, a coordinate system of the structured grid is defined. Among them, the horizontal rightward direction is defined as the positive direction of the x-axis, the vertically upward direction is defined as the positive direction of the z-axis, and the direction perpendicular to the x-axis and the z-axis and pointing to the back face of the cube is defined as the positive direction of the y-axis.

[0142] In one implementation, the size information of the structured grid includes: the number of grid cells n of the structured grid in the x-axis direction x , the number of grid cells n of the structured grid in the y-axis direction y and the number of grid cells n of the structured grid in the z-axis direction z . When the coefficient matrix of the sparse triangular equation set is a lower triangular matrix, the relationship between the subtask index number p and the coordinates (i, j, k) of the vertex of the structured grid satisfies:

[0143] p = i + j × n x + k × n x × n y

[0144] Exemplarily, for the sparse triangular equation system Wx = b, if the size of the coefficient matrix W is 27×27, the size of the solution vector x is 27×1, and the dimension of the right-hand vector b is 27×1, still taking the structured grid of the 3D19 template as an example, the structured grid of this sparse triangular equation system is the above-mentioned Figure 7 shown structured grid. The relationship between the coordinates of the points in the structured grid and the indices of the subtasks can be seen in Figure 10 shown ( Figure 10 schematically showing the coordinates of some points). According to the above calculation formula, the index of the subtask corresponding to each point can be calculated. For example, for the vertex (1, 1, 1) in the structured grid, the index of the subtask corresponding to this vertex calculated according to the calculation formula is 13.

[0145] S602. Analyze the dependency relationship between the first subtask and other subtasks among multiple subtasks, and determine the dependency rule.

[0146] In the embodiments of the present application, analyzing the dependency relationship between the first subtask and other subtasks among multiple subtasks is to analyze the distribution of the multiple subtasks indicated by the structured grid information in the structured grid, and determine which subtasks among the other subtasks are the tasks that the first subtask depends on.

[0147] In one implementation manner, in combination with Figure 6 , as Figure 9 shown, the above S602 is implemented through S6021 - S6022.

[0148] S6021. Analyze at least one other vertex that the first vertex corresponding to the first subtask in the structured grid of the grid computing task depends on.

[0149] In the embodiments of the present application, for any two vertices in the structured grid, such as the first vertex and the second vertex, in the structured grid, the positional relationship between at least one vertex that the first vertex depends on and the first vertex is the first positional relationship, and the positional relationship between at least one vertex that the second vertex depends on and the second vertex is the second positional relationship. According to the known characteristics of the structured grid, the first positional relationship is the same as the second positional relationship.

[0150] It can be understood that when the grid computing task is determined, the association relationship between the points in the structured grid is also determined. For example, when the coefficient matrix of the sparse triangular equation system is determined, the association relationship between each point in the structured grid of this sparse triangular equation system is known.

[0151] Continue to refer to the above Figure 10, for the assumption that the first vertex corresponding to the first task is point 13, and this point 13 depends on 9 points, namely point 1, point 3, point 4, point 5, point 7, point 9, point 10, point 11, and point 12. Thus, it can be known that subtask 13 depends on subtask 1, subtask 3, subtask 4, subtask 5, subtask 7, subtask 9, subtask 10, subtask 11, and subtask 12. According to the conclusion of the dependency relationship between multiple subtasks of the sparse triangular equation set introduced in the above embodiments: for subtask i, the index numbers of the subtasks on which subtask i depends are equal to the column numbers of the non-zero elements in the i-th row of the coefficient matrix. Then, the column indices of the non-zero elements in the 13th row of the coefficient matrix are 1, 3, 4, 5, 7, 9, 10, 11, 12, and 13.

[0152] For the sake of convenience in description, the positive direction of the x-axis of the coordinate system of the structured grid is simply referred to as the right side, the negative direction of the x-axis is simply referred to as the left side, the positive direction of the y-axis is simply referred to as the rear side, the negative direction of the y-axis is simply referred to as the front side, the positive direction of the z-axis is simply referred to as the upper side, and the negative direction of the z-axis is simply referred to as the lower side. In the structured grid of the 3D19 template, the points on the right side depend on the points on the left side, the points on the upper side depend on the points on the lower side, and the points on the rear side depend on the points on the front side.

[0153] Among the 9 points on which point 13 depends, point 1 is located in the front lower part of point 13. That is, starting from point 13, first move 1 unit along the front direction (referring to the side length of a grid cell of the cube), and then move 1 unit along the lower direction. Similarly, the positional relationships between point 13 and point 3, point 4, point 5, point 7, point 9, point 10, point 11, and point 12 can be analyzed, and the positional relationships are shown in Table 1 below.

[0154] Table 1

[0155] The point that point 13 depends on The positional relationship with point 13 Point 1 Front lower Point 3 Left lower Point 4 Directly below Point 5 Right lower Point 7 Rear lower Point 9 Left front Point 10 Directly in front Point 11 Right front Point 12 Directly to the left

[0156] S6022. Determine the dependency rule according to the positional relationship between the first vertex and at least one other vertex.

[0157] According to the positional relationship between the first vertex and at least one other vertex shown in the above Table 1, the dependency relationship between the first subtask and other subtasks can be obtained. The dependency relationship between subtasks is determined by the above positional relationship between points. Then, it can be understood that the above positional relationship can be regarded as the dependency rule that subtasks conform to.

[0158] S603. Analyze the dependency relationship between other subtasks among multiple subtasks according to the dependency rule.

[0159] Through the description of the above embodiments, according to the positional relationship shown in Table 1, it is possible to determine the points on which each point in the structured grid depends, that is, to determine the subtasks on which each subtask in the sparse triangular system of equations depends. It should be noted that for the points on the boundary surface of the structured grid, the number of points on which they depend is less than 9. For the above-mentioned sparse triangular system of equations, Table 2 shows the subtasks on which each subtask depends. The subtasks in Table 2 are the vertices in the structured grid. According to the point i corresponding to subtask i and the positional relationship analyzed above, one or more points on which the point i depends can be determined, and thus one or more subtasks on which the subtask i depends can be determined.

[0160] Table 2

[0161]

[0162] After analyzing the dependency relationships between the subtasks, it is possible to know which subtasks among the multiple subtasks are independent of each other, which subtasks are dependent on each other, and the number of layers of the dependency relationships.

[0163] In the embodiments of the present application, according to the structural characteristics of the structured grid, that is, the distribution of multiple subtasks, the dependency rules that the subtasks conform to are analyzed, and then the dependency relationships between the multiple subtasks are obtained. For the grid calculation task of solving the sparse triangular system of equations, there is no need to analyze the information of the coefficient matrix of the sparse triangular system of equations one by one to obtain the dependency relationships. It can be seen that the method for analyzing the subtask dependency relationships based on the structured grid in the present application can more quickly analyze and obtain the dependency relationships between multiple subtasks.

[0164] Optionally, in combination with Figure 9 , as Figure 11 shown, after S603, the method provided by the embodiments of the present application further includes S605.

[0165] S605. For each subtask among the multiple subtasks, delete the duplicate dependency relationships in the dependency relationships between the subtask and other subtasks, so as to prune the dependency relationships between the multiple subtasks.

[0166] In the embodiments of the present application, there may be redundancy in the dependency relationships between the subtasks. For example, among the at least one subtask on which a subtask depends, there is also one or more subtasks that depend on other subtasks among the at least one subtask. For example, for the above-mentioned subtask 13, the subtask 13 depends on subtask 1, subtask 3, subtask 4, subtask 5, subtask 7, subtask 9, subtask 10, subtask 11, and subtask 12. Among them, subtask 5 also depends on subtask 1 and subtask 4. Therefore, there are duplicate dependency relationships, that is, redundant dependency relationships. Pruning the redundant dependency relationships can simplify the dependency relationships and help with subsequent resource allocation.

[0167] Optionally, for each of the multiple subtasks, the duplicate dependencies in the dependencies between the deletion subtasks and other subtasks are specifically to delete the subtasks that are repeatedly depended on among the at least one subtask that each subtask depends on. For example, among the subtasks that subtask 13 depends on, at least subtask 1 and subtask 4 are repeatedly depended on.

[0168] In one implementation, for each of the multiple subtasks, the method of deleting the subtasks that are repeatedly depended on among the at least one subtask that each subtask depends on further includes step 1 and step 2.

[0169] Step 1: In the structural grid, for each vertex corresponding to a subtask, determine the positional relationship between the repeatedly depended-on vertex among the at least one vertex that the vertex depends on and this vertex.

[0170] Continuing with the above example for illustration, as Figure 12 shown, according to the rule of the dependency relationship of points in the structural grid of the 3D19 template, for subtask 13, that is, point 13 in the structural grid, among the points that point 13 depends on, points 1, 3, 4, 5, 9, and 10 are repeatedly depended on. Therefore, it is necessary to delete the corresponding subtasks among the subtasks that subtask 13 depends on, so that the dependency relationship of subtask 13 is trimmed to point 13 depending on point 7, point 11, and point 12. Combining the above Table 2 and Figure 12 , the positional relationship between the repeatedly depended-on points among the points that point 13 depends on and point 13 is: located in the front lower left, lower left, directly below, lower right, front left, and directly in front of point 13.

[0171] Step 2: According to the above positional relationship, determine the subtasks that are repeatedly depended on among the at least one subtask that each subtask depends on, and delete the repeatedly depended-on subtasks.

[0172] In the embodiment of the present application, in the structural grid, for all vertices in the structural grid, the positional relationship between the repeatedly depended-on vertex among the at least one vertex that the vertex depends on and this vertex is the same. Therefore, according to the positional relationship determined in step 1, the subtasks that are repeatedly depended on among the at least one subtask that each subtask depends on can be determined, and then the repeatedly depended-on subtasks can be deleted to obtain the at least one subtask that each subtask finally depends on. Table 3 exemplifies the dependency relationship after the dependency relationship between each subtask in Table 2 above is trimmed.

[0173] Table 3

[0174]

[0175]

[0176] Based on the pruned dependency relationships shown in Table 3, it can be analyzed that the dependency relationships between various subtasks include 13 layers, denoted as Each layer contains subtasks as Figure 13 shown.

[0177] In the process of pruning the dependency relationships between subtasks above, based on all vertices in the structural grid, and the structural feature that the positional relationship between the repeatedly dependent vertices and a vertex among at least one vertex on which the vertex depends is the same. After knowing the positional relationship, the repeatedly dependent vertices on which each vertex depends can be quickly determined according to the positional relationship, so as to more quickly achieve the pruning of dependency relationships.

[0178] S604. Allocate computing resources for multiple subtasks according to the dependency relationships between multiple subtasks, and execute multiple subtasks.

[0179] The computing resources allocated for multiple subtasks above refer to allocating multiple subtasks to different threads for running, and different threads process subtasks in parallel.

[0180] During the process of executing subtasks, based on the relationships between subtasks obtained from the above analysis, the execution order of multiple subtasks is determined as follows: Subtasks in the same layer are independent of each other and can run in parallel. There are dependency relationships between subtasks in different layers, and subtasks need to be run layer by layer.

[0181] For ease of understanding, taking Figure 14 to intuitively show the dependency relationships between multiple subtasks. Exemplarily, for Figure 14 the 27 subtasks shown, subtasks 0, 1, 2, 4, 5, 6, 7, 11, 8, 13, 14, 15, 16, 20, 17, 23, 25, 26 can be allocated to thread 1, and subtasks 3, 9, 10, 12, 18, 19, 21, 22, 24 can be allocated to thread 2. In this way, thread 1 and thread 2 can process subtasks in parallel, and subtasks are processed serially within a single thread.

[0182] When the grid computing task is to solve a sparse triangular equation system, the above execution of multiple subtasks specifically takes the coefficient matrix of the sparse triangular equation as the input, and multiple threads execute multiple subtasks of the sparse triangular equation according to the execution order of multiple subtasks to obtain the solution vector of the sparse triangular equation system.

[0183] For the sparse triangular equation system Wx = b, with the coefficient matrix W and the right-hand side vector b as known input information, the process of performing multiple subtasks of the sparse triangular equation system in the execution order of multiple subtasks to obtain the solution vector of the sparse triangular equation system specifically includes: each thread processes subtasks in parallel, and subtasks are processed serially within the same thread. Among them, the coefficient matrix W can be obtained based on three one-dimensional arrays in the CSR format of this coefficient matrix.

[0184] It should be understood that when a thread executes the corresponding subtask, according to other subtasks on which this subtask depends, the non-zero elements of the row corresponding to this subtask in the coefficient matrix, and the corresponding components in the right-hand side vector, the corresponding components in the solution vector of the triangular equation system are solved. Exemplarily, referring to the schematic diagram of the dependency relationship shown above Figure 14 When the thread executes subtask 13, the data required to execute subtask 13 are the 1st, 3rd, 4th, 5th, 7th, 9th, 10th, 11th, and 12th components in the solution vector, and b in the right-hand side vector. 13 .

[0185] The following takes one thread as an example to introduce the processing process of subtasks. As Figure 15 shown, the process of a thread (thread 1) executing subtasks includes the following S1501 - S1507.

[0186] S1501. Determine whether the current subtask depends on other subtasks.

[0187] When thread 1 executes the current subtask, first, according to the analyzed dependency relationship between subtasks, it is judged whether the current subtask depends on other subtasks. If the current subtask depends on other subtasks, then execute S1502; if the current subtask does not depend on other subtasks, then execute S1505 and subsequent steps.

[0188] For example, if the current subtask is subtask 0 in the above Figure 14 , according to the dependency relationship, it is determined that subtask 0 does not depend on other subtasks, then execute the following S1505 to execute this subtask to obtain the corresponding component x in the solution vector 0 . Also, for example, if the current subtask is subtask 4 in the above Figure 14 , it can be seen that subtask 4 depends on subtask 2 and subtask 3, then continue to execute S1502.

[0189] S1502. Determine whether there are unfinished subtasks among the subtasks on which the current subtask depends.

[0190] Since the current subtask depends on other subtasks and the execution of this subtask requires the data obtained after the execution of other subtasks is completed, it is necessary to determine whether the subtasks on which the current subtask depends have been completed. If there are no incomplete subtasks among the subtasks on which the current subtask depends (i.e., all the subtasks on which the current subtask depends have been executed), then execute S1505 to execute the current subtask; if there are incomplete subtasks among the subtasks on which the current subtask depends, then execute S1503 - S1504.

[0191] Optionally, the method for determining whether the subtasks on which the current subtask depends are completed is that thread 1 continuously reads the subtask completion flags of the dependent subtasks, and the subtask completion flag indicates whether the subtask has been executed.

[0192] S1503: Read the completion flags of the subtasks on which the current subtask depends.

[0193] Exemplarily, the completion flag of a subtask can be 0 or 1, where 1 indicates that the subtask has been completed and 0 indicates that the subtask is not completed.

[0194] S1504: Determine whether the dependent subtasks are completed based on the read subtask completion flags.

[0195] If the completion flag of the subtask on which the current subtask depends indicates that the subtask is not completed, then thread 1 returns to S1503, continues to wait and read the subtask completion flag; if the completion flag of the subtask on which the current subtask depends indicates that the subtask has been completed, then it is determined that the dependent subtask has been executed.

[0196] It should be noted that thread 1 needs to determine whether each of the subtasks on which the current subtask depends has been executed, and until there are no incomplete subtasks among all the subtasks on which the current subtask depends (i.e., all have been executed), execute S1505 to solve the component of the solution vector corresponding to this subtask.

[0197] S1505: Execute the current subtask to obtain the component corresponding to the solution vector.

[0198] S1506: Modify the completion flag of the current subtask.

[0199] After thread 1 executes a subtask, it needs to modify the completion flag of this subtask so that when subsequent other subtasks depend on this subtask, thread 1 can determine whether the subtask has been executed based on the subtask completion flag (i.e., the process described in S1502 above).

[0200] After thread 1 executes the current subtask, thread 1 continues to execute S1507 to determine whether thread 1 ends.

[0201] S1507. Determine whether the layer where the current subtask is located is the last layer of this thread.

[0202] It can be understood that in the embodiments of the present application, the dependency relationships between subtasks are hierarchical, and the thread executes subtasks in the order of layers. If the layer where the current subtask is located is the last layer of the dependency relationship, then the execution of thread 1 ends; if the layer where the current subtask is located is not the last layer of this thread, then thread 1 continues to loop and execute the above S1501 - S1507 to execute the subtasks of the next layer.

[0203] Each thread executes the corresponding subtasks according to the above processing process. After all threads have finished execution, the solution vector of the sparse triangular equation system can be obtained.

[0204] Generally, when storing the coefficient matrix in memory, the non - zero elements of each row are stored in sequence according to the order of the rows of the coefficient matrix. When a thread runs a certain subtask, it needs to read the non - zero elements of the row corresponding to this subtask in the coefficient matrix from memory (for example, when executing subtask 13, it needs to read the non - zero elements of the 13th row of the coefficient matrix). For a thread, the serial numbers of the subtasks it executes may not be continuous. Correspondingly, this thread needs to read the non - zero elements of the corresponding rows of the coefficient matrix from memory in a jump - like manner, which will cause the thread to be unable to access memory continuously during the process of executing subtasks (that is, discontinuous memory access), resulting in waste of memory bandwidth.

[0205] Continuing with the above Figure 14 Taking the above - shown example as an example, thread 1 executes each subtask in the order of subtask serial numbers 0, 1, 2, 4, 5, 6, 7, 11, 8, 13, 14, 15, 16, 20, 17, 23, 25, 26. It can be seen that when thread 1 executes subtask 0, it reads the non - zero elements of the 0th row of the coefficient matrix; after thread 1 finishes executing subtask 0, it executes subtask 1 and reads the non - zero elements of the 1st row of the coefficient matrix; after thread 1 finishes executing subtask 1, it executes subtask 2 and reads the non - zero elements of the 2nd row of the coefficient matrix. It can be seen that during the process of thread 1 executing subtasks 0, 1, and 2, the serial numbers of the subtasks are continuous, so there is no problem of discontinuous memory access. When thread 1 finishes executing subtask 2, thread 1 needs to execute subtask 4 and reads the non - zero elements of the 4th row of the coefficient matrix. In this case, thread 1 needs to skip the non - zero elements of the 3rd row of the coefficient matrix in memory and then read the non - zero elements of the 4th row, resulting in a phenomenon of discontinuous memory access.

[0206] Optionally, for the case where there are discontinuous memory accesses, in the embodiments of the present application, after analyzing the dependency relationships between multiple subtasks, the task processing method provided by the embodiments of the present application further includes: for subtasks without dependency relationships, adjusting the storage locations of the data of the tasks so that the storage locations of the data of tasks that can be executed in parallel are continuous. For example, for a first subtask and a second subtask that are executed in parallel and have no dependency relationships, adjust the storage location of the data of the first subtask or the data of the second subtask so that the storage locations of the data of the first subtask and the second subtask are continuous.

[0207] Exemplarily, taking Figure 14 the subtasks 7 and 11 executed by thread 1 in as an example, there is no dependency relationship between subtask 7 and subtask 11, and thread 1 can execute subtask 7 and subtask 11 in parallel. Then, adjust the positions of the data in the 7th row and / or the 11th row, and the 7th column and / or the 11th column of the sparse matrix so that the storage locations of the data of subtask 7 and subtask 11 are continuous.

[0208] In one implementation, for the solution of a sparse triangular equation system, the data of the subtasks are the data in the coefficient matrix of the sparse triangular equation system. The above method for adjusting the storage locations of the data of the tasks may include: rearranging the coefficient matrix of the sparse triangular equation system according to the execution order of the multiple subtasks and the computing resources allocated to the multiple subtasks, so that the memory access is continuous during the process of the thread executing the subtasks, and the memory bandwidth is fully utilized. Specifically, the method for rearranging the coefficient matrix includes: for one thread, rearranging both the rows and columns of the coefficient matrix according to the subtask serial numbers, and different threads rearrange the corresponding rows and columns in the coefficient matrix according to the thread serial numbers.

[0209] The following illustrates the rearrangement process of the coefficient matrix through an example. For the above Figure 14 example shown, subtasks 0, 1, 2, 4, 5, 6, 7, 11, 8, 13, 14, 15, 16, 20, 17, 23, 25, 26 correspond to thread 1, and subtasks 3, 9, 10, 12, 18, 19, 21, 22, 24 correspond to thread 2. Therefore, for the rows of the coefficient matrix, arrange them in the order of row indexes 0, 1, 2, 4, 5, 6, 7, 11, 8, 13, 14, 15, 16, 20, 17, 23, 25, 26, 3, 9, 10, 12, 18, 19, 21, 22, 24.

[0210] For the columns of the coefficient matrix, they also need to be rearranged in the order of the arrangement order of the rows.

[0211] For example, the original fourth row in the coefficient matrix is rearranged to the third row, and correspondingly, the fourth column in the coefficient matrix is rearranged to the third column; the original fifth row in the coefficient matrix is rearranged to the fourth row, and correspondingly, the fifth column in the coefficient matrix is rearranged to the fourth column; the original sixth row in the coefficient matrix is rearranged to the fifth row, and correspondingly, the sixth column in the coefficient matrix is rearranged to the fifth column; the columns of the coefficient matrix are rearranged according to this rule.

[0212] It should be noted that in the embodiments of the present application, the rearrangement of the coefficient matrix is a corresponding rearrangement of both the rows and columns of the coefficient matrix, which is only for optimizing the memory access continuity and will not change the dependency relationship between subtasks and the execution order of subtasks. Therefore, it will not affect the solution result.

[0213] It can be understood that during the execution of subtasks, it is necessary to access the components at the corresponding positions of the solution vector according to the column indices of the non-zero elements in the coefficient matrix. For example, for the component x 13 , the component x 13 depends on the components x 1 , x 3 , x 4 , x 5 , x 7 , x 9 , x 10 , x 11 , x 12 , when solving the component x 13 , it is necessary to read the components it depends on, and the positions of these components are determined based on the column indices (1, 3, 4, 5, 7, 9, 10, 11, 12) of the non-zero elements corresponding to the 13th row of the component x 13 in the coefficient matrix.

[0214] In the embodiments of the present application, taking the structured grid of the 3D19 template as an example, the way to store the coefficient matrix in a structured manner is A[nx, ny, nz, 0:9], and this storage method stores the structured coefficient matrix (which can be called a structured matrix). Among them, n x , n y , n z is the size information of the structured grid. When n x = 3, n y = 3, n z = 3, the size of this structured matrix is 27×10, and 0:9 indicates that the column indices of the structured matrix range from 0 to 9, with 10 elements.

[0215] The above-mentioned structured matrix is related to the structured grid of the grid computing task, so it can be called a structured matrix. Corresponding to the structured grid, the number of rows 27 of the structured matrix corresponds to 27 points in the structured grid, and the number of columns 10 of the structured matrix corresponds to the number of points in a point association group in the structured grid.

[0216] Reference Figure 16 , taking the p-th row of the structured matrix as an example, the p-th row corresponds to the p-th point in the structured grid (the value of p is related to the coordinates of the p-th point in the structured grid, see the above embodiments). According to the characteristics of the structured grid, each point on the non-boundary surface depends on 9 other points. Therefore, among the 10 elements in the p-th row, the element with column index 9 corresponds to the p-th point, and the elements with column indices 0-8 correspond to the other points that the p-th point in the structured grid depends on. Assuming the coordinates of the p-th point are (i, j, k), the coordinates of the points corresponding to the elements with column indices 0-8 in the structured grid are successively: (i, j-1, k-1), (i-1, j, k-1), (i, j, k-1), (i+1, j, k-1), (i, j+1, k-1), (i-1, j-1, k), (i, j-1, k), (i+1, j-1, k), (i-1, j, k). The relationship of the coordinates reflects the positional relationship between the p-th point and the 9 points it depends on, that is, the positional relationship obtained in the above embodiments. These 9 points are successively located in the front lower left, lower left, directly below, lower right, rear lower left, front left, directly in front, front right, and directly left of point p.

[0217] Since there is a mathematical relationship between the coordinates of a point and the index of the point (i.e., the index of the row) (p = i + j×n x +k×n x ×n y ), according to the coordinates of each point corresponding to the p-th row above, the column indices of the non-zero elements in the p-th row can be calculated. Taking the 13th row as an example, the coordinates (i, j, k) of point 13 in the structured grid are (1, 1, 1). The positions (i.e., the column indices of the non-zero elements) of the non-zero elements in the 13th row of the coefficient matrix can be calculated according to the coordinates of each point in the 13th row, which are 1, 3, 4, 5, 7, 9, 10, 11, 12, 13 respectively.

[0218] Figure 17 A simple schematic of the structure of the structured matrix is shown. The elements at the shaded positions in the matrix represent the stored non-zero elements. The column numbers of these non-zero elements are calculated through the coordinates of the points. The numerical values indicated at the shaded positions in the figure represent the column numbers of the non-zero elements. The values of the non-zero elements are stored at the positions of the non-zero elements in the structured matrix.

[0219] As above, according to the coordinates of the points in the structured grid corresponding to the elements in the coefficient matrix, the column coordinates of the non-zero elements in the coefficient matrix can be directly calculated. Based on the column coordinates, the relevant components can be read. Compared with the CSR format, during the access process of the solution vector, there is no need for secondary indexing (i.e., first reading the column index from the array d and then reading the component based on the column index), which can reduce the memory access volume, speed up the memory access speed, and improve the computational memory access ratio.

[0220] For the content described above, refer to Figure 18 , the solution process of the sparse triangular equation system may include a dependency analysis stage, a data preparation stage, and a calculation stage. Among them, the dependency analysis stage refers to analyzing the dependencies between multiple subtasks based on the structural grid of the sparse triangular equation system to obtain the dependencies of the multiple subtasks. The data preparation stage refers to the process of adjusting the positions of the data in the coefficient matrix (or the structured coefficient matrix) according to the analyzed dependencies. The calculation stage refers to solving the sparse triangular equation system based on the dependencies between the subtasks, the coefficient matrix after position adjustment, and other input information to obtain the solution vector.

[0221] Optionally, the idea of solving the sparse triangular equation system can also be used to solve the matrix factorization problem in the ILU algorithm. According to the description of the above embodiments, in the ILU algorithm, a matrix is decomposed into an upper triangular matrix and a lower triangular matrix. For example, when decomposing the sparse matrix A, LU = A. It can be seen that U or L can be regarded as the solution vector, and A can be regarded as the right-hand vector. Therefore, the process of solving U or L is essentially also the solution process of the sparse triangular equation system. Therefore, the method provided in the embodiments of the present application can also be used for ILU decomposition.

[0222] In the task processing method provided by the embodiments of the present application, since a vertex in the structural grid of the grid computing task corresponds to a subtask of the grid computing task, and the positional relationship between all vertices in the structural grid and the vertices on which they depend is the same. Therefore, after analyzing the subtasks on which a subtask depends based on the structural grid, according to the same positional relationship, the other subtasks on which multiple subtasks depend can be quickly analyzed, that is, the dependencies between multiple subtasks can be quickly analyzed through the structural grid of the grid computing task. In this way, the grid computing task can be quickly processed, improving the efficiency of business processing.

[0223] Furthermore, in the task processing method provided by the embodiments of the present application, instead of analyzing one by one based on the array stored in the CSR format, the dependencies between the subtasks are analyzed according to the structural grid, which can significantly reduce the memory access volume and improve the calculation-to-memory access ratio.

[0224] It can be understood that the above method is executed by a computing device, which, in order to implement the above functions, includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the method steps of each example described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0225] The embodiments of the present application can divide the functional modules of the above computing device according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0226] In the case of dividing each functional module corresponding to each function, Figure 19 FIG. shows a possible structural schematic diagram of the computing device involved in the above embodiments. The computing device includes an acquisition module 1901, an analysis module 1902, and a processing module 1903.

[0227] Among them, the acquisition module 1901 is used to execute S601 in the above method embodiment; the analysis module 1902 is used to execute S602 (including S6021 - S6022), S603, and S1501 in the above method embodiment; the processing module 1903 is used to execute S604, S1505, and S1506 in the above method embodiment.

[0228] Optionally, the computing device provided by the embodiments of the present application further includes a cropping module 1904 and an adjustment module 1905. Among them, the cropping module 1904 is used to execute S605 in the above method embodiment; the adjustment module 1905 is used to adjust the storage location of the data of the first subtask or the data of the second subtask for the first subtask and the second subtask that have no dependency relationship.

[0229] Each module of the above computing device can also be used to execute other actions in the above method embodiment. All relevant contents of each step involved in the above method embodiment can be cited in the function description of the corresponding functional module, and will not be elaborated here.

[0230] In the case of adopting an integrated unit, Figure 20Another possible structural schematic diagram of the computing device involved in the above embodiments is shown. As Figure 20 shown, the computing device provided by the embodiments of the present application may include: a processing module 2001 and a communication module 2002. The processing module 2001 may be used to control and manage the actions of the computing device. For example, the processing module 2001 may be used to support the Figure 19 acquisition module 1901, analysis module 1902, processing module 1903, clipping module 1904, and adjustment module 1905 in the above to execute corresponding steps, and / or for other processes of the technologies described herein. The communication module 2002 may be used to support the communication of the communication device with other network entities. As Figure 20 shown, the communication device may further include a storage module 2003 for storing computer instructions and data (such as a coefficient matrix).

[0231] Among them, the processing module 2001 may be a processor or a controller (for example, the processing module 2001 may be the Figure 5 processor 501 in the above), and the above processing module 2001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and so on. The communication module 2002 may be a communication interface (for example, the communication module 2002 may be the Figure 5 communication interface 503 in the above). The storage module 2003 may be a memory (for example, the storage module 2003 may be the Figure 5 memory 502 in the above). When the processing module 2001 is a processor, the communication module 2002 is a communication interface, and the storage module 2003 is a memory, the processor, transceiver, and memory may be connected through a bus.

[0232] For more details on the functions implemented by the modules included in the above computing device, please refer to the descriptions in the foregoing method embodiments, and will not be repeated here. Each embodiment in this specification is described in a progressive manner. For the same or similar parts between each embodiment, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments.

[0233] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in the form of a computer program product in whole or in part. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (such as a floppy disk, a magnetic disk, a magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid state drive (SSD)), etc.

[0234] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. The specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0235] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in an electrical, mechanical, or other form.

[0236] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0237] In addition, each functional unit in various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0238] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk, or optical disc, etc., various media that can store program codes.

[0239] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A task processing method, characterized in that, it includes: Obtain a grid computing task in an application, where the grid computing task includes multiple subtasks, and the dependency relationships between the multiple subtasks conform to dependency rules; Analyze the dependency relationships between a first subtask and other subtasks among the multiple subtasks, and determine the dependency rules; Analyze the dependency relationships between other subtasks among the multiple subtasks according to the dependency rules; Allocate computing resources to the multiple subtasks according to the dependency relationships between the multiple subtasks, and execute the multiple subtasks.

2. The method according to claim 1, characterized in that, the grid computing task is to solve a triangular equation by using matrix multiplication.

3. The method according to claim 1 or 2, characterized in that, the method further includes: For each subtask among the multiple subtasks, delete the duplicate dependency relationships in the dependency relationships between the subtask and other subtasks, so as to trim the dependency relationships between the multiple subtasks.

4. The method according to any one of claims 1 to 3, characterized in that, the method further includes: For a first subtask and a second subtask that are executed in parallel and have no dependency relationship between them, adjust the storage location of the data of the first subtask or the data of the second subtask, so that the storage locations of the data of the first subtask and the data of the second subtask are continuous.

5. The method according to any one of claims 1 to 4, characterized in that, the grid of the grid computing task is a structured grid, and multiple vertices of the structured grid correspond to the multiple subtasks one by one; The analyzing the dependency relationships between a first subtask and other subtasks among the multiple subtasks, and determining the dependency rules includes: Analyze at least one other vertex on which a first vertex corresponding to the first subtask depends; Determine the dependency rules according to the positional relationship between the first vertex and the at least one other vertex.

6. The method according to claim 5, characterized in that, the structured grid is a three-dimensional structured grid, and the relationship between the index number p of the subtask of the grid computing task and the coordinates (i, j, k) of the vertex of the structured grid satisfies: p = i + j×n x + k×n x ×n y where n x represents the number of grid cells of the structural grid in the x-axis direction, and n y represents the number of grid cells of the structural grid in the y-axis direction.

7. A task processing device, characterized in that, it includes: An acquisition module, an analysis module, and a processing module; The acquisition module is used to obtain a grid computing task in an application, where the grid computing task includes multiple subtasks, and the dependency relationships between the multiple subtasks conform to dependency rules; The analysis module is used to analyze the dependency relationships between a first subtask and other subtasks among the multiple subtasks, and determine the dependency rules; The analysis module is further used to analyze the dependency relationships between other subtasks among the multiple subtasks according to the dependency rules; The processing module is used to allocate computing resources to the multiple subtasks according to the dependency relationships between the multiple subtasks, and execute the multiple subtasks.

8. The task processing device according to claim 7, characterized in that, the grid computing task is to solve a triangular equation by using matrix multiplication.

9. The task processing device according to claim 7 or 8, wherein, it further comprises a pruning module; the pruning module is configured to, for each of the multiple subtasks, delete duplicate dependencies in the dependencies between the subtask and other subtasks, so as to prune the dependencies between the multiple subtasks.

10. The task processing device according to any one of claims 7 to 9, wherein, it further comprises an adjustment module; the adjustment module is configured to, for a first subtask and a second subtask that are executed in parallel and have no dependencies between them, adjust the storage location of the data of the first subtask or the data of the second subtask, so that the storage locations of the data of the first subtask and the data of the second subtask are consecutive.

11. The task processing device according to any one of claims 7 to 9, wherein, the grid of the grid computing task is a structured grid, the analysis module is specifically configured to analyze at least one other vertex on which the first vertex corresponding to the first subtask depends; and determine the dependency rule according to the positional relationship between the first vertex and the at least one other vertex.

12. The task processing device according to claim 11, wherein, the structured grid is a three-dimensional structured grid, and the relationship between the index number p of the subtask of the grid computing task and the coordinates (i, j, k) of the vertex of the structured grid satisfies: p = i + j×n x + k×n x ×n y where n x represents the number of grid cells of the structural grid in the x-axis direction, and n y represents the number of grid cells of the structural grid in the y-axis direction.

13. A computing device, wherein, it comprises a memory and at least one processor connected to the memory, the memory is used to store computer program code, the computer program code includes computer instructions, and when the computer instructions are executed by the at least one processor, the processor executes the method according to any one of claims 1 to 6.

14. A computer-readable storage medium, wherein, it stores computer instructions, and when the computer instructions run on a computer, they execute the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Task processing method and apparatus

    EP4807585A1

  • Task processing method and apparatus

    WO2025118892A1