Program code distribution method and device and computing equipment

By identifying and allocating the first and second types of computing codes in HPC tasks, and using data flow graphs for load balancing and resource allocation, the problem of low computing resource allocation efficiency in high-performance computing tasks is solved, and the effect of improving execution efficiency is achieved.

CN120045303APending Publication Date: 2025-05-27CHENGDU HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311592646.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In high-performance computing (HPC) tasks, the calculation of program code is large and complex, resulting in low execution efficiency and it is difficult to effectively allocate computing resources to improve execution efficiency.

Method used

By identifying the first type of calculation code with large amount of calculation and the second type of calculation code with small amount of calculation in the program code, each segment of the first type of calculation code is assigned to each segment of the calculation code, and load balancing and data flow diagram are established based on the calculation amount and data transfer amount to reduce data transfer across the calculation unit.

Benefits of technology

The execution efficiency of program code is improved, and the utilization rate and overall computing efficiency of the computing unit are improved through load balancing and reducing data handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045303A_ABST
    Figure CN120045303A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a program code distribution method and device and computing equipment, and belongs to the technical field of computers. The method comprises the steps of obtaining a program code used for executing a high performance computing (HPC) task in response to a program allocation instruction; at least one section of first calculation code for executing a first type of calculation in the HPC task is determined in the program code, and the calculation amount corresponding to the first type of calculation is larger than the calculation amount corresponding to a second type of calculation except the first type of calculation in the HPC task. In a plurality of computing units for executing the program code, a computing unit is allocated to each segment of the first computing code, respectively. By adopting the method and the device, each section of the first calculation code with large calculation amount in the program code can be allocated to the same calculation unit to be executed, the data transmission amount of the first calculation code across the calculation units can be reduced, and the execution efficiency of the program code can be further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and particularly to a method, an apparatus, and a computing device for allocating program code. Background Art

[0002] With the development of computer technologies and the improvement of the performance of computing hardware, the computational amount of program code is increasing, and the computational complexity is also getting higher. In particular, program code involving high-performance computing (HPC) includes a large number of complex mathematical calculations. Therefore, how to improve the execution efficiency of program code becomes particularly important. Summary of the Invention

[0003] Embodiments of the present application provide a method, an apparatus, and a computing device for allocating program code, which can allocate different computing units to program code to improve the execution efficiency of the program code. The corresponding technical solutions are as follows:

[0004] In a first aspect, a method for allocating program code is provided. The method can be executed by a processor and includes: in response to a program allocation instruction, obtaining program code for performing a high-performance computing (HPC) task. Determining at least one segment of first computing code for performing a first type of computation in the HPC task in the program code, where the computational amount corresponding to the first type of computation is greater than the computational amount corresponding to a second type of computation other than the first type of computation in the HPC task. Allocating a computing unit to each segment of the first computing code among a plurality of computing units for executing the program code.

[0005] The first type of computation is a computation (such as FFT computation, stencil computation, etc.) composed of basic computations (such as addition, subtraction, multiplication, etc.), and belongs to the key computation of the HPC task. The second type of computation is the remaining basic computation in the HPC task other than the first type of computation. The computing unit may be a processor core included in a computing device for executing the program code. The computing device for executing the program code and the computing device for executing the method for allocating the program code provided in the present application may be the same device or different devices.

[0006] In the solution shown in the present application, the first computing code for performing the first type of computation included in the program code can be identified. When allocating the computing units for executing the program code, each segment of the first computing code can be allocated to the same computing unit for execution. In this way, a large amount of data transfer across computing units generated when the first computing code with a large computational amount is executed on multiple computing units can be avoided, and thus the execution efficiency of the program code can be improved.

[0007] In an implementable manner, among multiple computing units for executing program code, a computing unit is allocated to each segment of first computing code respectively, including: determining a computing unit for executing each segment of first computing code among multiple computing units based on the computing amount corresponding to each segment of first computing code and the data transfer amount corresponding to each pair of adjacent first computing codes in the computing order.

[0008] In the solution shown in this application, by referring to the computing amount corresponding to each segment of first computing code and allocating a computing unit to the first computing code, load balancing of multiple computing units can be achieved, and the overall execution efficiency of the program code can be improved. Considering the data transfer amount corresponding to each pair of adjacent first computing codes in the computing order and allocating a computing unit to the first computing code, multiple segments of first computing code with a relatively large data transfer amount between them can be allocated to the same computing unit for execution, which can reduce data transfer across computing units and thus improve the computing efficiency of the program code.

[0009] In an implementable manner, the above method further includes: determining at least one segment of second computing code for executing second-type computing. For each segment of second computing code, a computing unit for executing the second computing code is determined among multiple computing units based on the data transfer amount corresponding to the second computing code and the computing code adjacent in the computing order.

[0010] In the solution shown in this application, since the computing amount of the second-type computing is small, when allocating a computing unit to the second computing code, the computing amount of the second computing code can be ignored, and only the data transfer amount between each segment of second computing code and each segment of program code already allocated to a computing unit is considered to allocate a computing unit to the second computing code. This can reduce data transfer across computing units and thus improve the computing efficiency of the program code.

[0011] In an implementable manner, before allocating a computing unit to each segment of first computing code, it further includes: determining the first computing code as a first-type node and the second computing code as a second-type node. Based on the computing amount respectively corresponding to each segment of first computing code, determining the node weight of each first-type node respectively corresponding to each segment of first computing code. Based on the data transfer amount of each pair of adjacent computing codes in the computing order, determining the edge weight between each pair of nodes corresponding to each pair of computing codes. Based on the first-type nodes and the corresponding node weights, the second-type nodes, and the edge weights corresponding to each pair of nodes, establishing a data flow graph corresponding to the program code.

[0012] In the solution shown in this application, by establishing a data flow graph corresponding to the program code, based on the data flow graph, the program code allocation method provided in this application can be implemented, and thus the execution efficiency of the program code on multiple computing units can be improved.

[0013] In an implementable manner, based on the amount of computation corresponding to each segment of the first computation code and the data transfer amount corresponding to each pair of adjacent first computation codes in the computation order, determining the computation units for executing each segment of the first computation code among multiple computation units includes: based on the node weights of each first type of node and the edge weights between each pair of first type of nodes in the data flow graph, dividing the multiple first type of nodes into multiple subgraphs, where the number of the multiple subgraphs is the same as the number of the multiple computation units. Assigning the first computation codes corresponding to the first type of nodes in the same subgraph to the same computation unit.

[0014] In an implementable manner, based on the node weights of each first type of node and the edge weights between each pair of first type of nodes in the data flow graph, dividing the multiple first type of nodes into multiple subgraphs includes: selecting multiple target nodes from the first type of nodes and dividing the multiple target nodes into different subgraphs, where the number of the multiple target nodes is equal to the number of the multiple subgraphs. For each remaining first type of node in the data flow graph after the target nodes are selected, based on the sum of the node weights of the first type of nodes included in each subgraph and the edge weights between the first type of node and the first type of nodes included in each subgraph, determining the subgraph to which the first type of node is added.

[0015] In an implementable manner, based on the data transfer amount corresponding to the second computation code and the computation codes adjacent in the computation order, determining the computation units for executing the second computation code among multiple computation units includes: based on the sum of the edge weights between the second type of nodes and the adjacent nodes in each subgraph, determining the subgraph to which the second type of node is added. Assigning the second computation codes corresponding to the second type of nodes in the same subgraph to the computation unit to which the first computation code corresponding to the first type of nodes in the same subgraph is assigned.

[0016] In an implementable manner, after determining at least one segment of the first computation code for executing the first type of computation in the HPC task in the program code, it further includes: obtaining the third computation code corresponding to the first type of computation from the code library, where the execution efficiency of the third computation code is higher than that of the first computation code; replacing the first computation code in the program code with the third computation code.

[0017] In the solution shown in this application, the optimized third computation codes corresponding to each first type of computation can be pre-stored in the code library. After identifying the first computation code included in the program code, the first computation code for implementing the first type of computation in the code library can be replaced with the corresponding third computation code. Since the third computation code is the optimized code and has a higher efficiency in implementing the first type of computation compared to the first computation code, the execution efficiency of the program code on multiple computation units can be further improved.

[0018] In one possible implementation, determining at least one segment of first calculation code for performing a first type of calculation in HPC tasks in program code includes: identifying, in the program code, calculation code that meets the calculation rules corresponding to the first type of calculation, and determining the identified calculation code as the first calculation code.

[0019] In one possible implementation, the first calculation code is added with comments corresponding to the first type of calculation; determining at least one segment of first calculation code for performing a first type of calculation in HPC tasks in program code includes: identifying, in the program code, comments corresponding to the first type of calculation, and determining the calculation code corresponding to the comments as the first calculation code.

[0020] In one possible implementation, the first type of calculation includes at least one of template stencil calculation, fast Fourier transform (FFT) calculation, and vector / matrix calculation.

[0021] In a second aspect, there is provided an apparatus for allocating program code, the apparatus including:

[0022] An acquisition module, configured to acquire program code for performing a high-performance computing (HPC) task in response to a program allocation instruction.

[0023] A determination module, configured to determine at least one segment of first calculation code for performing a first type of calculation in HPC tasks in the program code, where the calculation amount corresponding to the first type of calculation is greater than the calculation amount corresponding to a second type of calculation other than the first type of calculation in the HPC task.

[0024] An allocation module, configured to allocate a calculation unit to each segment of the first calculation code respectively among multiple calculation units for executing the program code.

[0025] In one possible implementation, the allocation module is configured to: determine, among multiple calculation units, a calculation unit for executing each segment of the first calculation code based on the calculation amount corresponding to each segment of the first calculation code and the data transfer amount corresponding to each pair of adjacent first calculation codes in the calculation order.

[0026] In one possible implementation, the determination module is further configured to: determine at least one segment of second calculation code for performing the second type of calculation.

[0027] The allocation module is further configured to: for each segment of the second calculation code, determine, among multiple calculation units, a calculation unit for executing the second calculation code based on the data transfer amount corresponding to the second calculation code and the adjacent calculation code in the calculation order.

[0028] In an implementable manner, the above-mentioned device further includes a graph construction module, which is used to: determine the first computing code as the first type of node, and determine the second computing code as the second type of node; based on the computing amount corresponding to each segment of the first computing code, determine the node weight of each segment of the first computing code corresponding to the first type of node; based on the data transfer amount between each pair of computing codes adjacent in the computing order, determine the edge weight between each pair of nodes corresponding to each pair of computing codes; based on the first type of node and the corresponding node weight, the second type of node, and the edge weight corresponding to each pair of nodes, construct a data flow graph corresponding to the program code.

[0029] In an implementable manner, the allocation module is used to: based on the node weight of each first type of node in the data flow graph and the edge weight between each pair of first type of nodes, divide multiple first type of nodes into multiple subgraphs, where the number of multiple subgraphs is the same as the number of multiple computing units; allocate the first computing code corresponding to the first type of node in the same subgraph to the same computing unit.

[0030] In an implementable manner, the allocation module is used to: select multiple target nodes from the first type of nodes, and divide the multiple target nodes into different subgraphs, where the number of multiple target nodes is equal to the number of multiple subgraphs; for each remaining first type of node in the data flow graph after selecting the target nodes, based on the sum of the node weights of the first type of nodes included in each subgraph and the edge weight between the first type of node and the first type of nodes included in each subgraph, determine the subgraph to which the first type of node is added.

[0031] In an implementable manner, the allocation module is used to: based on the sum of the edge weights between the second type of node and the adjacent nodes in each subgraph, determine the subgraph to which the second type of node is added; allocate the second computing code corresponding to the second type of node in the same subgraph to the computing unit to which the first computing code corresponding to the first type of node in the same subgraph is allocated.

[0032] In an implementable manner, the above-mentioned device further includes a replacement module, which is used to: obtain the third computing code corresponding to the first type of computing from the code library, and the execution efficiency of the third computing code is higher than that of the first computing code; replace the first computing code in the program code with the third computing code.

[0033] In an implementable manner, the determination module is used to: identify the computing code that satisfies the computing rule corresponding to the first type of computing in the program code, and determine the identified computing code as the first computing code.

[0034] In an implementable manner, the first computing code is added with an annotation corresponding to the first type of computing; the determination module is used to: identify the annotation corresponding to the first type of computing in the program code; determine the computing code corresponding to the annotation as the first computing code.

[0035] In one possible implementation, the first type of calculation includes at least one of stencil calculation, fast Fourier transform (FFT) calculation, and vector / matrix calculation.

[0036] In a third aspect, a computing device is provided, which includes a processor and a memory. The processor is configured to execute instructions stored in the memory, so that the computing device executes the method as described in the first aspect.

[0037] In a fourth aspect, a computer program product containing instructions is provided. When the instructions are run on a computing device, the computing device is caused to execute the method as described in the first aspect.

[0038] In a fifth aspect, a computer-readable storage medium is provided, which includes computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes the method as described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is a schematic structural diagram of a data flow graph provided by an embodiment of the present application;

[0040] Figure 2 is a schematic structural diagram of a computing device provided by an embodiment of the present application;

[0041] Figure 3 is a flowchart of a method for allocating program code provided by an embodiment of the present application;

[0042] Figure 4 is a flowchart of a method for establishing a data flow graph provided by an embodiment of the present application;

[0043] Figure 5 is a schematic structural diagram of a data flow graph provided by an embodiment of the present application;

[0044] Figure 6 is a structural diagram of an apparatus for allocating program code provided by an embodiment of the present application; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0046] A data flow graph (DFG) is a business logic expression method that effectively decouples the algorithm logic and resource scheduling of program code. Figure 1 is a schematic structural diagram of a data flow graph provided by an embodiment of the present application.

[0047] As Figure 1As shown, each node in the data flow diagram represents the calculation of data by the program code. For example, one node corresponds to an addition, multiplication, or division in the program code. If there is an edge between two nodes, it means there is data transfer between the program codes corresponding to these two nodes, and the direction of the corresponding edge is the direction of data transfer. For example, Figure 1 the node a in Figure 1 represents a step of addition in the program code, and the node b represents a step of multiplication in the program code. The edge ab between the node a and the node b can be used to represent that the result data of the addition operation corresponding to the node a will be input into the multiplication operation represented by the node b. Through the data flow diagram, the algorithm logic and resource scheduling involved in the program code can be displayed in a visual way. Technical personnel can further optimize the program code according to the structural characteristics of the data flow diagram generated from the program code. For example, adjust the order of calculation steps in the program code, or merge some calculation steps, so as to optimize the program code and improve the execution efficiency of the program code.

[0048] Currently, to generate a data flow diagram for program code, it is necessary to identify each section of calculation code included in the program code, determine each identified section of calculation code as a node in the data flow diagram, and then determine the edges and corresponding directions between the nodes according to the input-output relationships between the sections of calculation code. Since currently when generating the data flow diagram, the identified calculation codes all correspond to basic calculations such as addition, subtraction, multiplication, and division. This results in a large number of nodes in the generated data flow diagram, increasing the difficulty of analyzing and optimizing the program code, and there is no obvious focus in the data flow diagram, which cannot provide optimization ideas for technical personnel. Especially in High Performance Computing (HPC), because the calculations involved in the program code are more complex and the amount of calculation is larger. Therefore, the current data flow diagram has not been applied to the optimization of program code in the HPC field.

[0049] Figure 2 is a schematic diagram of a computing device provided by an embodiment of the present application. This computing device can be used to implement the program code allocation method provided by an embodiment of the present application. As Figure 2 shown, the computing device 200 includes: a bus 202, a processor 204, a memory 206, and a communication interface 208. Among them, the processor 204, the memory 206, and the communication interface 208 communicate with each other through the bus 202. The computing device 200 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 200.

[0050] The bus 202 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 2 it is represented by only one line in Figure 2 , but it does not mean that there is only one bus or one type of bus. The bus 202 can include a path for transmitting information between various components of the computing device 200 (for example, the memory 206, the processor 204, and the communication interface 208).

[0051] The processor 204 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0052] The memory 206 can include volatile memory, such as random access memory (RAM). The memory 206 can also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0053] Executable program code is stored in the memory 206, and the processor 204 executes the executable program code, which can implement the program code allocation method provided in the embodiments of the present application. For example, in response to a program allocation instruction, program code for performing a high-performance computing (HPC) task is obtained. At least one segment of first calculation code for performing the first type of calculation in the HPC task is determined in the program code. In a plurality of calculation units for executing the program code, a calculation unit is allocated for each segment of the first calculation code, etc.

[0054] The communication interface 208 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 200 and other devices or a communication network.

[0055] Figure 3 is a flowchart of a method for allocating program code provided in the embodiments of the present application. This method can be performed by the above-mentionedFigure 2 is executed by the computing device shown. Refer to Figure 3 , the method for allocating the program code includes:

[0056] Step 301, in response to a program allocation instruction, the computing device obtains program code for performing a high-performance computing (HPC) task.

[0057] Among them, the HPC tasks executed by the program code can be one or more, or the program code can also execute part of the calculations in an HPC task. The computing device that executes the method for allocating the program code provided in the embodiments of the present application and the computing device that executes the program code of the HPC task can be the same computing device or different computing devices.

[0058] In the case of the same computing device, the computing unit for allocating the program code is the multiple computing units included in the computing device. For example, the computing unit can be the processing core of the processor in the computing device. In the case of different computing devices, the computing unit for allocating the program code can be multiple computing devices, multiple processors, or multiple processor cores. For example, if the program code is executed by a computing device cluster, the computing unit can be the computing devices included in the computing device cluster, or if the program code is executed by other computing devices, the computing unit can be the processor or processor core in the computing device.

[0059] In one example, the method provided in the embodiments of the present application can be implemented by an application program running in the computing device. In the present application, this application program can be referred to as a code allocation application. The code allocation application can also be added to an existing compiler as a plug-in of the compiler. In implementation, the user can input the program code of the computing unit to be allocated in the code allocation application and set the number of computing units. Then, trigger a program allocation instruction in the code allocation application. In response to the program allocation instruction triggered in the code allocation application, the computing device can further execute the method for allocating the program code provided in the embodiment, that is, execute steps 302 to 303.

[0060] Step 302, determine at least one segment of first calculation code that executes the first type of calculation in the HPC task in the program code, where the calculation amount corresponding to the first type of calculation is greater than the calculation amount corresponding to the second type of calculation other than the first type of calculation in the HPC task.

[0061] After obtaining the program code, the code allocation application can identify the calculation code included in the program code. In this embodiment, the calculations included in the program code can be divided into a first type of calculation and a second type of calculation. Among them, the first type of calculation can be a complex calculation composed of basic calculations (such as addition, subtraction, multiplication, and division). For example, this complex calculation can be used to implement algorithms involved in HPC tasks, such as sorting algorithms, hash algorithms, etc. The second type of calculation is the remaining basic calculations in the HPC task other than the first type of calculation. Therefore, the amount of calculation of the first type of calculation is greater than the amount of calculation corresponding to the second type of calculation, and the first type of calculation can be predefined by technicians. For example, the first type of calculation can be FFT calculation, stencil calculation, vector / matrix calculation, etc. involved in HPC tasks.

[0062] In one example, the definition of the first type of calculation can be to set the function name of the calculation function corresponding to the first type of calculation, or to preset the calculation rules corresponding to the first type of calculation, such as the calculation order of the basic calculations included in the first type of calculation. After defining the first type of calculation, the first calculation code used to execute the first type of calculation included in the program code can be identified according to the function name, calculation rules, etc. corresponding to the defined first type of calculation.

[0063] The processing of identifying the first calculation code in the program code can at least include the following two methods:

[0064] Identification method 1: Identify the calculation code that satisfies the calculation rules corresponding to the first type of calculation in the program code, and determine the identified calculation code as the first calculation code.

[0065] In implementation, the code allocation application can identify the first calculation code included in the program code according to the calculation rules corresponding to the defined first type of calculation. For example, the code allocation application identifies the calculation type of the basic calculation corresponding to each line of calculation code in the program code, and then determines whether the calculation types and calculation orders corresponding to multiple adjacent calculation codes meet the calculation rules of the first type of calculation. If it meets, the multiple adjacent calculation codes can be determined as the first calculation code. If it does not meet, each line of calculation code in the multiple adjacent lines of calculation code is the second calculation code.

[0066] In one example, the code allocation application can construct an abstract syntax tree (AST) corresponding to the program code, and then traverse the abstract syntax tree. Each node in the abstract syntax tree can represent variables, function calls, calculation expressions, etc. in the program code. During the traversal, if it is determined that the structure formed by the nodes in the abstract syntax tree is the same as the node structure corresponding to the first type of calculation in the abstract syntax tree, it can be determined that the first type of calculation included in the program code is recognized. Further, according to the TaskLet mechanism (the soft interrupt latency mechanism in the linux interrupt handling mechanism) or by analyzing the code semantics, the program code corresponding to the node corresponding to the first type of calculation in the abstract syntax tree, that is, the first program code, can be determined. The node structure corresponding to the first type of calculation in the abstract syntax tree can be set according to the calculation rules of the first type of calculation.

[0067] Recognition method 2: Identify the annotation corresponding to the first type of calculation in the program code, and determine the calculation code corresponding to the annotation as the first calculation code.

[0068] During the process of writing the program code, the technician can add the annotation corresponding to the first type of calculation to each segment of the first calculation code in the program code. For example, an annotation can be added to the starting line of each segment of the first calculation code, or an annotation can also be added to each line of each segment of the first calculation code. In implementation, after the code allocation application obtains the program code, it can identify the annotation corresponding to the first type of calculation in the program code. After the annotation is recognized, the calculation code corresponding to the recognized annotation can be determined as the first calculation code.

[0069] In addition, after identifying the first calculation code in the program code, the remaining calculation code of each line can be determined as the second calculation code corresponding to the second type of calculation.

[0070] In the embodiments of the present application, the calculations included in the program code are divided into the first type of calculation with a large amount of calculation and the second type of calculation with a small amount of calculation. The first calculation code is identified according to the first type of calculation, and the second calculation code is identified according to other calculations. Since the first type of calculation is composed of multiple segments of calculation code for implementing basic calculations, in the embodiments of the present application, compared with identifying the calculation code of each basic calculation, the number of identified calculation codes is reduced, and the workload of optimizing the program code can be reduced. In addition, since the calculation amount of the first calculation code is large, an optimization direction is provided for subsequent optimization of the program code, that is, the first calculation code can be focused on for optimization.

[0071] Step 303: In multiple calculation units for executing the program code, allocate a calculation unit for each segment of the first calculation code respectively.

[0072] In the embodiments of the present application, each segment of the first calculation code can be assigned to a calculation unit for execution. In this way, it can be avoided that a segment of the first calculation code is assigned to multiple calculation units for execution, and thus a large amount of data transfer across calculation units generated during the execution of each segment of the first calculation code on multiple calculation units can be avoided, which can improve the execution efficiency of the first calculation code.

[0073] In one example, calculation units can be assigned to each segment of the first calculation code according to the calculation amount of the first calculation code and the data transfer amount between the first calculation code and other first calculation codes adjacent in the calculation order.

[0074] Among them, the calculation amount of the first calculation code can be determined by the complexity of the first type of calculation implemented by the first calculation code and the amount of data to be calculated. For example, the product of the complexity and the amount of data can be determined as the calculation amount corresponding to the first calculation code. The complexity of the first type of calculation is related to the calculation type of the first type of calculation. In one example, it can be determined by the corresponding relationship between the calculation type of the first type of calculation set and the complexity. The amount of data to be calculated can be determined according to the amount of data of the input data of the first calculation code and the data type of the input data. In one example, the product of the number and the number of bits occupied by the input data of the corresponding data type can be determined as the amount of data calculated by the first calculation code.

[0075] The calculation order in the program code refers to the execution order of multiple segments of the first calculation code and multiple segments of the second calculation code. For two calculation codes adjacent in the calculation order, the output data of the calculation code with the earlier calculation order includes the input data of the calculation code with the later calculation order. Among the output data of the calculation code with the earlier calculation order, the amount of data of the input data of the calculation code with the later calculation order included is the data transfer amount corresponding to the two calculation codes adjacent in the calculation order.

[0076] After determining the calculation amount corresponding to each segment of the first calculation code and the data transfer amount of each pair of calculation codes adjacent in the calculation order, calculation units can be assigned to each segment of the first calculation code in sequence. The process of assigning calculation units to the first calculation code can be as follows:

[0077] The code allocation application can allocate multiple segments of the first calculation code to multiple calculation units according to the calculation amount of the first calculation code and the data transfer amount between each pair of adjacent first calculation codes in the calculation order, so that the sum of the calculation amounts of the multiple segments of calculation code executed by each calculation unit is basically the same. If the difference in the sum of the calculation amounts of the first calculation code allocated to multiple calculation units is within a preset difference range, load balancing of multiple calculation units can be achieved. The data transfer amount between each segment of the first calculation program and other first calculation programs in the same calculation unit is greater than the data transfer amount with the first calculation programs allocated to other calculation units, which can reduce the data transfer amount between calculation units and thus improve the calculation efficiency. For example, when the calculation unit is a processor core, allocating two segments of the first calculation code with a large data transfer amount to the same processor core means that the data that needs to be transferred between the two segments of the first calculation code does not require additional transfer. However, if the two segments of the first calculation code are allocated to different processor cores, the data that needs to be transferred between the two segments of the first calculation code needs to be transferred across cores, which will affect the execution efficiency of the program code.

[0078] Among them, if there is a second calculation code between two adjacent first calculation codes in the calculation order, the data transfer amount between the two first calculation codes is equal to the sum of the data transfer amounts between each segment of the calculation code between the two first calculation codes. For example, if there are also a second calculation code B and a second calculation code C between the first calculation code A and the first calculation code D that are adjacent in the calculation order, the data transfer amount between the first calculation code A and the first calculation code D is equal to the sum of the data transfer amount between the first calculation code A and the second calculation code B, the data transfer amount between the second calculation code B and the second calculation code C, and the data transfer amount between the second calculation code C and the first calculation code D.

[0079] In an example, the process of allocating a calculation unit to the second calculation code may include: for each segment of the second calculation code, determining the calculation unit that executes each segment of the second calculation code among multiple calculation units based on the data transfer amount between the second calculation code and the adjacent calculation code in the calculation order.

[0080] After allocating multiple segments of first calculation code to multiple calculation units, second calculation code can be allocated to the multiple calculation units. In one example, since the calculations corresponding to the second calculation code are all basic calculations with small computational amounts, when allocating the second calculation code, the computational amount of the second calculation code can be ignored, and only the data transfer amount between the second calculation code and the calculation code adjacent to it in the calculation order is considered. For example, for each segment of the second calculation code, it can be determined whether there is a calculation code adjacent to the second calculation code in the calculation order in each calculation unit. For the existing adjacent calculation codes, it can be further determined which calculation code among the adjacent calculation codes has the largest data transfer amount with the second calculation code, and then the calculation unit to which the calculation code with the largest data transfer amount is allocated is determined as the calculation unit allocated to the second calculation code.

[0081] After determining the calculation units that execute the first calculation code and the second calculation code respectively in the program code, program code of the calculation units can be bound to each segment of the first calculation code and the second calculation code in the program code. For example, when the calculation unit is a processor core, program code for binding the cores to the first calculation code and the second calculation code can be added to the program code.

[0082] In the embodiments of the present application, on the one hand, the calculation programs in the program code are divided into first calculation code with large computational amounts and second calculation code with small computational amounts, thus reducing the number of divided calculation codes and reducing the difficulty of allocating calculation units to the calculation codes. On the other hand, according to the computational amount and data transfer amount of the first calculation code, allocating the calculation unit to the first calculation code first can achieve load balancing of the calculation units, and allocating the first calculation codes with a large data transfer amount between them to the same calculation unit can avoid data transmission across calculation units and improve the execution efficiency of the calculation code. On the further hand, for the second calculation code with small computational amounts, the calculation unit can be allocated to the second calculation code only considering the data transfer amount between the second calculation codes, which can further avoid data transmission across calculation units and further improve the execution efficiency of the calculation code.

[0083] In an implementable manner, the embodiments of the present application also provide a method for optimizing the first type of calculation, including: obtaining the third calculation code corresponding to the first type of calculation from the code library, and replacing the first calculation code in the program code with the third calculation code.

[0084] Among them, a code library can be set in the code allocation application, and the code library can store the optimized third calculation codes corresponding to each first type of calculation, and the execution efficiency of the third calculation code is higher than that of the first calculation code.

[0085] Since the first calculation code written by technicians may have redundancies in aspects such as the calculation process and has low execution efficiency. Therefore, in the embodiments of the present application, after the code allocation application identifies the first calculation code, it can determine the calculation category of the first type of calculation corresponding to the first calculation code. Among them, the calculation category can be the algorithm implemented by the first type of calculation, for example, it can be FFT calculation, stencil calculation, matrix multiplication calculation, etc. After determining the calculation category corresponding to the first calculation code, then obtain the third calculation code optimized for the execution of this calculation category in the code library, such as the optimized calculation code corresponding to FFT calculation and stencil calculation. The optimized third calculation code can be the calculation code after optimizing the calculation process of the first type of calculation, and has higher execution efficiency than the calculation code written by technicians. Therefore, after determining the third calculation code corresponding to the first calculation code, the first calculation code in the program code can be replaced with the third calculation code. It should be noted that before replacing the third calculation code, the variable names of the variables in the third calculation code can be kept the same as the variable names in the corresponding first calculation code, so that the third calculation code can be executed normally in the program code. For example, the variable names included in the first calculation code can be determined, and then the variable names in the optimized calculation code can be changed to the variable names included in the corresponding first calculation code. In this way, by replacing the first calculation code in the program code with the optimized third calculation code, the execution efficiency of the program code can be further improved.

[0086] The embodiments of the present application also provide a method for establishing a data flow graph. This method for establishing a data flow graph can be combined with the above-mentioned program code allocation method to implement the allocation of program code. Figure 4 This is the method for establishing a data flow graph provided by the embodiments of this period. This method can also be executed by the above Figure 2 shown computing device. See Figure 4 This method for establishing a data flow graph includes:

[0087] Step 401, obtain the program code of the computing unit to be allocated.

[0088] Step 402, determine the first calculation code corresponding to the first type of calculation included in the program code and the second calculation code corresponding to other calculations.

[0089] Among them, the processing of step 401 and step 402 can refer to the processing of step 301 and step 302 above, and will not be introduced in detail here.

[0090] Step 403, determine the first calculation code as the first type of node and the second calculation code as the second type of node.

[0091] After identifying the first calculation code and the second calculation code included in the program code, the first calculation code can be determined as the first type of node in the data flow diagram, and the second calculation code can be determined as the second type of node in the data flow diagram.

[0092] Step 404: Determine the node weights of the first type of nodes corresponding to each segment of the first calculation code based on the amount of calculation corresponding to each segment of the first calculation code.

[0093] In one example, the amounts of calculation corresponding to each first calculation code can be normalized, and then the normalized amount of calculation can be determined as the node weight of the first type of node.

[0094] Step 405: Determine the edge weights between each pair of nodes corresponding to each pair of calculation codes based on the data transfer amount between each pair of calculation codes adjacent in the calculation order.

[0095] In one example, the data transfer amounts corresponding to each pair of calculation codes adjacent in the calculation order can be normalized, and then the normalized data transfer amount can be determined as the edge weight between each pair of nodes corresponding to each pair of calculation codes. Among them, the edge weights between each pair of nodes include the edge weights between the first type of nodes and the first type of nodes, the edge weights between the first type of nodes and the second type of nodes, and the edge weights between the second type of nodes and the second type of nodes.

[0096] Step 406: Establish a data flow diagram corresponding to the program code based on the first type of nodes and the corresponding node weights, the second type of nodes, and the edge weights corresponding to each pair of nodes.

[0097] Before determining the first type of nodes, the second type of nodes, the node weights of the first type of nodes, and the edge weights between each node in the data flow diagram, a data flow diagram can be established. As Figure 5 shown, Figure 5 is a schematic diagram of a data flow diagram provided by an embodiment of the present application. In Figure 5 it, the nodes with larger areas identify the first type of nodes, the nodes with smaller areas identify the second type of nodes, and the pointing direction of the edges between the nodes represents the data flow direction between the nodes. Since the first type of calculation corresponding to the first calculation code is composed of basic calculations, the first type of nodes are actually formed by merging the nodes corresponding to multiple basic calculations. Therefore, determining one node for one segment of the first calculation code can reduce the number of nodes included in the data flow diagram and reduce the difficulty for technicians to analyze the data flow diagram. Moreover, the first type of calculation is generally the key calculation of the program code, and the calculation process composed of multiple first calculation codes is the main calculation process in the program code. Therefore, Figure 5In the data flow diagram shown, the first type of nodes A, the first type of nodes B, and the first type of nodes C can display the main architecture of the corresponding program code, and thus can provide a clear optimization direction for technicians.

[0098] In the data flow diagram provided in the embodiment of the present application, the computational expressions (nodes) and data scheduling (edges) in the program code can be represented in a visual manner, and thus the program code can be optimized in terms of both computational expressions and data scheduling. For computational expressions, the code corresponding to each computational node can be separately adapted to different hardware, thereby improving the portability of the program code between different hardwares. For data scheduling, some data scheduling processes can be merged and adjusted, thereby reducing the amount of data transfer in the program code process and improving the execution efficiency of the program code.

[0099] In implementation, the computational units can be allocated to the first computational code and the second computational code by dividing the data flow diagram into subgraphs. Among them, the number of subgraphs divided is the same as the number of computational units. For multiple segments of computational code corresponding to multiple nodes divided into the same subgraph, they are multiple segments of computational code allocated to the same computational unit.

[0100] In an implementable manner, the process of allocating computational units to the first computational code based on the data flow diagram may include: dividing multiple first type of nodes into multiple subgraphs based on the node weights of each first type of node and the edge weights between each pair of first type of nodes in the data flow diagram. In an example, the difference in the sum of the node weights of the first type of nodes included in multiple subgraphs is within a preset difference range, and the edge weights between the first type of nodes in the same subgraph are greater than the edge weights between the first type of nodes in this subgraph and the nodes in other subgraphs.

[0101] The following is an exemplary method for dividing the first type of nodes into subgraphs provided by the present application, as follows:

[0102] Select multiple target nodes from the first type of nodes and divide the multiple target nodes into different subgraphs. For each remaining first type of node in the data flow diagram after selecting the target nodes, determine the subgraph to which the first type of node is added based on the sum of the node weights of the first type of nodes included in each subgraph and the edge weights between the first type of node and the first type of nodes included in each subgraph.

[0103] In implementation, multiple first-type nodes can be randomly selected as target nodes in the data flow graph. The number of the target nodes is equal to the number of subgraphs to be partitioned, that is, equal to the number of computing units. In one example, there may be no edges between the selected target nodes. After selecting the target nodes, the remaining first-type nodes in the data flow graph can be traversed. For each traversed first-type node, the subgraph to which the first-type node joins can be determined according to the sum of the node weights of the first-type nodes included in each subgraph and the sum of the edge weights between the first-type node and the first-type nodes included in each subgraph.

[0104] In one example, the remaining first-type nodes can be partitioned into each subgraph by a one-way linear deterministic greedy (LDG) algorithm. The first objective function corresponding to the LDG algorithm can be g(v c ,P i )=|P i ∩N(v c *Com)|(1-|P i *(αCal / C cal +βMem / C mem )|), where v c is the traversed first-type node, P i refers to the i-th subgraph, |P i ∩N(v c *Com)| represents the sum of the edge weights between the first-type nodes adjacent to v c in the i-th subgraph and v c . 1-|P i *(αCal / C cal +βMem / C mem )| represents the remaining resource amount in the computing nodes corresponding to the i-th subgraph. α and β are preset coefficients, αCal / C cal +βMem / C mem represents the computing amount occupied by the first computing code already allocated in the corresponding computing unit, C cal represents the computing complexity, and C mem represents the memory occupancy. After calculating the values of the first objective function corresponding to each first-type node and each subgraph, the subgraph corresponding to the maximum value can be determined as the subgraph to which the first-type node joins.

[0105] In a realizable manner, the process of allocating computing units for the first computing code based on the data flow graph may include: determining the subgraph to which the second-type node joins according to the sum of the edge weights between the second-type node and the adjacent nodes in each subgraph. Allocate the second computing code corresponding to the second-type nodes in the same subgraph to the computing unit to which the first computing code corresponding to the first-type nodes in the same subgraph is allocated.

[0106] In implementation, after partitioning the first type of nodes in the data flow graph into each sub-graph, the remaining second type of nodes in the data flow graph can be traversed. For each traversed second type of node, the sum of the edge weights between the second type of node and each node included in each sub-graph determines the sub-graph to which the second type of node is added. That is, the sub-graph with the largest corresponding weight sum is determined as the sub-graph to which the second type of node is added.

[0107] The following is an exemplary method for partitioning the second type of nodes into sub-graphs provided by this application, as follows:

[0108] In one example, after partitioning the first type of nodes into each sub-graph through the LDG algorithm, the remaining second type of nodes can be further partitioned into each sub-graph through the LDG algorithm. The second objective function corresponding to the LDG algorithm can be g(v n ,P i )=|P i ∩N(v n *Com)|. Wherein, v n is the traversed second type of node, P i refers to the i-th sub-graph, and |P i ∩N(v n *Com)| represents the sum of the edge weights between the second type of nodes adjacent to v n in the i-th sub-graph and v n . After calculating the values of the second objective function corresponding to each second type of node and each sub-graph, the sub-graph corresponding to the maximum value can be determined as the sub-graph to which the second type of node is added.

[0109] In the embodiments of this application, by establishing a data flow graph to partition the first type of nodes and the second type of nodes into sub-graphs, the allocation of the first calculation code and the second calculation code in the program code can be completed. First partitioning the first type of nodes with large computational amounts into sub-graphs can balance the load of the computing units, and since the number of the first type of nodes is small, it can also reduce the difficulty and computational amount of partitioning the first type of nodes into sub-graphs. The edge weights between nodes in the same sub-graph are greater than the weights between nodes in other sub-graphs, so that the data transmission across computing units can be reduced, and thus the overall execution efficiency of the program code can be improved.

[0110] The embodiments of this application also provide an apparatus for allocating program code, as Figure 6 shown. This apparatus can be the computing device that executes the allocation of the program code. Referring to Figure 6 , this apparatus includes:

[0111] An acquisition module 610, configured to acquire program code for performing a high-performance computing (HPC) task in response to a program allocation instruction, and is specifically used to implement the acquisition function in the above-mentioned step 301 and implicit steps.

[0112] A determination module 620, configured to determine at least one section of first calculation code for performing a first type of calculation in the HPC task in the program code, where the amount of calculation corresponding to the first type of calculation is greater than the amount of calculation corresponding to a second type of calculation other than the first type of calculation in the HPC task, and is specifically used to implement the determination function in the above-mentioned step 302 and implicit steps.

[0113] An allocation module 630, configured to allocate a calculation unit to each section of the first calculation code respectively among multiple calculation units for executing the program code, and is specifically used to implement the allocation function in the above-mentioned step 303 and implicit steps.

[0114] In an implementable manner, the allocation module 630 is configured to: determine a calculation unit for executing each section of the first calculation code among multiple calculation units based on the amount of calculation corresponding to each section of the first calculation code and the data transfer amount corresponding to each pair of adjacent first calculation codes in the calculation order.

[0115] In an implementable manner, the determination module is further configured to: determine at least one section of second calculation code for performing the second type of calculation.

[0116] The allocation module 630 is further configured to: for each section of the second calculation code, determine a calculation unit for executing the second calculation code among multiple calculation units based on the data transfer amount corresponding to the second calculation code and the adjacent calculation code in the calculation order.

[0117] In an implementable manner, the above-mentioned apparatus further includes a graph construction module, configured to: determine the first calculation code as a first type of node and the second calculation code as a second type of node; determine the node weight of each section of the first calculation code corresponding to each section of the first calculation code based on the amount of calculation corresponding to each section of the first calculation code; determine the edge weight between each pair of nodes corresponding to each pair of calculation codes based on the data transfer amount between each pair of adjacent calculation codes in the calculation order; establish a data flow graph corresponding to the program code based on the first type of node and the corresponding node weight, the second type of node, and the edge weight corresponding to each pair of nodes.

[0118] In an implementable manner, the allocation module 630 is configured to: divide multiple first type of nodes into multiple subgraphs based on the node weight of each first type of node in the data flow graph and the edge weight between each pair of first type of nodes, where the number of multiple subgraphs is the same as the number of multiple calculation units; allocate the first calculation code corresponding to the first type of nodes in the same subgraph to the same calculation unit.

[0119] In an implementable manner, the allocation module 630 is configured to: select a plurality of target nodes from the first type of nodes, and divide the plurality of target nodes into different subgraphs, where the number of the plurality of target nodes is equal to the number of the plurality of subgraphs; for each of the remaining first type of nodes in the data flow graph after the target nodes are selected, determine the subgraph to which the first type of node is added based on the sum of the node weights of the first type of nodes included in each subgraph and the edge weights between the first type of node and the first type of nodes included in each subgraph.

[0120] In an implementable manner, the allocation module 630 is configured to: determine the subgraph to which the second type of node is added based on the sum of the edge weights between the second type of node and the adjacent nodes in each subgraph; allocate the second calculation code corresponding to the second type of nodes in the same subgraph to the calculation unit to which the first calculation code corresponding to the first type of nodes in the same subgraph is allocated.

[0121] In an implementable manner, the above device further includes a replacement module, configured to: obtain a third calculation code corresponding to the first type of calculation from a code library, and the execution efficiency of the third calculation code is higher than that of the first calculation code; replace the first calculation code in the program code with the third calculation code.

[0122] In an implementable manner, the determination module 620 is configured to: identify, in the program code, the calculation code that satisfies the calculation rule corresponding to the first type of calculation, and determine the identified calculation code as the first calculation code.

[0123] In an implementable manner, the first calculation code is added with an annotation corresponding to the first type of calculation; the determination module 620 is configured to: identify, in the program code, the annotation corresponding to the first type of calculation; and determine the calculation code corresponding to the annotation as the first calculation code.

[0124] In an implementable manner, the first type of calculation includes at least one of template stencil calculation, fast Fourier transform (FFT) calculation, and vector / matrix calculation.

[0125] The division of modules in the embodiments of the present application is illustrative. It is only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present application, each functional module may be integrated in one processor, may exist independently physically, or two or more modules may be integrated into one module. The above integrated module may be implemented in the form of hardware or in the form of a software functional module. In addition, the program code allocation device provided in the above embodiments and the embodiments of the program code allocation method belong to the same concept. The specific implementation process is detailed in the method embodiments and will not be elaborated here.

[0126] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a terminal device (which may be a personal computer, a mobile phone, or a network device, etc.) or a processor to execute all or part of the steps of the method in each embodiment of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0127] An embodiment of this application also provides a computer program product containing instructions. The computer program product can be a software or program product containing instructions that can run on a power management device or be stored in any available medium. When the computer program product runs on the power management device, it causes at least one computing device to execute the allocation method of the program code provided in the embodiment of this application.

[0128] An embodiment of this application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions that instruct the computing device to execute the allocation method of the program code provided in the embodiment of this application.

[0129] In this application, terms such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions and effects. It should be understood that there is no logical or temporal dependency between "first" and "second", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms "first" and "second" to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first calculation code can be called the second calculation code, and similarly, the second calculation code can be called the first calculation code. The first calculation code and the second calculation code can both be collectively referred to as the calculation code, and in some cases, they can be separate and different calculation codes.

[0130] In this application, the term "at least one" means one or more, and the term "a plurality of" means two or more.

[0131] The above description is only a specific embodiment of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

1. A method for allocating program code, characterized in that, the method includes: In response to a program allocation instruction, obtain program code for performing high-performance computing (HPC) tasks; Determine at least one segment of first calculation code in the program code for performing the first type of calculation in the HPC task, where the calculation amount corresponding to the first type of calculation is greater than the calculation amount corresponding to the second type of calculation other than the first type of calculation in the HPC task; Among multiple calculation units for executing the program code, allocate a calculation unit for each segment of the first calculation code respectively.

2. The method according to claim 1, characterized in that, the step of allocating a calculation unit for each segment of the first calculation code respectively among multiple calculation units for executing the program code includes: Based on the calculation amount corresponding to each segment of the first calculation code and the data transfer amount corresponding to each pair of adjacent first calculation codes in the calculation order, determine the calculation unit for executing each segment of the first calculation code among the multiple calculation units.

3. The method according to claim 1 or 2, characterized in that, the method further includes: Determine at least one segment of second calculation code for performing the second type of calculation; For each segment of the second calculation code, based on the data transfer amount corresponding to the second calculation code and the adjacent calculation code in the calculation order, determine the calculation unit for executing the second calculation code among the multiple calculation units.

4. The method according to claim 3, characterized in that, before allocating a calculation unit for each segment of the first calculation code, it further includes: Determine the first calculation code as the first type of node and the second calculation code as the second type of node; Based on the calculation amount corresponding to each segment of the first calculation code respectively, determine the node weight of the first type of node corresponding to each segment of the first calculation code; Based on the data transfer amount of each pair of adjacent calculation codes in the calculation order, determine the edge weight between each pair of nodes corresponding to each pair of calculation codes; Based on the first type of node and the corresponding node weight, the second type of node, and the edge weight corresponding to each pair of nodes, establish a data flow graph corresponding to the program code.

5. The method according to claim 4, characterized in that, the step of determining the calculation unit for executing each segment of the first calculation code among the multiple calculation units based on the calculation amount corresponding to each segment of the first calculation code and the data transfer amount corresponding to each pair of adjacent first calculation codes in the calculation order includes: Based on the node weight of each first type of node in the data flow graph and the edge weight between each pair of first type of nodes, divide multiple first type of nodes into multiple subgraphs, where the number of the multiple subgraphs is the same as the number of the multiple calculation units; Allocate the first calculation code corresponding to the first type of node in the same subgraph to the same calculation unit.

6. The method according to claim 5, characterized in that, the step of dividing multiple first type of nodes into multiple subgraphs based on the node weight of each first type of node in the data flow graph and the edge weight between each pair of first type of nodes includes: Select multiple target nodes from the first type of nodes, and divide the multiple target nodes into different subgraphs, where the number of the multiple target nodes is equal to the number of the multiple subgraphs; For each remaining first type of node in the data flow graph after selecting the target nodes, determine the subgraph to which the first type of node belongs based on the sum of the node weights of the first type of nodes included in each subgraph and the edge weights between the first type of node and the first type of nodes included in each subgraph.

7. The method according to any one of claims 4 to 6, characterized in that, The determining, in the multiple computing units, the computing unit that executes the second computing code based on the data transfer amount corresponding to the second computing code and the computing code adjacent in the computing order includes: Determine the subgraph to which the second type of node belongs based on the sum of the edge weights between the second type of node and the adjacent nodes in each subgraph; Allocate the second computing codes corresponding to the second type of nodes in the same subgraph to the computing units to which the first computing codes corresponding to the first type of nodes in the same subgraph are allocated.

8. The method according to any one of claims 1 to 7, characterized in that, After determining at least one first computing code that executes the first type of computation in the HPC task in the program code, further includes: Obtain a third computing code corresponding to the first type of computation from a code library, where the execution efficiency of the third computing code is higher than that of the first computing code; Replace the first computing code in the program code with the third computing code.

9. The method according to any one of claims 1 to 8, characterized in that, The determining at least one first computing code that executes the first type of computation in the HPC task in the program code includes: Identify, in the program code, computing codes that meet the computing rules corresponding to the first type of computation, and determine the identified computing codes as the first computing codes.

10. The method according to any one of claims 1 to 8, characterized in that, The first computing code is added with annotations corresponding to the first type of computation; The determining at least one first computing code that executes the first type of computation in the HPC task in the program code includes: Identify, in the program code, annotations corresponding to the first type of computation; Determine the computing codes corresponding to the annotations as the first computing codes.

11. The method according to any one of claims 1 to 10, characterized in that, The first type of computation includes at least one of template stencil computation, fast Fourier transform FFT computation, and vector / matrix computation.

12. An apparatus for allocating program code, characterized in that, The apparatus includes: An obtaining module, configured to obtain program code for executing a high performance computing HPC task in response to a program allocation instruction; A determining module, configured to determine at least one first computing code that executes the first type of computation in the HPC task in the program code, where the computation amount corresponding to the first type of computation is greater than the computation amount corresponding to the second type of computation other than the first type of computation in the HPC task; An allocation module, configured to allocate a computing unit for each segment of the first computing code respectively among a plurality of computing units for executing the program code.

13. The apparatus according to claim 12, wherein, the allocation module is configured to: Based on the computing amount corresponding to each segment of the first computing code and the data transfer amount corresponding to each pair of adjacent first computing codes in the computing order, determine the computing unit for executing each segment of the first computing code among the plurality of computing units.

14. The apparatus according to claim 12 or 13, wherein, the determination module is further configured to: determine at least one segment of the second computing code for executing the second type of computing; the allocation module is further configured to: for each segment of the second computing code, based on the data transfer amount corresponding to the second computing code and the computing code adjacent in the computing order, determine the computing unit for executing the second computing code among the plurality of computing units.

15. The apparatus according to claim 14, wherein, the apparatus further includes a graph construction module, configured to: Determine the first computing code as the first type of node and the second computing code as the second type of node; Based on the computing amount respectively corresponding to each segment of the first computing code, determine the node weight of the first type of node corresponding to each segment of the first computing code; Based on the data transfer amount of each pair of adjacent computing codes in the computing order, determine the edge weight between each pair of nodes corresponding to each pair of computing codes; Based on the first type of node and the corresponding node weight, the second type of node, and the edge weight corresponding to each pair of nodes, construct a data flow graph corresponding to the program code.

16. The apparatus according to claim 15, wherein, the allocation module is configured to: Based on the node weight of each first type of node in the data flow graph and the edge weight between each pair of first type of nodes, divide a plurality of first type of nodes into a plurality of subgraphs, where the number of the plurality of subgraphs is the same as the number of the plurality of computing units; Allocate the first computing code corresponding to the first type of node in the same subgraph to the same computing unit.

17. The apparatus according to claim 16, wherein, the allocation module is configured to: Select a plurality of target nodes from the first type of nodes and divide the plurality of target nodes into different subgraphs, where the number of the plurality of target nodes is equal to the number of the plurality of subgraphs; For each remaining first type of node in the data flow graph after selecting the target nodes, based on the sum of the node weights of the first type of nodes included in each subgraph and the edge weight between the first type of node and the first type of nodes included in each subgraph, determine the subgraph to which the first type of node belongs.

18. The apparatus according to any one of claims 15 to 17, wherein, the allocation module is configured to: Based on the sum of the edge weights between the second type of node and the adjacent nodes in each subgraph, determine the subgraph to which the second type of node belongs; Allocate the second computing code corresponding to the second type of node in the same subgraph to the computing unit to which the first computing code corresponding to the first type of node in the same subgraph is allocated.

19. The device according to any one of claims 12 to 18, wherein, the device further comprises a replacement module for: obtaining the third calculation code corresponding to the first type of calculation from a code library, the execution efficiency of the third calculation code being higher than that of the first calculation code; replacing the first calculation code in the program code with the third calculation code.

20. The device according to any one of claims 12 to 19, wherein, the determining module is used for: identifying, in the program code, calculation codes that meet the calculation rules corresponding to the first type of calculation, and determining the identified calculation codes as the first calculation code.

21. The device according to any one of claims 12 to 19, wherein, the first calculation code is added with annotations corresponding to the first type of calculation; the determining module is used for: identifying, in the program code, annotations corresponding to the first type of calculation; determining the calculation code corresponding to the annotation as the first calculation code.

22. The device according to any one of claims 12 to 21, wherein, the first type of calculation includes at least one of template stencil calculation, fast Fourier transform (FFT) calculation, and vector / matrix calculation.

23. A computing device, wherein, the computing device comprises a processor and a memory; the processor is configured to execute instructions stored in the memory, so that the computing device executes the method according to claims 1 to 11.

24. A computer program product comprising instructions, wherein, when the instructions are run on a computing device, the computing device is caused to execute the method according to claims 1 to 11.

25. A computer-readable storage medium, wherein, it includes computer program instructions, and when the computer program instructions are executed by a computing device, the computing device executes the method according to claims 1 to 11.