Method and device for executing business code and storage medium
By dividing the computing cluster into parallel execution blocks for computation and communication/memory access, the problem of low execution efficiency of business code in the computing cluster is solved, achieving efficient resource utilization and performance improvement.
Patent Information
- Application Number
- CN202410605439.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2025-11-18
AI Technical Summary
The low execution efficiency of business code in the computing cluster is due to the fact that after a node executes the first block of computational code, it needs to wait for a long time before it can execute the second block of computational code, resulting in wasted resources and extended execution time.
By dividing the business code based on computation tags and communication memory access tags, the scheduling information instructs the node to execute the communication memory access code block while executing the first computation code block, thereby reducing the waiting time of the second computation code block and executing the sub-computation code block and sub-communication memory access code block in parallel, thus optimizing the execution order of the code blocks.
It improved the execution efficiency of business code, reduced resource waste, and enhanced the computing performance of the computing cluster.
Smart Images

Figure CN120973483A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communications, and in particular to a method, apparatus, and storage medium for executing business code. Background Technology
[0002] Computing clusters typically consist of multiple nodes. Currently, business logic can be run using computing clusters. Business logic usually includes multiple computation code blocks and corresponding communication code blocks for each computation code block. Running business logic on a computing cluster means that multiple nodes in the computing cluster run the computation code blocks and communication code blocks that comprise the business logic.
[0003] For each computational code block, referred to as the first computational code block for ease of explanation, the business code includes a second computational code block that requires the execution result of the first computational code block. The computing cluster uses at least one first node to execute the first computational code block and at least one second node to execute the second computational code block. After at least one first node has executed the first computational code block and obtained the execution result, at least one first node executes the communication code block corresponding to the first computational code block. At least one first node sends the execution result to at least one second node by executing the communication code block. At least one second node executes the second computational code block based on the execution result.
[0004] After the first calculation code block is executed, at least one second node needs to wait for a long time before it can execute the second calculation code block, which reduces the execution efficiency of the business code. Summary of the Invention
[0005] This application provides a method, apparatus, and storage medium for executing business code, so as to improve the execution efficiency of business code. The technical solution is as follows:
[0006] Firstly, this application provides a method for executing business code. The method executes business code, which includes a computation marker for identifying computation code blocks and a communication memory access marker for identifying communication memory access code blocks. In the method, the business code is divided based on the computation marker and communication memory access marker to obtain a first computation code block, a second computation code block, and a communication memory access code block corresponding to the first computation code block. The second computation code block is a code block that requires the execution result of the first computation code block. Scheduling information is sent to at least one first node. The scheduling information includes the first computation code block and the communication memory access code block. The scheduling information instructs at least one first node to execute the first computation code block and the communication memory access code block, wherein the start time of executing the communication memory access code block is later than the start time of executing the first computation code block, and the start time of executing the communication memory access code block is earlier than the end time of executing the first computation code block. The communication memory access code block is used to send the execution result obtained during the execution of the first computation code block to at least one second node. The at least one second node is a node that needs to execute the second computation code block, and the at least one first node and the at least one second node are nodes in a computing cluster.
[0007] Since the scheduling information indicates that the start time of executing the communication memory access code block by at least one first node is later than the start time of executing the first computation code block, and the start time of executing the communication memory access code block is earlier than the end time of executing the first computation code block, at least one first node will also execute the communication memory access code block at the same time while executing the first computation code block, thereby reducing the waiting time of the second computation code block and improving the efficiency of executing business code.
[0008] By reducing the waiting time of the second computation code block, the idle time of at least one second node is also reduced, avoiding the long-term waste of resources of at least one second node and improving the computational performance of the computing cluster in executing business code.
[0009] In one possible implementation, the first computation code block is decomposed into M sub-computation code blocks, and the communication memory access code block is decomposed into M sub-communication memory access code blocks, where M is an integer greater than 1, and the M sub-computation code blocks correspond to the M sub-communication memory access code blocks. The first sub-computation code block is sent to at least one first node, which executes it. After at least one first node has executed the i-th sub-computation code block, a first scheduling command is sent to at least one first node. The first scheduling command includes the j-th sub-computation code block and the i-th sub-communication memory access code block corresponding to the i-th sub-computation code block, where i is a positive integer less than M, and j = i + 1. The first scheduling command instructs at least one first node to execute the j-th sub-computation code block and the i-th sub-communication memory access code block in parallel. The i-th sub-communication memory access code block is used to send the execution result of the i-th sub-computation code block to at least one second node. This allows the computational behavior implemented by executing the j-th sub-computation code block to mask the communication and memory access behavior implemented by executing the i-th sub-communication and memory access code block, thereby significantly improving the efficiency of executing business code.
[0010] In another possible implementation, at least one first node includes at least one computing node and at least one communication node, and the first scheduling command is used to simultaneously instruct at least one computing node to execute the j-th sub-computation code block and at least one communication node to execute the i-th sub-communication memory access code block.
[0011] In another possible implementation, i is less than M-1, and the execution time of at least one first node for the i-th sub-communication memory access code block is earlier than or equal to the execution time of the j-th sub-computation code block. When at least one first node completes the j-th sub-computation code block, a second scheduling command is sent to at least one first node. The second scheduling command includes the (j+1)-th sub-computation code block and the j-th sub-communication memory access code block corresponding to the j-th sub-computation code block. The second scheduling command instructs at least one first node to execute the (j+1)-th sub-computation code block and the j-th sub-communication memory access code block in parallel. The j-th sub-communication memory access code block is used to send the execution result of the j-th sub-computation code block to at least one second node. This allows the computational behavior implemented by executing the (j+1)-th sub-computation code block to mask the communication memory access behavior implemented by executing the j-th sub-communication memory access code block, thereby significantly improving the efficiency of executing business code.
[0012] In another possible implementation, i is less than M-1, and at least one first node completes the execution of the i-th sub-communication memory access code block later than the execution of the j-th sub-computation code block. When at least one first node completes the execution of the i-th sub-communication memory access code block, a second scheduling command is sent to at least one first node. The second scheduling command includes the (j+1)-th sub-computation code block and the j-th sub-communication memory access code block corresponding to the j-th sub-computation code block. The second scheduling command instructs at least one first node to execute the (j+1)-th sub-computation code block and the j-th sub-communication memory access code block in parallel. The j-th sub-communication memory access code block is used to send the execution result of the j-th sub-computation code block to at least one second node. This allows the computational behavior implemented by executing the (j+1)-th sub-computation code block to mask the communication memory access behavior implemented by executing the j-th sub-communication memory access code block, thereby significantly improving the efficiency of executing business code.
[0013] In another possible implementation, the i-th sub-communication memory access code block is also used to save the execution result obtained from executing the i-th sub-computation code block to the storage nodes included in the computing cluster.
[0014] In another possible implementation, multiple computation code blocks and multiple communication / memory access code blocks are obtained. The multiple computation code blocks include a first computation code block and at least one computation code block that has a dependency on the first computation code block. The multiple communication / memory access code blocks include a communication / memory access code block corresponding to the first computation code block and at least one communication / memory access code block corresponding to at least one computation code block. Based on the configuration information of the computing cluster, the execution time of each computation code block and the execution time of each communication / memory access code block are obtained. Based on the execution time of each computation code block and the execution time of each communication / memory access code block, the value of M is determined.
[0015] Since M is derived based on the execution time of each computation code block and the execution time of each communication / memory access code block, each computation code block includes a first computation code block and at least one computation code block that depends on the first computation code block. Each communication / memory access code block includes a communication / memory access code block corresponding to the first computation code block and at least one communication / memory access code block corresponding to at least one computation code block. This ensures that the value of M is adapted to the first computation code block. Therefore, decomposing the first computation code block and the first communication / memory access code block based on M optimizes the performance of the computation cluster in executing the first computation code block and the first communication / memory access code block.
[0016] Secondly, this application provides an apparatus for executing business code, used to execute the method in the first aspect or any possible implementation of the first aspect. Specifically, the apparatus includes a unit for executing the method in the first aspect or any possible implementation of the first aspect.
[0017] Thirdly, this application provides a computing device cluster, the computing device cluster including at least one computing device, each computing device including a processor and a memory;
[0018] The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method of the first aspect or any possible implementation thereof.
[0019] Fourthly, this application provides a computer program product containing instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method of the first aspect or any possible implementation thereof.
[0020] Fifthly, this application provides a computer-readable storage medium including computer program instructions that, when executed by a cluster of computing devices, perform the method of the first aspect or any possible implementation thereof.
[0021] In a sixth aspect, this application provides a chip including a memory and a processor, wherein the memory is used to store computer instructions, and the processor is used to retrieve and execute the computer instructions from the memory to perform the method in the first aspect or any possible implementation of the first aspect. Attached Figure Description
[0022] Figure 1 This is a schematic diagram illustrating the execution of computational and communication memory access behaviors provided in an embodiment of this application;
[0023] Figure 2 This is a schematic diagram of the structure of a computing cluster provided in an embodiment of this application;
[0024] Figure 3 This is a schematic diagram of another computing cluster structure provided in an embodiment of this application;
[0025] Figure 4 This is a flowchart of a method for executing business code provided in an embodiment of this application;
[0026] Figure 5 This is a schematic diagram of a business code provided in an embodiment of this application;
[0027] Figure 6This is a schematic diagram of an execution calculation code block and a communication memory access code block provided in an embodiment of this application;
[0028] Figure 7 This is a schematic diagram of a sub-computation code block and a sub-communication memory access code block provided in an embodiment of this application;
[0029] Figure 8 This application provides a schematic diagram of the scheduling order of executing sub-computation code blocks and sub-communication memory access code blocks according to an embodiment of the present application;
[0030] Figure 9 This application provides an embodiment of another scheduling order diagram for executing sub-computation code blocks and sub-communication memory access code blocks;
[0031] Figure 10 This is a schematic diagram of a device structure for executing business code provided in an embodiment of this application;
[0032] Figure 11 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application;
[0033] Figure 12 This is a schematic diagram of a cluster structure for executing business code provided in an embodiment of this application;
[0034] Figure 13 This is a schematic diagram of another cluster structure for executing business code provided in an embodiment of this application. Detailed Implementation
[0035] In the field of cluster computing, computing clusters can be used to execute business code. Business code includes numerous computational actions. For each computational action, the business code may also include corresponding communication and memory access actions; alternatively, the business code may not include the corresponding communication and memory access actions. Optionally, communication and memory access actions include communication actions and / or memory access actions, etc.
[0036] The business logic code includes numerous computational actions, which may be executed on different nodes within the computing cluster. For each computational action within the business logic code, it is designated as a "first computational action." Each first computational action may correspond to a communication / memory access action. After executing the first computational action and obtaining its result, some nodes in the computing cluster execute the corresponding communication / memory access action to transmit and / or save the execution result of the first computational action. For other computational actions that require the execution result of the first computational action, another set of nodes in the computing cluster executes these other computational actions. Upon receiving the execution result of the first computational action, these other nodes execute the other computational actions based on that result.
[0037] Optionally, if the communication-memory access behavior corresponding to the first computational behavior includes a communication behavior, some nodes, after obtaining the execution result of the first computational behavior, execute the communication behavior corresponding to the first computational behavior to transmit the execution result of the first computational behavior. For other computational behaviors that require the execution result of the first computational behavior, other nodes, after receiving the execution result of the first computational behavior, execute other computational behaviors based on the execution result of the first computational behavior. And / or, if the communication-memory access behavior corresponding to the first computational behavior includes a memory access behavior, some nodes, after obtaining the execution result of the first computational behavior, execute the memory access behavior corresponding to the first computational behavior to save the execution result of the first computational behavior. For other computational behaviors that require the execution result of the first computational behavior, after some nodes save the execution result of the first computational behavior, other nodes execute other computational behaviors based on the execution result of the first computational behavior.
[0038] The time between the end of the execution of the first computational action and the start of the execution of other computational actions is the waiting time required to execute other computational actions.
[0039] See Figure 1 If the first computational action, the corresponding communication memory access action, and other computational actions are executed serially, the waiting time required for other computational actions includes the execution time required for the communication memory access action.
[0040] Optionally, when the communication memory access behavior includes the communication behavior corresponding to the first computation behavior, the waiting time required to execute other computation behaviors includes the execution time required to execute the communication behavior, and / or, when the communication memory access behavior includes the memory access behavior corresponding to the first computation behavior, the waiting time required to execute other computation behaviors includes the execution time required to execute the memory access behavior. That is, the waiting time required to execute other computation behaviors includes the execution time required to execute the communication behavior and / or the execution time required to execute the memory access behavior.
[0041] Therefore, the longer the execution time required to perform the communication and memory access behavior after the computing cluster has completed the first computing behavior, the longer the waiting time required to perform other computing behaviors will be. This not only leads to a long period of waste of computing resources used to perform other computing behaviors, but also prolongs the execution time of business code and reduces the efficiency of business code execution.
[0042] In summary, the business logic code includes numerous computational actions. For one part of these computational actions, each action within that part has a corresponding communication / memory access action. For another part of the computational actions, there are no corresponding communication / memory access actions.
[0043] In some embodiments, the business code includes at least one computation code block, each of which implements computation behavior. Optionally, the business code may also include at least one communication memory access code block, each of which implements communication behavior and / or memory access behavior. Optionally, for each communication memory access code block, the communication memory access code block includes a communication code block and / or a memory access code block, the communication code block implementing communication behavior and the memory access code block implementing memory access behavior.
[0044] As computing clusters grow larger, the amount of communication and memory access operations they can execute in business code increases. The proportion of the total time required to execute these communication and memory access operations relative to the total time required to execute business code may also increase. This proportion will become a major bottleneck affecting the improvement of computing cluster performance. To address this, any of the following embodiments can be used to reduce this proportion.
[0045] See Figure 2 This application provides a computing cluster 200, which can be used to execute business code. The computing cluster 200 includes a management node 201, multiple computing nodes 202, and multiple communication nodes 203.
[0046] The management node 201 can communicate with each computing node 202 and each communication node 203. The management node 201 can manage and / or control each computing node 202, and / or manage and / or control each communication node 203.
[0047] For at least one of multiple computing nodes 202, the management node 201 can issue computing actions to the at least one computing node 202, and the at least one computing node 202 can execute the computing actions. Optionally, the management node 201 can issue computing code blocks to the at least one computing node 202 to implement the computing actions, and the at least one computing node 202 can execute the computing actions by executing the computing code blocks.
[0048] For at least one of the multiple communication nodes 203, the communication memory access behavior corresponding to the computation behavior includes a communication behavior. The management node 201 can issue a communication behavior to at least one communication node 203, and the at least one communication node 203 executes the communication behavior so that the at least one communication node 203 sends the execution result obtained in the process of executing the computation behavior. Optionally, the management node 201 can issue a communication code block to at least one communication node 203, and the at least one communication node 203 executes the communication behavior by executing the communication code block.
[0049] For each computing node 202, the computing node 202 can communicate with one or more communication nodes 203, which are used to send the execution results obtained by the computing node 202 in performing the computing behavior. Optionally, the computing node 202 and the one or more communication nodes 203 are different nodes, or the one or more communication nodes 203 are integrated on the computing node 202.
[0050] For each computational action, there is a corresponding communication action used to send the execution result obtained from the computational action. Management node 201 can first issue a computational action to at least one computational node 202, which then executes the computational action. After the start time of at least one computational node 202 executing the computational action, and before the end time of at least one computational node 202 completing the computational action, management node 201 can first issue a communication action to at least one communication node 203, which then executes the communication action to reduce the waiting time for other computational actions that require execution results.
[0051] In some embodiments, see Figure 3 The computing cluster 200 also includes multiple storage nodes 204. The management node 201 can communicate with each storage node 204 and can manage and / or control each storage node 204.
[0052] For at least one of multiple storage nodes 204, the communication memory access behavior corresponding to the computation behavior includes a memory access behavior. The management node 201 can issue a memory access behavior to at least one storage node 204, and the at least one storage node 204 executes the memory access behavior so that the at least one storage node 204 saves the execution result obtained in the process of executing the computation behavior. Optionally, the management node 201 can issue a memory access code block to at least one storage node 204 to implement the memory access behavior, and the at least one storage node 204 executes the memory access behavior by executing the memory access code block.
[0053] For each computational action, there is a corresponding memory access action, which is used to save the execution result obtained from the computational action. Management node 201 can first issue a computational action to at least one computational node 202, and at least one computational node 202 executes the computational action. After the start time of at least one computational node 202 executing its computational action, and before the end time of at least one computational node 202 completing its computational action, management node 201 can first issue a memory access action to at least one storage node 204, and at least one storage node 204 executes the memory access action, thereby reducing the waiting time for other computational actions that require execution results.
[0054] For each compute node 202, the compute node 202 can communicate with one or more storage nodes 204, which are used to store the execution results obtained by the compute node 202 in performing the computation. Optionally, the compute node 202 and the one or more storage nodes 204 are different nodes, or the one or more storage nodes 204 are integrated on the compute node 202.
[0055] See Figure 4 This application provides a method 400 for executing business code, the method 400 being applied to... Figure 2 or Figure 3 In the computing cluster 200 shown, the method 400 includes the following process.
[0056] Step 401: The management node obtains the business code to be executed. The business code includes a calculation tag for identifying the calculation code block and a communication memory access tag for identifying the communication memory access code block.
[0057] In some embodiments, the communication memory access tag used to identify the communication memory access code block may include a communication tag used to identify the communication code block, and / or a memory access tag used to identify the memory access code block.
[0058] The business logic code comprises multiple code blocks, each containing one or more code statements. These code blocks include computation code blocks and communication / memory access code blocks. The computation code blocks implement computational behavior, and the communication / memory access code blocks implement communication / memory access behavior. Optionally, the communication / memory access code blocks include communication code blocks and / or memory access code blocks, and the communication / memory access behavior includes communication behavior and / or memory access behavior. The communication code blocks implement communication behavior, and the memory access code blocks implement memory access behavior.
[0059] In step 401, the management node receives the business code to be executed and displays a skeleton interface. The skeleton interface includes the business code, allowing technicians to add markers to each code block included in the business code. For computation code blocks included in the business code, the markers added to the computation code blocks are computation markers used to identify the computation code blocks. For communication code blocks included in the business code, the markers added to the communication code blocks are communication markers used to identify the communication code blocks. For memory access code blocks included in the business code, the markers added to the memory access code blocks are memory access markers used to identify the memory access code blocks. Then, the management node retrieves the marked business code from the skeleton interface.
[0060] In some embodiments, the management node can use a hotspot function analysis tool to skeletonize the business code to identify the hotspot functions included in the business code, treat the hotspot functions as a code block, and mark the hotspot functions when displaying the business code in the skeleton interface, so as to facilitate technical personnel to identify the code block.
[0061] When branching modules (primarily if...else statements) exist in the business logic code, the hotspot function analysis tool uses calculation tags to mark them and indicate the probability of entering the branch. The code parsing module in the hotspot function analysis tool will then calculate the time cost of the branch based on the probability. When calculation modules (primarily for loops or matrix calculation formulas) exist in the business logic code, they are marked using calculation tags (such as compute).
[0062] For example, management nodes can be displayed in the skeleton interface as follows: Figure 5 The business code shown includes computation code blocks, memory access code blocks, and communication code blocks. Technical personnel can add a computation marker "compute" to the computation code blocks included in the skeleton interface, a memory access marker "malloc" to the memory access code blocks, and a communication marker "keep" to the communication code blocks. The management node then retrieves the marked business code from the skeleton interface.
[0063] Step 402: The management node divides the business code based on the computation tags and communication memory access tags included in the business code, resulting in multiple code blocks included in the business code, where the multiple code blocks in the business code include multiple computation code blocks and multiple communication memory access code blocks.
[0064] In some embodiments, for each computation code block, the computation code block may correspond to one or more of the plurality of communication memory access code blocks.
[0065] In some embodiments, the communication memory access code block corresponding to the computation code block includes a communication code block and / or a memory access code block, etc.
[0066] For example, for Figure 5 The business code shown includes a computation flag "compute", a memory access flag "malloc", and a communication flag "keep". The management node divides the business code based on these three flags, resulting in computation code blocks, memory access code blocks, and communication code blocks.
[0067] Step 403: The management node analyzes the multiple computation code blocks included in the business code to obtain the relationship between the multiple computation code blocks included in the business code.
[0068] For ease of explanation, any one of the multiple computation code blocks included in the business code is referred to as the first computation code block. At least one of the multiple code blocks included in the business code may have a dependency relationship with the first computation code block.
[0069] For example, suppose that among the multiple code blocks included in the business logic, there is a second computation code block that requires the execution result of the first computation code block. Then, the second computation code block is a code block that has a dependency relationship with the first computation code block. Also suppose that among the multiple code blocks included in the business logic, there is a third computation code block that requires the execution result of the second computation code block. Then, the third computation code block is a code block that has a dependency relationship with both the second and first computation code blocks.
[0070] The second computation code block may include a first call code statement for invoking the first computation code block. This first call code statement may be used to invoke and execute the first computation code block. Upon completion of the first computation code block, the execution result of the first computation code block is returned to the second computation code block, and then the second computation code block is executed based on that execution result.
[0071] Similarly, for the third computation code block, it may include a second calling code statement for invoking the second computation code block. This second calling code statement may be used to invoke and execute the second computation code block. Upon completion of the second computation code block, the execution result of the second computation code block is returned to the third computation code block, and then the third computation code block is executed based on that result.
[0072] In some embodiments, the operation of obtaining the relationship between the plurality of computation code blocks can be as follows: For the third computation code block, the management node detects whether the third computation code block includes a call statement for invoking the computation code block. If the third computation code block is detected to include a second call statement for invoking the second computation code block, it is determined that the third computation code block is a code block that has a dependency relationship with the second computation code block. Similarly, the management node detects whether the second computation code block includes a call statement for invoking the computation code block. If the second computation code block is detected to include a first call statement for invoking the first computation code block, it is determined that the second computation code block is a code block that has a dependency relationship with the first computation code block. Thus, the first computation code block, and the second and third computation code blocks that have dependencies on the first computation code block, are obtained.
[0073] In some embodiments, the abstract syntax tree (AST) of the business code includes the calling relationships between computation code blocks and other computation code blocks that the computation code block needs to call. The operation to obtain the relationships between these multiple computation code blocks can be as follows: the management node parses the business code to obtain the AST of the business code; for the third computation code block, it determines, based on the AST, the second computation code block that the third computation code block needs to call, and based on the AST, it determines the first computation code block that the second computation code block needs to call. Thus, the first computation code block, and the second and third computation code blocks that have dependencies on the first computation code block, are obtained.
[0074] In some embodiments, for any computation code block in the business code, the management node will also determine the communication memory access code block corresponding to the computation code block from the plurality of code blocks, that is, determine the communication code block and / or memory access code block corresponding to the computation code block.
[0075] For example, a first computation code block may include a call statement for invoking a first communication code block corresponding to the first computation code block and / or a call statement for invoking a first memory access code block corresponding to the first computation code block. Thus, the management node detects whether the first computation code block includes a call statement for invoking the first communication code block. If the management node detects that the first computation code block includes a call statement for invoking the first communication code block, it determines the first communication code block called by that call statement as the communication code block corresponding to the first computation code block. Similarly, the management node detects whether the first computation code block includes a call statement for invoking the first memory access code block. If the management node detects that the first computation code block includes a call statement for invoking the first memory access code block, it determines the first memory access code block called by that call statement as the memory access code block corresponding to the first computation code block. In this way, a first communication memory access code block corresponding to the first computation code block is obtained, and the first communication memory access code block includes the first communication code block and / or the first memory access code block. Similarly, the management node obtains the second communication code block and / or the second memory access code block corresponding to the second computation code block in the same way as described above, and obtains the third communication code block and / or the third memory access code block corresponding to the third computation code block, that is, it obtains the second communication memory access code block (including the second communication code block and / or the second memory access code block) corresponding to the second computation code block, and the third communication memory access code block (including the third communication code block and / or the third memory access code block) corresponding to the third computation code block.
[0076] In some embodiments, the Abstract Syntax Tree (AST) of the business code includes the calling relationship between a computation code block and a communication code block that the computation code block needs to call, and / or the calling relationship between the computation code block and a memory access code block that the computation code block needs to call. For a first computation code block, the management node determines the first communication code block that the first computation code block needs to call based on the AST of the business code, and / or determines the first memory access code block that the first computation code block needs to call based on the AST of the business code. Thus, the first communication memory access code block corresponding to the first computation code block is obtained, and the first communication memory access code block includes the first communication code block and / or the first memory access code block. Similarly, the management node obtains the second communication code block and / or the second memory access code block corresponding to the second computation code block in the same manner as described above, and obtains the third communication code block and / or the third memory access code block corresponding to the third computation code block, that is, it obtains the second communication memory access code block corresponding to the second computation code block (including the second communication code block and / or the second memory access code block), and the third communication memory access code block corresponding to the third computation code block (including the third communication code block and / or the third memory access code block).
[0077] Next, see Figure 6 For the first computation code block and the corresponding first communication memory access code block, the management node sends scheduling information to at least one first node in the computing cluster. The scheduling information includes the first computation code block and the first communication memory access code block. The scheduling information is used to instruct at least one first node to execute the first computation code block and the first communication memory access code block. The start time of executing the first communication memory access code block is later than the start time of executing the first computation code block, and the start time of executing the first communication memory access code block is earlier than the end time of executing the first computation code block.
[0078] The at least one first node includes at least one computing node, and may also include at least one communication node and / or at least one storage node. The scheduling information can be scheduling commands (such as the first scheduling command and / or the second scheduling command shown below), and the management node can send the scheduling information to the at least one first node through the following steps.
[0079] Step 404: The management node obtains a set of code blocks, which includes multiple computation code blocks and multiple communication memory access code blocks. The multiple computation code blocks include a first computation code block and at least one computation code block that has a dependency relationship with the first computation code block. The multiple communication memory access code blocks include a first communication memory access code block corresponding to the first computation code block and at least one communication memory access code block corresponding to the at least one computation code block.
[0080] For any computation code block included in the business code, i.e., for the first computation code block, at least one computation code block that has a dependency relationship with the first computation code block is obtained from the multiple code blocks included in the business code, and the multiple computation code blocks are obtained (such as including the first computation code block, the second computation code block and the third computation code block mentioned above).
[0081] The first communication memory access code block corresponding to the first calculation code block is obtained from the multiple code blocks included in the business code, and at least one communication memory access code block corresponding to the at least one calculation code block is obtained, resulting in multiple communication memory access code blocks (such as including the aforementioned first communication memory access code block, second communication memory access code block, and third memory access communication code block). The resulting set of code blocks includes the aforementioned first calculation code block, second calculation code block, third calculation code block, first communication memory access code block, second communication memory access code block, and third communication memory access code block.
[0082] In some embodiments, a management node can acquire multiple sets of code blocks. Computational code blocks and communication / memory access code blocks within the same set can form an event stream, where each code block is an event within that event stream. Different sets of code blocks constitute different event streams, and these event streams can be executed in parallel.
[0083] The following operations can be performed on each set of code blocks.
[0084] Step 405: The management node obtains the configuration information of the computing cluster, and based on the configuration information of the computing cluster, obtains the execution time of each computing code block and the execution time of each communication memory access code block in the code block set.
[0085] In some embodiments, the configuration information of the computing cluster includes one or more of the following: the network topology of the computing cluster, the number of computing nodes, the number of communication nodes, the number of storage nodes, configuration information of computing nodes, configuration information of communication nodes, or configuration information of storage nodes, etc.
[0086] Optionally, the network topology of the computing cluster is used to describe the connection relationships between multiple nodes in the computing cluster. For example, the network topology of the computing cluster is used to describe the connection relationships between multiple computing nodes, multiple communication nodes, and multiple storage nodes in the computing cluster.
[0087] Optionally, the configuration information of the computing node includes one or more of the following: the computing power of the computing node, the number of processor cores, or the processor clock speed. The computing node may be a processor or a computing device including a processor, which may be a graphics processing unit (GPU), a neural network processing unit (NPU), a central processing unit (CPU), or a field-programmable gate array (FPGA), etc.
[0088] Optionally, the configuration information of the communication node includes one or more of the following: latency, bandwidth, or network interface card (NIC) configuration information of the communication node. The communication node may be a switch corresponding to the computing node or a NIC integrated on the computing node.
[0089] Optionally, the configuration information of a storage node includes one or more of the following: memory configuration information, hard disk configuration information, or cache configuration information of the storage node. Optionally, the memory configuration information of the storage node includes the read / write bandwidth and / or read / write latency of the storage node's memory; the hard disk configuration information of the storage node includes the read / write bandwidth and / or read / write latency of the storage node's hard disk; and the cache configuration information of the storage node includes the read / write bandwidth and / or read / write latency of the storage node's cache.
[0090] In step 405, the management node can receive the configuration information of the computing cluster input by the technician, and determine the computing nodes, communication nodes and storage nodes included in the computing cluster based on the number of computing nodes, the number of communication nodes and the number of storage nodes included in the configuration information of the computing cluster.
[0091] For multiple computation code blocks in the code block set, one or more computation nodes are allocated to each computation code block based on the configuration information of the computation nodes. For multiple communication and memory access code blocks in the code block set, one or more communication nodes and / or one or more storage nodes are allocated to each communication and memory access code block based on the configuration information of the communication nodes, the configuration information of the storage nodes, and the network topology. Optionally, for each computation code block and its corresponding communication and memory access code block, there is a connection relationship between the one or more computation nodes allocated to the computation code block and the one or more communication nodes and / or one or more storage nodes allocated to the communication and memory access code block.
[0092] Based on the configuration information of the one or more computing nodes, the one or more communication nodes, the one or more storage nodes, and the dependencies between multiple computing code blocks in the code block set, the start and end execution times of each computing code block and each communication / memory access code block in the code block set are determined. Based on the start and end execution times of each computing code block, the execution duration of each computing code block and each communication / memory access code block is obtained.
[0093] In some embodiments, for each communication memory access code block, the communication memory access code block includes a communication code block and / or a memory access code block, and the execution time of the communication memory access code block is equal to the sum of the execution time of the communication code block and the execution time of the memory access code block.
[0094] Step 406: The management node determines the value of M based on the execution time of each computation code block and the execution time of each communication memory access code block. M is the number of sub-computation code blocks into which the first computation code block is divided, and M is an integer greater than or equal to 1.
[0095] In step 406, the management node determines the value of M based on the execution time of each computation code block and the execution time of each communication memory access code block using the following first formula.
[0096] The first formula is: M = Max{C1 / min(Cx), D1 / min(Dx), 1};
[0097] In the first formula, C1 is the execution time of the first computation code block, Cx is the execution time of the xth computation code block in the code block set, D1 is the execution time of the first communication memory access code block corresponding to the first computation code block, Dx is the execution time of the communication memory access code block corresponding to the xth computation code block in the code block set, x is an integer greater than 0, and x is less than or equal to the total number of computation code blocks included in the code block set.
[0098] min(Cx) represents the minimum execution time selected from the execution times of multiple computation code blocks in the code block set, and min(Dx) represents the minimum execution time selected from the execution times of multiple communication memory access code blocks in the code block set. For the three values C1 / min(Cx), D1 / min(Dx) and 1, M is the maximum value among the three values, and M is an integer greater than or equal to 1.
[0099] Since M is derived based on the execution time of each computation code block and the execution time of each communication / memory access code block, each computation code block includes a first computation code block and at least one computation code block that depends on the first computation code block. Each communication / memory access code block includes a communication / memory access code block corresponding to the first computation code block and at least one communication / memory access code block corresponding to at least one computation code block. This ensures that the value of M is adapted to the first computation code block. Therefore, decomposing the first computation code block and the first communication / memory access code block based on M optimizes the performance of the computation cluster in executing the first computation code block and the first communication / memory access code block.
[0100] For example, the code block set includes a first computation code block, a second computation code block, a third computation code block, a first communication memory access code block, a second communication memory access code block, and a third communication memory access code block. x can be equal to 1, 2, or 3. min(Cx) is the minimum execution time among the execution times of the first, second, and third computation code blocks. C1 / min(Cx) is the first ratio between the execution time of the first computation code block and this minimum execution time. min(Dx) is the minimum execution time among the execution times of the first, second, and third communication memory access code blocks. D1 / min(Dx) is the second ratio between the execution time of the first communication memory access code block and this minimum execution time. Therefore, M is the maximum value among the first ratio, the second ratio, and 1.
[0101] Step 407: The management node decomposes the first computation code block into M sub-computation code blocks and the first communication memory access code block into M sub-communication memory access code blocks, with each of the M sub-computation code blocks corresponding to the M sub-communication memory access code blocks.
[0102] In some embodiments, the execution time of each sub-computation code block may be equal, and the execution time of each sub-communication memory access code block may be equal.
[0103] For the second computation code block that requires the execution result of the first computation code block and the second communication memory access code block corresponding to the second computation code block, the number N of sub-blocks that need to be decomposed into the second computation code block can be obtained in step 406, where N is an integer greater than or equal to 1. Then, the second computation code block is decomposed into N sub-computation code blocks and the second communication memory access code block is decomposed into N sub-communication memory access code blocks in this step. The N sub-computation code blocks correspond one-to-one with the N sub-communication memory access code blocks.
[0104] For example, see Figure 7Assuming M=3, the management node decomposes the first computation code block into sub-computation code block 11, sub-computation code block 12, and sub-computation code block 13. It also decomposes the first communication memory access code block into sub-communication memory access code block 21, sub-communication memory access code block 22, and sub-communication memory access code block 23.
[0105] In some embodiments, the management node can obtain the execution duration of M sub-computation code blocks and the execution duration of M sub-communication and memory access code blocks. Based on the execution durations of the M sub-computation code blocks and the M sub-communication and memory access code blocks, it determines the start execution time of the M sub-computation code blocks and the start execution time of the M sub-communication and memory access code blocks. Based on the start execution times of the M sub-computation code blocks and the M sub-communication and memory access code blocks, the scheduling order corresponding to the first computation code block is obtained.
[0106] Optionally, see Figure 8 In the scheduling order shown, when the execution time of the y-th sub-computation code block is greater than or equal to the execution time of the y-th sub-communication memory access code block corresponding to the y-th sub-computation code block, the end execution time of the y-th sub-computation code block is the start execution time of the y-th sub-communication memory access code block, and also the start execution time of the (y+1)-th sub-computation code block, where y = 1, 2, ..., M-1. The end execution time of the M-th sub-computation code block is the start execution time of the M-th sub-communication memory access code block. The end execution time of the M-th sub-communication memory access code block is the start execution time of the 1st sub-computation code block in the second computation code block.
[0107] Optionally, see Figure 9 In the scheduling order shown, when the execution time of the z-th sub-computation code block is less than the execution time of the z-th sub-communication memory access code block corresponding to the z-th sub-computation code block, the end execution time of the first sub-computation code block is the start execution time of the first sub-communication memory access code block, and also the start execution time of the second sub-computation code block. The end execution time of the z-th sub-communication memory access code block is the start execution time of the z-th sub-communication memory access code block, and also the start execution time of the (z+1)-th sub-computation code block, where z = 2, ..., M-1. The end execution time of the (M-1)-th sub-communication memory access code block is the start execution time of the M-th sub-communication memory access code block. The end execution time of the M-th sub-communication memory access code block is the start execution time of the first sub-computation code block in the second computation code block.
[0108] In some embodiments, the first communication memory access code block includes a first communication code block, which is decomposed into M sub-communication code blocks, i.e., each sub-communication memory access code block includes one sub-communication code block. And / or, the first communication memory access code block includes a first memory access code block, which is decomposed into M sub-communication memory access code blocks, i.e., each sub-communication memory access code block includes one sub-communication code block and / or one sub-memory code block.
[0109] Step 408: The management node sends the first sub-computation code block to at least one first node. The at least one first node executes the first sub-computation code block. After the at least one first node finishes executing the i-th sub-computation code block, the management node sends a first scheduling command to the at least one first node. The first scheduling command includes the j-th sub-computation code block and the i-th sub-communication memory access code block corresponding to the i-th sub-computation code block, where i is a positive integer less than M and j = i + 1.
[0110] The first scheduling command is used to instruct the at least one first node to execute the j-th sub-computation code block and the i-th sub-communication memory access code block in parallel. The i-th sub-communication memory access code block is used to send the execution result obtained from executing the i-th sub-computation code block to at least one second node. The at least one second node is the node that needs to execute the second computation code block. The at least one first node and the at least one second node are nodes in the computing cluster.
[0111] In step 408, the management node sends the first sub-computation code block to at least one first node, which executes the first sub-computation code block. Then, after at least one first node has executed the i-th sub-computation code block, a first scheduling command is sent to at least one first node based on the scheduling order corresponding to the first computation code block. The first scheduling command includes the j-th sub-computation code block and the i-th sub-communication memory access code block corresponding to the i-th sub-computation code block, where i is a positive integer less than M, and j = i + 1.
[0112] In some embodiments, at least one first node includes at least one computing node allocated for the first computing code block, and at least one first node further includes at least one communication node and / or at least one storage node allocated for the first communication memory access code block. At least one computing node receives and executes the first sub-computation code block. After at least one computing node has executed the i-th sub-computation code block, the management node sends a first scheduling command including the j-th sub-computation code block to at least one computing node based on the scheduling order corresponding to the first computing code block, and sends a first scheduling command including the i-th sub-communication memory access code block corresponding to the i-th sub-computation code block to at least one communication node and / or at least one storage node.
[0113] At least one computing node receives a first scheduling command including the j-th sub-computation code block and executes the j-th sub-computation code block. At least one communication node and / or at least one storage node receives a first scheduling command including the ith sub-communication memory access code block, executes the ith sub-communication memory access code block to send the execution result obtained from executing the ith sub-computation code block to at least one second node, and / or saves the execution result obtained from executing the ith sub-computation code block.
[0114] Optionally, the operation of at least one communication node and / or at least one storage node receiving the first scheduling command can be as follows: If the i-th sub-communication memory access code block includes the i-th sub-communication code block, at least one communication node receives the first scheduling command including the i-th sub-communication code block, executes the i-th sub-communication code block, and sends the execution result obtained from executing the i-th sub-computation code block to at least one second node. And / or, if the i-th sub-communication memory access code block includes the i-th sub-memory access code block, at least one storage node receives the first scheduling command including the i-th sub-memory access code block, executes the i-th sub-memory access code block, and saves the execution result obtained from executing the i-th sub-computation code block.
[0115] In some embodiments, see Figure 8 If i is less than M-1, the time it takes for at least one first node to complete the execution of the i-th sub-communication memory access code block is earlier than or equal to the time it takes to complete the execution of the j-th sub-computation code block. In other words, the execution time required to execute the i-th sub-communication memory access code block is less than or equal to the execution time required to execute the j-th sub-computation code block.
[0116] In this scenario, the order in which the management node issues sub-computation code blocks and sub-communication memory access code blocks can be as follows: When at least one first node completes the execution of the first sub-computation code block and obtains its execution result (i=1, j=2), the management node sends a first scheduling command to at least one first node, including the second sub-computation code block and the first sub-communication memory access code block corresponding to the first sub-computation code block. This first scheduling command instructs the at least one first node to execute the second sub-computation code block and the first sub-communication memory access code block in parallel. The first sub-communication memory access code block is used to send the execution result obtained from executing the first sub-computation code block to at least one second node, and / or to save the execution result obtained from executing the first sub-computation code block.
[0117] When at least one first node has completed the execution of the j-th sub-computation code block, and i is an integer greater than 1 and less than M, then j = i + 1. The management node sends a second scheduling command to at least one first node. The second scheduling command includes the (j+1)-th sub-computation code block and the j-th sub-communication memory access code block corresponding to the j-th sub-computation code block. Specifically, the second scheduling command instructs at least one first node to execute the (j+1)-th sub-computation code block and the j-th sub-communication memory access code block in parallel. The j-th sub-communication memory access code block is used to send the execution result of the j-th sub-computation code block to at least one second node, and / or to save the execution result of the j-th sub-computation code block.
[0118] Optionally, see Figure 8 At least one first node includes at least one computing node, and at least one communication node and / or at least one storage node. The order in which the sub-computation code block and the sub-communication memory access code block are issued can be detailed as follows: the management node sends the first sub-computation code block to at least one computing node, and at least one computing node executes the first sub-computation code block.
[0119] When at least one computing node completes the execution of the first sub-computation code block (i=1, j=2), the management node sends a first scheduling command including the second sub-computation code block to at least one computing node, and a first scheduling command including the first sub-communication memory access code block corresponding to the first sub-computation code block to at least one communication node and / or at least one storage node. At least one computing node executes the second sub-computation code block, and at least one communication node and / or at least one storage node executes the first sub-communication memory access code block, to send the execution result obtained from executing the first sub-computation code block to at least one second node, and / or to save the execution result obtained from executing the first sub-computation code block.
[0120] When at least one computing node completes the execution of the j-th sub-computation code block (where i is an integer greater than 1 and less than M, and j = i + 1), the management node sends a second scheduling command including the (j+1)-th sub-computation code block to at least one computing node, and a second scheduling command including the j-th sub-communication memory access code block to at least one communication node and / or at least one storage node. At least one computing node receives the second scheduling command including the (j+1)-th sub-computation code block and executes it. At least one communication node and / or at least one storage node receives the second scheduling command including the j-th sub-communication memory access code block, executes it, and sends the execution result obtained from executing the j-th sub-computation code block to at least one second node, and / or saves the execution result obtained from executing the j-th sub-computation code block.
[0121] For the Mth sub-communication memory access code block, the management node sends a second scheduling command, including the Mth sub-communication memory access code block, to at least one communication node and / or at least one storage node after at least one computing node has completed executing the Mth sub-computation code block. Upon receiving the second scheduling command including the Mth sub-communication memory access code block, at least one communication node and / or at least one storage node executes the Mth sub-communication memory access code block to send the execution result obtained from executing the Mth sub-computation code block to at least one second node, and / or saves the execution result obtained from executing the Mth sub-computation code block.
[0122] When at least one communication node and / or at least one storage node has completed executing the Mth sub-communication memory access code block, the management node executes the second computation code block and the second communication memory access code block in the same manner as the first computation code block and the first communication memory access code block described above. That is, for the first sub-computation code block included in the second computation code block, when at least one communication node and / or at least one storage node has completed executing the Mth sub-communication memory access code block in the first communication memory access code block, the management node sends the first sub-computation code block included in the second computation code block to at least one second node (the computation node within the at least one second node). For example... Figure 8 As shown, in this case, the waiting time of the second calculation code block needs to be equal to the execution time of the Mth sub-communication memory access code block, thereby reducing the waiting time of the second calculation code block.
[0123] In some embodiments, see Figure 9 If i is less than M-1, at least the time it takes for the first node to complete the i-th sub-communication memory access code block is later than the time it takes to complete the j-th sub-computation code block. In other words, the execution time required to execute the i-th sub-communication memory access code block is greater than the execution time required to execute the j-th sub-computation code block.
[0124] In this scenario, the order in which the management node issues sub-computation code blocks and sub-communication memory access code blocks can be as follows: When at least one first node completes the execution of the first sub-computation code block and obtains its execution result (i=1, j=2), the management node sends a first scheduling command to at least one first node, including the second sub-computation code block and the first sub-communication memory access code block corresponding to the first sub-computation code block. The first scheduling command instructs at least one first node to execute the second sub-computation code block and the first sub-communication memory access code block in parallel. The first sub-communication memory access code block is used to send the execution result obtained from executing the first sub-computation code block to at least one second node, and / or to save the execution result obtained from executing the first sub-computation code block.
[0125] When at least one first node finishes executing the i-th sub-communication memory access code block (where i is an integer greater than 1 and less than M, j = i + 1), the management node sends a second scheduling command to at least one first node. The second scheduling command includes the (j+1)-th sub-computation code block and the j-th sub-communication memory access code block corresponding to the j-th sub-computation code block. Specifically, the second scheduling command instructs at least one first node to execute the (j+1)-th sub-computation code block and the j-th sub-communication memory access code block in parallel. The j-th sub-communication memory access code block is used to send the execution result of the j-th sub-computation code block to at least one second node, and / or to save the execution result of the j-th sub-computation code block.
[0126] Optionally, see Figure 9 At least one first node includes at least one computing node, and at least one communication node and / or at least one storage node. The order in which the sub-computation code block and the sub-communication memory access code block are issued can be detailed as follows: the management node sends the first sub-computation code block to at least one computing node, and at least one computing node executes the first sub-computation code block.
[0127] When at least one computing node completes the execution of the first sub-computation code block (i=1, j=2), the management node sends a first scheduling command including the second sub-computation code block to at least one computing node, and a first scheduling command including the first sub-communication memory access code block corresponding to the first sub-computation code block to at least one communication node and / or at least one storage node. At least one computing node executes the second sub-computation code block, and at least one communication node and / or at least one storage node executes the first sub-communication memory access code block, to send the execution result obtained from executing the first sub-computation code block to at least one second node, and / or to save the execution result obtained from executing the first sub-computation code block.
[0128] When at least one communication node and / or at least one storage node completes the execution of the i-th sub-communication memory access code block (where i is an integer greater than 1 and less than M, and j = i + 1), the management node sends a second scheduling command including the (j+1)-th sub-computation code block to at least one computing node, and also sends a second scheduling command including the j-th sub-communication memory access code block to at least one communication node and / or at least one storage node. At least one computing node receives the second scheduling command including the (j+1)-th sub-computation code block and executes the (j+1)-th sub-computation code block. At least one communication node and / or at least one storage node receives the second scheduling command including the j-th sub-communication memory access code block, executes the j-th sub-communication memory access code block, and sends the execution result obtained from executing the j-th sub-computation code block to at least one second node, and / or saves the execution result obtained from executing the j-th sub-computation code block.
[0129] For the Mth sub-communication memory access code block, when at least one communication node and / or at least one storage node has completed executing the (M-1)th sub-communication memory access code block, the management node sends a second scheduling command including the Mth sub-communication memory access code block to at least one communication node and / or at least one storage node. Upon receiving the second scheduling command including the Mth sub-communication memory access code block, at least one communication node and / or at least one storage node executes the Mth sub-communication memory access code block to send the execution result obtained from executing the Mth sub-computation code block to at least one second node, and / or saves the execution result obtained from executing the Mth sub-computation code block.
[0130] When at least one communication node and / or at least one storage node has completed executing the Mth sub-communication memory access code block, the management node executes the second computation code block and the second communication memory access code block in the same manner as the first computation code block and the first communication memory access code block described above. That is, for the first sub-computation code block included in the second computation code block, when at least one communication node and / or at least one storage node has completed executing the Mth sub-communication memory access code block, the management node sends the first sub-computation code block included in the second computation code block to at least one second node (the computation node within the at least one second node). For example... Figure 9 As shown, in this case, the waiting time of the second computation code block needs to be equal to the time between the end execution time of the Mth sub-computation code block and the end execution time of the Mth sub-communication memory access code block, thereby reducing the waiting time of the second computation code block.
[0131] In this embodiment, the management node divides the business code into multiple code blocks, including a first computation code block and a corresponding first communication memory access code block. The first computation code block is further decomposed into M sub-computation code blocks, and the first communication memory access code block is also decomposed into M sub-communication memory access code blocks. The management node sends the first sub-computation code block to at least one first node, which executes the first sub-computation code block. After at least one first node completes the execution of the i-th sub-computation code block, a first scheduling command is sent to at least one first node. The first scheduling command includes the j-th sub-computation code block and the i-th corresponding sub-communication memory access code block, where i is a positive integer less than M, and j = i + 1. The first scheduling command instructs at least one first node to execute the j-th sub-computation code block and the i-th sub-communication memory access code block in parallel. The i-th sub-communication memory access code block sends the execution result obtained from executing the i-th sub-computation code block to at least one second node. By using the computation behavior corresponding to the j-th sub-computation code block to mask the communication and memory access behavior corresponding to the i-th sub-communication and memory access code block, the waiting time of the second computation code block that requires the execution result of the first computation code block is reduced, thereby improving the efficiency of executing business code and reducing the wasted time of computing resources in the computing cluster.
[0132] See Figure 10 This application provides an apparatus 1000 for executing business code, which can be deployed in... Figure 2 or Figure 3 On the management node of the computing cluster 200 shown, or deployed on Figure 4 On the management node of method 400 shown. The device 1000 is used to execute business code, which includes a calculation tag for identifying a calculation code block and a communication memory access tag for identifying a communication memory access code block. The device 1000 includes:
[0133] The processing unit 1001 is used to divide the business code based on the calculation mark and communication memory access mark included in the business code, to obtain a first calculation code block, a second calculation code block and a communication memory access code block corresponding to the first calculation code block included in the business code, wherein the second calculation code block is a code block that needs the execution result of the first calculation code block;
[0134] The sending unit 1002 is configured to send scheduling information to at least one first node. The scheduling information includes a first computation code block and the communication memory access code block. The scheduling information is used to instruct at least one first node to execute the first computation code block and the communication memory access code block. The start time of executing the communication memory access code block is later than the start time of executing the first computation code block, and the start time of executing the communication memory access code block is earlier than the end time of executing the first computation code block.
[0135] The communication memory access code block is used to send the execution result obtained during the execution of the first computation code block to at least one second node. The at least one second node is a node that needs to execute the second computation code block. The at least one first node and the at least one second node are nodes in the computing cluster.
[0136] Optionally, for the detailed implementation process of the processing unit 1001 obtaining the first calculation code block, the second calculation code block, and the communication memory access code block corresponding to the first calculation code block, see [link to relevant documentation]. Figure 4 The relevant content in step 402 of method 400 shown will not be described in detail here.
[0137] Optionally, for details of the implementation process of sending scheduling information by the sending unit 1002 to at least one first node, please refer to [link to relevant documentation]. Figure 4 The relevant content in step 408 of method 400 shown will not be described in detail here.
[0138] Optionally, the processing unit 1001 is further configured to decompose the first computation code block into M sub-computation code blocks and decompose the communication memory access code block into M sub-communication memory access code blocks, where M is an integer greater than 1, and the M sub-computation code blocks correspond to the M sub-communication memory access code blocks.
[0139] The sending unit 1002 is used to send the first sub-computation code block to at least one first node, and the at least one first node is used to execute the first sub-computation code block; after the at least one first node has executed the i-th sub-computation code block, a first scheduling command is sent to the at least one first node, the first scheduling command includes the j-th sub-computation code block and the i-th sub-communication memory access code block corresponding to the i-th sub-computation code block, where i is a positive integer less than M, and j = i + 1;
[0140] The first scheduling command is used to instruct at least one first node to execute the j-th sub-computation code block and the i-th sub-communication memory access code block in parallel. The i-th sub-communication memory access code block is used to send the execution result obtained from executing the i-th sub-computation code block to at least one second node.
[0141] Optionally, the processing unit 1001 decomposes the first computation code block into M sub-computation code blocks, and the detailed implementation process of decomposing the communication memory access code block into M sub-communication memory access code blocks can be found in [link to documentation]. Figure 4 The relevant content in step 404 of method 400 shown will not be described in detail here.
[0142] Optionally, the sending unit 1002 sends the first sub-computation code block to at least one first node. For details on the implementation of sending the first scheduling command to at least one first node, please refer to [link to relevant documentation]. Figure 4 The relevant content in step 408 of method 400 shown will not be described in detail here.
[0143] Optionally, at least one first node includes at least one computing node and at least one communication node, and the first scheduling command is used to simultaneously instruct at least one computing node to execute the j-th sub-computation code block and at least one communication node to execute the i-th sub-communication memory access code block.
[0144] Optionally, where i is less than M-1, and at least one first node completes the execution of the i-th sub-communication memory access code block earlier than or equal to the execution of the j-th sub-computation code block, the sending unit 1002 is further configured to:
[0145] When at least one first node has finished executing the j-th sub-computation code block, a second scheduling command is sent to at least one first node. The second scheduling command includes the (j+1)-th sub-computation code block and the j-th sub-communication memory access code block corresponding to the j-th sub-computation code block.
[0146] The second scheduling command is used to instruct at least one first node to execute the (j+1)th sub-computation code block and the jth sub-communication memory access code block in parallel. The jth sub-communication memory access code block is used to send the execution result obtained from executing the jth sub-computation code block to at least one second node.
[0147] Optionally, the sending unit 1002 sends the first sub-computation code block to at least one first node, and the detailed implementation process of sending the second scheduling command to at least one first node is described in [link to documentation]. Figure 4 The relevant content in step 408 of method 400 shown will not be described in detail here.
[0148] Optionally, where i is less than M-1, and at least one first node completes the execution of the i-th sub-communication memory access code block later than the execution of the j-th sub-computation code block, the sending unit 1002 is further configured to:
[0149] When at least one first node has completed the execution of the i-th sub-communication memory access code block, a second scheduling command is sent to at least one first node. The second scheduling command includes the j+1-th sub-computation code block and the j-th sub-communication memory access code block corresponding to the j-th sub-computation code block.
[0150] The second scheduling command is used to instruct at least one first node to execute the (j+1)th sub-computation code block and the jth sub-communication memory access code block in parallel. The jth sub-communication memory access code block is used to send the execution result obtained from executing the jth sub-computation code block to at least one second node.
[0151] Optionally, the i-th sub-communication memory access code block is also used to save the execution result obtained from executing the i-th sub-computation code block to the storage nodes included in the computing cluster.
[0152] Optionally, the processing unit 1001 is further configured to:
[0153] Obtain multiple computation code blocks and multiple communication memory access code blocks. The multiple computation code blocks include a first computation code block and at least one computation code block that has a dependency relationship with the first computation code block. The multiple communication memory access code blocks include a communication memory access code block corresponding to the first computation code block and at least one communication memory access code block corresponding to at least one computation code block.
[0154] Based on the configuration information of the computing cluster, the execution time of each computing code block in multiple computing code blocks and the execution time of each communication memory access code block in multiple communication memory access code blocks are obtained.
[0155] The value of M is determined based on the execution time of each computation code block and the execution time of each communication memory access code block.
[0156] Optionally, for details on the implementation process of the processing unit 1001 acquiring multiple computation code blocks and multiple communication memory access code blocks, please refer to [link to relevant documentation]. Figure 4 The relevant content in step 404 of method 400 shown will not be described in detail here.
[0157] Optionally, for details on how processing unit 1001 obtains the execution time of each computation code block and the execution time of each communication memory access code block from the multiple computation code blocks, please refer to [link to relevant documentation]. Figure 4 The relevant content in step 405 of method 400 shown will not be described in detail here.
[0158] Optionally, the processing unit 1001 determines the value of M based on the execution time of each computation code block and the execution time of each communication memory access code block. For details on this process, please refer to [link to relevant documentation]. Figure 4 The relevant content in step 406 of method 400 shown will not be described in detail here.
[0159] In this embodiment, because the scheduling information sent by the sending unit indicates that the start time of executing the communication memory access code block by at least one first node is later than the start time of executing the first computation code block, and the start time of executing the communication memory access code block is earlier than the end time of executing the first computation code block, at least one first node will simultaneously execute the first communication memory access code block while executing the first computation code block. This reduces the waiting time of the second computation code block and improves the efficiency of executing business code. Because the waiting time of the second computation code block is reduced, the idle time of at least one second node is also reduced, avoiding prolonged waste of resources by at least one second node and improving the computational performance of the computing cluster in executing business code.
[0160] See Figure 11 This application provides a computing device 1100. For example, the computing device 1100 may be... Figure 2 or Figure 3 The management node in the computing cluster 200 shown, or the computing device 1100, may be... Figure 4 The management node, etc. in method 400 shown.
[0161] like Figure 11 As shown, the computing device 1100 includes a bus 1102, a processor 1104, a memory 1106, and a communication interface 1108. The processor 1104, the memory 1106, and the communication interface 1108 communicate with each other via the bus 1102. The computing device 1100 can be a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 1100.
[0162] Bus 1102 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 11 The bus 1102 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 1102 may include a path for transmitting information between various components of the computing device 1100 (e.g., processor 1104, memory 1106, communication interface 1108).
[0163] The processor 1104 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0164] The memory 1106 may include volatile memory, such as random access memory (RAM). The memory 1106 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0165] See Figure 11 The memory 1106 stores executable program code, and the processor 1104 executes the executable program code to implement the following respectively. Figure 10 The processing unit 1001 and the sending unit 1002 in the illustrated device 1000 perform their functions to implement the method provided in any of the above embodiments. That is, the memory 1106 stores instructions for executing the method provided in any of the above embodiments. Alternatively,
[0166] The communication interface 1108 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 1100 and other devices or communication networks.
[0167] This application also provides a cluster for executing business code. The cluster for executing business code includes at least one computing device. This computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smartphone.
[0168] like Figure 12 As shown, the cluster executing the business code includes at least one computing device 1100. The memory 1106 of one or more computing devices 1100 in the cluster executing the business code may store the same instructions for executing the methods provided in any of the above embodiments.
[0169] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the cluster executing the business code may also store partial instructions for executing the method of executing the business code described above. In other words, a combination of one or more computing devices 1100 can jointly execute instructions for executing the method provided in any of the above embodiments.
[0170] In some possible implementations, one or more computing devices in the cluster executing business code can be connected via a network. This network can be a wide area network (WAN), a local area network (LAN), or similar. Figure 13 One possible implementation is shown. For example... Figure 13 As shown, the two computing devices 1100A and 1100B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device.
[0171] In some possible implementations, the memory 1106 in the computing device 1100A stores the execution of, for example Figure 10 Instructions for the function of the processing unit 1001 in the illustrated embodiment. Meanwhile, the memory 1106 in the computing device 1100B stores instructions for executing such... Figure 10 Instructions for the function of the sending unit 1002 in the illustrated embodiment.
[0172] It should be understood that Figure 13 The functions of computing device 1100A shown can also be performed by multiple computing devices 1100. Similarly, the functions of computing device 1100B can also be performed by multiple computing devices 1100.
[0173] This application also provides another type of cluster for executing business code. The connection relationships between the computing devices in the cluster for executing business code can be similarly referenced. Figure 13 The connection method of the cluster executing the business code is different. In addition, the memory 1106 of one or more computing devices 1100 in the cluster executing the business code can store the same instructions for executing the methods provided in any of the above embodiments.
[0174] In some possible implementations, the memory 1106 of one or more computing devices 1100 in the cluster executing the business code may also store partial instructions for executing the methods provided in any of the above embodiments. In other words, a combination of one or more computing devices 1100 can jointly execute instructions for executing the methods provided in any of the above embodiments.
[0175] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform the methods provided in any of the above embodiments.
[0176] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct a computing device to perform the method provided in any of the above embodiments.
[0177] All information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in this application are authorized by the user or fully authorized by all parties, and the collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0178] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0179] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A method for executing business code, characterized in that, The method is used to execute business code, which includes a computation tag for identifying computation code blocks and a communication memory access tag for identifying communication memory access code blocks, including: The business code is divided based on the calculation tags and communication memory access tags included in the business code to obtain a first calculation code block, a second calculation code block, and a communication memory access code block corresponding to the first calculation code block. The second calculation code block is a code block that requires the execution result of the first calculation code block. Send scheduling information to at least one first node. The scheduling information includes the first computation code block and the communication memory access code block. The scheduling information is used to instruct the at least one first node to execute the first computation code block and the communication memory access code block. The start time of executing the communication memory access code block is later than the start time of executing the first computation code block, and the start time of executing the communication memory access code block is earlier than the end time of executing the first computation code block. The communication memory access code block is used to send the execution result obtained during the execution of the first computation code block to at least one second node. The at least one second node is a node that needs to execute the second computation code block. The at least one first node and the at least one second node are nodes in the computing cluster.
2. The method as described in claim 1, characterized in that, Sending scheduling information to at least one first node includes: The first computation code block is decomposed into M sub-computation code blocks, and the communication memory access code block is decomposed into M sub-communication memory access code blocks, where M is an integer greater than 1, and the M sub-computation code blocks correspond to the M sub-communication memory access code blocks; Send a first sub-computation code block to the at least one first node, wherein the at least one first node is used to execute the first sub-computation code block; After the at least one first node has completed the execution of the i-th sub-computation code block, a first scheduling command is sent to the at least one first node. The first scheduling command includes the j-th sub-computation code block and the i-th sub-communication memory access code block corresponding to the i-th sub-computation code block, where i is a positive integer less than M and j = i + 1. The first scheduling command is used to instruct the at least one first node to execute the j-th sub-computation code block and the i-th sub-communication memory access code block in parallel, and the i-th sub-communication memory access code block is used to send the execution result obtained from executing the i-th sub-computation code block to at least one second node.
3. The method as described in claim 2, characterized in that, The at least one first node includes at least one computing node and at least one communication node, and the first scheduling command is used to simultaneously instruct the at least one computing node to execute the j-th sub-computation code block and the at least one communication node to execute the i-th sub-communication memory access code block.
4. The method as described in claim 2 or 3, characterized in that, If i is less than M-1, and the time when the at least one first node completes the execution of the i-th sub-communication memory access code block is earlier than or equal to the time when it completes the execution of the j-th sub-computation code block, then sending scheduling information to the at least one first node further includes: When the at least one first node finishes executing the j-th sub-computation code block, a second scheduling command is sent to the at least one first node. The second scheduling command includes the (j+1)-th sub-computation code block and the j-th sub-communication memory access code block corresponding to the j-th sub-computation code block. The second scheduling command is used to instruct the at least one first node to execute the (j+1)th sub-computation code block and the jth sub-communication memory access code block in parallel, and the jth sub-communication memory access code block is used to send the execution result obtained from executing the jth sub-computation code block to the at least one second node.
5. The method as described in claim 2 or 3, characterized in that, If i is less than M-1, and the execution time of the at least one first node of the i-th sub-communication memory access code block is later than the execution time of the j-th sub-computation code block, the method further includes: When the at least one first node finishes executing the i-th sub-communication memory access code block, a second scheduling command is sent to the at least one first node. The second scheduling command includes the (j+1)-th sub-computation code block and the j-th sub-communication memory access code block corresponding to the j-th sub-computation code block. The second scheduling command is used to instruct the at least one first node to execute the (j+1)th sub-computation code block and the jth sub-communication memory access code block in parallel, and the jth sub-communication memory access code block is used to send the execution result obtained from executing the jth sub-computation code block to the at least one second node.
6. The method according to any one of claims 2-5, characterized in that, The i-th sub-communication memory access code block is also used to save the execution result obtained by executing the i-th sub-computation code block to the storage nodes included in the computing cluster.
7. The method according to any one of claims 2-6, characterized in that, Before decomposing the first computation code block into M sub-computation code blocks and the communication memory access code block into M sub-communication memory access code blocks, the method further includes: Obtain multiple computation code blocks and multiple communication memory access code blocks. The multiple computation code blocks include the first computation code block and at least one computation code block that has a dependency relationship with the first computation code block. The multiple communication memory access code blocks include the communication memory access code block corresponding to the first computation code block and at least one communication memory access code block corresponding to the at least one computation code block. Based on the configuration information of the computing cluster, the execution time of each computing code block and the execution time of each communication memory access code block in the plurality of computing code blocks are obtained. The value of M is determined based on the execution time of each computation code block and the execution time of each communication memory access code block.
8. An apparatus for executing business code, characterized in that, The apparatus is used to execute business code, the business code including a calculation tag for identifying a calculation code block and a communication memory access tag for identifying a communication memory access code block, the apparatus comprising: The processing unit is configured to divide the business code based on the calculation mark and communication memory access mark included in the business code, to obtain a first calculation code block, a second calculation code block and a communication memory access code block corresponding to the first calculation code block, wherein the second calculation code block is a code block that requires the execution result of the first calculation code block; A sending unit is configured to send scheduling information to at least one first node. The scheduling information includes the first computation code block and the communication memory access code block. The scheduling information is configured to instruct the at least one first node to execute the first computation code block and the communication memory access code block, wherein the start time of executing the communication memory access code block is later than the start time of executing the first computation code block, and the start time of executing the communication memory access code block is earlier than the end time of executing the first computation code block. The communication memory access code block is used to send the execution result obtained during the execution of the first computation code block to at least one second node. The at least one second node is a node that needs to execute the second computation code block. The at least one first node and the at least one second node are nodes in the computing cluster.
9. The apparatus as claimed in claim 8, characterized in that, The processing unit is further configured to decompose the first computation code block into M sub-computation code blocks and decompose the communication memory access code block into M sub-communication memory access code blocks, where M is an integer greater than 1, and the M sub-computation code blocks correspond to the M sub-communication memory access code blocks. The sending unit is configured to send a first sub-computation code block to the at least one first node, and the at least one first node is configured to execute the first sub-computation code block; after the at least one first node has executed the i-th sub-computation code block, a first scheduling command is sent to the at least one first node, the first scheduling command including the j-th sub-computation code block and the i-th sub-communication memory access code block corresponding to the i-th sub-computation code block, where i is a positive integer less than M, and j = i + 1; The first scheduling command is used to instruct the at least one first node to execute the j-th sub-computation code block and the i-th sub-communication memory access code block in parallel, and the i-th sub-communication memory access code block is used to send the execution result obtained from executing the i-th sub-computation code block to at least one second node.
10. The apparatus as claimed in claim 9, characterized in that, The at least one first node includes at least one computing node and at least one communication node, and the first scheduling command is used to simultaneously instruct the at least one computing node to execute the j-th sub-computation code block and the at least one communication node to execute the i-th sub-communication memory access code block.
11. The apparatus as claimed in claim 9 or 10, characterized in that, If i is less than M-1, and the time when the at least one first node completes the execution of the i-th sub-communication memory access code block is earlier than or equal to the time when it completes the execution of the j-th sub-computation code block, the sending unit is further configured to: When the at least one first node finishes executing the j-th sub-computation code block, a second scheduling command is sent to the at least one first node. The second scheduling command includes the (j+1)-th sub-computation code block and the j-th sub-communication memory access code block corresponding to the j-th sub-computation code block. The second scheduling command is used to instruct the at least one first node to execute the (j+1)th sub-computation code block and the jth sub-communication memory access code block in parallel, and the jth sub-communication memory access code block is used to send the execution result obtained from executing the jth sub-computation code block to the at least one second node.
12. The apparatus as claimed in claim 9 or 10, characterized in that, If i is less than M-1, and the execution time of the at least one first node of the i-th sub-communication memory access code block is later than the execution time of the j-th sub-computation code block, the sending unit is further configured to: When the at least one first node finishes executing the i-th sub-communication memory access code block, a second scheduling command is sent to the at least one first node. The second scheduling command includes the (j+1)-th sub-computation code block and the j-th sub-communication memory access code block corresponding to the j-th sub-computation code block. The second scheduling command is used to instruct the at least one first node to execute the (j+1)th sub-computation code block and the jth sub-communication memory access code block in parallel, and the jth sub-communication memory access code block is used to send the execution result obtained from executing the jth sub-computation code block to the at least one second node.
13. The apparatus according to any one of claims 9-12, characterized in that, The i-th sub-communication memory access code block is also used to save the execution result obtained by executing the i-th sub-computation code block to the storage nodes included in the computing cluster.
14. The apparatus according to any one of claims 9-13, characterized in that, The processing unit is further configured to: Obtain multiple computation code blocks and multiple communication memory access code blocks. The multiple computation code blocks include the first computation code block and at least one computation code block that has a dependency relationship with the first computation code block. The multiple communication memory access code blocks include the communication memory access code block corresponding to the first computation code block and at least one communication memory access code block corresponding to the at least one computation code block. Based on the configuration information of the computing cluster, the execution time of each computing code block and the execution time of each communication memory access code block in the plurality of computing code blocks are obtained. The value of M is determined based on the execution time of each computation code block and the execution time of each communication memory access code block.
15. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-7.
16. A computer-readable storage medium, characterized in that, Includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1-7.
17. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster performs the method as described in any one of claims 1-7.