Distributed task migration method with consistent computing node states, I / O device control method, electronic device, and readable storage medium
By transmitting static and dynamic state information of tasks between computing nodes and restoring the task state on the target node, the problem of inconsistent state during task migration is solved, the resource utilization and execution efficiency of the computing node cluster are improved, and the decoupling of tasks from I/O devices is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTHWESTERN POLYTECHNICAL UNIV
- Filing Date
- 2023-01-06
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies fail to effectively guarantee the consistency of states before and after task migration, resulting in limited reliability and efficiency of task migration between computing nodes.
By transmitting static and dynamic status information of a task from the source computing node to the target computing node, the target computing node restores the task status based on this information, enabling the migration task to execute normally on the target node. Furthermore, by controlling I/O devices through a shared information channel, the task and I/O are decoupled.
It improves the resource utilization and execution efficiency of the computing node cluster, enhances computing power and reliability, and enables seamless migration of tasks between different computing nodes and decoupling of I/O device control.
Smart Images

Figure CN116521357B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of cross-platform computing node task migration, and relates to a distributed task migration method with consistent computing node states. Background Technology
[0002] With the widespread expansion of distributed networks, virtualization technology has developed rapidly. Task migration, a key feature of virtualization, refers to the movement of an executing task from one computing node in a distributed computing system to another. Tasks may require resources unavailable at their current location; task migration involves reassigning these tasks to resource-rich computing nodes within the network before execution. The consistency issue in task migration is crucial for achieving effective and correct migration. Therefore, for cross-platform task migration, a migration method that maintains state consistency before and after migration needs to be researched.
[0003] Currently, research on mission migration mainly focuses on resource allocation and performance consumption during the mission migration process, with cost reduction as an effective objective. The paper "MEC Service Management and Mission Migration Optimization Method in LEO Satellite Networks, Liu Zhihui, Wei Huan, Yin Jie, Wang Junyi, Jin Shichao, Dong Tao, Space-Ground Integrated Information Network, 2016, Vol. 3, No. 3" studies the MEC service management model of LEO satellite networks to support the implementation and execution of remote sensing satellite mission offloading to communication satellite constellations, achieving efficient collaboration between remote sensing and communication satellites. The paper "Algorithm for Mission Migration and Resource Allocation in Low-Earth Orbit Satellite Collaborative Edge Computing, Song Zhengyu, Hao Yuanyuan, Sun Xin, Acta Electronica Sinica, 2022, Vol. 50, No. 3" studies the mission migration and resource allocation problem of low-Earth orbit satellite collaborative edge computing based on inter-satellite links, providing edge computing services to users in remote areas. It adopts a partial mission migration mechanism, establishes an optimization problem with the goal of minimizing the weighted total energy consumption of ground users, and proposes an algorithm for mission migration and resource allocation in low-Earth orbit satellite collaborative edge computing. The above methods mainly consider indicators such as total energy consumption and processing cost of task migration, but do not consider the issue of ensuring consistency of state before and after task migration. Summary of the Invention
[0004] To address the issue of consistency in task migration states across different computing nodes in a computing system, in a first aspect, a distributed task migration method for ensuring consistent computing node states, according to some embodiments of this application, is used for the source computing node, including...
[0005] The source compute node sends a task migration request to the target compute node;
[0006] The source computing node receives the response information sent by the target computing node;
[0007] Task management agent of source compute node Identify the target computing node to be migrated;
[0008] The source compute node sends static status information of the tasks to be migrated to the target compute node, whereby the static status information is the task management agent of the target compute node. The information upon which the migration task is created on the target computing node to be migrated;
[0009] The source compute node sends dynamic status information of the task to be migrated to the target compute node, and the dynamic status information is the task management agent of the target compute node. The information on which the created migration task is based to restore the target computing node to its state before migration.
[0010] According to some embodiments of the distributed task migration method for ensuring consistent computing node states in this application, in the step of the source computing node issuing a task migration request to the target computing node, the task migration request is made by the task management agent of the source computing node. The task management agent of the source computing node was detected. The task queue contains tasks awaiting migration from the task management agent of the source compute node. issue;
[0011] In the step of the source computing node receiving response information from the target computing node, the response information is the response information sent by the target computing node to the source computing node in response to the task migration request, and the response information includes the load information of the target computing node.
[0012] Task management agent of source compute node In the step of determining the target compute node to be migrated, the target compute node to be migrated is the task management agent of the source compute node. The target computing node with the minimum load determined based on the load information in the response information;
[0013] In the step of the source compute node sending the static status information of the task to be migrated to the target compute node, the static status information is provided by the task management agent of the source compute node. The task to be migrated is obtained and sent to the target computing node to be migrated;
[0014] In the step of the source compute node sending the dynamic status information of the task to be migrated to the target compute node, the dynamic status information is provided by the task management agent of the source compute node. The task to be migrated is obtained and sent to the target computing node to be migrated.
[0015] In a second aspect, a distributed task migration method for ensuring consistent computing node states according to some embodiments of this application is used for a target computing node, including...
[0016] The target compute node receives a task migration request from the source compute node;
[0017] The target computing node sends a response message to the source computing node;
[0018] The target compute node to be migrated receives the static status information of the tasks to be migrated, and the task management agent of the target compute node to be migrated... A migration task is created on the target computing node to be migrated based on the static status information.
[0019] Task management agent for the target compute node to be migrated Receive the dynamic status information of the task to be migrated, and the task management agent of the target computing node to be migrated Based on the dynamic status information of the task to be migrated, the created migration task is restored to its pre-migration state on the target computing node to be migrated.
[0020] According to some embodiments of the distributed task migration method for ensuring consistent computing node states in this application, in the step of the target computing node receiving a task migration request issued by the source computing node, the task migration request is made by the task management agent of the source computing node. The task management agent TMAS of the source compute node was detected to have tasks awaiting migration in its task queue. The task migration request was issued;
[0021] In the step of the target computing node sending response information to the source computing node, the response information is the response information sent by the target computing node to the source computing node in response to the task migration request, and the response information includes the load information of the target computing node.
[0022] The target compute node to be migrated receives the static status information of the task to be migrated, and the task management agent of the target compute node to be migrated... In the step of creating a migration task on the target compute node to be migrated based on the static state information, wherein the static state information is generated by the task management agent of the source compute node. The task to be migrated is obtained and sent to the target compute node, which is the task management agent of the source compute node. The target computing node with the minimum load determined based on the load information in the response information;
[0023] Task management agent for the target compute node to be migrated Receive the dynamic status information of the task to be migrated, and the task management agent of the target computing node to be migrated The created migration task is restored in the target compute node recovery step based on the dynamic status information of the task to be migrated, wherein the dynamic status information is provided by the task management agent of the source compute node. The task to be migrated is obtained and sent to the target computing node to be migrated.
[0024] According to the distributed task migration method for ensuring consistent computing node states in some embodiments of this application, the recovered migration task is performed by the task management agent of the target computing node to be migrated. Migration tasks are added to the task set of the target computing node to be migrated.
[0025] According to the distributed task migration method with consistent computing node states in some embodiments of this application, the load is obtained based on the following formula:
[0026]
[0027] in , , , These are the current load size, CPU utilization, memory utilization, and number of tasks of the target computing node, respectively. , , These represent the proportions of CPU utilization, memory utilization, and task number in the compute node load, respectively, and all three are constants not greater than 1. .
[0028] According to some embodiments of the distributed task migration method with consistent computing node states, the format of the response information is as follows: Where id represents the identifier of the compute node, and load represents the load of the compute node;
[0029] The format of the dynamic status information ,in Indicates the task owner. Indicates the task ID. Indicates permissions. This indicates that the file handle is open. Indicates disk information, Represents the program stack. Represents the program register.
[0030] On a third-party level, the I / O device control method according to some embodiments of this application includes...
[0031] The migration task restored from the task set of the target computing node to be migrated sends a command to the shared information channel (SIC) of the target computing node to control the I / O device, and the shared information channel (SIC) of the target computing node to be migrated receives the command.
[0032] The I / O controller reads the instruction from the shared information channel (SIC) and forwards it to the corresponding I / O device according to the DT field in the instruction.
[0033] The I / O device receives the instruction and performs control according to the instruction;
[0034] The format of the instruction , DT represents the serial number and the device type. Indicates the length of the attribute. Represents a list of attributes; This represents the actual instruction stream that the task sends to the I / O controller. Indicates the length of the marking instruction sequence. This indicates whether the instruction has been completed.
[0035] In a fourth aspect, an electronic device according to some embodiments of this application includes: one or more processors, a memory, and one or more programs; wherein the one or more programs are stored in the memory, and the one or more programs include instructions that, when executed by the electronic device, cause the electronic device to perform the technical solutions of the first aspect of the embodiments of this application and any possible design of the first aspect.
[0036] In a fifth aspect, a computer-readable storage medium according to some embodiments of the present application includes a computer program that, when run on an electronic device, causes the electronic device to perform the technical solutions of the first aspect of the present application and any possible design of the first aspect.
[0037] For the technical effects that may be achieved in the above aspects, please refer to the description of the technical effects that may be achieved by the various possible solutions in the first and second aspects above, which will not be repeated here.
[0038] Beneficial Effects: The various solutions of the embodiments of the present invention migrate the migration task of the source computing node in the computing system to other computing nodes within the same communication domain for execution. This can improve the resource utilization and execution efficiency of the computing node cluster. By cooperating with other computing nodes to execute the migration task, the computing node cluster gains more powerful computing capabilities and higher reliability. Furthermore, the solutions of the present invention enable continued control of I / O devices after the task is migrated to other computing nodes, achieving decoupling between tasks and I / O, thus demonstrating applicability. Additional aspects and advantages of the present invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0039] Figure 1 This is a flowchart of the task migration method in the embodiment.
[0040] Figure 2 This is a diagram illustrating a specific implementation scenario of cross-platform task migration in the example.
[0041] Figure 3 This is the instruction sequence format in the embodiment. Detailed Implementation
[0042] The embodiments of this application are described in detail below with reference to the accompanying drawings, examples of which are illustrated in the drawings. This application provides a method, electronic device, and computer-readable storage medium to solve the problem of consistency in task migration states between different computing nodes in a computing system. The method, electronic device, and computer-readable storage medium are based on the same technical concept. Since the principles by which the method, electronic device, and computer-readable storage medium solve the problem are similar, the implementation of the device and method can refer to each other, and repeated details will not be repeated.
[0043] In one embodiment, this disclosure describes a distributed task migration method with consistent computing node states. The method is based on the idea of state consistency, which transmits the static task state information and dynamic task state information of the task to be migrated from the source computing node to the target computing node. The target computing node restores the task locally according to the two states of the migrated task, so that the migrated task continues to be executed normally on the target computing node.
[0044] According to the proposed solution, this invention migrates the task to be migrated to other computing nodes within the same communication domain for execution, thereby improving the resource utilization and execution efficiency of the computing node cluster. By cooperating with other computing nodes to execute the migration task, the computing node cluster gains more powerful computing capabilities and higher reliability. Furthermore, this method ensures that the task can continue to control I / O devices after being migrated to other computing nodes, thus achieving decoupling between tasks and I / O and demonstrating good applicability.
[0045] To better understand this invention, the technical terms involved are defined or described herein. Those skilled in the art can understand this invention through the following definitions or descriptions. Computing nodes are divided into two types: source computing nodes. and target computing nodes Both the source and target computing nodes contain several tasks, and the task set of each computing node is defined as follows: , Indicates the first One task, and Represents the source compute node The same applies to each task. Represents the target computing node Tasks are categorized into two types on each compute node: portable tasks and non-portable tasks. Portable tasks are those that can be migrated to other compute nodes for continued execution when they fail to execute normally. Non-portable tasks are local tasks, meaning they can only be executed on the current compute node.
[0046] In this invention, each computing node deploys a Task Management Agent (TMA), which is used to detect whether any of the K traversable tasks on the computing node are experiencing execution failures, excessive load, or other situations that prevent the task from executing normally on the current computing node. If the task... If a situation arises where the task cannot be executed normally, TMA will... As a task to be migrated, it sends task migration requests to other compute nodes. During the task migration process, the TMA can receive response information returned by other compute nodes to the source compute node, and can transmit and receive static and dynamic task status information of the migrated task. The Shared Information Channels (SICs) introduced in this invention are a local storage space of the compute node, denoted as SIC, used to store instructions issued by the task, and the I / O controller can read instructions from them and forward them to its hardware devices to complete hardware control.
[0047] In this invention, the dynamic status information sent by the source computing node The format is ,in Indicates the task owner. Indicates the task ID. Indicates permissions. This indicates that the file handle is open. Indicates disk information, Represents the program stack. This represents the program register, which allows the target compute node to resume the migration task locally based on the dynamic status information transmitted from the source compute node.
[0048] In this invention, the response information returned by the computing node after receiving the task migration request The format is id is the identifier of the compute node, used to uniquely identify a compute node. load is defined as the load of the compute node. The load of a compute node can be measured by factors such as its CPU utilization, system idle time and memory utilization.
[0049] In this invention, the instructions sent by the task control I / O device to the shared information channel The format is , DT is the sequence number, used to uniquely identify a command; DT is the device type, used to specify the hardware device to which this command is sent. It is the attribute length, used to represent the number of hardware attributes; It is a list of attributes used to represent the attributes of the hardware; This is the actual instruction stream that the task sends to the I / O controller. The hardware device uses this field to complete hardware device control. Each hardware device has its own instruction set. Mark the length of the instruction sequence; Used to indicate whether this instruction has been completed;
[0050] The distributed task migration method with consistent computing node states described in this invention specifically includes the following steps:
[0051] Step 1: Initializing the Task Management Agent:
[0052] Task management agent on the source compute node Create a migration task queue And the response information set RIS, for the migration task queue Initialize it, and set its length to... The source computing node task set All K( (N and K are both positive integers greater than 0) transferable tasks are added sequentially. In the middle, the response information set RIS is initialized to empty. It must not be empty. If it is empty, it means that the source computing node is an isolated computing node and cannot perform task migration in or out with other computing nodes. Isolated computing nodes do not meet the requirements of this invention.
[0053] Step 2: Check the migration task queue and perform migration for the tasks to be migrated.
[0054] Step 2.1 For migration task queue The tasks in the process are checked, and if any tasks are detected that need to be migrated, a task is selected from the tasks to be migrated. Proceed to step 2.2; otherwise, delay. After a time right Retest.
[0055] Step 2.2 To m other target computing nodes within the same communication domain Send a task migration request and wait for a response. The target compute node that receives the task migration request sends to Return response information , Response information will be received Add to the response information set in sequence .
[0056] Step 2.3 If the response information set If the value is empty, meaning no target computing node responds to the task migration request, then the task is considered... If the current condition is not transferable, return to step 2.1. Otherwise, iterate through the previous steps. Select the response information with the least load. The identifier is The target compute node is used as the target compute node for migration. .
[0057] Step 2.4 Save task First, extract the task's static state information, i.e., the binary executable file, and send it to the task management agent of the target compute node to be migrated. , After receiving the static status information of the migration task, create a new task on the target compute node. Once created, the task will be suspended. If the task... Creation successful. The task Dynamic status information Send to The target computing node is based on Restore it to its state before the migration; otherwise, from Delete it and return to step 2.3.
[0058] Step 2.5 The source compute node will From the task set Delete it, similarly. Will From the migration task queue Delete, the target compute node will create a new task Join its task set .
[0059] Step 3: Task migration complete, control I / O devices:
[0060] Step 3.1 If the task During operation, it is necessary to control I / O devices, and the commands issued by them... Required Figure 3 The format shown is loaded into the Shared Information Channel (SIC); otherwise, return to step two.
[0061] Step 3.2 The I / O controller reads from the SiC It then forwards the command to the corresponding I / O device based on the DT field in the instruction, thus completing device control.
[0062] Compared with existing technologies, the state-consistent distributed task migration method of the present invention is a task migration method based on user space, which does not require kernel modification and can migrate tasks that cannot be executed normally to other computing nodes for execution. In addition, when the migration task needs to control I / O devices during the operation, the instructions it issues can be loaded into the shared information channel, without having to consider the reconstruction of network communication after migration, thus achieving decoupling of tasks and I / O.
[0063] To better understand the technical solution of this invention, an example of implementing the method according to this invention is provided. In this specific example, a task migration scenario is constructed. There are M+1 computing nodes in the communication domain, one of which is the source computing node, and the other M computing nodes are the target computing nodes. Each computing node has a shared information channel (SIC), and each SIC is identified by a computing node identifier. The computing nodes in the communication domain need to control I / O devices through an I / O controller. The I / O controller is responsible for managing the I / O devices and can forward instructions to the I / O devices to achieve specific control of the devices.
[0064] In this instance, a distributed task migration method that ensures consistent compute node states, such as... Figure 1 As shown, firstly, Create a migration task queue and the response information set RIS, initialize If RIS is empty, add the transferable tasks from the source compute node sequentially. , Detection If a task is detected to be migrated, select one task from the list of tasks to be migrated. Send task migration requests to other compute nodes; otherwise, delay. Then, a re-test was performed. Next, other compute nodes received the task migration request and... Send response information, Add all received response information to the RIS. If the RIS is empty, the task is not currently available for migration. Otherwise, select the computing node with the least load from the RIS as the target computing node for migration. The static task status information and dynamic task status information are transmitted sequentially to the target computing node. The target computing node resumes the migration task based on the two task statuses. If the migration task requires control of I / O devices during execution, the instructions are sent to the shared information channel, and the I / O controller reads the new instructions from the shared information channel and forwards them to the I / O devices to complete device control. Otherwise, the migration process ends.
[0065] like Figure 2 The diagram illustrates a specific scenario for task migration. The specific steps of this invention are as follows:
[0066] Step 1: Initializing the Task Management Agent:
[0067] Task management agent on the source compute node Create a migration task queue And the response information set RIS, for the migration task queue Initialize it, and set its length to... The source computing node task set All K transferable tasks in the process are added sequentially. In the middle, the response information set RIS is initialized to empty.
[0068] Step 2: Check the migration task queue and perform migration for the tasks to be migrated.
[0069] Step 2.1 For migration task queue The tasks in the process are checked, and if any tasks are detected that need to be migrated, a task is selected from the tasks to be migrated. Proceed to step 2.2; otherwise, delay. After a time right Retest.
[0070] Step 2.2 To m other target computing nodes within the same communication domain Send a task migration request and wait for a response. The target compute node that receives the task migration request sends to Return response information , Response information will be received Add to the response information set in sequence .
[0071] Step 2.3 If the response information set If the value is empty, meaning no target computing node responds to the task migration request, then the task is considered... If the current condition is not transferable, return to step 2.1. Otherwise, iterate through the previous steps. Select the response information with the least load. The identifier is The target compute node is used as the target compute node for migration. .
[0072] In this embodiment, a method for calculating load is provided, and the formula is as follows:
[0073]
[0074] in , , , Target computing nodes Current load size, CPU utilization, memory utilization, and number of tasks. , , These represent the proportions of CPU utilization, memory utilization, and task number in the compute node load, respectively, and all three are constants not greater than 1. .
[0075] According to the aforementioned scheme , , These metrics respectively represent the proportion of CPU utilization, memory utilization, and task count in the compute node load. Increasing these proportions increases the importance of the metrics in the load, for example... =0.5 =0.3 =0.2 indicates that the CPU utilization accounts for the majority of the load. In different computing systems, the load considerations need to be adjustable under certain requirements. The load determination method described in this invention accurately selects parameters that can affect the actual load and determines the load in a proportionally adjustable manner. This not only accurately reflects the load but also allows the load representation to be adjusted according to the actual situation.
[0076] Step 2.4 Save task First, extract the task's static state information, i.e., the binary executable file, and send it to the task management agent of the target compute node to be migrated. , After receiving the static status information of the migration task, create a new task on the target compute node. Once created, the task will be suspended. If the task... Creation successful. The task Dynamic status information Send to The target computing node is based on Restore it to its state before the migration; otherwise, from Delete it and return to step 2.3.
[0077] Step 2.5 The source compute node will From the task set Delete it, similarly. Will From the migration task queue Delete, the target compute node will create a new task Join its task set .
[0078] Step 3: Task migration complete, control I / O devices:
[0079] Step 3.1 If the task During operation, it is necessary to control I / O devices, and the commands issued by them... Required Figure 3 The format shown is loaded into the shared information channel of the target compute node during migration. If correct, proceed to step two; otherwise, return to step two.
[0080] Step 3.2 I / O controller from Read from The instruction is forwarded to the corresponding hardware device based on the DT field in the instruction, and the hardware device completes the specific control of the device based on the Order field in the instruction.
[0081] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0082] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0083] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0084] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A distributed task migration method with consistent computing node states, characterized in that, Used for source compute nodes, including The source compute node sends a task migration request to the target compute node; The source compute node receives response information from the target compute node; wherein, the response information is the response information sent by the target compute node to the source compute node in response to the task migration request, and the response information includes the load information of the target compute node; Task management agent for source compute nodes The target compute node to be migrated is determined; the source compute node sends static status information of the tasks to be migrated to the target compute node, wherein the static status information is the task management agent of the target compute node. The information upon which the migration task is created on the target computing node to be migrated; The source compute node sends dynamic status information of the task to be migrated to the target compute node, and the dynamic status information is the task management agent of the target compute node. The information on which the created migration task is based to restore the target computing node to its state before migration; The load is obtained based on the following formula: in , , , These are the current load size, CPU utilization, memory utilization, and number of tasks of the target computing node, respectively. , , These represent the proportions of CPU utilization, memory utilization, and the number of tasks in the compute node load, respectively, and all three are constants not greater than 1. ; The format of the response information Where id represents the identifier of the compute node, and load represents the load of the compute node; The format of the dynamic status information ,in Indicates the task owner. Indicates the task ID. Indicates permissions. This indicates that the file handle is open. Indicates disk information, Represents the program stack. Represents the program register.
2. The distributed task migration method with consistent computing node states according to claim 1, characterized in that, In the step of the source compute node issuing a task migration request to the target compute node, the task migration request is made by the task management agent of the source compute node. The task management agent of the source computing node was detected. The task queue contains tasks awaiting migration from the task management agent of the source compute node. issue; Task management agent for source compute nodes In the step of determining the target compute node to be migrated, the target compute node to be migrated is the task management agent of the source compute node. The target computing node with the minimum load is determined based on the load information in the response information; In the step of the source compute node sending the static status information of the task to be migrated to the target compute node, the static status information is provided by the task management agent of the source compute node. The task to be migrated is obtained and sent to the target computing node to be migrated; In the step of the source compute node sending the dynamic status information of the task to be migrated to the target compute node, the dynamic status information is provided by the task management agent of the source compute node. The task to be migrated is obtained and sent to the target computing node to be migrated.
3. The distributed task migration method with consistent computing node states according to any one of claims 1-2, characterized in that, The restored migration task is performed by the task management agent of the target compute node to be migrated. Migration tasks are added to the task set of the target computing node to be migrated.
4. A distributed task migration method with consistent computing node states, characterized in that, For target computing nodes, including The target compute node receives a task migration request from the source compute node; The target computing node sends a response message to the source computing node; wherein, the response message is the response message sent by the target computing node to the source computing node in response to the task migration request, and the response message includes the load information of the target computing node; The target compute node to be migrated receives the static status information of the tasks to be migrated, and the task management agent of the target compute node to be migrated... A migration task is created on the target computing node to be migrated based on the static status information. Task management agent for the target compute node to be migrated Receive the dynamic status information of the task to be migrated, and the task management agent of the target computing node to be migrated Based on the dynamic status information of the task to be migrated, the created migration task is restored to its state before migration on the target computing node to be migrated; The load is obtained based on the following formula: in , , , These are the current load size, CPU utilization, memory utilization, and number of tasks of the target computing node, respectively. , , These represent the proportions of CPU utilization, memory utilization, and task number in the compute node load, respectively, and all three are constants not greater than 1. ; The format of the response information Where id represents the identifier of the compute node, and load represents the load of the compute node; The format of the dynamic status information ,in Indicates the task owner. Indicates the task ID. Indicates permissions. This indicates that the file handle is open. Indicates disk information, Represents the program stack. Represents the program register.
5. The distributed task migration method with consistent computing node states according to claim 4, characterized in that, In the step where the target compute node receives a task migration request from the source compute node, the task migration request is made by the task management agent of the source compute node. The task management agent TMAS of the source compute node was detected to have tasks awaiting migration in its task queue. The task migration request was issued; The target compute node to be migrated receives the static status information of the task to be migrated, and the task management agent of the target compute node to be migrated... In the step of creating a migration task on the target compute node to be migrated based on the static state information, wherein the static state information is generated by the task management agent of the source compute node. The task to be migrated is obtained and sent to the target compute node, which is the task management agent of the source compute node. The target computing node with the minimum load is determined based on the load information in the response information; Task management agent for the target compute node to be migrated Receive the dynamic status information of the task to be migrated, and the task management agent of the target computing node to be migrated The created migration task is restored in the target compute node recovery step based on the dynamic status information of the task to be migrated, wherein the dynamic status information is provided by the task management agent of the source compute node. The task to be migrated is obtained and sent to the target computing node to be migrated.
6. The distributed task migration method with consistent computing node states according to any one of claims 4-5, characterized in that, The restored migration task is performed by the task management agent of the target compute node to be migrated. Migration tasks are added to the task set of the target computing node to be migrated.
7. A method for acquiring migration task control I / O devices, characterized in that, The distributed task migration method with consistent computing node states as described in any one of claims 1-6; include The migration task restored from the task set of the target computing node to be migrated sends a command to the shared information channel (SIC) of the target computing node to control the I / O device, and the shared information channel (SIC) of the target computing node to be migrated receives the command. The I / O controller reads the instruction from the shared information channel (SIC) and forwards it to the corresponding I / O device according to the DT field in the instruction. The I / O device receives the instruction and performs control according to the instruction; The format of the instruction , DT represents the serial number and the device type. Indicates the length of the attribute. Represents a list of attributes; This represents the actual instruction stream that the task sends to the I / O controller. Indicates the length of the marking instruction sequence. This indicates whether the instruction has been completed.
8. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the steps of the method as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
System and method for supporting common transaction identifier (xid) optimization and transaction affinity based on resource manager (rm) instance awareness in transactional environment
CN106255956A
Dynamic adaptive IO load balancing method for wide-area high-performance computing environment
CN110213351A