Task execution method and device, equipment, medium and program product

By creating container groups based on historical images stored by the work nodes in the accelerator card-free mode, the resource waste problem caused by the need to provide additional CPU nodes in the prior art is solved, and resource utilization and task execution efficiency are improved.

CN120179387APending Publication Date: 2025-06-20DAWNING INT INFORMATION IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510238215.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art requires additional CPU nodes that do not set up acceleration cards during script debugging, resulting in wasting node resources.

Method used

The task execution request of the terminal device is obtained through the management node, and the work node that performs the historical task is determined based on the identification of the historical task. In the accelerator card mode, a container group is created based on the historical image stored by the work node to perform tasks.

Benefits of technology

The tasks can be executed in accelerating card-free mode without the help of additional CPU nodes, which improves node resource utilization, reduces task execution costs, and improves task execution efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179387A_ABST
    Figure CN120179387A_ABST
Patent Text Reader

Abstract

The invention provides a task execution method and device, equipment, a medium and a program product. The method comprises the steps that a management node obtains a task execution request sent by terminal equipment; the task execution request comprises an identifier of a historical task; the historical task is a task executed in a historical container group in an accelerator card mode; the management node determines a working node for executing the historical task according to the identifier of the historical task; wherein the working node comprises an accelerator card; the management node determines available resource information corresponding to the acceleration card-free mode; the management node controls the working node to create a container group based on a historical mirror image stored by the working node according to the available resource information, and controls the working node to execute a task in the container group in a mode without an accelerator card; wherein the historical mirror image is a mirror image used for creating the historical container group. According to the method, the working nodes for setting the acceleration card can be controlled, the task is executed in the mode without the acceleration card, and the node resource utilization rate is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic devices, and in particular, to a task execution method, apparatus, device, medium, and program product. Background Art

[0002] A node equipped with an acceleration card can provide computing services for users based on the container group service of Kubernetes. The cost of acceleration card resources is relatively high, and users need to pay as long as they start a container group to execute tasks.

[0003] In related technologies, during script debugging, a central processing unit (CPU) node without an acceleration card can be used to start a container group to execute tasks in order to reduce costs.

[0004] However, in related technologies, it is necessary to additionally provide a CPU node without an acceleration card for script debugging, resulting in a problem of wasting node resources in the methods of related technologies. Summary of the Invention

[0005] This application provides a task execution method, apparatus, device, medium, and program product, which can control a node equipped with an acceleration card to execute tasks in a no-acceleration-card mode, improving the utilization rate of node resources.

[0006] In a first aspect, this application provides a task execution method, which is applied to a management node. The method includes:

[0007] Obtain a task execution request sent by a terminal device; the task execution request includes an identifier of a historical task; the historical task is a task executed in a historical container group in an acceleration card mode;

[0008] Determine a working node for executing the historical task according to the identifier of the historical task; where the working node includes an acceleration card;

[0009] Determine available resource information corresponding to the no-acceleration-card mode;

[0010] According to the available resource information, control the working node to create a container group based on a historical image stored in the working node, and control the working node to execute a task in the container group in a no-acceleration-card mode; where the historical image is an image used to create the historical container group.

[0011] In this solution, the management node can obtain a task execution request sent by a terminal device; the task execution request includes the identifier of a historical task; the historical task is a task executed in a historical container group in the accelerated card mode. The management node can determine the working node (including the accelerated card) that executed the historical task based on the identifier of the historical task, and determine the available resource information corresponding to the non-accelerated card mode. The management node can control the working node to create a container group based on the historical image (the image used to create the historical container group) stored by the working node according to the available resource information, and control the working node to execute the task in the container group in the non-accelerated card mode. By the above method of controlling the node with an accelerated card to execute the task in the container group in the non-accelerated card mode based on the available resource information corresponding to the non-accelerated card mode, the task without an accelerated card can be executed without relying on an additional CPU node, improving the utilization rate of node resources. In addition, by the above method of reusing the historical image for constructing the historical container group, the container group can be constructed without pulling the image from the image repository, improving the construction efficiency of the container group, and further improving the task execution efficiency, thus improving the user experience.

[0012] In one implementation, determining the available resource information corresponding to the non-accelerated card mode includes:

[0013] Determine a resource grouping according to the identifier of the historical task; among them, the working node that executed the historical task belongs to the resource grouping;

[0014] Determine the corresponding available resource information according to the resource grouping and the non-accelerated card mode; among them, the available resource information includes the number of central processing units (CPUs) used and the size of the memory used.

[0015] In this solution, the available resource information of different resource groupings is different in the non-accelerated card mode. The management node can determine the resource grouping to which the working node that executed the historical task belongs based on the identifier of the historical task, and then determine the available resource information (the number of CPUs used and the size of the memory used) in the non-accelerated card mode based on the resource grouping. In addition, by setting the number of CPUs used and the size of the memory used, on the one hand, it avoids the situation that the task execution process in the non-accelerated mode occupies too much CPU and memory, resulting in the task execution process in the accelerated card mode being affected; on the other hand, it ensures that the working node with an accelerated card can utilize the available resource information to execute the task in the non-accelerated card mode, reducing the task execution cost.

[0016] In one implementation, according to the available resource information, controlling the working node to create a container group based on the historical image stored by the working node and controlling the working node to execute the task in the container group in the non-accelerated card mode includes:

[0017] Obtain resource limit information; wherein, the resource limit information indicates that the acceleration card is invisible in the container group;

[0018] According to the resource limit information and the available resource information, control the worker node to create a container group based on the historical image, and control the worker node to execute tasks in the container group in the acceleration cardless mode.

[0019] In this solution, based on the resource limit information indicating that the acceleration card is invisible in the container group, the container group created by the worker node is configured to be invisible to the acceleration card. When the container group is configured to be invisible to the acceleration card, the worker node cannot use the acceleration card when executing tasks in the container group. That is to say, the worker node can execute tasks in the container group in the acceleration cardless mode. Through the above method, it is ensured that the worker node equipped with an acceleration card can execute tasks in the acceleration cardless mode, reducing the task execution cost.

[0020] In one implementation, before obtaining the resource limit information, the method further includes:

[0021] Obtain the remaining resource information corresponding to the worker node; wherein, the remaining resource information includes the number of remaining CPUs;

[0022] When it is determined that the number of used CPUs does not exceed the number of remaining CPUs, it is determined that the worker node meets the container group creation condition.

[0023] In this solution, the management node can determine whether the worker node meets the container group creation condition by comparing the number of used CPUs and the number of remaining CPUs. Through the above method, it is possible to avoid the situation where when the container group executes tasks in the acceleration cardless mode, it directly occupies the CPU based on the number of used CPUs, resulting in excessive CPU occupancy and affecting other tasks executed in the acceleration card mode.

[0024] In one implementation, when it is determined that the number of used CPUs does not exceed the number of remaining CPUs, determining that the worker node meets the container group creation condition includes:

[0025] When the number of used CPUs does not exceed the number of remaining CPUs, obtain the status of the device plugin in the worker node; wherein, the device plugin is used to obtain the resource information of multiple components of the worker node for the container group to use the components based on the resource information; the multiple components include the acceleration card;

[0026] When it is determined that the status of the device plugin is the running state, it is determined that the worker node meets the container group creation condition.

[0027] In this solution, when the number of CPUs used by the management node does not exceed the number of remaining CPUs, the management node can obtain the status of the device plugin in the worker node. When it is determined that the status of the device plugin is the running state, it is determined that the worker node meets the container group creation condition. In the above manner, on the one hand, it is possible to avoid the situation where when a container group executes a task in the non-acceleration card mode, it directly occupies the CPU based on the number of CPUs used, resulting in excessive CPU occupancy and affecting other tasks executed in the acceleration card mode. On the other hand, it is possible to avoid the situation where the device plugin fails to run, resulting in the inability to create a container group.

[0028] In one implementation, before determining the worker node that executes the historical task according to the identifier of the historical task, the method further includes:

[0029] Obtain the status of the historical task according to the identifier of the historical task;

[0030] Determine that the status of the historical task is not the running state.

[0031] In this solution, based on the historical task in the running state (a task executed in the historical container group in the acceleration card mode), non-acceleration card execution is not supported. Therefore, the management node can obtain the status of the historical task. When the management node determines that the status of the historical task is not the running state, it can determine the worker node that executes the historical task according to the identifier of the historical task, so as to control the worker node to re-execute the task in the container group in the non-acceleration card mode.

[0032] In a second aspect, an embodiment of the present application provides a task execution device, which is applied to a management node. The device includes:

[0033] An obtaining module, configured to obtain a task execution request sent by a terminal device; the task execution request includes an identifier of a historical task; the historical task is a task executed in a historical container group in the acceleration card mode;

[0034] A processing module, configured to determine a worker node that executes the historical task according to the identifier of the historical task; wherein, the worker node includes an acceleration card;

[0035] The processing module is further configured to determine available resource information corresponding to the non-acceleration card mode;

[0036] The processing module is further configured to control the worker node to create a container group based on a historical image stored in the worker node according to the available resource information, and control the worker node to execute a task in the container group in the non-acceleration card mode; wherein, the historical image is an image used to create the historical container group.

[0037] The task execution device provided by the embodiment of the present application can execute the technical solution shown in the above method embodiment. The implementation principle and beneficial effects are similar and will not be elaborated here.

[0038] In one implementation manner, the processing module is specifically configured to:

[0039] Determine a resource group according to the identifier of the historical task; wherein, the working nodes executing the historical task belong to the resource group;

[0040] Determine the corresponding available resource information according to the resource group and the accelerator card - free mode; wherein, the available resource information includes the number of central processing units (CPUs) used and the size of the memory used.

[0041] The task execution device provided by the embodiment of the present application can execute the technical solution shown in the above method embodiment. The implementation principle and beneficial effects are similar and will not be elaborated here.

[0042] In one implementation manner, the processing module is specifically configured to:

[0043] Obtain resource limit information; wherein, the resource limit information indicates that the accelerator card is invisible in the container group;

[0044] Control the working nodes to create a container group based on the historical image according to the resource limit information and the available resource information, and control the working nodes to execute tasks in the container group in the accelerator card - free mode.

[0045] The task execution device provided by the embodiment of the present application can execute the technical solution shown in the above method embodiment. The implementation principle and beneficial effects are similar and will not be elaborated here.

[0046] In one implementation manner, the processing module is further configured to:

[0047] Obtain the remaining resource information corresponding to the working node; wherein, the remaining resource information includes the number of remaining CPUs;

[0048] When it is determined that the number of CPUs used does not exceed the number of remaining CPUs, determine that the working node meets the container group creation condition.

[0049] The task execution device provided by the embodiment of the present application can execute the technical solution shown in the above method embodiment. The implementation principle and beneficial effects are similar and will not be elaborated here.

[0050] In one implementation manner, the processing module is specifically configured to:

[0051] When the number of CPUs used does not exceed the number of remaining CPUs, obtain the status of the device plugin in the worker node; where the device plugin is used to obtain the resource information of multiple components in the worker node for the container group to use the components based on the resource information; the multiple components include acceleration cards.

[0052] When it is determined that the status of the device plugin is the running state, determine that the worker node meets the container group creation condition.

[0053] The task execution device provided by the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and its implementation principle and beneficial effects are similar, and will not be elaborated here.

[0054] In one implementation, the processing module is further configured to:

[0055] Obtain the status of the historical task according to the identifier of the historical task;

[0056] Determine that the status of the historical task is not the running state.

[0057] The task execution device provided by the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and its implementation principle and beneficial effects are similar, and will not be elaborated here.

[0058] In a third aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0059] The memory stores computer-executable instructions;

[0060] The processor executes the computer-executable instructions stored in the memory to implement the method as in the first aspect.

[0061] The electronic device provided by the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and its implementation principle and beneficial effects are similar, and will not be elaborated here.

[0062] In a fourth aspect, the embodiments of the present application provide a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the method as in the first aspect.

[0063] When the computer-executable instructions stored in the computer-readable storage medium provided by the embodiments of the present application are executed by a processor, they can implement the technical solutions shown in the above method embodiments, and its implementation principle and beneficial effects are similar, and will not be elaborated here.

[0064] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method as in the first aspect.

[0065] When the computer program in the computer program product provided by the embodiments of the present application is executed by a processor, it can implement the technical solutions shown in the above method embodiments. The implementation principles and beneficial effects are similar and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0067] Figure 1 It is a schematic diagram of an application scenario provided by the embodiments of the present application;

[0068] Figure 2 It is a flowchart of a task execution method provided by the embodiments of the present application Figure 1 ;

[0069] Figure 3a It is a flowchart of a task execution method provided by the embodiments of the present application Figure 2 ;

[0070] Figure 3b It is a schematic diagram of uploading a configuration file provided by the embodiments of the present application;

[0071] Figure 3c It is a flowchart of applying for container group resources provided by the embodiments of the present application;

[0072] Figure 4 It is the third flowchart of a task execution method provided by the embodiments of the present application;

[0073] Figure 5 It is a schematic diagram of a task execution device provided by the embodiments of the present application;

[0074] Figure 6 It is a structural diagram of an electronic device provided by the present application.

[0075] Through the above drawings, the clear embodiments of the present application have been shown, and there will be more detailed descriptions later. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0076] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0077] Glossary of terms:

[0078] Kubernetes: Abbreviated as K8s, which is formed by replacing the eight characters "ubernete" in the name with 8. It is an open-source system used to manage containerized applications on multiple hosts in a cloud platform. The goal of Kubernetes is to make it simple and efficient to deploy containerized applications. Kubernetes provides a mechanism for application deployment, scheduling, updating, and maintenance.

[0079] k8s-device-plugin: A device plugin that allows users to incorporate device resources (such as graphics processors, complex programmable logic devices, etc.) into the Kubernetes scheduler, thereby helping users better manage and utilize these device resources (components of nodes). Through this device plugin, device resources such as graphics processors on Kubernetes nodes can be directly invoked and used by container groups (tasks in container groups).

[0080] Jupyter Notebook: A web application that facilitates the creation and sharing of program documentation, supports real-time code, mathematical equations, visualization, and Markdown. The uses of Jupyter Notebook include data cleaning and transformation, numerical simulation, statistical modeling, machine learning, etc.

[0081] Container: A main body that provides services. An instance after a mirror is started is called a container; a container is one or a group of applications that run independently.

[0082] Pod (Container group): The smallest resource management component in Kubernetes and the smallest resource object for running containerized applications. A Pod represents a process running in the cluster.

[0083] Acceleration card: A hardware device specifically designed to improve computing performance. They usually integrate high-performance computing cores and a large amount of memory to accelerate complex computing tasks. Exemplarily, the acceleration card can be a graphics processing unit (GPU).

[0084] In a server cluster (cloud computing platform), a node equipped with an acceleration card can provide computing services for users based on the container group service of Kubernetes (abbreviated as K8S). Users can apply for acceleration card resources that meet their computing needs and execute tasks (such as model training, etc.).

[0085] However, the cost of acceleration card resources is relatively high, and users need to pay as long as they start a container group to execute tasks (such as Jupyter Notebook tasks).

[0086] In related technologies, a Central Processing Unit (CPU) node without an acceleration card can be used to provide an environment without an acceleration card. During the script debugging process, users can use the CPU node without an acceleration card to start a container group to execute tasks to reduce costs.

[0087] However, in related technologies, an additional CPU node without an acceleration card needs to be provided, resulting in a problem of wasting node resources in the methods of related technologies.

[0088] Based on the above technical problems, the technical concept of the embodiments of this application is as follows:

[0089] The management node can determine the working node (equipped with an acceleration card) that executes the historical task according to the identifier of the historical task (the task executed in the historical container group in the acceleration card mode). The management node can control the working node to create a container group based on the historical image (the image stored in the working node and used to create the historical container group) according to the available resource information corresponding to the non-acceleration card mode, and control the working node to execute the task in the container group in the non-acceleration card mode.

[0090] By controlling the working node equipped with an acceleration card to execute tasks in the container group in the non-acceleration card mode, the utilization rate of node resources is improved.

[0091] Next, a task execution method provided by the embodiments of this application will be described in detail.

[0092] For ease of understanding, first, in combination with Figure 1 the application scenario involved in the embodiments of this application will be described.

[0093] Figure 1 is a schematic diagram of an application scenario provided by the embodiments of this application.

[0094] As Figure 1 shown, this application scenario includes a server cluster 10 and a terminal device 20.

[0095] Among them, the server cluster 10 includes a management node 101 and multiple candidate working nodes. Exemplarily, Figure 1 4 candidate working nodes are shown, namely candidate working node 102, candidate working node 103, candidate working node 104, and candidate working node 105. It should be noted that one of the multiple candidate working nodes is the working node that executes the historical task in the embodiment of the present application.

[0096] Taking candidate working node 102 as an example, candidate working node 102 may include at least one CPU, at least one acceleration card, and memory. It should be noted that other candidate working nodes are similar to candidate working node 102 and will not be described repeatedly here.

[0097] It should also be noted that the terminal device 20 is communicatively connected to the management node 101, and the management node 101 is communicatively connected to multiple candidate working nodes.

[0098] It should be noted that Figure 1 is only a schematic diagram of the application scenario provided by the embodiment of the present application. The embodiment of the present application does not limit Figure 1 the actual form of various components included therein, nor does it limit Figure 1 the interaction method between the components therein. In the application of the solution, it can be set according to actual needs.

[0099] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0100] Figure 2 is a schematic flowchart of a task execution method provided by an embodiment of the present application Figure 1 as Figure 2 shown. The method includes the following steps:

[0101] S201: Obtain a task execution request sent by the terminal device.

[0102] In this embodiment, the management node may obtain a task execution request sent by the terminal device. In one implementation, the display interface of the terminal device may display a control for the non-acceleration card mode corresponding to the historical task. The terminal device may generate a task execution request in response to a user's operation on the control for the non-acceleration card mode and send the task execution request to the management node.

[0103] Among them, the task execution request may include an identifier of the historical task.

[0104] It should be noted that the historical task is a task executed in the historical container group in the acceleration card mode.

[0105] S202: Determine the working node for executing the historical task according to the identifier of the historical task.

[0106] In this embodiment, the management node can determine the working node for executing the historical task according to the identifier of the historical task. In one implementation, the management node can query the database according to the identifier of the historical task to determine the working node for executing the historical task.

[0107] It should be noted that the working node includes an acceleration card.

[0108] S203: Determine the available resource information corresponding to the non-acceleration card mode.

[0109] In this embodiment, the management node can determine the available resource information corresponding to the non-acceleration card mode.

[0110] In one implementation, the database of the management node can store the available resource information corresponding to the non-acceleration card mode. The management node can query the database to obtain the available resource information corresponding to the non-acceleration card mode.

[0111] In one implementation,

[0112] The management node can determine the resource group corresponding to the historical task according to the identifier of the historical task.

[0113] In one implementation, the management node can store the corresponding relationship between the identifier of the historical task and the resource group. The management node can query this corresponding relationship according to the identifier of the historical task to determine the resource group corresponding to the historical task.

[0114] In one implementation, the management node can determine the working node for executing the historical task according to the identifier of the historical task. The management node can determine the resource group to which the working node belongs and determine this resource group as the resource group corresponding to the historical task. It should be noted that one resource group can include at least one candidate working node. Exemplarily, the resource group to which the working node belongs includes candidate working node 102 (the working node for executing the historical task) and candidate working node 103.

[0115] The management node can determine the available resource information according to this resource group and the non-acceleration card mode. Among them, the available resource information includes the number of CPUs used and the size of the memory used. Exemplarily, the number of CPUs used can be 1. Exemplarily, the size of the memory used can be 4GB (or 1CB).

[0116] S204: Based on the available resource information, control the worker node to create a container group based on the historical image stored in the worker node, and control the worker node to execute the target task in the container group in the non-accelerator card mode.

[0117] In this embodiment, the management node can, based on the available resource information, control the worker node to create a container group (pod) based on the historical image. Herein, the historical image is the image stored in the worker node. The historical image is the image used to create the historical container group. It should be noted that the historical image is pulled by the worker node from the image repository when creating the historical container group.

[0118] The management node can control the worker node to execute a task in the container group in the non-accelerator card mode.

[0119] Beneficial effects of this embodiment: In this embodiment, the management node can obtain a task execution request sent by the terminal device; wherein, the task execution request includes the identifier of the historical task; the historical task is a task executed in the historical container group in the accelerator card mode. The management node can, according to the identifier of the historical task, determine the worker node (including the accelerator card) that executed the historical task, and determine the available resource information corresponding to the non-accelerator card mode. The management node can, based on the available resource information, control the worker node to create a container group based on the historical image (the image used to create the historical container group) stored in the worker node, and control the worker node to execute the task in the container group in the non-accelerator card mode. By the above method of controlling the node with an accelerator card to execute the task in the container group in the non-accelerator card mode, it is possible to execute the task in the non-accelerator card mode without relying on an additional CPU node, thereby improving the utilization rate of node resources.

[0120] Figure 3a It is a flowchart illustration of a task execution method provided by an embodiment of the present application Figure 2 as Figure 3a shown, this method includes the following steps:

[0121] S301: Obtain a task execution request sent by the terminal device.

[0122] In this embodiment, the management node can obtain a task execution request sent by the terminal device.

[0123] Wherein, the task execution request includes the identifier of the historical task. The historical task is a task executed in the historical container group in the accelerator card mode.

[0124] The specific implementation process is the same as that of S201 and will not be elaborated herein.

[0125] S302: Determine the worker node that executed the historical task according to the identifier of the historical task.

[0126] In this embodiment, the management node may determine the working node for executing the historical task according to the identifier of the historical task.

[0127] Among them, the working node includes an acceleration card.

[0128] The specific implementation process is the same as that of S202 and will not be elaborated here.

[0129] S303: Determine the resource group to which the working node for executing the historical task belongs according to the identifier of the historical task.

[0130] In this embodiment, the management node may divide multiple candidate working nodes in the server cluster into at least one resource group. Among them, one resource group may include one or more candidate working nodes.

[0131] The management node may determine the resource group to which the working node for executing the historical task belongs according to the identifier of the historical task.

[0132] S304: Determine the corresponding available resource information according to the resource group and the acceleration card-free mode.

[0133] In this embodiment, the available resource information of different resource groups in the acceleration card-free mode may be the same or different.

[0134] For the same resource group, the available resource information in the acceleration card mode and the acceleration card-free mode may be different.

[0135] The management node may determine the corresponding available resource information according to the resource group (the resource group to which the working node for executing the historical task belongs) and the acceleration card-free mode.

[0136] Among them, the available resource information includes the number of central processing units (CPUs) used and the size of the memory used.

[0137] S305: Obtain the resource limit information.

[0138] In this embodiment, the management node may obtain the configuration file stored by the management node.

[0139] Among them, the configuration file includes the resource limit information, and the resource limit information indicates that the acceleration card is invisible in the container group. The configuration file is pre-obtained by the management node and stored in the management node.

[0140] In one implementation,

[0141] The management node may determine whether the working node meets the container group creation condition.

[0142] The management node can obtain resource limit information when determining that the worker node meets the container group creation conditions.

[0143] Next, the process by which the management node determines whether the worker node meets the container group creation conditions will be described.

[0144] In one implementation,

[0145] The management node can obtain the remaining resource information corresponding to the worker node. Among them, the remaining resource information includes the number of remaining CPUs.

[0146] The management node can determine that the worker node meets the container group creation conditions when determining that the number of used CPUs does not exceed the number of remaining CPUs.

[0147] The management node can determine that the worker node does not meet the container group creation conditions when determining that the number of used CPUs exceeds the number of remaining CPUs.

[0148] In one implementation,

[0149] The management node can obtain the remaining resource information corresponding to the worker node. Among them, the remaining resource information includes the number of remaining CPUs.

[0150] The management node can determine that the worker node does not meet the container group creation conditions when determining that the number of used CPUs exceeds the number of remaining CPUs.

[0151] The management node can obtain the status of the device plugin in the worker node when determining that the number of used CPUs does not exceed the number of remaining CPUs. The device plugin is used to obtain the resource information of multiple components of the worker node for the container group to use the components based on the resource information; the multiple components include acceleration cards.

[0152] The management node can determine that the worker node meets the container group creation conditions when determining that the status of the device plugin is the running state.

[0153] The management node can determine that the worker node does not meet the container group creation conditions when determining that the status of the device plugin is not the running state.

[0154] Next, the process by which the management node pre-obtains and stores the configuration file will be described.

[0155] In one implementation,

[0156] The terminal device can obtain the configuration file uploaded by the user. Among them, the configuration file includes available resource information and resource limit information. In addition, the configuration file can also include the number of reserved CPUs (nocard(n)),

[0157] It should be noted that the available resource information includes the number of CPUs used and the amount of memory used.

[0158] The resource limit information is used to indicate that the acceleration card is invisible in the container group.

[0159] The maximum value of the number of remaining CPUs is the same as the number of reserved CPUs. In other words, when all the CPUs available for the non-acceleration card mode are not occupied, the number of remaining CPUs is the same as the number of reserved CPUs.

[0160] Exemplarily, Figure 3b is a schematic diagram of an uploaded configuration file provided by an embodiment of the present application. As Figure 3b shown, the configuration file uploaded by the user includes available resource information (the number of central processing units (CPUs) used is 1, and the amount of memory used is 4096 MB (4 GB)), and indication information (name: VISIBLE_DEVICES; value: none). Among them, the value of VISIBLE_DEVICES is none, indicating that the acceleration card is invisible in the container group. In addition, the configuration file may further include that the number of reserved CPUs is 10.

[0161] The management node can obtain the configuration file sent by the terminal device and perform storage processing on the configuration file.

[0162] In addition, it should be noted that the management node can also pre-install the K8S environment and device plugins (such as k8s-device-plugin) for each candidate working node in the server cluster. It should be noted that the working node executing the historical task is one of the multiple candidate working nodes included in the server cluster. It should also be noted that the device plugin is used to obtain the resource information of multiple components of the working node for the container group to use the components based on the resource information. The multiple components include acceleration cards, and the multiple components may also include CPUs and memory.

[0163] S306: According to the resource limit information and the available resource information, control the working node to create a container group based on the historical image, and control the working node to execute tasks in the container group in the non-acceleration card mode.

[0164] In this embodiment, the management node can control the working node to create a container group based on the historical image according to the resource limit information and the available resource information. Based on the fact that the resource limit information indicates that the acceleration card is invisible in the container group, the container group created by the working node is configured to be invisible to the acceleration card.

[0165] In the case where the container group is configured to be invisible to the acceleration card, the working node can execute tasks in the container group in the acceleration cardless mode. Understandably, based on the container group being configured to be invisible to the acceleration card, the working node cannot use the acceleration card when executing tasks in the container group.

[0166] It should also be noted that during the process of creating a container group by the working node, it is necessary to apply for container group resources through a device plugin.

[0167] Figure 3c This is a schematic diagram of a process for applying for container group resources provided by an embodiment of the present application. As Figure 3c shown, a device plugin and a container group management component (kubelet) are running on the working node, a server (kube-apiserver) is running on the management node, and a client (client) is running on the terminal device.

[0168] Step 1: The device plugin can send a registration request to the kubelet through a Unix socket (an inter-process communication mechanism). The registration request specifies the API version of the device plugin and identifies the resource type (identification of components) (ResourceName) obtained by the device plugin.

[0169] Step 2: The container group management component can obtain the resource information of multiple components included in the working node by calling the ListAndWatch method.

[0170] Step 3: The container group management component can update the resource information of the working node (resource information of multiple components) stored on the server.

[0171] Step 4: The server can obtain a container group (pod) creation request from the client.

[0172] Step 5: The server can schedule the container group to the working node.

[0173] Step 6: After receiving the scheduling information of the container group, the container group management component can check the components (resources) required by the container group and call the allocate interface of the device plugin to request components that the container group can use from the device plugin.

[0174] Advantages of this embodiment: In this embodiment, the management node can obtain a task execution request sent by a terminal device; the task execution request includes an identifier of a historical task; the historical task is a task executed in a historical container group in the accelerator card mode. The management node can determine a working node (including an accelerator card) that executed the historical task according to the identifier of the historical task, and determine a resource group to which the working node that executed the historical task belongs. The management node can determine corresponding available resource information according to the resource group and the non-accelerator card mode. The management node can control the working node to create a container group based on a historical image (an image used to create the historical container group) stored by the working node according to the obtained resource limit information and available resource information, and control the working node to execute the task in the container group in the non-accelerator card mode. Through the above method, the available resource information (the number of CPUs used and the size of the memory used) in the non-accelerator card mode can be accurately determined. In addition, by setting the number of CPUs used and the size of the memory used, on the one hand, it is avoided that the task execution process in the non-accelerator mode occupies too much CPU and memory, resulting in the task execution process in the accelerator card mode being affected; on the other hand, it is ensured that the working node equipped with an accelerator card can execute tasks in the non-accelerator card mode, reducing the task execution cost. In addition, based on the resource limit information indicating that the accelerator card is invisible in the container group, the created container group is invisible to the accelerator card, thereby ensuring that the working node equipped with an accelerator card can execute tasks in the non-accelerator card mode, reducing the task execution cost, and improving the node resource utilization rate.

[0175] Figure 4 FIG. 3 is a schematic flowchart of a task execution method provided by an embodiment of the present application. As Figure 4 shown, the method includes the following steps:

[0176] S401: Obtain a task execution request sent by a terminal device.

[0177] In this embodiment, the management node can obtain a task execution request sent by a terminal device.

[0178] Among them, the task execution request includes an identifier of a historical task; the historical task is a task executed in a historical container group in the accelerator card mode.

[0179] The specific implementation process is the same as that of S201 and will not be described in detail here.

[0180] S402: Obtain the status of the historical task according to the identifier of the historical task.

[0181] In this embodiment, the management node can obtain the status of the historical task according to the identifier of the historical task.

[0182] S403: Determine that the status of the historical task is not the running status.

[0183] In this embodiment, the management node can determine whether the status of the historical task is the running status.

[0184] When the management node determines that the status of the historical task is not the running status, it can execute S404;

[0185] When the management node determines that the status of the historical task is the running status, it ends.

[0186] S404: Determine the working node that executes the historical task according to the identifier of the historical task.

[0187] In this embodiment, the management node can determine the working node that executes the historical task according to the identifier of the historical task.

[0188] Among them, the working node includes an acceleration card.

[0189] The specific implementation process is the same as that of S202 and will not be elaborated here.

[0190] S405: Determine the resource grouping according to the identifier of the historical task.

[0191] In this embodiment, the management node can divide multiple candidate working nodes in the server cluster into at least one resource grouping. Among them, one resource grouping can include one or more candidate working nodes.

[0192] The management node can determine the resource grouping according to the identifier of the historical task. Among them, the working node that executes the historical task belongs to this resource grouping.

[0193] S406: Determine the corresponding available resource information according to the resource grouping and the acceleration card - free mode.

[0194] In this embodiment, for different resource groupings, the available resource information in the acceleration card - free mode can be the same or different.

[0195] For the same resource grouping, the available resource information in the acceleration card mode can be different from that in the acceleration card - free mode.

[0196] The management node can determine the corresponding available resource information according to the resource grouping (the resource grouping to which the working node that executes the historical task belongs) and the acceleration card - free mode.

[0197] Among them, the available resource information includes the number of CPUs used and the size of the memory used. Exemplarily, the number of CPUs used can be 10, and the size of the memory used can be 1GB or 4GB.

[0198] S407: Obtain the remaining resource information corresponding to the working node.

[0199] In this embodiment, the management node can obtain the remaining resource information corresponding to the working node.

[0200] Among them, the remaining resource information includes the number of remaining CPUs. It should be noted that the number of remaining CPUs is the number of unoccupied CPUs.

[0201] In addition, in one implementation, the remaining resource information may further include the size of the remaining memory.

[0202] S408: When it is determined that the number of used CPUs does not exceed the number of remaining CPUs, obtain the status of the device plugin in the working node.

[0203] In this embodiment, the management node can compare the number of used CPUs in the available resource information with the number of remaining CPUs in the remaining resource information.

[0204] The management node can end when it is determined that the number of used CPUs exceeds the number of remaining CPUs.

[0205] The management node can obtain the status of the device plugin (such as k8s - device - plugin) in the working node when it is determined that the number of used CPUs does not exceed the number of remaining CPUs.

[0206] It should be noted that the device plugin is used to obtain the resource information of multiple components of the working node for the container group to use the components based on the resource information. It should also be noted that the multiple components may include acceleration cards.

[0207] S409: When it is determined that the status of the device plugin is the running state, determine that the working node meets the container group creation condition.

[0208] In this embodiment, the management node can determine that the working node can obtain the resource information of multiple components through the device plugin for the working node to run the container group and can use the components based on the resource information when it is determined that the status of the device plugin is the running state. That is to say, the management node can determine that the working node meets the container group creation condition when it is determined that the status of the device plugin is the running state.

[0209] S410: Obtain the resource limit information.

[0210] In this embodiment, the management node can obtain the configuration file stored in the management node. Among them, the configuration file includes the resource limit information, and the resource limit information indicates that the acceleration card is invisible in the container group.

[0211] The specific implementation process is the same as that of S305 and will not be elaborated here.

[0212] S411: According to the resource limit information and available resource information, control the worker node to create a container group based on the historical image, and control the worker node to execute tasks in the container group in the non-accelerated card mode.

[0213] In this embodiment, the management node can control the worker node to create a container group based on the historical image according to the resource limit information and available resource information. It should be noted that based on the indication in the resource limit information that the acceleration card is invisible in the container group, the container group created by the worker node is configured to be invisible to the acceleration card. In the case where the container group is configured to be invisible to the acceleration card, the worker node can execute tasks in the container group in the non-accelerated card mode.

[0214] In addition, in one implementation,

[0215] The management node can determine the corresponding CPU unit price according to the resource grouping and non-accelerated mode.

[0216] The management node can obtain the CPU core hours.

[0217] The management node can determine the cost information according to the CPU unit price and CPU core hours.

[0218] It should be noted that the CPU core hours refer to the product of the number of CPUs used by the container group and the running duration during the execution of tasks by the container group.

[0219] Beneficial effects of this embodiment: Based on historical tasks in the running state (tasks executed in the historical container group in the accelerator card mode), execution without an accelerator card is not supported. Therefore, when the management node determines that the status of the historical task is not the running state, it can determine the worker node that executes the historical task according to the identifier of the historical task. The management node can determine the resource grouping to which the worker node that executes the historical task belongs. The management node can determine the corresponding available resource information according to this resource grouping and the accelerator cardless mode. When the number of CPUs used does not exceed the number of remaining CPUs and the device plugin of the worker node is in the running state, the management node can determine that the worker node meets the container group creation conditions. Through the above method, on the one hand, it can be avoided that when a container group executes a task in the accelerator cardless mode, it directly occupies the CPU based on the number of CPUs used, resulting in excessive CPU occupation and affecting other tasks executed in the accelerator card mode; on the other hand, it can be avoided that the device plugin is not running, resulting in the inability to create a container group. When the worker node meets the container group creation conditions, the management node can control the worker node to create a container group based on the historical image stored in the worker node (the image used to create the historical container group) according to the resource limit information and the available resource information, and control the worker node to execute the task in the container group in the accelerator cardless mode. Through the above method, tasks without an accelerator card can be executed without relying on an additional CPU node, improving the utilization rate of node resources. In addition, through the above method of reusing the historical image for constructing the historical container group, a container group can be constructed without pulling the image from the image repository, improving the construction efficiency of the container group, further improving the task execution efficiency, and thus improving the user experience.

[0220] Figure 5 Schematic diagram of a task execution device provided by an embodiment of the present application. The task execution device is applied to a management node, as Figure 5 shown, the task execution device 50 includes an acquisition module 51 and a processing module 52.

[0221] The acquisition module 51 is used to acquire a task execution request sent by a terminal device; the task execution request includes an identifier of a historical task; the historical task is a task executed in a historical container group in the accelerator card mode;

[0222] The processing module 52 is used to determine a worker node that executes the historical task according to the identifier of the historical task; wherein, the worker node includes an accelerator card;

[0223] The processing module 52 is further used to determine available resource information corresponding to the accelerator cardless mode;

[0224] The processing module 52 is further configured to control a working node to create a container group based on a historical image stored in the working node according to available resource information, and control the working node to execute a task in the container group in a mode without an acceleration card; wherein, the historical image is an image used to create a historical container group.

[0225] The task execution device provided by the embodiment of the present application can execute the technical solution shown in the above method embodiment, and its implementation principle and beneficial effects are similar, which will not be elaborated here.

[0226] In one implementation manner, the processing module 52 is specifically configured to:

[0227] Determine a resource grouping according to an identifier of a historical task; wherein, the working node that executes the historical task belongs to the resource grouping;

[0228] Determine corresponding available resource information according to the resource grouping and the mode without an acceleration card; wherein, the available resource information includes the number of central processing units (CPUs) used and the size of the memory used.

[0229] The task execution device provided by the embodiment of the present application can execute the technical solution shown in the above method embodiment, and its implementation principle and beneficial effects are similar, which will not be elaborated here.

[0230] In one implementation manner, the processing module 52 is specifically configured to:

[0231] Obtain resource limit information; wherein, the resource limit information indicates that the acceleration card is invisible in the container group;

[0232] Control the working node to create a container group based on the historical image according to the resource limit information and the available resource information, and control the working node to execute a task in the container group in a mode without an acceleration card.

[0233] The task execution device provided by the embodiment of the present application can execute the technical solution shown in the above method embodiment, and its implementation principle and beneficial effects are similar, which will not be elaborated here.

[0234] In one implementation manner, the processing module 52 is further configured to:

[0235] Obtain remaining resource information corresponding to the working node; wherein, the remaining resource information includes the number of remaining CPUs;

[0236] When it is determined that the number of CPUs used does not exceed the number of remaining CPUs, it is determined that the working node meets the container group creation condition.

[0237] The task execution device provided by the embodiment of the present application can execute the technical solution shown in the above method embodiment, and its implementation principle and beneficial effects are similar, which will not be elaborated here.

[0238] In one implementation, the processing module 52 is specifically configured to:

[0239] When the number of CPUs used does not exceed the number of remaining CPUs, obtain the status of the device plugin in the working node; wherein, the device plugin is used to obtain the resource information of multiple components in the working node for the container group to use the components based on the resource information; the multiple components include acceleration cards.

[0240] When it is determined that the status of the device plugin is the running state, determine that the working node meets the container group creation condition.

[0241] The task execution device provided by the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and its implementation principle and beneficial effects are similar, and will not be elaborated here.

[0242] In one implementation, the processing module 52 is further configured to:

[0243] Obtain the status of the historical task according to the identifier of the historical task;

[0244] Determine that the status of the historical task is not the running state.

[0245] The task execution device provided by the embodiments of the present application can execute the technical solutions shown in the above method embodiments, and its implementation principle and beneficial effects are similar, and will not be elaborated here.

[0246] Figure 6 It is a structural diagram of an electronic device provided by the present application. As Figure 6 shown, the electronic device 60 includes a processor 61 and a memory 62. Among them, the processor 61 is communicatively connected to the memory 62, and the memory 62 is used to store computer execution instructions; the processor 61 is configured to execute the technical solutions in any of the foregoing method embodiments by executing the computer execution instructions stored in the memory 62.

[0247] Optionally, the memory 62 can be either independent or integrated with the processor 61. Optionally, when the memory 62 is a device independent of the processor 61, the electronic device 60 may further include: a bus 63 for connecting the above devices.

[0248] This electronic device is used to execute the technical solutions in any of the foregoing method embodiments, and its implementation principle and technical effects are similar, and will not be elaborated here.

[0249] The embodiments of the present application further provide a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the technical solutions provided in any of the foregoing method embodiments.

[0250] The embodiments of the present application also provide a computer program product, including a computer program which, when executed by a processor, is used to implement the technical solutions provided by the foregoing method embodiments.

[0251] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0252] It should be further noted that although the steps in the flowchart are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limitation, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0253] It should be understood that the above device embodiments are illustrative, and the devices of the present application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units, modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed.

[0254] In addition, without special description, in each embodiment of the present application, each functional unit / module can be integrated in one unit / module, or each unit / module can exist physically alone, or two or more units / modules can be integrated together. The above integrated unit / module can be implemented in the form of hardware or in the form of a software program module.

[0255] When the integrated unit / module is implemented in the form of hardware, the hardware can be a digital circuit, an analog circuit, etc. The physical implementation of the hardware structure includes but is not limited to transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.

[0256] When the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. And the aforementioned memory includes: USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs, etc., all kinds of media that can store program codes.

[0257] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0258] Other embodiments of the present application will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only illustrative, and the true scope and spirit of the present application are pointed out by the following claims.

[0259] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A task execution method, characterized in that: Applied to a management node, the method comprises: Acquire a task execution request sent by a terminal device; the task execution request includes an identifier of a historical task; the historical task is a task executed in a historical container group in an accelerator card mode; Determine a working node for executing the historical task according to the identifier of the historical task; wherein the working node includes an accelerator card; Determine the available resource information corresponding to the mode without accelerator card; According to the available resource information, the working node is controlled to create a container group based on the historical image stored in the working node, and the working node is controlled to execute the task in the container group in a mode without an accelerator card; wherein the historical image is an image used to create the historical container group.

2. The method according to claim 1, characterized in that The determining of available resource information corresponding to the mode without an accelerator card includes: Determine a resource group according to the identifier of the historical task; wherein the working node executing the historical task belongs to the resource group; According to the resource grouping and the non-acceleration card mode, the corresponding available resource information is determined; wherein the available resource information includes the number of central processing units CPU used and the size of the used memory.

3. The method according to claim 2, characterized in that The step of controlling the working node to create a container group based on the historical image stored in the working node according to the available resource information, and controlling the working node to execute the task in the container group in a mode without an accelerator card, includes: Acquire resource restriction information; wherein the resource restriction information indicates that the accelerator card is not visible in the container group; According to the resource restriction information and the available resource information, the working node is controlled to create the container group based on the historical image, and the working node is controlled to execute the task in the container group in the non-acceleration card mode.

4. The method according to claim 3, characterized in that Before obtaining the resource limitation information, the method further includes: Obtaining the remaining resource information corresponding to the working node; wherein the remaining resource information includes the number of remaining CPUs; When it is determined that the number of used CPUs does not exceed the number of remaining CPUs, it is determined that the working node meets the container group creation condition.

5. The method according to claim 4, characterized in that The determining that the working node meets the container group creation condition when determining that the number of used CPUs does not exceed the number of remaining CPUs includes: When the number of used CPUs does not exceed the number of remaining CPUs, the state of the device plug-in in the working node is obtained; wherein the device plug-in is used to obtain resource information of multiple components of the working node, so that the container group can use the components based on the resource information; the multiple components include the accelerator card; When it is determined that the state of the device plug-in is a running state, it is determined that the working node meets the container group creation condition.

6. The method according to any one of claims 1 to 5, characterized in that: Before determining the working node for executing the historical task according to the identifier of the historical task, the method further includes: According to the identifier of the historical task, obtaining the status of the historical task; It is determined that the status of the historical task is not a running status.

7. A task execution device, characterized in that: Applied to a management node, the device comprises: An acquisition module, used for acquiring a task execution request sent by a terminal device; the task execution request includes an identifier of a historical task; the historical task is a task executed in a historical container group in an accelerator card mode; A processing module, used to determine a working node for executing the historical task according to the identifier of the historical task; wherein the working node includes an accelerator card; The processing module is further used to determine available resource information corresponding to the mode without an accelerator card; The processing module is also used to control the working node to create a container group based on the historical image stored in the working node according to the available resource information, and control the working node to execute the task in the container group in a mode without an accelerator card; wherein the historical image is an image used to create the historical container group.

8. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.

10. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.