A method, device, equipment and storage medium for allocating computing power resources
By configuring containers for each business and configuring computing resource quotas in the container, the problem of computing resource preemption between different businesses is solved, ensuring the operation isolation of instance tasks of the same business is improved, and the user experience is improved.
Patent Information
- Application Number
- CN202011395313.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-03
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-12-03
AI Technical Summary
During the AI model training process, the problem of computing resource seizing between different businesses has affected the user experience.
By configuring the corresponding container for each business and configuring the computing resource quota in each container, we ensure that the instance tasks of the same business run in the corresponding container, avoiding cross-service computing resource preemption.
It realizes the isolation of operation of instance tasks for the same business, avoids the problem of computing resource seizing between different businesses, and improves the user experience.
Smart Images

Figure CN112380020B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a computing resource allocation method, device, equipment and storage medium. Background Art
[0002] With the rapid development of science and technology, various advanced technologies are constantly emerging. Graphics Processing Unit (GPU) is becoming more and more popular due to its excellent computing power. GPU is often used for computing in various scenarios. For example, it is used for artificial intelligence (AI) model training.
[0003] Different businesses need to share computing resources during AI model training. For businesses using computing resources, the current solution is to provide a unified task layer. Multiple business instances are placed in a unified task layer queue waiting for scheduling. The scheduling strategy is basically first come first served or configure the priority of the business, and grab computing resources according to the priority.
[0004] However, this approach expands the scope of priority preemptive scheduling between different businesses, resulting in the problem of computing power resource preemption between different businesses, affecting user experience. Summary of the invention
[0005] In order to solve the above technical problems, the present application provides a computing power resource allocation method, device, equipment and storage medium to ensure that the operation of instance tasks belonging to the same business is isolated in the corresponding container. Even if priority preemption occurs during scheduling, it is only constrained within the container, avoiding the problem of computing power resource preemption between different businesses, which affects the operation of other businesses and further affects the user experience.
[0006] The embodiments of the present application disclose the following technical solutions:
[0007] In a first aspect, an embodiment of the present application provides a computing resource allocation method, where different business configurations correspond to containers, and each container configuration corresponds to a computing resource quota, and the method includes:
[0008] Converting the acquired first instance into an instance task, where the first instance belongs to the target business;
[0009] Determine the target container corresponding to the target service according to the correspondence between the service and the container;
[0010] If it is determined that there is no available computing power resource in the target container according to the computing power resource quota corresponding to the target container, controlling the instance task corresponding to the first instance to enter the task queue of the target container;
[0011] When it is detected that there are available computing resources in the target container, the target instance task is scheduled from the task queue to run.
[0012] In a second aspect, an embodiment of the present application provides a computing resource allocation device, wherein different business configurations correspond to containers, and each container is configured with a corresponding computing resource quota, and the device includes a conversion unit, a determination unit, an entry unit, and a scheduling unit:
[0013] The conversion unit is used to convert the acquired first instance into an instance task, where the first instance belongs to the target business;
[0014] The determining unit is used to determine the target container corresponding to the target business according to the correspondence between the business and the container;
[0015] The entry unit is configured to control the instance task corresponding to the first instance to enter the task queue of the target container if it is determined that there are no available computing resources in the target container according to the computing resource quota corresponding to the target container;
[0016] The scheduling unit is used to schedule the target instance task from the task queue to run when it is detected that there are available computing resources in the target container.
[0017] In a third aspect, an embodiment of the present application provides a device for allocating computing resources, wherein the electronic device includes a processor and a memory:
[0018] The memory is used to store program code and transmit the program code to the processor;
[0019] The processor is configured to execute the method described in the first aspect according to instructions in the program code.
[0020] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute the method described in the first aspect.
[0021] It can be seen from the above technical solution that the present application configures corresponding containers for different business configurations, thereby isolating different businesses through containers, and the instance tasks corresponding to each business run in the corresponding container without interfering with other tasks. The container may have the ability to manage and control computing resources, that is, configure the corresponding computing resource quota in each container to ensure the running quality of multiple instance tasks. In this way, when the submitted instance is obtained for the target business, taking the first instance as an example, the obtained first instance can be converted into an instance task, and the target container corresponding to the target business is determined according to the correspondence between the business and the container. If it is determined that there are no available computing resources in the target container according to the computing resource quota corresponding to the target container, the instance task corresponding to the first instance is controlled to enter the task queue of the target container, thereby realizing the use of the computing resources configured by the target container to run the first instance without preempting the computing resources of other businesses. When available computing resources are detected in the target container, the target instance task is scheduled from the task queue and put into operation. That is, the instance tasks belonging to the same business enter the task queue of the corresponding container and wait for scheduling. Once available computing resources appear in the container, the target instance task is scheduled from the task queue and put into operation, ensuring that the instance tasks belonging to the same business are isolated in the corresponding container. Even if priority preemption occurs during scheduling, it is only constrained within the container, avoiding the problem of computing resource preemption between different businesses, which affects the operation of other businesses and further affects the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technical members in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0023] Figure 1 A schematic diagram of the system architecture of a computing resource allocation method provided in an embodiment of the present application;
[0024] Figure 2 A flowchart of a computing resource allocation method provided in an embodiment of the present application;
[0025] Figure 3 A schematic diagram of a system architecture for allocating computing resources based on namespaces provided in an embodiment of the present application;
[0026] Figure 4 A schematic diagram of a process of priority preemption in a single namespace provided in an embodiment of the present application;
[0027] Figure 5 A flowchart of a method for expanding the capacity of a target container provided in an embodiment of the present application;
[0028] Figure 6 A flowchart of a computing resource allocation method provided in an embodiment of the present application;
[0029] Figure 7 A structural diagram of a computing resource allocation device provided in an embodiment of the present application;
[0030] Figure 8 A structural diagram of a terminal device provided in an embodiment of the present application;
[0031] Fig. 9 A structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0032] The embodiments of the present application are described below in conjunction with the accompanying drawings.
[0033] For businesses using computing resources, the current solution is to provide a unified task layer. Multiple business instances are placed in a unified task layer queue waiting for scheduling. The scheduling strategy is basically first come first served or configure the priority of the business, and occupy computing resources according to the priority.
[0034] Taking business A and B as an example, business A has instances a1, a2, a3, etc., and business B has instances b1, b2, b3, etc. If the instance tasks corresponding to the instances of business A and B are waiting for scheduling in the task layer queue, when the resource quota of business A shows that there are available computing resources, if the priority of business B is higher than that of business A, the available computing resources corresponding to business A will be preempted by the instances of business B, resulting in a situation where business A has resource quota but no resources, causing business A to be unable to run, affecting its user experience.
[0035] In order to solve the above technical problems, an embodiment of the present application provides a business configuration container (such as a namespace, a file system, a resource view, etc.) to isolate businesses. When an instance is submitted, resource preemption is performed only based on whether there are available resources in the container corresponding to the business to which it belongs, thereby constraining resource preemption within a queue at the business granularity and avoiding resource preemption between different businesses.
[0036] It should be noted that an embodiment of the present application provides a computing power resource allocation method, which can isolate non-operating businesses through containers. When an instance is submitted, resource preemption is performed only based on whether the container corresponding to the business to which it belongs has available computing power resources, thereby constraining the computing power resource preemption within the queue at the business granularity and avoiding resource preemption between different businesses.
[0037] The method provided in the embodiment of the present application relates to the field of cloud technology, such as the field of cloud computing. Cloud computing refers to the delivery and use mode of Internet technology (IT) infrastructure, which refers to obtaining required resources through the network in an on-demand, easily scalable manner; cloud computing in a broad sense refers to the delivery and use mode of services, which refers to obtaining required services through the network in an on-demand, easily scalable manner. Such services can be IT and software, Internet-related, or other services. Cloud computing is the product of the development and integration of traditional computer and network technologies such as grid computing, distributed computing, parallel computing, utility computing, network storage technologies, virtualization, and load balancing.
[0038] The method provided in the embodiment of the present application also involves the field of artificial intelligence (AI). Artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making.
[0039] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operating / interactive systems, mechatronics, and other technologies. Artificial intelligence software technologies mainly include computer vision technology, speech processing technology, natural language processing technology, machine learning / deep learning, and autonomous driving.
[0040] With the development of artificial intelligence, AI training needs to be run in more and more scenarios to train AI models. When running AI training, computing resources, such as GPU computing resources, need to be delivered. Therefore, the method provided in the embodiment of the present application is used to put AI training of the same business into the same container to achieve isolation between different businesses, thereby avoiding resource preemption between different businesses after submitting an instance of running AI training.
[0041] The method provided in the embodiment of the present application can be applied to a data processing device, which can be a server or a terminal device. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application.
[0042] Next, we will take the data processing device, which is a server, as an example to introduce the system architecture of the computing resource allocation method. Figure 1 , Figure 1 A schematic diagram of the system architecture of the computing resource allocation method provided in the embodiment of the present application. The system architecture includes a server 101, on which a computing platform can be deployed, which provides services for various businesses. This application mainly introduces the example of the computing platform providing services for business operation AI training.
[0043] The computing power platform needs to deliver computing power resources, such as GPU computing power resources, to run AI training for service businesses. To ensure isolation between businesses, different businesses are configured with corresponding containers, and each container is configured with a corresponding computing power resource quota.
[0044] Container is a virtualization technology in computer operating systems. This technology allows processes to run in a relatively independent and isolated environment, thereby simplifying the software deployment process, enhancing the portability and security of software, and improving the utilization of computing resources. Container technology is widely used in service scenarios in the field of cloud computing. In this embodiment, the container can be an independent file system, namespace, or resource view, etc. This application mainly introduces the container as the namespace.
[0045] Containers can have the ability to manage and control business computing resources, that is, within a single container, quotas are set for containers, and corresponding computing resource quotas are configured for containers to ensure the quality of multi-instance operation. This is called container quotas. If the container is a namespace, it becomes a namespace quota. Data processing devices such as servers can be equipped with GPU cards, central processing units (CPUs), memory, and disk space. Namespaces can be obtained by dividing the above GPU cards, CPUs, memory, and disk space.
[0046] A business is divided into multiple instances during AI training operation, and the multiple instances share the computing power resources in the container corresponding to the business. When a certain instance of the target business, such as the first instance, is submitted for operation, the server 101 can convert the acquired first instance into an instance task. The server 101 determines the target container corresponding to the target business based on the correspondence between the business and the container, so as to determine whether there are available computing power resources in the target container (i.e., whether the computing power resource quota has been fully used) according to the computing power resource quota corresponding to the target container. If it is determined that there are available computing power resources in the target container according to the computing power resource quota corresponding to the target container, the instance task corresponding to the first instance is executed. If it is determined that there are no available computing power resources in the target container according to the computing power resource quota corresponding to the target container, the server 101 controls the instance task corresponding to the first instance to enter the task queue of the target container and wait for operation. When available computing power resources are detected in the target container, the server 101 schedules the target instance task from the task queue for operation.
[0047] The target instance may be any instance task in the task queue, may be determined based on priority, may be confirmed based on the first-in-first-out principle, etc. In some cases, the target instance task may be the instance task corresponding to the first instance, etc.
[0048] Next, the computing power resource allocation method provided in the embodiment of the present application will be introduced in detail by mainly taking the container as a namespace as an example and combining with the accompanying drawings.
[0049] See also Figure 2 , Figure 2 A flowchart of a computing resource allocation method is shown, the method comprising:
[0050] S201: A data processing device converts an acquired first instance into an instance task, where the first instance belongs to a target business.
[0051] In this embodiment, corresponding containers can be configured for different services, and each container can be configured with a corresponding computing resource quota. In some cases, computing resource quotas can also be configured for containers according to the instances in the container, that is, each instance is configured with a corresponding computing resource quota, and the computing resource quota of the container is the sum of the computing resource quotas of the instances in the container.
[0052] The computing resource quota corresponding to each container and / or the computing resource quota of the instance can be recorded in a database (DB), see Figure 3As shown. As AI training progresses, computing resources may be gradually occupied, and the database can also record the size of the occupied computing resources or the size of the available computing resources at any time, so as to facilitate the subsequent confirmation of whether the computing resource quota has been fully used, that is, whether there are available computing resources in the container. Of course, it can also be recorded on the data processing device that executes the computing resource allocation method, which is not limited in this embodiment.
[0053] If a data processing device is set with the number of GPU cards, and the corresponding CPU, memory, and disk space, then when dividing the container (such as namespace), it can be divided equally according to the number of GPU cards, and the CPU, memory, and disk space corresponding to each GPU card are obtained by dividing equally according to the number of GPU cards. For example, if a GPU device has 8 GPUs, 96-core CPUs, 512G memory, and 4T disk space, it can be divided into 8 parts according to the GPU cards, and each CPU, memory, and disk space are 12 cores, 64G, and 500G respectively.
[0054] See also Figure 3 As shown, Figure 3 For example, if the data processing device is a GPU device and the container is a namespace, if the namespace corresponding to the business includes the namespace AD, each namespace is assigned a namespace quota. The computing power resource quota of the configured namespace can be represented by a quota card, for example Figure 3 In the example, the quota of namespace A is M card, the quota of namespace B is N card, the quota of namespace C is K card, and the quota of namespace D is L card. Each namespace can run instances of the corresponding business, and each instance is configured with a computing resource quota. For example, the instances running on namespace A include instances A1, A2, etc., the quota of instance A1 is a1 card, and the quota of instance A2 is a2 card; the instances running on namespace B include instances B1, B2, etc., the quota of instance B1 is b1 card, and the quota of instance B2 is b2 card; the instances running on namespace C include instances C1, C2, etc., the quota of instance C1 is c1 card, and the quota of instance C2 is c2 card; the instances running on namespace D include instances D1, D2, etc., the quota of instance D1 is d1 card, and the quota of instance D2 is d2 card.
[0055] When the first instance of the target service is submitted for execution, the data processing device may convert the first instance into an instance task for scheduling and execution. Figure 4 As shown, if the first instance is represented by A1, after the first instance is obtained, the first instance can be converted into an instance task, and the instance task corresponding to the first instance can be represented by instance task A1.
[0056] S202: The data processing device determines a target container corresponding to the target service according to a correspondence between services and containers.
[0057] S203: If the data processing device determines that there is no available computing power resource in the target container according to the computing power resource quota corresponding to the target container, control the instance task corresponding to the first instance to enter the task queue of the target container.
[0058] Since different services are configured with corresponding containers, the data processing device can determine the corresponding target container according to the target service to which the first instance belongs, so as to know whether its computing power resource quota can meet the operation of the service instance. In a possible implementation, the first instance obtained by the data processing device can have a corresponding service identifier, so as to determine the target container corresponding to the target service according to the service identifier and the corresponding relationship.
[0059] The computing power resource quota corresponding to the target container can be read by the data processing device from the database. In order to facilitate the determination of whether there are available computing power resources in the target container, the size of the occupied computing power resources can also be read from the database to determine whether the computing power resource quota corresponding to the target container has been fully used.
[0060] If the data processing device determines that there are available computing resources in the target container, the instance task corresponding to the first instance is executed, for example Figure 4 As shown, the instance task A1 corresponding to the first instance is put into the scheduler to run the instance task. If the data processing device determines that there is no available computing power resource in the target container, the instance task corresponding to the first instance is controlled to enter the task queue of the target container and wait.
[0061] Figure 4 It also includes other instances submitted for running, such as instance A2, instance A3, and instance A4, and the corresponding instance tasks are instance task A2, instance task A3, and instance task A4. According to the above steps S201-S203, it can be determined whether instance task A2, instance task A3, and instance task A4 are running or entering the task queue. Figure 4 For example, instance task A1 runs directly, and instance tasks A2, A3, and A4 enter the task queue and wait. When instance A1 has a new task running, if it is determined again that the computing power resource quota is full, instance task A1 corresponding to instance A1 will also enter the task queue and wait.
[0062] In some cases, the computing power resource quota of each container is configured according to the instance, and the computing power resource quota of the target container is the sum of the computing power resource quotas of the instances in the target container. If it is determined that there are available computing power resources in the target container, but since each instance is configured with a corresponding computing power resource quota, the available computing power resources are not necessarily the computing power resources used to run the first instance. In this case, the data processing device can further determine whether the computing power resource quota of the first instance is not fully used. If it is determined that the computing power resource quota of the first instance is fully used, execute step S203.
[0063] S204: When the data processing device detects that there are available computing resources in the target container, the target instance task is scheduled from the task queue to run.
[0064] The data processing device can periodically monitor whether there are available computing resources in the target container. When there are available computing resources, that is, the computing resource quota of the target container becomes free, the target instance task is selected from the task queue and put into operation. For example, the scheduler periodically polls the task queue, and when there are available computing resources, the scheduler selects the target instance task and puts it into operation. Figure 4 As shown, if the instance task Ax is determined as the target instance task according to the priority, the scheduler will give priority to selecting the instance task Ax for operation.
[0065] It should be noted that the instance tasks waiting in the task queue may include multiple ones, that is, in addition to the instance tasks corresponding to the first instance, other instance tasks may also be included. The scheduler can select which instance tasks to put into operation from the task queue based on different principles, such as the priority preemption principle and the first-in-first-out principle (that is, the instance tasks that enter the task queue first are scheduled to be put into operation by the scheduler first). In a possible implementation, taking the instance tasks corresponding to the first instance and the instance tasks corresponding to the second instance included in the task queue as an example, the second instance also belongs to the target business, and the implementation method of S204 can be to obtain the priority of the first instance and the priority of the second instance. If it is determined that the priority of the first instance is higher than the priority of the second instance, the instance task corresponding to the first instance is used as the target instance task, and the instance task corresponding to the first instance is scheduled from the task queue for operation. That is, the instance task corresponding to the first instance with a higher priority can seize computing resources, and the operator resource preemption occurs in the target container and does not affect other businesses.
[0066] In one possible implementation, the computing power resource quota of each container is configured according to the instance, that is, each instance is configured with a corresponding computing power resource quota. If the data processing device detects that available computing power resources appear in the target container, it can further identify which instance the available computing power resources correspond to. If it is determined that the available computing power resources are determined according to the computing power resource quota of the first instance, that is, the available computing power resources are the computing power resources of the first instance, indicating that the computing power resource quota of the first instance is free, the data processing device takes the instance task corresponding to the first instance as the target instance task, and schedules the instance task corresponding to the first instance from the task queue for operation.
[0067] However, the priorities of multiple instance tasks included in the task queue may be different. The higher the priority, the more important the corresponding instance task is, and the more likely it is to be run first to better meet user needs. Therefore, in this case, even if the available computing power resources are the computing power resources of a certain instance, indicating that the computing power resource quota of the instance is free, but because the task queue also includes other instance tasks, it is necessary to preempt the computing power resources in the target container according to the priority of the instance task, so as to run the instance tasks with higher priority first.
[0068] Taking the example of a task queue including an instance task corresponding to the first instance and an instance task corresponding to the second instance, even if the available computing power resources are determined according to the computing power resource quota of the second instance, that is, the available computing power resources are the computing power resources of the second instance, it means that the computing power resource quota of the second instance is free. The data processing device can further obtain the priority of the first instance and the priority of the second instance. If it is determined that the priority of the first instance is higher than the priority of the second instance, the instance task corresponding to the first instance is used as the target instance task, and the instance task corresponding to the first instance is scheduled from the task queue and put into operation, that is, the first instance has seized the computing power resources of the second instance.
[0069] Although the computing power resources of the second instance are preempted by the first instance, when the data processing device monitors whether there are available computing power resources in the target container, when it is detected that there are available computing power resources in the target container again, the instance task corresponding to the second instance is scheduled from the task queue and put into operation. That is, after the computing power resources of an instance are preempted, when available computing power resources appear again, the instance task corresponding to the instance is executed first, thereby avoiding the problem of the instance being unable to run due to the computing power resources of the instance being constantly preempted by other instances with higher priority in resource delivery, that is, avoiding the situation of instance starvation.
[0070] It can be seen from the above technical solution that the present application configures corresponding containers for different business configurations, thereby isolating different businesses through containers, and the instance tasks corresponding to each business run in the corresponding container without interfering with other tasks. The container may have the ability to manage and control computing resources, that is, configure the corresponding computing resource quota in each container to ensure the running quality of multiple instance tasks. In this way, when the submitted instance is obtained for the target business, taking the first instance as an example, the obtained first instance can be converted into an instance task, and the target container corresponding to the target business is determined according to the correspondence between the business and the container. If it is determined that there are no available computing resources in the target container according to the computing resource quota corresponding to the target container, the instance task corresponding to the first instance is controlled to enter the task queue of the target container, thereby realizing the use of the computing resources configured by the target container to run the first instance without preempting the computing resources of other businesses. When available computing resources are detected in the target container, the target instance task is scheduled from the task queue and put into operation. That is, the instance tasks belonging to the same business enter the task queue of the corresponding container and wait for scheduling. Once available computing resources appear in the container, the target instance task is scheduled from the task queue and put into operation, ensuring that the instance tasks belonging to the same business are isolated in the corresponding container. Even if priority preemption occurs during scheduling, it is only constrained within the container, avoiding the problem of computing resource preemption between different businesses, which affects the operation of other businesses and further affects the user experience.
[0071] In addition, multi-tenant isolation of the business is achieved through containers, and task scheduling is integrated according to the granularity of the business, which facilitates the operation of the computing power platform.
[0072] As the number of business instances in the container increases, the computing resource quota of the container may not be able to meet the business needs, so the container can be expanded. Taking the case where the target container needs to be expanded, the data processing device can receive an expansion application, which includes a quota expansion capacity, which is determined based on the newly added instances. The data processing device adjusts the computing resource quota of the target container based on the quota expansion capacity. Among them, the expansion application can be triggered by the operator on the computing platform.
[0073] When applying to expand the target container, since the computing resources used for the expansion can be obtained from the resource pool, the computing resources in the resource pool are also limited, and the computing resource quota of the target container cannot be expanded infinitely. Therefore, the flowchart of the method for expanding the target container can be found in Figure 5 As shown, the method includes:
[0074] S501: A data processing device receives a capacity expansion application.
[0075] S502: The data processing device determines whether the amount of idle resources in the resource pool meets the quota expansion capacity. If not, execute S503; if yes, execute S504.
[0076] S503: The data processing device expands the resource pool.
[0077] If it is determined that the amount of free resources in the resource pool does not meet the quota expansion capacity, the resource pool can be expanded first, and then the computing power resource quota of the target container can be adjusted according to the quota expansion capacity using the expanded resource pool, that is, S504 is executed to expand the computing power resource quota of the target container.
[0078] S504: The data processing device adjusts the computing resource quota of the target container according to the quota expansion capacity.
[0079] S505: The data processing device adjusts the amount of idle resources in the resource pool according to the quota expansion capacity.
[0080] After adjusting the computing resource quota of the target container, the amount of idle resources in the resource pool will decrease, so the amount of idle resources in the resource pool can be updated according to the quota expansion capacity, so that when the capacity is expanded again later, step S502 can be executed according to the accurate amount of idle resources.
[0081] The embodiment of the present application isolates the business through containers. When the number of instances in a business container increases, the operator of the computing power platform can evaluate the increase in computing power resources of the container (i.e., quota expansion), which solves the problem that it is difficult to evaluate quota expansion at the traditional unified task layer.
[0082] Next, the computing power resource allocation method provided in the embodiment of the present application will be introduced in combination with actual application scenarios. In the scenario where the business operation AI training requires GPU computing power resources, in order to avoid computing power resource preemption between businesses, namespaces can be configured for the business to isolate the businesses. Different namespaces share the physical resources of GPU computing power, but the logical quota of GPU computing power of the namespace is not shared. A business is divided into multiple instances during AI training operation. Under the premise that multiple instances share the GPU computing power resources in the same namespace, their respective computing power resource quotas are configured according to the instances. Based on this, taking the data processing device as a server as an example, when an instance, such as the first instance, is submitted for operation, it can be based on Figure 6 The method shown implements computing resource allocation, and the method includes:
[0083] S601: The server converts the acquired first instance into an instance task.
[0084] S602: If the server determines that there are available computing resources in the target namespace, the instance task is run.
[0085] Since the first instance belongs to the target business, this embodiment configures the business namespace to isolate the businesses. Each business has a corresponding relationship with the namespace. The first instance obtained by the server can have a corresponding business identifier, so as to determine the target container corresponding to the target business according to the business identifier and the corresponding relationship.
[0086] If the server determines that the computing power resource quota of the target namespace is not fully used, that is, there are available computing power resources, the instance task corresponding to the first instance can be directly run.
[0087] S603: If the server determines that there are no available computing resources in the target namespace, control the instance task corresponding to the first instance to enter the task queue of the target namespace.
[0088] If the server determines that the computing resource quota of the target namespace has been fully used, that is, there are no available computing resources, the instance task corresponding to the first instance is controlled to enter the task queue and wait.
[0089] S604: The server polls and monitors whether there are available computing resources in the target namespace.
[0090] S605: The server detects that available computing resources appear in the target namespace, and schedules the target instance task from the task queue to run.
[0091] There may be multiple instance tasks waiting in the task queue, that is, in addition to the instance task corresponding to the first instance, it may also include other instance tasks. The other instance tasks may enter the task queue before the instance task corresponding to the first instance, or may enter the task queue after the instance task corresponding to the first instance. This embodiment does not limit this.
[0092] The server can periodically poll to monitor whether available computing resources appear in the target namespace. If available computing resources are detected in the target namespace, the instance tasks in the task queue can be scheduled based on the priority preemption principle, thereby constraining the priority preemption in the target namespace.
[0093] For example, the task queue includes an instance task corresponding to the first instance and an instance task corresponding to the second instance. The second instance also belongs to the target business. If the server determines that the priority of the first instance is higher than the priority of the second instance, the instance task corresponding to the first instance can preempt the available computing power resources, and the server schedules the instance task corresponding to the first instance from the task queue and puts it into operation.
[0094] Of course, the instance tasks in the task queue can also be scheduled based on the first-in-first-out principle. For example, the task queue includes the instance tasks corresponding to the first instance and the instance tasks corresponding to the second instance. The second instance also belongs to the target business, and the instance tasks corresponding to the second instance enter the task queue after the instance tasks corresponding to the first instance. If the server detects that available computing resources appear in the target namespace, even if the available computing resources are determined by the computing resource quota of the second instance, that is, the available computing resources are the computing resources of the second instance, it means that the computing resource quota of the second instance is free. Since the instance tasks corresponding to the first instance enter the task queue first, the instance tasks corresponding to the first instance can be scheduled to run first, thereby realizing resource preemption in the target namespace.
[0095] The embodiment of the present application configures corresponding namespaces for different services, so that different services are isolated through namespaces, and the instance tasks corresponding to each service are run in the corresponding namespace without interfering with other tasks. The namespace can have the ability to control computing resources, that is, configure the corresponding computing resource quota in each namespace to ensure the running quality of multiple instance tasks. In this way, when the first instance submitted is obtained for the target service, the server can convert the obtained first instance into an instance task, and determine the target namespace corresponding to the target service according to the corresponding relationship between the service and the namespace. If it is determined that there are no available computing resources in the target namespace, the instance task corresponding to the first instance is controlled to enter the task queue of the target namespace, so as to realize the operation of the first instance using the computing resources configured by the target namespace without preempting the computing resources of other services. When it is monitored that there are available computing resources in the target namespace, the target instance task is scheduled from the task queue to run, ensuring that the running of the instance tasks belonging to the same service is isolated in the corresponding namespace, even if priority preemption occurs during scheduling, it is only constrained within the namespace, avoiding the problem of computing resource preemption between different services, and affecting the operation of other services, and then affecting the user experience.
[0096] based on Figure 2 The computing resource allocation method provided in the corresponding embodiment, the embodiment of the present application also provides a computing resource allocation device, different business configurations corresponding to the container, each container configuration corresponding to the computing resource quota, see Figure 7 , the device 700 includes a conversion unit 701, a determination unit 702, an entry unit 703 and a scheduling unit 704:
[0097] The conversion unit 701 is used to convert the acquired first instance into an instance task, where the first instance belongs to the target business;
[0098] The determining unit 702 is used to determine the target container corresponding to the target business according to the correspondence between the business and the container;
[0099] The entry unit 703 is used to control the instance task corresponding to the first instance to enter the task queue of the target container if it is determined that there is no available computing resource in the target container according to the computing resource quota corresponding to the target container;
[0100] The scheduling unit 704 is used to schedule the target instance task from the task queue to run when it is detected that there are available computing resources in the target container.
[0101] In a possible implementation manner, if the task queue further includes an instance task corresponding to a second instance, and the second instance belongs to the target service, the scheduling unit 704 is configured to:
[0102] Obtaining a priority of the first instance and a priority of the second instance;
[0103] If it is determined that the priority of the first instance is higher than the priority of the second instance, taking the instance task corresponding to the first instance as the target instance task;
[0104] The instance task corresponding to the first instance is scheduled from the task queue for operation.
[0105] In a possible implementation, the computing resource quota of each container is configured according to the instance, the computing resource quota of the target container is the sum of the computing resource quotas of the instances in the target container, and the determining unit 702 is further used to:
[0106] If it is determined that there are available computing resources in the target container according to the computing resource quota corresponding to the target container, determine whether the computing resource quota of the first instance is not fully used;
[0107] If the determining unit determines that the computing resource quota of the first instance has been fully used, the entry unit is triggered to execute the step of controlling the instance task corresponding to the first instance to enter the task queue of the target container.
[0108] In a possible implementation, the scheduling unit 704 is configured to:
[0109] If it is determined that the available computing power resources are determined according to the computing power resource quota of the first instance, the instance task corresponding to the first instance is used as the target instance task;
[0110] The instance task corresponding to the first instance is scheduled from the task queue for operation.
[0111] In a possible implementation, if the task queue also includes an instance task corresponding to a second instance, the second instance belongs to the target business, and if the available computing power resources are determined according to the computing power resource quota of the second instance, the scheduling unit 704 is used to:
[0112] Obtaining a priority of the first instance and a priority of the second instance;
[0113] If it is determined that the priority of the first instance is higher than the priority of the second instance, taking the instance task corresponding to the first instance as the target instance task;
[0114] The instance task corresponding to the first instance is scheduled from the task queue for operation.
[0115] In a possible implementation, after scheduling the instance task corresponding to the first instance from the task queue to be put into operation, the scheduling unit 704 is further configured to:
[0116] When it is again detected that there are available computing resources in the target container, the instance task corresponding to the second instance is scheduled from the task queue to run.
[0117] In a possible implementation manner, the device further includes a receiving unit and an adjusting unit:
[0118] The receiving unit is used to receive a capacity expansion application, wherein the capacity expansion application includes a quota capacity expansion, and the quota capacity expansion is determined according to a newly added instance;
[0119] The adjustment unit is used to adjust the computing resource amount of the target container according to the quota expansion capacity.
[0120] In a possible implementation manner, the determining unit 702 is further configured to:
[0121] Determine whether the amount of idle resources in the resource pool meets the quota expansion capacity;
[0122] If satisfied, trigger the adjustment unit to execute the step of adjusting the computing resource amount of the target container according to the quota expansion capacity;
[0123] The adjustment unit is further configured to adjust the amount of idle resources in the resource pool according to the quota expansion capacity.
[0124] In a possible implementation manner, the adjusting unit is further configured to:
[0125] If the determining unit 702 determines that the amount of idle resources in the resource pool does not meet the quota expansion capacity, the resource pool is expanded;
[0126] The computing resource amount of the target container is adjusted according to the quota expansion capacity by using the expanded resource pool.
[0127] The embodiment of the present application further provides a device for computing resource allocation, which may be a data processing device for executing a computing resource allocation method, and which may be a terminal device. For example, the terminal device is a smart phone:
[0128] Figure 8 The block diagram shows a partial structure of a smart phone related to the terminal device provided in the embodiment of the present application. Figure 8 The smartphone includes: a radio frequency (RF) circuit 810, a memory 820, an input unit 830, a display unit 840, a sensor 850, an audio circuit 860, a wireless fidelity (WiFi) module 870, a processor 880, and a power supply 890. The input unit 830 may include a touch panel 831 and other input devices 832, the display unit 840 may include a display panel 841, and the audio circuit 860 may include a speaker 861 and a microphone 862. Those skilled in the art will appreciate that Figure 8 The structure of the smartphone shown in the figure does not constitute a limitation of the smartphone, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0129] The memory 820 can be used to store software programs and modules. The processor 880 executes various functional applications and data processing of the smartphone by running the software programs and modules stored in the memory 820. The memory 820 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the smartphone (such as audio data, a phone book, etc.), etc. In addition, the memory 820 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0130] The processor 880 is the control center of the smartphone, which uses various interfaces and lines to connect various parts of the entire smartphone, and executes various functions of the smartphone and processes data by running or executing software programs and / or modules stored in the memory 820, and calling data stored in the memory 820. Optionally, the processor 880 may include one or more processing units; preferably, the processor 880 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 880.
[0131] In this embodiment, the processor 880 in the terminal device 800 may perform the following steps:
[0132] Converting the acquired first instance into an instance task, where the first instance belongs to the target business;
[0133] Determine the target container corresponding to the target service according to the correspondence between the service and the container;
[0134] If it is determined that there is no available computing power resource in the target container according to the computing power resource quota corresponding to the target container, controlling the instance task corresponding to the first instance to enter the task queue of the target container;
[0135] When it is detected that there are available computing resources in the target container, the target instance task is scheduled from the task queue to run.
[0136] The device may also include a server. The embodiment of the present application also provides a server. Fig. 9 As shown, Fig. 9 The structural diagram of the server 900 provided in the embodiment of the present application, the server 900 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 922 (for example, one or more processors) and a memory 932, and one or more storage media 930 (for example, one or more mass storage devices) storing application programs 942 or data 944. Among them, the memory 932 and the storage medium 930 can be temporary storage or permanent storage. The program stored in the storage medium 930 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 922 can be configured to communicate with the storage medium 930 and execute a series of instruction operations in the storage medium 930 on the server 900.
[0137] The server 900 may also include one or more power supplies 926, one or more wired or wireless network interfaces 950, one or more input and output interfaces 958, and / or one or more operating systems 941, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0138] In this embodiment, the central processor 922 in the server 900 may perform the following steps:
[0139] Converting the acquired first instance into an instance task, where the first instance belongs to the target business;
[0140] Determine the target container corresponding to the target service according to the correspondence between the service and the container;
[0141] If it is determined that there is no available computing power resource in the target container according to the computing power resource quota corresponding to the target container, controlling the instance task corresponding to the first instance to enter the task queue of the target container;
[0142] When it is detected that there are available computing resources in the target container, the target instance task is scheduled from the task queue to run.
[0143] According to one aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute the computing power resource allocation method described in the aforementioned embodiments.
[0144] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in various optional implementations of the above embodiments.
[0145] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0146] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0147] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0148] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0149] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (RandomAccess Memory, referred to as RAM), disk or optical disk and other media that can store program codes.
[0150] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, ordinary technical members in the art should understand that they can still modify the technical solutions recorded in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for allocating computing resources. It is characterized in that Different business configurations correspond to containers, each container is configured with a corresponding computing resource quota, and the computing resource quota corresponding to the container is associated with an instance in the container. The method includes: Converting the acquired first instance into an instance task, where the first instance belongs to the target business; Determine the target container corresponding to the target service according to the correspondence between the service and the container; If it is determined that there is no available computing power resource in the target container according to the computing power resource quota corresponding to the target container, controlling the instance task corresponding to the first instance to enter the task queue of the target container, the task queue including the instance task corresponding to the second instance, and the second instance belongs to the target business; If the available computing resources are determined according to the computing resource quota of the second instance, when it is detected that there are available computing resources in the target container, the priority of the first instance and the priority of the second instance are obtained; If it is determined that the priority of the first instance is higher than the priority of the second instance, taking the instance task corresponding to the first instance as the target instance task; Scheduling the instance task corresponding to the first instance from the task queue for operation; When it is again detected that there are available computing resources in the target container, the instance task corresponding to the second instance is scheduled from the task queue to run.
2. The method according to claim 1, It is characterized in that The computing resource quota of each container is configured according to the instance, the computing resource quota of the target container is the sum of the computing resource quotas of the instances in the target container, and the method further includes: If it is determined that there are available computing resources in the target container according to the computing resource quota corresponding to the target container, determine whether the computing resource quota of the first instance is not fully used; If it is determined that the computing power resource quota of the first instance has been fully used, a step of controlling the instance task corresponding to the first instance to enter the task queue of the target container is performed.
3. The method according to claim 2, It is characterized in that When it is detected that there are available computing resources in the target container, the target instance task is scheduled from the task queue to run, including: If it is determined that the available computing power resources are determined according to the computing power resource quota of the first instance, the instance task corresponding to the first instance is used as the target instance task; The instance task corresponding to the first instance is scheduled from the task queue for operation.
4. The method according to claim 1, It is characterized in that The method further comprises: receiving a capacity expansion application, wherein the capacity expansion application includes a quota capacity expansion, and the quota capacity expansion is determined according to a newly added instance; The computing resource amount of the target container is adjusted according to the quota expansion capacity.
5. The method according to claim 4, It is characterized in that Before adjusting the computing resource amount of the target container according to the quota expansion, the method further includes: Determine whether the amount of idle resources in the resource pool meets the quota expansion capacity; If satisfied, executing the step of adjusting the computing resource amount of the target container according to the quota expansion capacity; The amount of idle resources in the resource pool is adjusted according to the quota expansion capacity.
6. The method according to claim 5, It is characterized in that The method further comprises: If it is determined that the amount of idle resources in the resource pool does not meet the quota expansion capacity, the resource pool is expanded; The computing resource amount of the target container is adjusted according to the quota expansion capacity by using the expanded resource pool.
7. A computing resource allocation device, It is characterized in that Different business configurations correspond to containers, each container configuration corresponds to a computing resource quota, the computing resource quota corresponding to the container is associated with an instance in the container, and the device includes a conversion unit, a determination unit, an entry unit, and a scheduling unit: The conversion unit is used to convert the acquired first instance into an instance task, where the first instance belongs to the target business; The determining unit is used to determine the target container corresponding to the target business according to the correspondence between the business and the container; The entry unit is configured to control the instance task corresponding to the first instance to enter the task queue of the target container if it is determined that there are no available computing resources in the target container according to the computing resource quota corresponding to the target container; The scheduling unit is used to schedule the target instance task from the task queue to run when it is detected that there are available computing resources in the target container; If the task queue further includes an instance task corresponding to a second instance, the second instance belongs to the target business, and if the available computing power resources are determined according to the computing power resource quota of the second instance, the scheduling unit is configured to: Obtaining a priority of the first instance and a priority of the second instance; If it is determined that the priority of the first instance is higher than the priority of the second instance, taking the instance task corresponding to the first instance as the target instance task; Scheduling the instance task corresponding to the first instance from the task queue for operation; After scheduling the instance task corresponding to the first instance from the task queue to be put into operation, the scheduling unit is further used to: When it is again detected that there are available computing resources in the target container, the instance task corresponding to the second instance is scheduled from the task queue to run.
8. The device according to claim 7, It is characterized in that The computing resource quota of each container is configured according to the instance, the computing resource quota of the target container is the sum of the computing resource quotas of the instances in the target container, and the determining unit is further used for: If it is determined that there are available computing resources in the target container according to the computing resource quota corresponding to the target container, determine whether the computing resource quota of the first instance is not fully used; If the determining unit determines that the computing resource quota of the first instance has been fully used, the entry unit is triggered to execute the step of controlling the instance task corresponding to the first instance to enter the task queue of the target container.
9. The device according to claim 8, It is characterized in that The scheduling unit is used to: If it is determined that the available computing power resources are determined according to the computing power resource quota of the first instance, the instance task corresponding to the first instance is used as the target instance task; The instance task corresponding to the first instance is scheduled from the task queue for operation.
10. The device according to claim 7, It is characterized in that The device also includes a receiving unit and an adjusting unit: The receiving unit is used to receive a capacity expansion application, wherein the capacity expansion application includes a quota capacity expansion, and the quota capacity expansion is determined according to a newly added instance; The adjustment unit is used to adjust the computing resource amount of the target container according to the quota expansion capacity.
11. The device according to claim 10, It is characterized in that The determining unit is further configured to: Determine whether the amount of idle resources in the resource pool meets the quota expansion capacity; If satisfied, trigger the adjustment unit to execute the step of adjusting the computing resource amount of the target container according to the quota expansion capacity; The adjustment unit is further configured to adjust the amount of idle resources in the resource pool according to the quota expansion capacity.
12. The device according to claim 11, It is characterized in that The adjustment unit is also used for: If the determining unit determines that the amount of idle resources in the resource pool does not meet the quota expansion capacity, the resource pool is expanded; The computing resource amount of the target container is adjusted according to the quota expansion capacity by using the expanded resource pool.
13. A device for allocating computing resources, It is characterized in that The device comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the method according to any one of claims 1 to 6 according to the instructions in the program code.
14. A computer-readable storage medium, It is characterized in that The computer-readable storage medium is used to store program codes, and the program codes are used to execute the method according to any one of claims 1 to 6.
15. A computer program product, It is characterized in that The computer program product comprises instructions, which, when executed on a computer device, cause the computer device to perform the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Operation request method and device, electronic equipment and storage medium
CN109558446A
Virtual instance scheduling system and method based on cloud platform
CN109962940A
Task scheduling method and device, electronic equipment and computer readable storage medium
CN110837410A