Resource isolation method, distributed platform, computer device and storage medium
By introducing resource isolation methods into the distributed platform, we judge the resource requirements of task instances and allocate resources, the conflict problems in resource sharing between task instances are solved, and the resource utilization rate and task execution success rate are improved.
Patent Information
- Application Number
- CN201910541011.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-06-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2039-06-21
AI Technical Summary
In the prior art, there are conflicts in resource sharing between task instances, resulting in low resource utilization and inability to effectively avoid resource conflicts.
By introducing resource isolation methods into the distributed platform, we can obtain the process to be run in the task instance and determine whether it needs to consume target resources. If necessary, it is determined whether the sum of the first resource quantity and the second resource quantity is greater than the resource application quantity. If it exceeds the resource application, the identification information of the resource application failure will be returned, otherwise the target resource will be allocated.
It realizes the mutual isolation of resource applications between task instances, avoids resource conflicts, improves resource utilization, and enables task instances to be successfully executed.
Smart Images

Figure CN112114958B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of distributed platform technology, and in particular to a resource isolation method, a distributed platform, a computer device and a storage medium. Background Art
[0002] In order to improve the task processing capability and reliability of a single node, the prior art proposes a distributed platform that centrally manages the resources in several physical server nodes or virtual machine nodes to respond to task requests. In order to improve the resource utilization of the distributed platform, there is a related study in the prior art that over-divides the physical devices in the node so that the logical number of physical devices in the node is greater than the actual number of physical devices, thereby realizing shared scheduling of physical devices.
[0003] However, the inventors have discovered that the current method of super-dividing physical devices is to super-divide one physical device into two, so that one physical device is super-divided into two logical devices. When the two logical devices are assigned to two task instances, it is equivalent to two task instances sharing the same physical device. This shared scheduling method still has the following problems: in the prior art, some processes of a task instance will occupy all of some resources on a physical device, so when other task instances are assigned to the physical device, this type of resource conflict will occur.
[0004] Therefore, providing a resource scheduling method, distributed platform, computer device and storage medium to further improve resource utilization and reduce resource conflicts has become a technical problem that urgently needs to be solved in this field. Summary of the invention
[0005] The purpose of the present invention is to provide a resource isolation method, a distributed platform, a computer device and a storage medium, which are used to solve the above-mentioned technical problems existing in the prior art.
[0006] To achieve the above objective, the present invention provides a resource isolation method.
[0007] The resource isolation method includes: obtaining a process to be run in a task instance; judging whether the process to be run needs to consume target resources, wherein the task instance includes one or more processes; when the process to be run needs to consume target resources, judging whether the sum of a first resource amount and a second resource amount is greater than a resource application amount, wherein the first resource amount is the amount of target resources currently occupied by the task instance, the second resource amount is the amount of target resources required by the process to be run, and the resource application amount is the amount of target resources applied for by the task instance; when the sum of the first resource amount and the second resource amount is greater than the resource application amount, returning identification information indicating a failure in the resource application of the process to be run to the task instance; and when the sum of the first resource amount and the second resource amount is less than or equal to the resource application amount, allocating the target resource of the second resource amount to the process to be run.
[0008] Furthermore, the process to be run is a process that sends an interface call request, and the step of determining whether the process to be run needs to consume the target resources is specifically: determining whether the interface called by the interface call request is an interface for applying for the target resources; wherein, when the interface called by the interface call request is an interface for applying for the target resources, the process to be run needs to consume the target resources.
[0009] Furthermore, before the step of determining whether the sum of the first resource amount and the second resource amount is greater than the resource application amount, the resource isolation method includes: determining the physical device to which the target resource that the process to be run needs to consume belongs, and obtaining the first physical device, wherein the task instance applies for the target resource on at least two physical devices when it is executed, and the first physical device is one of the at least two physical devices; obtaining the amount of the target resource that the task instance is occupying on the first physical device, and obtaining the first resource amount; obtaining the amount of the target resource applied for by the task instance on the first physical device, and obtaining the resource application amount.
[0010] Furthermore, all processes of the task instance share a resource occupancy variable. Before the step of determining whether the sum of the first resource amount and the second resource amount is greater than the resource application amount, the resource isolation method also includes: reading the value of the resource occupancy variable to obtain the first resource amount.
[0011] Furthermore, the task instance creates target resources applied for on N physical devices, the resource occupancy variable is an array, the array includes at least N elements, and the amount of target resources on each physical device being occupied by the task instance is recorded by one element.
[0012] Furthermore, the target resource is a video memory resource of a GPU physical device.
[0013] Furthermore, the resource isolation method also includes: forwarding the interface call request when the process to be run does not need to consume the target resource; the step of allocating the second amount of target resources to the process to be run is specifically: forwarding the interface call request.
[0014] To achieve the above objectives, the present invention provides a distributed platform.
[0015] The distributed platform includes: a management node and several processing nodes, the processing nodes include target resources, a task creation device and a resource processing device, the resource processing device includes a resource isolation module and a resource management module, wherein: the management node is used to schedule tasks to the processing nodes according to the information of the target resources on each processing node; the task creation device is used to create a task instance when being scheduled on the processing node; the resource management module is used to report the information of the target resources on the processing node to the management node, and is also used to allocate target resources to the task instance; and the resource isolation module is used to execute any one of the resource isolation methods provided by the present invention.
[0016] Furthermore, the processing node includes a GPU physical device, the target resource is a video memory resource on the GPU physical device, the resource isolation module is an so library, and the resource management module is also used to: mount the GPU physical device and the resource isolation module to the task instance; set the environment variables of the task instance, wherein the environment variables include the resource application amount of the target resource and the dynamic library loading variable, and the value of the dynamic library loading variable is the resource isolation module.
[0017] To achieve the above objectives, the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0018] To achieve the above object, the present invention also provides a computer-readable storage medium on which a computer program is stored, and the computer program implements the steps of the above method when executed by a processor.
[0019] The resource isolation method, distributed platform, computer equipment and storage medium provided by the present invention respectively allocate the target resources required by each task instance, and the resource application amounts obtained by each task instance are isolated from each other. When the process of each task instance is to be run, it is first determined whether the process to be run needs to consume the target resources. When it needs to consume the target resources, it is then determined whether the remaining resource application amount of the task instance can meet the needs of the process to be run. If it cannot meet the needs, identification information representing the failure of the resource application of the process to be run is returned to the task instance, so that the task instance can perform process allocation according to the process allocation mechanism, and finally the task instance is successfully executed; if it can meet the needs, the target resources are allocated to the process to be run. The resource application amounts of different task instances are isolated from each other, and do not interfere with each other, thereby avoiding conflicts in resource applications between task instances. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1A flowchart of a resource isolation method provided by an embodiment of the present invention;
[0021] Figure 2 A schematic diagram of resource isolation provided by an embodiment of the present invention;
[0022] Figure 3 A block diagram of a resource isolation device provided by an embodiment of the present invention;
[0023] Figure 4 A block diagram of a distributed platform provided by an embodiment of the present invention;
[0024] Figure 5 and Figure 6 A schematic diagram of a business processing flow of a distributed platform provided by an embodiment of the present invention;
[0025] Figure 7 A schematic diagram of applying for resources for a task instance provided in an embodiment of the present invention; and
[0026] Figure 8 A hardware structure diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0028] Currently, in a distributed platform, a graphics processing unit (GPU), as a resource set in a physical server node or a virtual machine node in a cluster, is mainly scheduled in blocks.
[0029] In order to improve the utilization rate of GPU physical devices, existing research has also made some changes to the scheduling of GPU physical devices on distributed platforms to achieve shared scheduling of GPU physical devices. However, the inventors further discovered that for tasks such as machine learning, a process will occupy all the video memory resources of the GPU physical device by default, which can easily cause other processes to fail to apply for the video memory resources of the GPU physical device, thus causing resource conflicts between task instances.
[0030] In order to solve the above technical difficulties, the present invention proposes a resource isolation method, a distributed platform, a computer device and a storage medium. In the resource isolation method, for a task instance including one or more processes, a certain amount of target resources is allocated to the task instance, that is, each task instance will apply for a certain amount of target resources according to the needs of the task instance. Based on this, during the execution of the task instance, when a process in the task instance is in a waiting state, the waiting process is obtained, and it is determined whether the waiting process needs to consume the target resources. When the waiting process needs to consume the target resources, the first resource amount (that is, the amount of the target resource being occupied by the task instance) is determined. ) and the second resource amount (that is, the amount of target resources required by the process to be run) is greater than the resource application amount (that is, the amount of target resources applied for by the task instance). When the sum of the first resource amount and the second resource amount is greater than the resource application amount, identification information representing the failure of resource application is returned to the task instance. At this time, the task instance can adjust the first resource amount by controlling the process occupying the target resource to release the target resource, and re-run the process that receives the identification information representing the failure of resource application, so as to finally make the task instance run successfully; and when the sum of the first resource amount and the second resource amount is less than or equal to the resource application amount, the target resource of the second resource amount is allocated to the process to be run. It can be seen from this that within a task instance, each process shares the resource application amount; between each task instance, the process in each task instance only consumes the target resources of the task instance in which it is located, and the target resources between each other are isolated, which will not cause resource conflicts, thereby solving the technical problem of resource conflicts in resource sharing in the prior art.
[0031] In the above resource isolation method, when the target resource is a video memory resource of a GPU physical device, the conflict of video memory resources between task instances in GPU physical device resource scheduling can be resolved.
[0032] The resource isolation method, distributed platform, computer device and storage medium provided by the present invention will be described in detail below through specific embodiments. It should be noted that, for the convenience of description, the detailed description in the following embodiments is described by taking the video memory resources of the GPU physical device as an example, but the resource isolation method of the present invention is not limited to the video memory resources of the GPU physical device.
[0033] Embodiment 1
[0034] Embodiment 1 of the present invention provides a resource isolation method. In an application scenario, the execution subject of the resource isolation method can be a resource processing device in a processing node of a distributed platform. When the management node of the distributed platform schedules a task to a processing node, a task creation device in the processing node creates a task instance. During the execution of the task instance, the resource processing device allocates target resources required by each process of each task instance to achieve resource isolation between task instances. Specifically, Figure 1 A flowchart of a resource isolation method provided by an embodiment of the present invention, such as Figure 1 As shown, the resource isolation method provided in this embodiment includes the following steps S101 to S105.
[0035] Step S101: Obtain the process to be run in the task instance.
[0036] The task instance includes one or more processes, and the status of each process may include not running, waiting to run, running, and running completed. Among them, the not running process means that the process has not been started; the waiting to run process means that the process has been started but has not yet started to execute. In this step S101, the waiting to run process in the task instance is obtained, specifically including intercepting interface call requests, etc.
[0037] Step S102: Determine whether the process to be run needs to consume target resources.
[0038] In this step, whether the target resource needs to be consumed can be determined based on the process content of the process to be run. For example, if the process to be run is an interface call request, it can be determined whether the interface called by the interface call request is related to the allocation of the target resource.
[0039] When the process to be run needs to consume the target resource, the following step S103 is executed; when the process to be run does not need to consume the target resource, the following step is not processed.
[0040] Step S103: When the process to be run needs to consume the target resource, it is determined whether the sum of the first resource amount and the second resource amount is greater than the resource application amount.
[0041] Among them, the first resource amount is the amount of target resources occupied by the task instance, the second resource amount is the amount of target resources required by the process to be run, and the resource application amount is the amount of target resources applied for by the task instance. When the task instance is created, a certain amount of target resources is allocated, and the amount of target resources allocated here is the resource application amount; during the operation of the task instance, the amount of target resources occupied by the task instance is updated in real time according to the consumption of target resources by the process in the task instance, that is, the first resource amount is maintained in real time.
[0042] When the process to be run needs to consume target resources, first determine the amount of target resources required by the process to be run, that is, the second resource amount, and then determine whether the sum of the first resource amount and the second resource amount is greater than the resource application amount, that is, determine whether the current remaining target resources of the task instance meet the target resources required by the process to be run.
[0043] When the sum of the first resource amount and the second resource amount is greater than the resource application amount, execute the following step S104; when the sum of the first resource amount and the second resource amount is less than or equal to the resource application amount, execute the following step S105.
[0044] Step S104: Return identification information indicating that the resource application of the process to be run has failed to the task instance.
[0045] When the target resources currently remaining in the task instance cannot meet the target resources required by the process to be run, that is, the target resources can no longer be provided to the process in the task instance, the identification information of the failed application is returned to the task instance. After the task instance obtains the identification information representing the failed resource application, the internal process allocation mechanism can release the target resources by controlling the process occupying the target resources to adjust the target resources currently consumed by the task instance, and finally make each process in the task instance run successfully.
[0046] Step S105: Allocate the second amount of target resources to the process to be run.
[0047] If the target resources currently remaining in the task instance can also satisfy the target resources required by the process to be run, the target resources can be provided to the process to be run.
[0048] By adopting the resource isolation method provided by this embodiment, each task instance is respectively allocated the target resources it needs, and the resource application amounts obtained by each task instance are isolated from each other. When the process of each task instance is to be run, it is first determined whether the process to be run needs to consume the target resources. When it needs to consume the target resources, it is then determined whether the remaining resource application amount of the task instance can meet the needs of the process to be run. If it cannot meet the needs, identification information representing the failure of the resource application of the process to be run is returned to the task instance, so that the task instance can perform process allocation according to the process allocation mechanism, and finally the task instance is successfully executed; if it can meet the needs, the target resources are allocated to the process to be run. The resource application amounts of different task instances are isolated from each other and do not interfere with each other, thereby avoiding conflicts in resource applications between task instances.
[0049] Optionally, in one embodiment, the process to be run is a process that sends an interface call request, and step S102, that is, the step of determining whether the process to be run needs to consume target resources, is specifically: determining whether the interface called by the interface call request is an interface for consuming target resources; wherein, when the interface called by the interface call request is an interface for consuming target resources, it is determined that the process to be run needs to consume target resources.
[0050] Specifically, in the prior art, during the execution of the task instance, the interface call request will directly call the interface and return the result, while in the present invention, the interface call request will be intercepted, and after the interception, it will be determined whether the interface called by the interface call request is an interface for consuming the target resources. If it is an interface for consuming the target resources, it means that the process to be run needs to consume the target resources, and then the above steps S103 to S105 are executed.
[0051] The resource isolation method provided in this embodiment is adopted to intercept the interface call request and then judge whether the called interface is an interface for consuming the target resource, so as to determine whether the process to be run needs to consume the target resource.
[0052] Further optionally, when the process to be run does not need to consume the target resource, the interface call request is forwarded. In step S105, the step of allocating the target resource of the second resource amount to the process to be run is specifically: forwarding the interface call request.
[0053] Specifically, by judging whether the called interface is an interface for consuming the target resources, whether the process to be run needs to consume the target resources is determined. If the called interface is not an interface for consuming the target resources, that is, the process to be run does not need to consume the target resources, then the interface call request is directly forwarded, which means that the target resources of the second amount of resources are allocated to the process to be run; if the called interface is an interface for consuming the target resources, that is, the process to be run needs to consume the target resources, and at the same time the sum of the first amount of resources and the second amount of resources is less than or equal to the resource application amount, then the interface call request is directly forwarded, which means that the target resources of the second amount of resources are allocated to the process to be run.
[0054] Optionally, in one embodiment, before the step of determining whether the sum of the first resource amount and the second resource amount is greater than the resource application amount, the resource isolation method also includes: determining the physical device to which the target resource that the process to be run needs to consume belongs, and obtaining the first physical device, wherein the task instance applies for the target resource on at least two physical devices when it is executed, and the first physical device is one of the at least two physical devices; obtaining the amount of the target resource that the task instance is occupying on the first physical device, and obtaining the first resource amount; obtaining the amount of the target resource applied for by the task instance on the first physical device, and obtaining the resource application amount.
[0055] Specifically, when creating a task instance, target resources on multiple physical devices can be allocated to the task instance. In addition, the amount of target resources applied for by the same task instance on different physical devices is isolated from each other, and the amount of target resources occupied by the same task instance on different physical devices is also isolated from each other.
[0056] After obtaining the process to be run in the task instance, first determine the physical device to which the target resource to be consumed by the process to be run belongs. For example, if the process to be run is a process that sends an interface call request, the physical device to which the target resource to be consumed by the process to be run belongs can be determined by the identifier of the physical device requested by the interface call request. Here, the physical device is defined as the first physical device. After determining the first physical device, the first resource amount is obtained by obtaining the amount of target resources that the task instance is occupying on the first physical device, and the resource application amount is obtained by obtaining the amount of target resources applied for by the task instance on the first physical device. Then, it is determined whether the sum of the first resource amount and the second resource amount is greater than the resource application amount.
[0057] When the target resources currently remaining on the first physical device of the task instance cannot meet the target resources required by the process to be run, that is, the first physical device can no longer provide target resources to the process in the task instance, the identification information of the application failure is returned to the task instance. After the task instance obtains the identification information representing the failure of the resource application of the process to be run on the first physical device, the internal process allocation mechanism can control the process to be run to apply for the target resources on the second physical device, such as calling the interface on the second physical device, so as to finally make each process in the task instance run successfully.
[0058] By adopting the resource isolation method provided in this embodiment, the same task instance can apply for target resources on different physical devices, and the amount of target resources on different physical devices occupied by the task instance is isolated from each other. The amount of target resources applied for by the task instance on different physical devices is isolated from each other. When allocating resources to the process of the task instance, different physical devices can be controlled separately, making the resource allocation method more flexible.
[0059] Optionally, in one embodiment, all processes of the task instance share a resource occupancy variable. Before the step of determining whether the sum of the first resource amount and the second resource amount is greater than the resource application amount, the resource isolation method also includes: reading the value of the resource occupancy variable to obtain the first resource amount.
[0060] Specifically, in the present invention, the amount of target resources occupied by a task instance is recorded by sharing resource occupancy variables among all processes. When a process is assigned to a target resource, the shared resource occupancy variable is updated. When a process is executed and the target resource is released, the shared resource occupancy variable is updated, so that the shared resource occupancy variable can reflect the amount of target resources occupied by the task instance in real time. Based on this, the first resource amount can be obtained by reading the value of the resource occupancy variable. Processes in different task instances share different resource occupancy variables to achieve resource isolation between different task instances.
[0061] By adopting the resource isolation method provided in this embodiment, the target resources occupied by the task instance are recorded by sharing the resource occupation variable. Each time the first resource amount is obtained, it is only necessary to read the value of the resource occupation variable, and the processing logic is simple.
[0062] Optionally, in one embodiment, when a task instance is created, target resources on N physical devices are applied for, and the resource occupancy variable is an array, the array includes at least N elements, and the amount of target resources occupied by the task instance on each physical device is recorded by one element.
[0063] Specifically, the resource occupancy variable is implemented through an array structure. The number of elements in the array is at least equal to the number of physical devices allocated to the task instance. That is, when the task instance is created and applies for target resources on N physical devices, the array will accordingly include at least N elements. There can be a physical device corresponding to a different element, and different elements are used to record the amount of target resources on different physical devices that the task instance is occupying.
[0064] With the resource isolation method provided in this embodiment, the resource occupancy variable is implemented using an array structure, and each physical device can correspond to one element, which facilitates updating and reading of the resource occupancy variable.
[0065] Optionally, in one embodiment, the target resource is a video memory resource of a GPU physical device.
[0066] Specifically, the GPU physical device includes computing resources and video memory resources, etc. The target resources in the present invention can be the video memory resources of the GPU physical device. In addition, the Volta architecture series of GPU physical devices launched by NVIDIA supports the use of MPS hardware to accelerate the parallel execution of multiple processes under a single GPU physical device to improve resource utilization. At the same time, the usage percentage of threads (that is, computing resources) of the corresponding process can be controlled through environment variables. Therefore, in this embodiment, fine-grained isolation of GPU physical devices is added, including video memory resources and threads, to avoid the problem of multi-tasking video memory resource conflicts in machine learning scenarios and improve the resource utilization of GPU physical devices in distributed platforms. Specifically, Figure 2 A schematic diagram of resource isolation provided by an embodiment of the present invention, such as Figure 2 As shown in the figure, a processing node (GPU computing node) of the distributed platform includes two GPU physical devices, specifically physical GPU0 and physical GPU1, and two task instances are running in the processing node, namely task instance 1 and task instance 2. The data structure of the resource usage variable is an array of length 4, which records the usage of the video memory resources of each GPU physical device of the processing node by the task instance.
[0067] When task instance 1 is created, it applies for video memory resources on two GPU physical devices, and the video memory resources applied for each GPU physical device are 14GB (i.e., the resource application amount, which is less than or equal to the video memory resource size of a single GPU physical device, and the video memory resource sizes of different GPU physical devices may also be different), and the threads percentage of each GPU physical device is 80. N process tasks are running simultaneously in task instance 1, and the video memory resource application amount (i.e., the first resource amount) of all process tasks for GPU0 is stored in the array mem_used[4]. For example, mem_used[0]=14G, which means that the video memory resource application amount of all processes in the task instance for GPU 0 physical device of the processing node is 14GB, which has reached the maximum application video memory amount (i.e., resource application amount). Similarly, the video memory resource application amount of all processes in task instance 1 for GPU 1 physical device of the processing node is mem_used[1]=12G, indicating that the task instance can also apply for 2GB of video memory resources from physical device 1. The mem_used variable is stored in a shared memory manner, and the semaphore mechanism in the prior art is used to solve the problem of multi-process read and write conflicts, so as to provide multiple processes of the task instance with sharing the usage of the GPU physical device video memory resources of the task instance.
[0068] Task instance 2 has applied for 1 GPU physical device, corresponding to GPU physical device 1, and the video memory resource applied for the GPU physical device is 1GB (that is, the resource application amount), and the threads percentage of the GPU physical device is 5. The video memory used by all processes under the current task instance 2 is 1GB, that is, mem_used[1] = 1.
[0069] Embodiment 2
[0070] Corresponding to the above-mentioned embodiment 1, the second embodiment of the present invention provides a resource isolation module, and the corresponding technical features and technical effects are not described in detail in this embodiment, and the relevant parts can refer to the above-mentioned embodiment 1. Specifically, Figure 3 A block diagram of a resource isolation module provided in an embodiment of the present invention, such as Figure 3 As shown, the module includes a first acquisition unit 301 , a first judgment unit 302 , a second judgment unit 303 and a processing unit 304 .
[0071] Among them, the first acquisition unit 301 is used to acquire the process to be run in the task instance; the first judgment unit 302 is used to judge whether the process to be run needs to consume the target resources, wherein the task instance includes one or more processes; the second judgment unit 303 is used to judge whether the sum of the first resource amount and the second resource amount is greater than the resource application amount when the process to be run needs to consume the target resources, wherein the first resource amount is the amount of target resources occupied by the task instance, the second resource amount is the amount of target resources required by the process to be run, and the resource application amount is the amount of target resources applied by the task instance; the processing unit 304 is used to return identification information representing the failure of resource application of the process to be run to the task instance when the sum of the first resource amount and the second resource amount is greater than the resource application amount; and when the sum of the first resource amount and the second resource amount is less than or equal to the resource application amount, allocate the target resource of the second resource amount to the process to be run.
[0072] Optionally, in one embodiment, the process to be run is a process that sends an interface call request. When the first judgment unit 302 judges whether the process to be run needs to consume target resources, the specific steps performed are: judging whether the interface called by the interface call request is an interface for applying for target resources; wherein, when the interface called by the interface call request is an interface for applying for target resources, the process to be run needs to consume the target resources.
[0073] Optionally, in one embodiment, the resource isolation device further includes a determination unit and a second acquisition unit. The determination unit is used to determine the physical device to which the target resource to be consumed by the process to be run belongs before the second judgment unit 303 determines whether the sum of the first resource amount and the second resource amount is greater than the resource application amount, and obtain the first physical device, wherein the task instance applies for the target resource on at least two physical devices when executing, and the first physical device is one of the at least two physical devices; the second acquisition unit is used to obtain the amount of the target resource that the task instance is occupying on the first physical device, and obtain the first resource amount, and obtain the amount of the target resource applied for by the task instance on the first physical device, and obtain the resource application amount.
[0074] Optionally, in one embodiment, all processes of the task instance share a resource occupancy variable, and the resource isolation device also includes a reading unit, which is used to read the value of the resource occupancy variable to obtain the first resource amount before the step of determining whether the sum of the first resource amount and the second resource amount is greater than the resource application amount.
[0075] Optionally, in one embodiment, the task instance creates and applies for target resources on N physical devices, the resource occupancy variable is an array, the array includes at least N elements, and the amount of target resources on each physical device being occupied by the task instance is recorded by one element.
[0076] Optionally, in one embodiment, the target resource is a video memory resource of a GPU physical device.
[0077] Optionally, in one embodiment, the processing unit 304 is also used to forward the interface call request when the process to be run does not need to consume the target resources, and when allocating the second amount of target resources to the process to be run, the specific steps performed are: forwarding the interface call request.
[0078] Embodiment 3
[0079] Embodiment 3 of the present invention provides a distributed platform. Figure 4 A block diagram of a distributed platform provided by an embodiment of the present invention, such as Figure 4 As shown, the distributed platform includes: a management node 41 and several processing nodes 42, the processing node 42 includes a target resource 421, a task creation device 422 and a resource processing device 423, and the resource processing device 423 includes a resource isolation module 4231 and a resource management module 4232.
[0080] Among them, the management node 41 is used to schedule tasks to the processing nodes 42 according to the information of the target resources on each processing node 42; the task creation device 422 is used to create a task instance when it is scheduled on the processing node 42 where it is located; the resource management module 4232 is used to report the information of the target resources on the processing node 42 where it is located to the management node 41, and is also used to allocate target resources to the task instance; and the resource isolation module 4231 is used to execute any one of the resource isolation methods provided by the present invention.
[0081] Optionally, in one embodiment, the processing node 42 includes a GPU physical device, the target resource is a video memory resource on the GPU physical device, the resource isolation module 4231 is an so library, and the resource management module 4232 is also used to: mount the GPU physical device and the resource isolation module 4231 to the task instance; set the environment variables of the task instance, wherein the environment variables include the resource application amount of the target resource and the dynamic library loading variable, and the value of the dynamic library loading variable is the resource isolation module.
[0082] Figure 5 and Figure 6 A schematic diagram of a business processing flow of a distributed platform provided by an embodiment of the present invention. In an embodiment, as Figure 5 and Figure 6 As shown, in a distributed platform, the video memory resources and threads resources of the GPU physical device can be isolated.
[0083] Specifically, the GPU Memory Manage module in the processing node is a resource isolation module, which is a gpu_memory_manage.so library, which implements the memory resource isolation of the task instance. The thread resource isolation of the GPU physical device is implemented by the MPS Server (also known as the MPS server).
[0084] The resource management module is responsible for synchronizing the resource information of the GPU physical devices on the processing node to the management node, and is responsible for:
[0085] (1) Mount the GPU physical device;
[0086] (2) Mount the gpu_memory_manage.so library into the task instance at / tmp / gpu_memory_manage.so;
[0087] (3) Set some GPU resource isolation-related environment variables for the task instance, including:
[0088] MAX_MEM;
[0089] LD_PRELOAD;
[0090] CUDA_MPS_ACTIVE_THREAD_PERCENTAGE.
[0091] Before processing business on the distributed platform, module deployment is performed first, including:
[0092] 1) The GPU Memory Manage module (also known as the resource isolation module) is a gpu_memory_manage.so file library, which is deployed in the / tmp / directory of the processing node in the form of a file.
[0093] 2) MPS Server is deployed on each processing node in a service mode to implement multi-process task services for GPU physical devices and isolate the Threads resources of GPU physical devices.
[0094] 3) The resource management module is deployed on each processing node in a service mode.
[0095] The various steps of the business processing flow of the distributed platform are described as follows.
[0096] Step S.1: The resource management module initializes resource information and synchronizes resources to the scheduling module.
[0097] The management node includes a task management module and a scheduling module. For the GPU physical resources on the processing node, the resource management module collects the number of GPU physical devices on each processing node and the capacity of the video memory resources of each GPU physical device, and synchronizes them to the scheduling module of the management node in the distributed platform.
[0098] Step S.2: The scheduling module queries the task instance to be scheduled.
[0099] After the user submits a task on the distributed platform, the resource request parameters for requesting GPU physical resources are determined according to the type of task. For example, if the user submits a video processing task, the video memory resources, Threads resources, and the number of GPU physical devices required for the task can be preset, so the distributed platform can set the resource request parameters for the task according to the task type.
[0100] Alternatively, the request parameters of the task include resource request parameters, such as the amount of video memory resources on each GPU physical device, the number of GPU physical devices, and the percentage of Threads resources of each GPU physical device, for example:
[0101] …
[0102] -name:job n
[0103] resources:
[0104] limits:
[0105] nvidia.com / gpu:4
[0106] nvidia.com / gpu_threads:33
[0107] nvidia.com / gpu_mem:5461
[0108] From the above parameters, we can see that the number of GPU physical devices is 4, the percentage of Threads resources of each GPU physical device is 33, and the amount of video memory resources on each GPU physical device is 5461.
[0109] Step S.3: The task management module sends the task instance to be scheduled to the scheduling module.
[0110] Step S.4: The scheduling module performs scheduling and binds the scheduling node to the task instance.
[0111] Among them, after the scheduling module receives the task submitted by the user, it will use node pre-selection and optimization strategies to schedule the task according to the resource request parameters of the requested GPU physical resources and the GPU resource conditions of each processing node in the distributed platform.
[0112] Step S.5: The resource management module allocates resources to the task instance and updates the allocated resources to the scheduling module.
[0113] After the scheduling is completed and the processing node to which the task is assigned is obtained, the task creation device on the processing node creates a task instance. The resource management module will perform the following tasks according to the resource request parameters of the task instance requesting GPU physical resources:
[0114] 1) Mount the / tmp / gpu_memory_manage.so library in the processing node to the / tmp / gpu_memory_manage.so inside the task instance in read-only mode;
[0115] 2) Set the task instance's environment variable MAX_MEM to the amount of video memory resources on each GPU physical device;
[0116] 3) Set the task instance's environment variable CUDA_MPS_ACTIVE_THREAD_PERCENTA GE to the percentage of threads resources for each GPU physical device requested by the task instance;
[0117] 4) Set the LD_PRELOAD environment variable of the task instance to / tmp / gpu_memory_manage.so to enable the GPU memory isolation module.
[0118] Step S.6: The task instance applies for resources.
[0119] Among them, the process of task instance applying for resources is also the workflow of resource isolation module; Figure 7 A schematic diagram of applying for resources for a task instance provided in an embodiment of the present invention, such as Figure 7 As shown in the figure, when the process in the task instance uses the GPU physical device to call the CUDA (Compute Unified Device Architecture) API interface, it will be intercepted by the gpu_memory_manage.so library implemented by the GPU MemoryManage module. The GPU Memory Manage module will determine whether the interface called by the task instance is an interface related to video memory resource allocation. If not, it will call the original CUDA API and return the call result to the task instance. If the GPU Memory Manage module determines that the interface called by the task instance is an interface related to video memory allocation, it will first read the video memory resource allocation of the current task instance from the shared memory. The GPUMemory Manage module will determine whether the video memory resources currently used by the task instance have reached the value of the environment variable MAX_MEM. If the sum of the amount of video memory resources currently required and the value in the shared memory exceeds the maximum value of MAX_MEM, the GPUMemory Manage module will return the CUDA_ERROR_OUT_OF_MEMORY status code to the task instance.
[0120] Step S.7: The resource isolation module returns the resource application result to the program in the task instance.
[0121] Embodiment 4
[0122] This embodiment also provides a computer device, such as a smart phone, tablet computer, laptop computer, desktop computer, rack server, blade server, tower server or cabinet server (including an independent server or a server cluster composed of multiple servers) that can execute programs. Figure 8 As shown, the computer device 01 of this embodiment includes at least but is not limited to: a memory 011 and a processor 012 which can be interconnected via a system bus. Figure 8 It should be pointed out that Figure 8Only a computer device 01 having components memory 011 and processor 012 is shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0123] In this embodiment, the memory 011 (i.e., readable storage medium) includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 011 can be an internal storage unit of the computer device 01, such as a hard disk or memory of the computer device 01. In other embodiments, the memory 011 can also be an external storage device of the computer device 01, such as a plug-in hard disk equipped on the computer device 01, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the memory 011 can also include both the internal storage unit of the computer device 01 and its external storage device. In this embodiment, the memory 011 is generally used to store the operating system and various application software installed on the computer device 01, such as the program code of the resource isolation device of the second embodiment. In addition, the memory 011 can also be used to temporarily store various types of data that have been output or are to be output.
[0124] In some embodiments, the processor 012 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 012 is generally used to control the overall operation of the computer device 01. In this embodiment, the processor 012 is used to run the program code stored in the memory 011 or process data, such as a resource isolation method.
[0125] Embodiment 5
[0126] This embodiment also provides a computer-readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a disk, an optical disk, a server, an App application store, etc., on which a computer program is stored, and the program implements the corresponding function when executed by the processor. The computer-readable storage medium of this embodiment is used to store the resource isolation method, and implements the resource isolation method of embodiment 1 when executed by the processor.
[0127] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0128] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0129] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0130] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A resource isolation method, characterized in that: include: Get the process to be run in the task instance; Determine whether the process to be run needs to consume target resources, wherein the task instance includes one or more processes, and the target resource is a video memory resource of a GPU physical device; When the process to be run needs to consume the target resource, determine whether the sum of the first resource amount and the second resource amount is greater than the resource application amount, wherein the first resource amount is the amount of the target resource currently occupied by the task instance, the second resource amount is the amount of the target resource required by the process to be run, and the resource application amount is the amount of the target resource applied for by the task instance; When the sum of the first resource amount and the second resource amount is greater than the resource application amount, identification information indicating that the resource application of the process to be run has failed is returned to the task instance. After the task instance obtains the identification information indicating that the resource application has failed, the process allocation mechanism controls the process occupying the target resource to release the target resource, so as to adjust the amount of the target resource being occupied by the task instance, so that each process in the task instance runs successfully; and When the sum of the first resource amount and the second resource amount is less than or equal to the resource application amount, allocating the target resource of the second resource amount to the process to be run; Before the step of determining whether the sum of the first resource amount and the second resource amount is greater than the resource application amount, the resource isolation method includes: Determine the physical device to which the target resource to be consumed by the process to be run belongs, and obtain a first physical device, wherein the target resource on at least two physical devices is applied for when the task instance is executed, and the first physical device is one of the at least two physical devices; Obtaining the amount of the target resource on the first physical device that is being occupied by the task instance to obtain the first resource amount; The amount of the target resource applied for by the task instance on the first physical device is obtained to obtain the resource application amount.
2. The resource isolation method according to claim 1, characterized in that: The process to be run is a process that sends an interface call request, and the step of determining whether the process to be run needs to consume target resources is specifically as follows: Determine whether the interface called by the interface call request is an interface for applying for the target resource; Wherein, when the interface called by the interface call request is an interface for applying for the target resource, the process to be run needs to consume the target resource.
3. The resource isolation method according to claim 1, characterized in that: All processes of the task instance share the resource occupancy variable. Before the step of determining whether the sum of the first resource amount and the second resource amount is greater than the resource application amount, the resource isolation method further includes: The value of the resource occupancy variable is read to obtain the first resource amount.
4. The resource isolation method according to claim 3, characterized in that: The task instance creates the target resource applied for on N physical devices, the resource occupancy variable is an array, the array includes at least N elements, and the amount of the target resource on each physical device being occupied by the task instance is recorded by one of the elements.
5. The resource isolation method according to claim 2, characterized in that: The resource isolation method further includes: forwarding the interface call request when the process to be run does not need to consume the target resource; The step of allocating the target resource of the second resource amount to the process to be run is specifically: forwarding the interface call request.
6. A distributed platform, characterized in that: The distributed platform includes: a management node and several processing nodes, the processing node includes a target resource, a task creation device and a resource processing device, the resource processing device includes a resource isolation module and a resource management module, wherein: The management node is used to schedule tasks to the processing nodes according to the information of the target resources on each of the processing nodes; The task creation means is used to create a task instance when being scheduled on the processing node; The resource management module is used to report the information of the target resource on the processing node to the management node, and is also used to allocate the target resource to the task instance; and The resource isolation module is used to execute the resource isolation method described in any one of claims 1 to 5.
7. The distributed platform according to claim 6, characterized in that: The processing node includes a GPU physical device, the target resource is a video memory resource on the GPU physical device, the resource isolation module is an so library, and the resource management module is further used for: Mounting the GPU physical device and the resource isolation module to the task instance; The environment variables of the task instance are set, wherein the environment variables include the resource application amount of the target resource and a dynamic library loading variable, and the value of the dynamic library loading variable is the resource isolation module.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Allocation method and device for resource pool
CN102761469A
CPU (central processing unit) resource allocation method and device and electronic equipment
CN105988872A
PaaS cloud implementation method based on containers
CN106445515A
Resource shared using method and system based on preemptive scheduling and equipment
CN108769254A
Task processing method, device and system
CN109471727A