Resource allocation method and system for containerized development environment
By monitoring the resource usage data of container instances and judging their status, recycling the freeable resources of low-load instances and allocating them to instances with lower optimization levels, the resource waste caused by Kubernetes pre-approval resource allocation is solved and the resource utilization rate is improved.
Patent Information
- Application Number
- CN202411666607.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-11-21
AI Technical Summary
In the scenario of large-scale deployment of containerized development environment instances, Kubernetes adopts pre-applied resource allocation methods, unused or low-load resources cannot be reused by other instances, resulting in waste of resources.
By monitoring the resource usage data of the container instance, it determines its status (idle, low load, or normal operation), and recycles resources in a low load state and allocates them to instances with lower optimization levels.
It improves the resource utilization rate of the containerized development environment and avoids resource waste. At the same time, when resource demand increases or recovers, resources can be withdrawn from low-priority instances, with less impact.
Smart Images

Figure CN119201358B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a resource allocation method and system for a containerized development environment. Background Art
[0002] Containerized development environment is an important cloud-based development tool. Commonly used tools include Jupyter NoteBook and VS-Code, which allow developers to write, debug, and deploy code online through a browser. The development and compilation process is carried out in the cloud, and cloud resources can be fully utilized for code debugging and running. Especially in scenarios such as machine learning, big data analysis, deep learning, and large models, high-performance GPUs, large memory, and high-performance CPUs are often required for code development, debugging, and running. The containerized development environment provides pre-configured tools and resources, allowing developers to quickly start and run machine learning projects without having to worry about environment configuration and resource management issues.
[0003] With the increasing development and maturity of technologies such as cloud computing and cloud native, managing cloud resources in a cloud native way has become a technical trend. The traditional way of running online development environments based on physical machines and virtual machines is gradually being replaced by containerization, that is, one online development environment corresponds to one container in the cloud. Multiple online development environments are managed and scheduled in a unified manner through container orchestration tools such as Kubernetes. The containerized online development environment has quickly become the mainstream method in the industry because it can automatically adjust computing resources according to demand to avoid resource waste; the pre-configured development environment avoids complex local environment configuration and other advantages.
[0004] In the scenario of large-scale deployment of containerized development environment instances, multiple online development environment instances are generally created through Kubernetes and allocated to different R&D engineers. Kubernetes applies for resources in a pre-application manner, that is, the resources of each online development environment instance are already allocated by Kubernetes at the time of application. Even if some online development environment instances are not used in certain time periods, the resources of these online development environment instances cannot be used by other online development environment instances.
[0005] In actual operation, the resources of an online development environment instance are often not used or fully used. However, since Kubernetes applies for resources in a pre-application manner, these resources cannot be released even if they are not used, and other instances cannot use these resources, resulting in serious waste of resources. Summary of the invention
[0006] In view of the above technical problems existing in the prior art, the present invention provides a resource allocation method and system for a containerized development environment to improve the resource utilization rate of the containerized development environment.
[0007] The present invention discloses a resource allocation method for a containerized development environment, comprising the following steps: obtaining resource usage data of a first instance; obtaining a state of the first instance based on the resource usage data, the state comprising an idle state, a low-load state or a normal operating state; if the state is a low-load state, calculating and reclaiming releasable resources of the first instance; obtaining a second instance requesting resources; from the second instance, screening a third instance whose priority is lower than a fourth threshold; and allocating the reclaimed resources to the third instance.
[0008] Preferably, the method for determining the state includes:
[0009] Collect resource usage and network traffic;
[0010] Calculate the variance of historical network traffic;
[0011] Determining whether the variance is less than a first threshold;
[0012] If so, calculate the historical mean of any dimension of resource usage;
[0013] Determine whether the historical average is less than a second threshold;
[0014] If it is less than a second threshold, the state is an idle state;
[0015] If not or greater than the second threshold, determine whether the historical average is less than a third threshold, and the third threshold is greater than the second threshold;
[0016] If the historical average is less than the third threshold, the state is a low load state;
[0017] If the historical average is greater than the third threshold, the state is a normal operating state.
[0018] Preferably, the third threshold value is calculated in the following manner: request * A1, wherein request represents the requested resource amount, and A1 represents the coefficient.
[0019] Preferably, the calculation method of the releasable resources is:
[0020] Rel i = request i – Use i – Buffer i
[0021] Buffer i= A2 * request i
[0022] Among them, Rel i Represented as a releasable resource of dimension i, request i Expressed as the requested resource amount of dimension i, Use i Represented as the historical mean of dimension i, Buffer i It is represented as the resource reservation amount of dimension i, and A2 is represented as the reservation coefficient.
[0023] Preferably, the method of allocating the recovered resources to the third instance includes:
[0024] Based on the Webhook mechanism, intercept the creation request of the third instance;
[0025] According to the creation request, the reclaimed resources are injected into the third instance, and the requested resource amount of the third instance is modified accordingly.
[0026] Preferably, the present invention also includes a method for resource recovery:
[0027] Determining whether the first instance is in a normal operating state;
[0028] If so, the resources of the third instance are released and allocated to the first instance.
[0029] Preferably, if the state is an idle state, after the first instance is saved as a first image, the first instance is stopped and resources of the first instance are released.
[0030] Preferably, the method for node resource management is:
[0031] Obtaining information about the first instance of releasable resources reported;
[0032] Obtain the reallocatable resources of the node where the first instance is located, where the reallocatable resources are calculated as follows: Nod = Nod0 + Rel
[0033] Nod represents the reallocatable resources of the node, Nod0 represents the already allocable resources on the node, and Rel represents the releasable resources of the first instance.
[0034] The amount of reallocatable resources of the node is updated through the interface.
[0035] The present invention also provides a system for implementing the above method, including a monitoring module, an instance status analysis module, a resource recovery module and a resource allocation module, wherein the monitoring module is used to obtain resource usage data of a first instance; the instance status analysis module is used to obtain the status of the first instance based on the resource usage data; the resource recovery module is used to calculate and recover the releasable resources of the first instance when the instance status is a low load state; the resource allocation module is used to obtain a second instance requesting resources; from the second instance, a third instance having a priority lower than a fourth threshold is screened; and the recovered resources are allocated to the third instance.
[0036] Preferably, the system also includes a resource recovery module and a node resource management module, wherein the resource recovery module is used to release the resources of the third instance and allocate the resources to the first instance when the first instance resumes normal operation; the node resource management module is used to obtain information about releasable resources of the instance, calculate reallocatable resources of the node, and update the amount of reallocatable resources of the node through an interface.
[0037] Compared with the prior art, the beneficial effects of the present invention are: the releasable resources of the first instance are reallocated and allocated to the third instance with a lower optimization level, which improves resource utilization on the one hand; on the other hand, when the resource demand of the first instance increases and recovers, the reallocated resources can be withdrawn from the third instance at any time without affecting the main or important instances. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 is a flow chart of a resource allocation method for a containerized development environment of the present invention;
[0039] Figure 2 It is a flow chart of the method for determining the instance status;
[0040] Figure 3 It is a system logic block diagram of the present invention. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0042] The present invention is further described in detail below in conjunction with the accompanying drawings:
[0043] A resource allocation method for a containerized development environment, such as Figure 1 As shown, the following steps are included:
[0044] Step 101: Obtain resource usage data of a first instance.
[0045] Step 102: Obtain the state of the first instance according to the resource usage data, where the state includes an idle state, a low-load state, or a normal operating state.
[0046] Step 103: If the state is an idle state, release the resources of the first instance.
[0047] More specifically, after saving the first instance as the first image, stop the first instance and release the resources of the first instance. Set IDLE_STATE=LEISURE to stop the first instance directly. The specific logic is to change the replicas of the Deployment corresponding to the first instance to 0. In order to ensure that the code, files, dependencies, etc. in the online development environment are not lost when subsequent users open and use the online development environment, the online development environment container will be automatically saved as an image when it is stopped. The next time the online development environment is opened, the saved image will be used by default to start. These resources do not require additional processing. When a new instance needs to be created and run, k8s will automatically use these resources to schedule the instance.
[0048] Step 104: If the state is a normal operating state, continuously monitor the resource usage of the first instance.
[0049] Step 105: If the state is a low-load state, calculate and reclaim the releasable resources of the first instance.
[0050] Step 106: Obtain a second instance of the requested resource.
[0051] Step 107: Filter out instances whose priorities are lower than a fourth threshold from the second instances and record them as third instances.
[0052] Step 108: Allocate the reclaimed resources to the third instance.
[0053] The releasable resources of the first instance are reallocated and allocated to the third instance with a lower optimization level. On the one hand, resource utilization is improved; on the other hand, when the resource demand of the first instance increases / recovers, the reallocated resources can be withdrawn from the third instance at any time without affecting the main or important instances. This enables k8s to use the recovered resources to schedule and run low-priority instances.
[0054] like Figure 2 In step 102, the method for determining the instance status includes the following steps:
[0055] Step 201: Collect resource usage and network traffic.
[0056] For the online development environment instances running in the cluster, collect their network traffic data and resource usage data in real time, and report the data to the Promethues monitoring system. Resource usage data collection can be performed using the cAdvisor component, which is based on the open source cAdvisor component. It expands the collection of container GPU, NPU and other resource usage based on a plug-in approach, and the plug-in approach facilitates the subsequent expansion of other resources. Taking the GPU plug-in as an example, its core is to mobilize the CUDA driver-related interface to obtain the real-time computing power usage and video memory usage of the GPU bound to the container, and aggregate them to form the container's GPU usage monitoring data. cAdvisor collects the CPU, memory, GPU / NPU resource dimensions of each container on the node and reports them to the Promethues monitoring system for storage. Network traffic data collection is implemented in SideCar mode. When all online development environment instances are created, the platform automatically intercepts API-Server requests for online development environment instance creation based on K8swebhook, and automatically injects the net-exportor SideCar container into the online development environment instance Pod, and exposes the access address of the SideCar container to the user. This ensures that all user operations on the online development environment instance are actually access to the net-exportor SideCar container. After the net-exportor SideCar container receives the user's operation request, it forwards the operation request to the online development environment instance container in the same Pod for processing and response, and on the other hand, it converts the request into network traffic data and reports it to the Promethues monitoring system. The network traffic data reported to the Promethues monitoring system only needs the number of requests per minute, and the detailed information of the request itself does not need to be reported.
[0057] Step 202: Calculate the variance of historical network traffic.
[0058] Historical data represents the instance operation status over the past period of time. The definition of the past period of time is determined by the automatic recovery time parameter specified by the user when creating the online development environment instance, that is, the value of the environment variable AUTO_RECOVERY_TIME. The variance judgment method is introduced to avoid the influence of some systematic timed network requests on the judgment.
[0059] Step 203: Determine whether the variance is less than or equal to a first threshold.
[0060] If yes, it is considered that the first instance has not received an operation request, and step 204 is executed: the historical average of any dimension of resource usage is calculated. For example, when it is equal to 0, it is considered that there has been no user operation in the past period of time.
[0061] Step 205: Determine whether the historical average is less than a second threshold.
[0062] The second threshold may be the initial resource usage of the instance, which is the resource usage data when the container just enters the running state, but is not limited thereto. More specifically, it is required that the history of all dimensions is less than the corresponding second threshold.
[0063] If it is less than the second threshold, step 206 is executed: the state is idle. For example, if the average of the resource usage data in the CPU, memory, GPU / NPU resource dimensions over a period of time is less than the initial resource usage, it is determined and marked that the online development environment is in idle state, that is, the environment variable IDLE_STATE=LEISURE is set.
[0064] If not or greater than the second threshold, execute step 207: determine whether the historical average is less than a third threshold, and the third threshold is greater than the second threshold.
[0065] If the historical average is less than the third threshold, step 208 is executed: the state is a low load state. At this time, the instance is in a low load state in terms of CPU, memory, and GPU / NPU resource dimensions, and the environment variable IDLE_STATE=LOWLOAD is set.
[0066] If the historical average is greater than the third threshold, step 209 is executed: the state is a normal operation state, and the environment variable IDLE_STATE=NORMAL is set.
[0067] The third threshold is calculated as: request * A1, wherein request represents the requested resource amount, and A1 represents a coefficient, which can be set based on experience. For example, A1 can be 0.3, 0.5, 0.7, etc., but is not limited thereto.
[0068] In a specific embodiment, an environment variable with a switch state is set for the instance. The environment variable is specified when the instance is created, AUTO_RECOVERY_SWITCH: ON, representing that the switch is turned on; AUTO_RECOVERY_SWITCH: OFF, representing that the switch is turned off. If the switch state is off, the state of the online development environment instance is directly marked as a normal operating state, that is, the environment variable IDLE_STATE=NORMAL is set; if the switch state is on, steps 201-109 are executed to determine the state of the instance.
[0069] In step 105, the calculation method of the releasable resources is:
[0070] Rel i = request i – Use i – Buffer i
[0071] Buffer i = A2 * request i
[0072] Among them, Rel i Represented as a releasable resource of dimension i, request i Expressed as the requested resource amount of dimension i, Use i Represented as the historical mean of dimension i, Buffer i It is represented as the resource reservation amount of dimension i, and A2 is represented as the reservation coefficient (Buffer), also known as the buffer coefficient.
[0073] More specifically, the Buffer coefficient of CPU and GPU / NPU computing resources is set to a smaller value, with the default value of 10%; the Buffer coefficient of memory and GPU / NPU video memory resources is set to a larger value, with the default value of 20%.
[0074] In step 108, the method of allocating the recovered resources to the third instance includes:
[0075] Based on the Webhook mechanism, intercept the creation request of the third instance;
[0076] According to the creation request, the recycled resources are injected into the third instance, and the requested resource amount of the third instance is modified accordingly. If the recycled resources have satisfied the request of the third instance, the requested resource amount of the third instance is set to 0 to ensure that k8s can use compressed resources to schedule and run low-priority online development environment instances. If the request amount of the third instance cannot be fully satisfied, the allocated resource amount will be deducted from the request.
[0077] The present invention also includes a first example resource recovery method:
[0078] Step 901: Determine whether the first instance has recovered and is in normal operation.
[0079] If so, execute step 902: release the resources of the third instance and allocate the resources to the first instance.
[0080] If not, execute step 903: continuously monitor the resource usage and network usage of the first instance.
[0081] The present invention also includes a method for node resource management:
[0082] Step 911: Obtain the reported information of the releasable resources of the first instance.
[0083] Step 912: obtaining the reallocatable resources of the node where the first instance is located, wherein the reallocatable resources are calculated as follows:
[0084] Nod = Nod0 + Rel
[0085] Nod represents the reallocatable resources of the node, Nod0 represents the already allocable resources on the node, and Rel represents the releasable resources of the first instance.
[0086] Step 913: Update the reallocatable resource amount of the node through the interface.
[0087] More specifically, call the update interface of the api-server to update the following reallocatable resources: CPU compression resources: compressed_cpu; memory compression resources: compressed_mem; GPU / NPU compressed computing power resources: compressed_gpu / compressed_npu; GPU / NPU compressed video memory resources: compressed_gpu_mem / compressed_npu_mem.
[0088] The present invention also provides a system for implementing the above method, Figure 3 As shown, it includes a monitoring module 1, an instance status analysis module 2, a resource recovery module 3 and a resource allocation module 4.
[0089] The monitoring module 1 is used to obtain resource usage data of the first instance;
[0090] The instance status analysis module 2 is used to obtain the status of the first instance according to the resource usage data;
[0091] The resource recovery module 3 is used to calculate and recover the releasable resources of the first instance when the instance state is in a low load state;
[0092] The resource allocation module 4 is used to obtain a second instance requesting resources; filter a third instance having a priority lower than a fourth threshold from the second instance; and allocate the recovered resources to the third instance.
[0093] The system also includes a resource recovery module 5 and a node resource management module 6.
[0094] The resource recovery module 5 is used to release the resources of the third instance and allocate the resources to the first instance when the first instance recovers to a normal operating state;
[0095] The node resource management module 6 is used to obtain the information of the releasable resources of the instance, calculate the reallocatable resources of the node, and update the amount of the reallocatable resources of the node through the interface.
[0096] The present invention uses network request data and resource usage to comprehensively judge the idle state of the online development environment instance, avoiding the unreasonable situations such as "although there is no user request, there is a program running in the online development environment" judged as idle by simply using network request data; "the user is developing, there is no program running, so the resource usage is small" judged as idle by simply using resource usage data. Seizing the characteristics that the scheduled system request is regular and the user's operation request is irregular, the variance of the network traffic data is calculated to judge whether there is a user operation, and the effect is very obvious.
[0097] A low-load state is introduced to cover various scenarios such as idle instances caused by user rest and low-load instances caused by user development. Limited resource recovery methods are provided for different scenarios, which greatly improves resource utilization.
[0098] In specific applications, the method and system of the present invention are used in the MLP platform of a certain automobile enterprise and a certain government, effectively improving the resource utilization rate of computing resources by more than 20%. Without expanding computing resources, 30% more online development environment instances can be supported and newly built, achieving real cost reduction and efficiency improvement.
[0099] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A resource allocation method for a containerized development environment, characterized in that: The following steps are involved: Obtain resource usage data of the first instance; Obtaining a state of the first instance according to the resource usage data, the state comprising an idle state, a low-load state, or a normal operating state; If the state is a low-load state, calculating and reclaiming releasable resources of the first instance; Obtaining a second instance of the requested resource; From the second instances, filter out third instances whose priorities are lower than a fourth threshold; and allocating the recovered resources to the third instance; The method for determining the state includes: Collect resource usage and network traffic; Calculate the variance of historical network traffic; Determining whether the variance is less than a first threshold; If so, calculate the historical mean of any dimension of resource usage; Determine whether the historical average is less than a second threshold; If it is less than a second threshold, the state is an idle state; If not or greater than the second threshold, determine whether the historical average is less than a third threshold, and the third threshold is greater than the second threshold; If the historical average is less than the third threshold, the state is a low load state; If the historical average is greater than the third threshold, the state is a normal operating state.
2. The resource allocation method according to claim 1, characterized in that: The third threshold value is calculated in the following manner: request * A1, wherein request represents the requested resource amount, and A1 represents the coefficient.
3. The resource allocation method according to claim 1, characterized in that: The calculation method for releasable resources is: Rel i = request i – Use i – Buffer i Buffer i = A2 * request i Among them, Rel i Represented as a releasable resource of dimension i, request i Expressed as the requested resource amount of dimension i, Use i Represented as the historical mean of dimension i, Buffer i It is represented as the resource reservation amount of dimension i, and A2 is represented as the reservation coefficient.
4. The resource allocation method according to claim 1, characterized in that: The method of allocating the recovered resources to the third instance includes: Based on the Webhook mechanism, intercept the creation request of the third instance; According to the creation request, the reclaimed resources are injected into the third instance, and the requested resource amount of the third instance is modified accordingly.
5. The resource allocation method according to claim 1, characterized in that: It also includes methods for resource recovery: Determining whether the first instance is in a normal operating state; If so, the resources of the third instance are released and allocated to the first instance.
6. The resource allocation method according to claim 1, characterized in that: If the state is an idle state, after the first instance is saved as a first image, the first instance is stopped and resources of the first instance are released.
7. The resource allocation method according to claim 1, characterized in that: It also includes methods for node resource management: Obtaining information about the first instance of releasable resources reported; Obtaining reallocatable resources of the node where the first instance is located, wherein the reallocatable resources are calculated as follows: Nod = Nod0 + Rel Wherein, Nod represents the reallocatable resources of the node, Nod0 represents the already allocable resources on the node, and Rel represents the releasable resources of the first instance; The amount of reallocatable resources of the node is updated through the interface.
8. A system for implementing the resource allocation method according to any one of claims 1 to 7, characterized in that: It includes monitoring module, instance status analysis module, resource recovery module and resource allocation module. The monitoring module is used to obtain resource usage data of the first instance; The instance status analysis module is used to obtain the status of the first instance according to the resource usage data; The resource recovery module is used to calculate and recover the releasable resources of the first instance when the instance state is in a low load state; The resource allocation module is used to obtain a second instance requesting resources; filter a third instance having a priority lower than a fourth threshold from the second instance; and allocate the recovered resources to the third instance.
9. The system according to claim 8, characterized in that It also includes resource recovery module and node resource management module. The resource recovery module is used to release the resources of the third instance and allocate the resources to the first instance when the first instance recovers to a normal operating state; The node resource management module is used to obtain the information of the releasable resources of the instance, calculate the reallocatable resources of the node, and update the amount of the reallocatable resources of the node through the interface.
Citation Information
Patent Citations
Method and system for managing calculation examples in cloud platforms
CN103761147A
Fairness job scheduling method based on reservation mechanism
CN112506634A