Container management system control method and device and storage medium
By analyzing the operation type and content of dynamic resource requirements, calculating the tenant's real-time resource allocation and adjusting the resource usage quota, the problem of uneven resource usage in multi-tenant shared cluster resources is solved, and dynamic scheduling and efficient utilization of tenant resources are achieved.
Patent Information
- Application Number
- CN202511188611.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-25
AI Technical Summary
In scenarios where multiple tenants share cluster resources, traditional static resource restriction solutions cannot effectively respond to tenants' dynamic business needs, causing some tenants to over-occupy resources and affecting the normal operation of other tenants' applications.
By analyzing the operation type and content of dynamic resource requirements, calculating the real-time resource allocation of tenants, and adjusting the resource scheduling strategy of cluster nodes, it ensures dynamic adjustment of resource usage quotas.
It realizes dynamic management of tenant resources, meets the changing resource demands of tenants due to business operations, and improves resource utilization efficiency and the flexibility of the container management system.
Smart Images

Figure CN120723474A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of container management, and in particular to a control method, device, and storage medium for a container management system. Background Art
[0002] Kubernetes cluster resource management involves allocating, scheduling, monitoring, and limiting compute, storage, network, and other resources within a cluster. Its core goal is to ensure efficient resource utilization, avoid resource contention, and guarantee stable application operation, while also enabling fair sharing of cluster resources among multiple tenants and teams. The core manageable resources in Kubernetes are compute resources such as CPUs and memory, but Kubernetes also supports extended resources such as GPUs, FPGAs (Field-Programmable Gate Arrays), and storage IOPS (Input / Output Operations Per Second).
[0003] However, when multiple tenants share cluster resources, the traditional static resource restriction solution is prone to some tenants over-occupying CPU, memory, GPU and other resources in scenarios with dynamic tenant business needs, resulting in restrictions on the operation of other tenants' applications. Summary of the Invention
[0004] The main purpose of this application is to provide a control method, device and storage medium for a container management system, aiming to solve the technical problem that in scenarios with dynamic business needs of tenants, some tenants may over-occupy resources, resulting in restricted operation of other tenants' applications.
[0005] To achieve the above objectives, the present application provides a method for controlling a container management system, the method comprising: In response to a trigger operation caused by dynamic resource demand, the operation type and operation content of the trigger operation are parsed, wherein the operation type includes deploying a new application, performing replica operations on an existing application, updating an application image version, and adjusting application runtime resource request parameters; Obtaining an initial resource quota allocated when the tenant corresponding to the trigger operation registers, the initial resource quota including a first resource allocation amount, a second resource allocation amount, a third resource allocation amount, and a fourth resource allocation amount; Determining a resource demand change corresponding to the triggering operation according to the operation type and the operation content; Substitute the used resources, initial resource quota, and resource demand change into the preset resource allocation formula to calculate the tenant's real-time resource allocation; The resource usage quota of the tenant is adjusted according to the real-time resource allocation amount, and the resource scheduling policy of the cluster node is synchronously updated.
[0006] In one embodiment, the step of parsing the operation type and operation content of the triggering operation in response to the dynamic resource demand triggering operation includes: If the operation type is to deploy a new application, the operation content is parsed to obtain the preset number of replicas, the CPU request amount, memory request amount, GPU request amount, and storage request amount required for each replica; If the operation type is to scale replicas of an existing application, the operation content is parsed to obtain the current number of replicas, the target number of replicas, and the fixed resource configuration information of a single replica. The fixed resource configuration information includes the baseline amount of CPU, memory, GPU, and storage occupied by a single replica. If the operation type is to update the application image version, the operation content is parsed to obtain the resource requirement parameters of the current image version and the resource requirement parameters of the new version image. The resource requirement parameters include the CPU upper limit, memory upper limit, GPU usage threshold, and storage IOPS requirement; If the operation type is to adjust the resource request parameters of the application runtime, the operation content is parsed to obtain the resource type to be adjusted and the corresponding request amount adjustment value, where the resource type includes at least one of CPU, memory, GPU, and storage.
[0007] In one embodiment, the first resource is a CPU, the second resource is a memory, the third resource is a GPU, and the fourth resource is a storage. Before the step of obtaining the initial resource quota allocated when the tenant corresponding to the triggering operation is registered, the following steps are included: Receive a registration request from a tenant, parse the registration request to obtain tenant type information, business scenario description, and estimated resource requirements, where the tenant type information includes individual tenant, enterprise tenant, or priority tenant, and the business scenario description includes compute-intensive, storage-intensive, or network-intensive; The basic quota coefficient is determined based on the preset resource quota benchmark table and tenant type information. The basic quota coefficient for enterprise tenants is higher than that for individual tenants, and the basic quota coefficient for priority tenants is higher than that for ordinary tenants of the same type. Match the corresponding resource weight matrix based on the business scenario description. In compute-intensive scenarios, the weight of CPU and GPU is higher than that of memory and storage. In storage-intensive scenarios, the weight of storage is higher than that of CPU, GPU, and memory. Combining the basic quota coefficient, the resource weight matrix, and the estimated resource demand, the first resource allocation amount, the second resource allocation amount, the third resource allocation amount, and the fourth resource allocation amount are calculated to determine the tenant's initial resource quota; The initial resource quota is associated with the tenant identifier and stored, and a resource upper limit threshold for the tenant is set, where the resource upper limit threshold is a preset multiple of the initial resource quota and does not exceed a preset proportion of the total available resources of the cluster.
[0008] In one embodiment, the step of determining the resource demand change corresponding to the triggering operation based on the operation type and the operation content includes: If the operation type is to deploy a new application, the resource demand changes for CPU, memory, GPU, and storage are calculated based on the parsed CPU request, memory request, GPU request, and storage request for a single replica and the preset number of replicas. If the operation type is to scale replicas of an existing application, the resource demand change for each resource type is calculated based on the difference between the current number of replicas and the target number of replicas, combined with the fixed resource configuration information for a single replica. If the operation type is to update the application image version, calculate the difference in resource requirement parameters between the new version image and the current version image, and use the difference as the resource requirement change for each resource type; If the operation type is to adjust the resource request parameters when the application is running, the adjusted value of the request amount of each resource type obtained by parsing will be directly used as the corresponding resource demand change.
[0009] In one embodiment, the real-time resource allocation amount is obtained by initial resource quota, used resource amount, resource type weight, resource demand change amount and tenant priority factor.
[0010] In one embodiment, after the step of adjusting the tenant's resource usage quota according to the real-time resource allocation amount and synchronously updating the resource scheduling policy of the cluster node, the following steps are included: For CPU, memory, GPU, and storage, calculate the ratio of used resources to the initial resource quota to obtain the utilization rate of each resource type; Based on the preset health assessment model, the utilization rate of each resource type is mapped to the corresponding health value; Determine the health weight of each resource type, with CPU and GPU having higher weights than memory and storage. The overall health of the tenant is calculated by calculating the health value of each resource type and configuring the health weight.
[0011] In one embodiment, after the step of adjusting the tenant's resource usage quota according to the real-time resource allocation amount and synchronously updating the resource scheduling policy of the cluster node, the following steps are included: Determine the initial GPU resource limits for the computing GPU, rendering GPU, and general-purpose GPU, respectively, based on the initial resource quotas corresponding to the GPUs. Get the tenant's initial GPU resource limit usage; Extract the tenant's historical GPU resource consumption data, including daily peak usage, average usage, and resource request frequency, as well as current GPU resource usage and request queue length. Generates dynamic correction factors based on daily peak usage, average usage, resource request frequency, current GPU resource usage, and request queue length; Adjust the GPU resource limit of tenants based on dynamic correction factors.
[0012] In one embodiment, before the step of parsing the operation type and operation content of the triggering operation in response to the dynamic resource demand, the step includes: In response to a virtual cluster isolation instruction, determining a virtual cluster instance of the tenant from which the virtual cluster isolation instruction originated, the virtual cluster instance comprising a virtual API server, a virtual controller manager, and a virtual scheduler, wherein the virtual cluster instance and the physical cluster implement isolated scheduling of container resources through a resource mapping layer; Deploy the virtual control plane components in a dedicated namespace within the underlying physical Kubernetes cluster. The step of parsing the operation type and operation content of the triggering operation in response to the dynamic resource demand includes: In response to a click operation on the virtual control plane component, determining an API interface; If the API interface belongs to the virtual cluster instance associated with the tenant identifier of the tenant, the step of parsing the operation type of the triggering operation and the operation content of the triggering operation is performed.
[0013] In addition, to achieve the above-mentioned objectives, the present application also provides a control device for a container management system, wherein the control device for the container management system includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the control method for the container management system as described above.
[0014] In addition, to achieve the above-mentioned purpose, the present application also provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores a program for implementing the control method of the container management system. The program for implementing the control method of the container management system is executed by a processor to implement the steps of the control method of the container management system as described above.
[0015] The present application provides a control method for a container management system. The present application first analyzes the operation type and operation content of the trigger operation in response to a trigger operation of dynamic resource demand, wherein the operation type includes deploying a new application, performing a replica operation on an existing application, updating an application image version, and adjusting application runtime resource request parameters; obtains the initial resource quota allocated when the tenant corresponding to the trigger operation is registered, wherein the initial resource quota includes a first resource allocation amount, a second resource allocation amount, a third resource allocation amount, and a fourth resource allocation amount; determines the resource demand change corresponding to the trigger operation based on the operation type and the operation content; substitutes the used resource amount, the initial resource quota, and the resource demand change into a preset resource ratio formula to calculate the tenant's real-time resource allocation amount; adjusts the tenant's resource usage quota based on the real-time resource allocation amount, and simultaneously updates the resource scheduling policy of the cluster node. That is, the present application determines the resource demand change by distinguishing different operation types and operation contents, and then calculates the tenant's real-time resource allocation amount based on the resource demand change amount and the initial resource quota allocated when the tenant is registered, thereby adjusting the resource scheduling policy of the cluster node and the tenant's resource usage quota. It solves the technical problem that in the scenario of dynamic business needs of tenants, some tenants may over-occupy resources, resulting in restricted operation of other tenants' applications, and achieves the technical effect of dynamic resource scheduling for tenants. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 A flowchart of the first embodiment of the control method for the container management system of the present application is provided; Figure 2 A flowchart of the seventh embodiment of the control method for the container management system of the present application is provided; Figure 3 A flowchart of an eighth embodiment of the control method for a container management system of the present application is provided; Figure 4 This is a schematic diagram of the hardware architecture involved in an embodiment of the control device of the container management system of this application.
[0019] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0020] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not intended to limit the present application.
[0021] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0022] Currently, when multiple tenants share cluster resources, traditional static resource restriction solutions can easily lead to some tenants over-occupying CPU, memory, GPU and other resources in scenarios with dynamic tenant business needs, resulting in restrictions on the operation of other tenants' applications.
[0023] The main solution of this application is to determine the change in resource demand by distinguishing different operation types and operation contents. Then, combining the change in resource demand with the initial resource quota allocated when the tenant registered, the tenant's real-time resource allocation is calculated, and the resource scheduling strategy of the cluster nodes and the tenant's resource usage quota are adjusted. This solves the technical problem that in scenarios with dynamic tenant business needs, some tenants are prone to excessive use of CPU, memory, GPU and other resources, resulting in the restricted operation of other tenants' applications, and achieves the technical effect of dynamic resource scheduling for tenants.
[0024] It should be noted that the execution entity of this embodiment can be a container management system, a computing service device with data processing, network communication, and program execution capabilities, such as a tablet computer, personal computer, or mobile phone, or a container management system control device capable of implementing the above functions. This embodiment does not specifically limit this. The following uses the container management system as an example to illustrate this embodiment and the following embodiments.
[0025] Based on this, the first embodiment of the present application proposes a control method for a container management system, please refer to Figure 1 The control method of the container management system includes steps S10 to S50: Step S10, in response to the triggering operation of dynamic resource demand, parsing the operation type and operation content of the triggering operation, wherein the operation type includes deploying a new application, performing scaling operations on existing applications, updating the application image version, and adjusting the application runtime resource request parameters.
[0026] In this embodiment, a dynamic resource demand trigger refers to an operation performed by a tenant during its use of the container management system that can cause a change in its resource demand. The operation type categorizes the triggering operation, distinguishing different types of resource demand changes. The operation content is the specific information contained in the triggering operation, such as the application identifier and resource parameters involved in the operation.
[0027] As an optional implementation, the container management system collects tenant operation requests in real time and initiates the parsing process when it detects an operation that meets preset trigger conditions. For operations such as deploying new applications, the system parses the application name, the image used, the number of required replicas, and the specific requested values for each resource. For operations such as scaling replicas, the system parses the target application ID, the current number of replicas, and the target number of replicas. For operations such as updating an application image version, the system parses the application ID, the current image version, and the new image version. For operations such as adjusting application runtime resource request parameters, the system parses the application ID, the resource type to be adjusted, and the specific values after the adjustment.
[0028] As another optional implementation, the container management system includes a dedicated operation log collection module that regularly analyzes collected tenant operation logs and selects records of operations triggered by dynamic resource demands. Using pre-set parsing rules, the module extracts the operation type and content from the log records. For example, a log entry such as "Deploy application: Application A, Image X, Number of replicas 2, CPU requesting 1 core" can be identified and parsed to identify the operation type as deploying a new application, and the operation content includes information such as Application A, Image X, Number of replicas 2, and CPU requesting 1 core.
[0029] Step S20: obtaining an initial resource quota allocated when the tenant corresponding to the triggering operation registers, wherein the initial resource quota includes a first resource allocation amount, a second resource allocation amount, a third resource allocation amount, and a fourth resource allocation amount.
[0030] In this embodiment, the initial resource quota is the initial amount of resources allocated to a tenant based on their information when registering with the container management system. The first, second, third, and fourth resource allocations correspond to different types of resource allocations, such as CPU, memory, GPU, and storage allocations, respectively.
[0031] As an optional implementation, after tenant registration is complete, the container management system associates the allocated initial resource quota with the tenant ID and stores it in a database. When access is needed, the initial resource quota record for the tenant is directly queried from the database based on the tenant ID corresponding to the triggering operation, and the first through fourth resource allocations are retrieved.
[0032] As another optional implementation, the container management system implements a resource quota cache module to cache frequently accessed tenant initial resource quotas in memory. When access is needed, the system first checks whether the tenant's initial resource quota exists in the cache. If so, it retrieves the quota directly from the cache. If not, it queries the database and stores the query result in the cache for faster subsequent access.
[0033] Step S30: determining the resource demand change corresponding to the triggering operation according to the operation type and the operation content.
[0034] In this embodiment, the resource demand change refers to the amount of change in the tenant's demand for various resources due to the execution of the triggering operation, which can be positive or negative. A positive value indicates an increase in demand, and a negative value indicates a decrease in demand.
[0035] As an optional implementation, corresponding calculation rules are preset for different operation types. For the operation of deploying a new application, the new demand for various resources is calculated based on the single-copy resource request and the number of copies in the operation content, that is, the resource demand change = single-copy resource request × number of copies; for the operation of scaling copies, the resource demand change = (target number of copies - current number of copies) × single-copy resource amount is calculated based on the difference between the target number of copies and the current number of copies and the single-copy resource amount; for the operation of updating the application image version, the difference between the resource demand of the new image and the current image is calculated as the resource demand change; for the operation of adjusting the resource request parameters during application runtime, the difference between the adjusted parameter value and the original parameter value is used as the resource demand change.
[0036] As another optional implementation, the container management system implements a resource demand assessment model that combines the operation type, operation content, and historical resource usage data to estimate changes in resource demand. For example, when deploying a new application, the model not only considers the resource requirements per replica and the number of replicas, but also considers the actual resource consumption of similar applications in similar scenarios. This model then modifies the calculated change in resource demand to produce a more realistic result.
[0037] Step S40 , substituting the used resource amount, the initial resource quota, and the resource demand change into a preset resource allocation formula to calculate the tenant's real-time resource allocation.
[0038] In this embodiment, used resources refer to the amount of various resources currently occupied by a tenant. Real-time resource allocations refer to the real-time quotas of various resources currently allocated to a tenant, as determined by the container management system. The preset resource allocation formula is a mathematical formula used to calculate real-time resource allocations, taking into account factors such as used resources, initial resource quotas, and changes in resource demand.
[0039] As an optional implementation, the default resource allocation formula is: real-time resource allocation = initial resource quota - used resource quota + change in resource demand. The corresponding data for CPU, memory, GPU, storage, and other resource types are substituted into the formula to obtain the real-time resource allocation for each resource type. An upper limit is also set for the real-time resource allocation to ensure it does not exceed the maximum resource limit set by the cluster for the tenant.
[0040] As another optional implementation, the preset resource allocation formula is: real-time resource allocation = (initial resource quota - used resource quota) × resource type weight + change in resource demand × tenant priority factor. The resource type weight is set based on the resource's importance, and the tenant priority factor is determined by the tenant's level. During the calculation, the real-time resource allocation for each resource type is substituted into the formula, ensuring that it does not exceed the maximum resource limit and is not less than the currently required resource amount.
[0041] Step S50: adjusting the tenant's resource usage quota according to the real-time resource allocation amount, and synchronously updating the resource scheduling policy of the cluster node.
[0042] In this embodiment, resource usage quota refers to the maximum limit of various resources that the container management system allows tenants to use. Cluster node resource scheduling policy refers to the rules and methods followed by the container management system when allocating cluster node resources to containers.
[0043] As an optional implementation, the container management system directly modifies the tenant's corresponding resource usage quota record based on the calculated real-time resource allocation, updating the original quota value to the real-time resource allocation. Simultaneously, a new resource scheduling policy is generated that prioritizes scheduling the tenant's containers to nodes with sufficient remaining resources to meet the real-time resource allocation, and considers node load balancing during the scheduling process.
[0044] As an alternative implementation, when adjusting resource usage quotas, the container management system performs a pre-verification check to see if the adjusted quota will adversely affect the overall cluster resource allocation. If the check passes, the adjustment is executed. A dynamic update mechanism is used for cluster node resource scheduling policies. Scheduling weights are regularly recalculated based on real-time resource allocations and node resource status, allowing the scheduling policy to adapt to resource changes and improve resource utilization.
[0045] For example, when an enterprise tenant registers with a container management system, it receives an initial resource quota of 4 CPU cores, 8 GB of memory, 1 GPU card, and 100 GB of storage. The tenant triggers a new application deployment operation, requesting application B using image Y with 2 replicas. For each replica, the system requests 1 CPU core, 2 GB of memory, 0.5 GPU card, and 10 GB of storage. After analyzing the operation type and content, the system determines the change in resource requirements as 2 CPU cores, 4 GB of memory, 1 GPU card, and 20 GB of storage. Assuming the tenant's current used resources are 1 CPU core, 2 GB of memory, 0.3 GPU card, and 15 GB of storage, the system calculates the real-time resource allocation using the formula: real-time resource quota = initial resource quota - used resources + change in resource requirements. The resulting real-time resource allocation is 5 CPU cores, 10 GB of memory, 1.7 GPU card, and 105 GB of storage. The system adjusts the tenant's resource usage quota based on the real-time resource allocation and updates the cluster node scheduling policy to schedule the container of application B to a node with sufficient remaining resources.
[0046] This embodiment triggers operations in response to tenants' dynamic resource demands, accurately analyzes operation information, calculates real-time resource allocation based on initial resource quotas, used resources, and changes in resource demands, and then adjusts resource usage quotas and scheduling strategies. This achieves dynamic management of tenant resources, can promptly meet changes in tenant resource demands due to business operations, and improves resource utilization efficiency and the flexibility of the container management system.
[0047] Based on any of the above embodiments, in the second embodiment of the present application, step S10 includes: Step S11: If the operation type is to deploy a new application, the operation content is parsed to obtain the preset number of replicas, the CPU request amount, memory request amount, GPU request amount, and storage request amount required for a single replica.
[0048] In this embodiment, the preset number of replicas refers to the number of container instances that a tenant pre-sets when deploying a new application. The CPU request required for a single replica is the amount of CPU resources required to run a single container instance. The memory request required for a single replica is the amount of memory resources required to run a single container instance. The GPU request required for a single replica is the amount of GPU resources required to run a single container instance (if the application requires GPU support). The storage request required for a single replica is the amount of storage resources required to run a single container instance.
[0049] As an optional implementation, the container management system provides a visual deployment interface where tenants enter information for deploying a new application, including the number of replicas to be configured and the CPU, memory, GPU, and storage requirements for each replica. The system then directly reads these input values through the interface form parsing module and uses them as the parsed operation content.
[0050] As another alternative implementation, a tenant submits a request to deploy a new application through an API. The request parameters are passed in JSON format, including fields for the preset number of replicas and the requested amount of each resource. The system's API parsing module parses the JSON parameters, extracts the corresponding values, and completes the parsing of the operation content.
[0051] In step S12, if the operation type is to perform a replica scaling operation on an existing application, the operation content is parsed to obtain the current number of replicas, the target number of replicas, and the fixed resource configuration information of a single replica, wherein the fixed resource configuration information includes the baseline amount of CPU, memory, GPU, and storage occupied by a single replica.
[0052] In this embodiment, the current number of replicas refers to the number of container instances running for the existing application before the replica scaling operation is performed. The target number of replicas refers to the number of container instances the tenant expects to achieve through the replica scaling operation. The fixed resource configuration information for a single replica refers to the baseline values of various resources that are fixedly occupied by a single container instance during operation and serves as the basis for calculating changes in resource requirements during replica scaling operations.
[0053] As an optional implementation, the system queries the application management database for the current running information of the existing application, obtaining the current number of replicas and the fixed resource configuration of a single replica. Simultaneously, it reads the target number of replicas contained in the replica scaling operation instructions submitted by the tenant, integrating them to obtain the complete operation details.
[0054] As another optional implementation, the system monitors the scaling events of the application. When a scaling operation is detected, the data collection process is triggered. By interacting with the container runtime interface, the current number of replicas and the fixed configuration information of single replica resources are obtained, and then the target number of replicas is parsed from the operation instructions to complete the parsing of the operation content.
[0055] In step S13, if the operation type is to update the application image version, the operation content is parsed to obtain the resource requirement parameters of the current image version and the resource requirement parameters of the new version image. The resource requirement parameters include CPU upper limit, memory upper limit, GPU usage threshold and storage IOPS requirement.
[0056] In this embodiment, the resource requirement parameters of the current image version refer to the resource requirement limit parameters for various types of resources that the image currently used by the application places on during runtime. The resource requirement parameters of the new version image refer to the resource requirement limit parameters for various types of resources that the image to be updated places on during runtime. The CPU upper limit refers to the maximum amount of CPU resources allowed to be used when the container is running. The memory upper limit refers to the maximum amount of memory resources allowed to be used when the container is running. The GPU utilization threshold refers to the maximum proportion limit of GPU resources used by the container. The storage IOPS requirement refers to the number of input / output operations per second that the container requires for the storage device.
[0057] As an optional implementation, the system retrieves metadata for the current and new image versions from the image repository. This metadata includes descriptions of resource requirements. The metadata parsing module extracts parameters such as CPU limits, memory limits, GPU usage thresholds, and storage IOPS requirements as the parsed operation content.
[0058] As another optional implementation, when a tenant submits a request to update an application image version, the system requires the tenant to upload a resource requirement configuration file for the new image version. The system parses this configuration file and combines it with the resource requirement parameters of the current image version retrieved from the system to determine the operation details.
[0059] Step S14: If the operation type is to adjust the resource request parameters during application runtime, the operation content is parsed to obtain the resource type to be adjusted and the corresponding request amount adjustment value, wherein the resource type includes at least one of CPU, memory, GPU, and storage.
[0060] In this embodiment, the resource type to be adjusted refers to the type of resource whose requested amount the tenant wishes to change, such as CPU, memory, etc. The requested amount adjustment value refers to the specific value by which the tenant wishes to increase or decrease the requested amount of the resource type, with a positive number indicating an increase and a negative number indicating a decrease.
[0061] As an optional implementation, tenants can select the resource type to be adjusted and enter the corresponding adjustment value in the resource adjustment interface during application runtime. The system's interface interaction module records these selections and inputs and directly interprets them as the resource type to be adjusted and the requested adjustment value.
[0062] As another optional implementation, the tenant submits a resource parameter adjustment request command using a command line tool. The command includes the resource type and adjustment value parameters. The system's command parsing module parses the command and extracts the corresponding resource type and requested adjustment value as the operation content.
[0063] For example, the tenant performed four operations: first, deploying a new application, entering the preset number of copies 2 through the visual interface, the single-copy CPU request of 1 core, the memory request of 2GB, the GPU request of 0.5 cards, and the storage request of 10GB, and the system parsing to obtain the corresponding operation content; second, performing scale-up and scale-down operations on existing applications, the system obtains the current number of copies 3 and the fixed configuration of single-copy resources (CPU 1 core, memory 2GB, etc.) from the application management database, and parses out the target number of copies 5 in the tenant's instructions; third, updating the application image version, the system parses out the current image CPU upper limit of 2 cores, memory upper limit of 4GB, etc., as well as the corresponding parameters of the new version image from the image warehouse metadata; fourth, adjusting the application runtime resource request parameters, the tenant selects the CPU type on the interface, enters the adjustment value of 0.5 cores, and the system parses out that the resource type to be adjusted is CPU, and the request adjustment value is 0.5 cores.
[0064] This embodiment accurately extracts key resource information from various operation contents by formulating corresponding parsing methods for different operation types, providing an accurate data basis for subsequent calculation of changes in resource demand, and improving the accuracy and efficiency of the container management system's response to tenants' dynamic resource demands.
[0065] Based on any of the above embodiments, in the second embodiment of the present application, before step S20, the following steps are included: Step S21, receive the tenant's registration request, parse the registration request to obtain tenant type information, business scenario description and estimated resource requirements, where the tenant type information includes personal tenant, enterprise tenant or priority tenant, and the business scenario description includes computing intensive, storage intensive or network intensive.
[0066] In this embodiment, tenant type information is a classification identifier used to distinguish tenant attributes. Different tenant types will have different resource allocations. The business scenario description is the tenant's description of the application operation scenario, which is used to determine the focus of resource allocation. The estimated resource demand is the tenant's estimated demand for various resources based on its business situation. The first resource is CPU, the second resource is memory, the third resource is GPU, and the fourth resource is storage.
[0067] As an optional implementation, the container management system provides an online registration form with options for selecting the tenant type, inputting a description of the business scenario, and filling in estimated resource requirements. After the tenant submits the registration request, the system extracts the relevant information through a form parsing module.
[0068] As another optional implementation, the tenant submits a registration request by uploading a registration information document. The system uses a document parsing tool to parse the document to extract tenant type information, business scenario description, and estimated resource requirements.
[0069] Step S22: Determine a basic quota coefficient based on a preset resource quota benchmark table and tenant type information, wherein the basic quota coefficient of an enterprise tenant is higher than that of an individual tenant, and the basic quota coefficient of a priority tenant is higher than that of an ordinary tenant of the same type.
[0070] In this embodiment, the resource quota benchmark table is a pre-set table containing reference values for base quota coefficients corresponding to different tenant types. The base quota coefficient is used to adjust the baseline resource amount when calculating the initial resource quota. The larger the coefficient, the more resources are allocated.
[0071] As an optional implementation, the resource quota benchmark table specifies that the base quota coefficient for individual tenants is 1.0, for enterprise tenants is 1.5, for priority individual tenants is 1.2, and for priority enterprise tenants is 1.8. The system matches the corresponding base quota coefficient directly from the table based on the parsed tenant type information.
[0072] As another optional implementation method, the system sets up a basic quota coefficient calculation model, inputs additional parameters such as tenant type information and tenant size, and the model calculates the basic quota coefficient through an algorithm. For example, corporate tenants are divided into grades according to the number of employees, and different grades correspond to different coefficients.
[0073] Step S23, matching the corresponding resource weight matrix according to the business scenario description, wherein the weight ratio of CPU and GPU in computing-intensive scenarios is higher than that of memory and storage, and the weight ratio of storage in storage-intensive scenarios is higher than that of CPU, GPU and memory.
[0074] In this embodiment, the resource weight matrix is a matrix composed of weight values of various types of resources (CPU, memory, GPU, storage), and the weight value reflects the importance of the resource in the corresponding business scenario.
[0075] As an optional implementation, the system presets three resource weight matrices: compute-intensive (CPU:GPU:memory:storage) = 4:3:2:1; storage-intensive (storage:CPU:memory:GPU) = 4:2:2:2; and network-intensive (memory:GPU:CPU:storage) = 3:3:2:2. The system directly matches the corresponding matrix based on the business scenario description.
[0076] As another optional implementation, the system allows administrators to customize the resource weight matrix. After the tenant selects a business scenario, the system associates it with the customized matrix set by the administrator based on the scenario.
[0077] In step S24, the first resource allocation amount, the second resource allocation amount, the third resource allocation amount and the fourth resource allocation amount are calculated based on the basic quota coefficient, the resource weight matrix and the estimated resource demand to determine the tenant's initial resource quota, where the first resource is the CPU, the second resource is the memory, the third resource is the GPU and the fourth resource is the storage.
[0078] In this embodiment, the first resource allocation is the initial allocation for CPU, the second resource allocation is the initial allocation for memory, the third resource allocation is the initial allocation for GPU, and the fourth resource allocation is the initial allocation for storage. The initial resource quota is the sum of these four resource allocations, which constitutes the resource allocation amount for the tenant.
[0079] As an optional implementation, the system first determines a baseline amount for each resource based on estimated resource demand. The system then multiplies the baseline amount by the basic quota coefficient and the corresponding resource's weight in the weight matrix to determine the allocated amount for each resource. For example, if the estimated CPU demand is 10 cores, the basic quota coefficient is 1.5, and the weight is 0.4, the CPU allocation is 10 × 1.5 × 0.4 = 6 cores.
[0080] As another optional implementation, the system sets the resource allocation calculation formula: Resource allocation = (estimated resource demand × basic quota coefficient) × resource weight + scenario basic resource quantity, where the scenario basic resource quantity is the minimum guaranteed quantity of various resources under different business scenarios to ensure that tenants obtain basic resources.
[0081] Step S25 : The initial resource quota is associated with the tenant identifier and stored, and a resource upper limit threshold for the tenant is set, where the resource upper limit threshold is a preset multiple of the initial resource quota and does not exceed a preset proportion of the total available resources of the cluster.
[0082] In this embodiment, the tenant identifier is a string or code used to uniquely identify the tenant. The resource upper limit threshold is set by the system for the tenant, which is the maximum limit allowed for the tenant to use resources to prevent the tenant from over-occupying resources.
[0083] As an optional implementation, the system associates the initial resource quota with the tenant ID and stores it in the tenant resource table in the database. The resource upper limit threshold is set to twice the initial resource quota, but not exceeding 10% of the total available resources in the cluster. Database constraints ensure that the threshold setting is effective.
[0084] As an alternative implementation, the system uses a distributed storage system to store tenant resource information, distributively associating initial resource quotas with tenant identifiers. The resource upper limit threshold is dynamically adjusted based on cluster load, set to 2.5 times the initial quota when load is low and 1.8 times when load is high, with no limit exceeding 8% of the total available cluster resources.
[0085] For example, an enterprise tenant submits a registration request. The tenant type is analyzed as enterprise, the business scenario is described as compute-intensive, and the estimated resource requirements are 20 CPU cores, 30 GB of memory, 5 GPU cards, and 100 GB of storage. The system matches the enterprise tenant's base quota coefficient with 1.5 from the resource quota benchmark table. Based on the compute-intensive scenario, the resource weight matrix is matched as 4:3:2:1 for CPU:GPU:Memory:Storage. The resource allocations are calculated as follows: CPU = 20 × 1.5 × (4 / 10) = 12 cores, Memory = 30 × 1.5 × (2 / 10) = 9 GB, GPU = 5 × 1.5 × (3 / 10) = 2.25 cards, and Storage = 100 × 1.5 × (1 / 10) = 15 GB. The initial resource quota is associated with the enterprise tenant ID and stored, and the resource upper limit is set to twice the initial quota, but not exceeding 10% of the total available resources in the cluster.
[0086] This embodiment receives tenant registration requests and parses key information, determines the basic quota coefficient based on the tenant type, matches the resource weight matrix according to the business scenario, and finally calculates the initial resource quota and sets the upper limit threshold. This makes the initial resource allocation more in line with the actual needs of the tenant and the characteristics of the business scenario. It not only guarantees the basic resource needs of the tenant, but also prevents resource abuse through the upper limit threshold, laying a reasonable foundation for subsequent dynamic resource adjustment.
[0087] Based on any of the above embodiments, in the fourth embodiment of the present application, step S30 includes: In step S31, if the operation type is to deploy a new application, based on the CPU request amount, memory request amount, GPU request amount, storage request amount and the preset number of replicas obtained by parsing the required single copy, the resource demand changes of CPU, memory, GPU and storage are calculated respectively.
[0088] In this embodiment, the resource request amount required for a single copy refers to the application amount for various resources when a single application copy is running. The preset number of copies is the total number of application copies that the tenant expects to deploy, and the change in resource demand is the number of various resource demands added due to the deployment of new applications.
[0089] As an optional implementation method, the system extracts the CPU, memory, GPU, storage request amount and preset number of copies of a single copy from the parsed operation content, and calculates them separately through multiplication operations, that is, the change in CPU resource demand = single copy CPU request amount × preset number of copies. The change in resource demand for memory, GPU, and storage is calculated in the same way.
[0090] As another optional implementation, the system sets a resource redundancy factor and multiplies the basic calculation result by this factor to account for possible resource fluctuations. For example, if the redundancy factor is set to 1.2, the change in CPU resource demand = the CPU request per replica × the preset number of replicas × 1.2. The calculation method for other resource types is the same.
[0091] In step S32, if the operation type is to scale replicas of an existing application, the resource demand change of each resource type is calculated based on the difference between the current number of replicas and the target number of replicas obtained by analysis and the fixed resource configuration information of a single replica.
[0092] In this embodiment, the current number of replicas is the number of running replicas of the application before the operation, the target number of replicas is the desired number of replicas after the operation, a positive difference indicates expansion, a negative difference indicates contraction, and the baseline quantity is the fixed quantity of each resource occupied by a single replica. The change in resource demand for each resource type is calculated using the formula "Resource demand change = (target number of replicas - current number of replicas) × baseline quantity."
[0093] As an optional implementation, the system directly substitutes the current number of replicas, the target number of replicas, and the baseline resource requirements for each replica into the formula to calculate the change in resource requirements for CPU, memory, GPU, and storage. If the target number of replicas is greater than the current number of replicas, the result is positive, indicating that more resources are needed; otherwise, the result is negative, indicating that resources need to be reduced.
[0094] As another optional implementation, the system introduces a step coefficient during calculation. When the number of replicas changes significantly, such as when the difference exceeds a preset threshold, the calculation result is adjusted in steps. For example, if the difference is within 5, the coefficient is 1, and if it exceeds 5, the coefficient is 0.9, to avoid over-allocation or under-allocation of resources.
[0095] Step S33: If the operation type is to update the application image version, the resource requirement parameter difference between the new image version and the current image version is calculated, and the difference is used as the resource requirement change of each resource type.
[0096] In this embodiment, the resource requirement parameter is a restriction or requirement indicator for various resources when the image is running, such as a CPU upper limit, a memory upper limit, etc. The difference is the difference in resource requirement parameters between the new version and the current version. A positive number indicates an increase in resource demand, and a negative number indicates a decrease.
[0097] As an optional implementation, the system extracts parameters such as the CPU upper limit, memory upper limit, GPU usage threshold, and storage IOPS requirement of the current image and the new version image, and calculates the difference one by one, that is, the change in CPU resource requirement = the new version CPU upper limit - the current version CPU upper limit. The difference calculation method for other resource types is the same, and these differences are the corresponding resource requirement changes.
[0098] As another optional implementation, the system performs a rationality check on the calculated difference. If the absolute value of the difference for a certain type of resource exceeds a preset range (such as the CPU difference exceeds 50% of the CPU upper limit of the current version), an early warning is issued and manual confirmation is prompted. After confirmation, the difference is used as the change in resource demand.
[0099] Step S34: If the operation type is to adjust the resource request parameters during application runtime, the request adjustment value of each resource type obtained by parsing is directly used as the corresponding resource demand change.
[0100] In this embodiment, the resource type request quantity adjustment value is a specific value that the tenant uses to adjust the request quantity of various types of resources when the application is running. A positive number indicates an increase in the request quantity, and a negative number indicates a decrease in the request quantity. The change in resource demand is consistent with the adjustment value.
[0101] As an optional implementation, the system directly assigns the parsed request adjustment values of resource types such as CPU, memory, GPU, and storage to the corresponding resource demand changes without the need for additional calculations.
[0102] As another optional implementation, the system limits the range of the adjustment value. If the adjustment value exceeds a certain proportion of the initial quota of the resource type (such as 30%), it is truncated proportionally, and the truncated value is used as the change in resource demand, and the adjustment log is recorded.
[0103] For example, the tenant performs four operations: when deploying a new application, the CPU request for a single copy is 2 cores, memory 3GB, GPU 0.5 card, and storage 10GB. The preset number of copies is 4, and the calculated change in CPU resource demand is 2×4=8 cores, memory = 3×4=12GB, GPU = 0.5×4=2 cards, and storage = 10×4=40GB; for scaling copies of existing applications, the current number of copies is 2, the target number of copies is 5, the CPU baseline for a single copy is 1 core, and the change in CPU resource demand is (5-2)×1=3 cores; for updating the application image version, the current version has a CPU upper limit of 4 cores, and the new version has a CPU upper limit of 6 cores, and the change in CPU resource demand is 6-4=2 cores; for adjusting runtime resource parameters, the CPU request adjustment value is 1 core, and the change in CPU resource demand is 1 core.
[0104] This embodiment ensures the accuracy of resource demand changes by formulating precise resource demand change calculation methods for different operation types, provides a reliable basis for subsequent resource allocation adjustments, and helps the container management system to more reasonably schedule resources to meet the dynamic needs of tenants.
[0105] Based on any of the above embodiments, in embodiment 5 of the present application, the real-time resource allocation amount is obtained through the initial resource quota, the used resource amount, the resource type weight, the resource demand change and the tenant priority factor.
[0106] Specifically, the resource allocation formula is: real-time resource allocation amount = (initial resource quota - used resource amount) × resource type weight + resource demand change amount × tenant priority factor.
[0107] In this embodiment, the amount of used resources is the amount of each type of resources currently actually occupied by the tenant; the initial resource quota is the initial amount of each type of resources allocated by the system when the tenant registers; the change in resource demand is the change in the tenant's demand for each type of resource due to the triggering operation; the resource type weight is the weight value set according to the importance of different resources in the business scenario, and the sum of the weights of each resource type is 1; the tenant priority factor is a coefficient set according to the tenant level, which is used to reflect the priority differences of different tenants in resource allocation.
[0108] As an optional implementation, the system assigns resource type weights to CPU, memory, GPU, and storage, for example, 0.4 for CPU, 0.3 for GPU, 0.2 for memory, and 0.1 for storage. Priority factors are also assigned to different tenants, such as 1.0 for individual tenants, 1.2 for enterprise tenants, and 1.5 for priority tenants. During the calculation, the used resource amount, initial resource quota, and change in resource demand corresponding to each resource type are substituted into the formula to obtain the real-time resource allocation for each resource type.
[0109] As an alternative implementation, the system allows administrators to dynamically adjust resource type weights and tenant priority factors based on cluster resource usage. For example, when GPU resources are scarce in the cluster, the GPU resource type weight can be reduced; when a tenant's business is urgent, its priority factor can be temporarily increased. The adjusted weights and factors are substituted into the calculation formula.
[0110] For example, an enterprise tenant (with a priority factor of 1.2) has an initial resource quota of 10 CPU cores, 4 cores in use, and a change in CPU resource demand due to new application deployment of 3 cores. The CPU resource type weight is 0.4. Substituting these data into the formula, the real-time CPU resource allocation is calculated as (10-4) × 0.4 + 3 × 1.2 = 2.4 + 3.6 = 6 cores. Similarly, the real-time resource allocation for memory, GPU, and storage can be calculated for the tenant.
[0111] This embodiment introduces a resource allocation formula that combines resource type weights and tenant priority factors. This allows the calculation of real-time resource allocation to take into account both the importance of the resources themselves and the priority differences of tenants. This makes resource allocation more reasonable and accurate, better adapting to different business scenarios and tenant needs, and improving the flexibility and effectiveness of resource scheduling in the container management system.
[0112] Based on any of the above embodiments, in the sixth embodiment of the present application, after step S50, the following steps are included: In step A10, for CPU, memory, GPU, and storage, the ratio of used resources to the initial resource quota is calculated to obtain the utilization rate of each resource type.
[0113] In this embodiment, the amount of used resources is the amount of each type of resources currently actually occupied by the tenant, the initial resource quota is the initial amount of each type of resources obtained when the tenant registers, and the utilization rate is the percentage of the amount of used resources to the initial resource quota, which is used to reflect the degree of saturation of the use of each type of resources.
[0114] As an optional implementation, the system obtains the used resources of CPU, memory, GPU, and storage from the resource monitoring module, extracts the corresponding initial resource quota from the tenant resource configuration table, and calculates them separately through division operations, that is, CPU utilization = (CPU used resources ÷ CPU initial resource quota) × 100%. The calculation method for memory, GPU, and storage utilization is similar.
[0115] As another optional implementation, the system sets a data collection cycle (e.g., every 5 minutes), collects the used resource amount and initial resource quota in each cycle, calculates the utilization rate of each resource type, and takes the average value of multiple cycles as the final utilization rate to reduce the impact of data fluctuations.
[0116] Step A20 : Mapping the utilization rate of each resource type to a corresponding health value based on a preset health evaluation model.
[0117] In this embodiment, the health assessment model is pre-set by the system and is a rule model used to convert resource utilization into a health score. The health value is a numerical value reflecting the health status of resource utilization, and the range can be set to 0-100 points. The higher the score, the better the health status.
[0118] As an optional implementation, the health assessment model sets the following health values: 90-100 points for utilization ≤ 60%; 70-89 points for utilization 60% < ≤ 80%; 50-69 points for utilization 80% < ≤ 90%; and 0-49 points for utilization > 90%. The system matches the corresponding health value range directly from the model based on the utilization of each resource type, and then uses the middle value as the health value for that resource type.
[0119] As another optional implementation, the health assessment model uses a linear mapping method, where a utilization rate of 0% corresponds to 100 points, a utilization rate of 100% corresponds to 0 points, and the health value is calculated according to a linear proportion for intermediate utilization rates, that is, health value = 100-(utilization rate × 100). For example, when the utilization rate is 70%, the health value is 30 points.
[0120] Step A30: Determine the health weight of each resource type configuration, where the weights of CPU and GPU are higher than those of memory and storage.
[0121] In this embodiment, the health weight is a coefficient that reflects the importance of each resource type in the overall health assessment. The higher the weight, the greater the impact of the health status of the resource type on the overall health.
[0122] As an optional implementation, the system presets health weights: CPU is 0.35, GPU is 0.3, memory is 0.2, storage is 0.15, and the sum of the weights is 1, which reflects that the importance of CPU and GPU is higher than memory and storage.
[0123] As another optional implementation, the system allows weights to be adjusted according to business scenarios. In compute-intensive scenarios, the GPU weight can be increased to 0.35, the CPU weight to 0.35, and the storage weight to 0.25 in storage-intensive scenarios, while ensuring that the CPU and GPU weights are always higher than memory and storage.
[0124] In step A40 , the overall health of the tenant is calculated using the health value of each resource type and the configured health weight.
[0125] Specifically, the overall health of the tenant is calculated by: Overall health = Σ (health value × configured health weight).
[0126] In this embodiment, the overall health is an overall health score of tenant resource usage obtained by integrating the health status of various resources, and is used to intuitively reflect the overall health level of tenant resource usage.
[0127] As an optional implementation method, the system multiplies the health value of each resource type by the corresponding health weight, and then adds the products, that is, overall health = CPU health value × CPU weight + memory health value × memory weight + GPU health value × GPU weight + storage health value × storage weight, to obtain the final result and retain one decimal place.
[0128] As another optional implementation, the system normalizes the calculation results and limits the overall health level to a range of 0-100 points. If the calculation result exceeds 100 points, 100 points will be taken, and if it is less than 0 points, 0 points will be taken to ensure the rationality of the score.
[0129] For example, a tenant has 6 CPU cores in use, an initial resource quota of 10 cores, a utilization rate of 60%, and a mapping health value of 90. 12 GB of memory is in use, an initial quota of 20 GB, a utilization rate of 60%, and a health value of 90. Three GPUs are in use, an initial quota of 5 GPUs, a utilization rate of 60%, and a health value of 90. 80 GB of storage is in use, an initial quota of 100 GB, a utilization rate of 80%, and a health value of 80. Set the health weights as: CPU 0.35, GPU 0.3, memory 0.2, and storage 0.15. The overall health score is 90 × 0.35 + 90 × 0.2 + 90 × 0.3 + 80 × 0.15 = 31.5 + 18 + 27 + 12 = 88.5.
[0130] This embodiment calculates the utilization rate, health value, and overall health of each resource type in steps, and distinguishes the importance of resources by weight. This can comprehensively reflect the health status of tenant resource usage, provide a basis for subsequent resource adjustments in the system, help to discover resource usage risks in advance, and ensure the stable operation of tenant services.
[0131] Based on any of the above embodiments, in the seventh embodiment of the present application, refer to Figure 2 , after step S50, including: Step S61 : determining initial GPU resource limits corresponding to the computing GPU, the rendering GPU, and the general-purpose GPU, respectively, based on the initial resource quotas corresponding to the GPUs.
[0132] In this embodiment, the portion of the initial resource quota corresponding to the GPU is the initial total amount of GPU resources obtained when the tenant registers; a computing GPU is a GPU type suitable for high-performance computing scenarios, a rendering GPU is a GPU type suitable for graphics rendering scenarios, and a general-purpose GPU is a GPU type that takes into account multiple scenarios; the initial GPU resource limit is the initial usage limit set for different types of GPUs.
[0133] As an optional implementation method, the system extracts the total GPU quota from the initial resource quota and allocates it according to the business scenario selected by the tenant when registering. For example, in computing-intensive scenarios, computing GPUs account for 60%, general-purpose GPUs account for 30%, and rendering GPUs account for 10%, thereby determining the initial resource limits of each type of GPU.
[0134] As another optional implementation, the system presets a basic allocation ratio for each type of GPU: computing type: rendering type: general type = 5:2:3. Combined with the initial total GPU quota, the system calculates the initial resource limit for each type of GPU, that is, the initial limit of a certain type = total quota × corresponding ratio.
[0135] Step S62: Obtain the tenant's initial GPU resource usage limit.
[0136] The initial GPU resource limit usage rate = (current GPU usage ÷ initial GPU resource limit) × 100%.
[0137] In this embodiment, the current GPU usage is the total amount of various GPU resources currently actually occupied by the tenant; the initial GPU resource limit is the sum of various GPU resource limits determined in step S61; and the initial GPU resource limit utilization rate is an indicator reflecting the proportion of current GPU resource usage to the initial limit.
[0138] As an optional implementation method, the system collects the computing, rendering, and general GPU resources currently used by the tenant in real time and sums them up to obtain the current GPU usage, substitutes it into the formula to calculate the usage rate, and retains two decimal places.
[0139] As another optional implementation, the system calculates the usage rate by GPU type and takes the average as the initial GPU resource limit usage rate. For example, if the computing type usage rate is 80%, the rendering type is 60%, and the general type is 70%, then the average usage rate is 70%.
[0140] Step S63: extract the daily peak usage rate, average usage rate, and resource request frequency corresponding to the tenant's historical GPU resource consumption data, and collect the current GPU resource usage and request queue length.
[0141] In this embodiment, historical GPU resource consumption data is a record of tenants' use of GPU resources over a period of time in the past; daily peak usage is the highest proportion of GPU usage per day in the historical data; average usage is the average proportion of GPU usage in the historical data; resource request frequency is the number of times a tenant applies for GPU resources per unit time; current GPU resource usage is the actual amount of GPU resources currently occupied; and request queue length is the number of GPU resource requests waiting to be allocated.
[0142] As an optional implementation, the system extracts historical GPU data from the database for the past 30 days, calculates the daily peak usage and average usage, counts the number of resource requests per hour as the request frequency, and obtains the current usage and request queue length through the resource monitoring module.
[0143] As another optional implementation, the system processes data separately by GPU type, such as extracting the historical peak usage of computing GPUs and the average usage of rendering GPUs, and then integrates various types of data for subsequent calculations.
[0144] Step S64 , generating a dynamic correction factor based on the daily peak usage rate, the average usage rate, the resource request frequency, the current GPU resource usage, and the request queue length.
[0145] In this embodiment, the dynamic correction factor is a coefficient calculated based on multiple GPU resource usage indicators, and is used to dynamically adjust GPU resource limits to make resource allocation more in line with actual needs.
[0146] As an optional implementation, the system sets weights for each indicator, with peak usage accounting for 30%, average usage accounting for 25%, request frequency accounting for 20%, current usage accounting for 15%, and request queue length accounting for 10%. After standardizing each indicator, the weighted sum is then mapped to a dynamic correction factor in the range of 0.8-1.5, such as a comprehensive score of 0.6 corresponds to a factor of 1.2.
[0147] As another optional implementation, the system sets a threshold judgment rule. When the average utilization rate is ≤50% and the request queue length is 0, the factor is 0.9; when the average utilization rate is 50%-80% and the queue length is ≤2, the factor is 1.1; when the average utilization rate is >80% or the queue length is >2, the factor is 1.3.
[0148] Step S65: Adjust the tenant's GPU resource limit according to the dynamic correction factor.
[0149] Exemplarily, the adjustment formula is: revised GPU resource limit = initial GPU resource limit × dynamic correction factor, where the revised GPU resource limit does not exceed the GPU resource upper limit set by the cluster for the tenant, nor is it lower than the tenant's current GPU resource usage.
[0150] In this embodiment, the revised GPU resource limit is the GPU resource limit that the tenant can use after the adjustment; the GPU resource upper limit set by the cluster for the tenant is the maximum limit set by the system to prevent the tenant from over-occupying GPU resources; and the current GPU resource usage is the actual amount of GPU resources occupied by the tenant during the adjustment.
[0151] As an optional implementation, the system multiplies the initial GPU resource limit by the dynamic correction factor to obtain a corrected value. If the value exceeds the upper limit, the upper limit is used; if it is lower than the current usage, the current usage is used to ensure reasonable adjustments.
[0152] As another optional implementation, the system calculates the revised GPU resource limit by GPU type, such as computational GPU = initial computational limit × dynamic correction factor, and then aggregates the various limits to obtain the total revised limit, while ensuring that all limits meet the upper and lower limit requirements.
[0153] For example, a tenant's initial GPU resource quota is 10 GPUs. Based on a compute:rendering:general use ratio of 5:2:3, the initial limits are determined to be 5, 2, and 3 GPUs, respectively. Currently, 6 GPUs are in use, and the initial limit utilization is (6 ÷ 10) × 100% = 60%. Historical data from the past 30 days shows: daily peak utilization of 75%, average utilization of 60%, request frequency of 3 per hour, current usage of 6 GPUs, and request queue length of 1. A dynamic correction factor of 1.1 is generated based on the threshold rule. The adjusted GPU resource limit is 10 × 1.1 = 11 GPUs. This value is within the cluster's upper limit of 15 GPUs and is higher than the current usage of 6 GPUs. The final adjusted limit is 11 GPUs.
[0154] This embodiment achieves precise adjustment of GPU resource limits by subdividing GPU types and combining historical and real-time data to generate dynamic correction factors. This not only meets the actual usage needs of tenants, but also avoids resource waste, thereby improving the utilization efficiency and scheduling flexibility of GPU resources.
[0155] Based on any of the above embodiments, in the eighth embodiment of the present application, refer to Figure 3 , in response to a trigger operation of dynamic resource demand, before the step of parsing the operation type of the trigger operation and the operation content of the trigger operation, including: Step B10, in response to the virtual cluster isolation instruction, determines the virtual cluster instance of the tenant from which the virtual cluster isolation instruction comes, the virtual cluster instance includes a virtual API server, a virtual controller manager and a virtual scheduler, and the virtual cluster instance and the physical cluster implement isolated scheduling of container resources through a resource mapping layer.
[0156] In this embodiment, the virtual cluster isolation instruction is an instruction used to start or configure the virtual cluster isolation mechanism; the virtual cluster instance is a virtual cluster unit created separately for a tenant, which contains core virtual components that implement cluster functions; the virtual API server is a component that processes API requests within the virtual cluster; the virtual controller manager is a component that manages controllers within the virtual cluster; the virtual scheduler is a component responsible for resource scheduling within the virtual cluster; the resource mapping layer is an intermediate layer that connects the virtual cluster and the physical cluster to implement resource mapping and isolation scheduling.
[0157] As an optional implementation method, after receiving the virtual cluster isolation instruction, the system parses the tenant identifier carried in the instruction, and searches for the corresponding virtual cluster instance from the cluster instance list based on the tenant identifier. The instance includes a preset virtual API server, virtual controller manager and virtual scheduler, and establishes an isolated scheduling relationship with the physical cluster through the resource mapping layer.
[0158] As another optional implementation, after receiving the isolation instruction, the system first verifies the legitimacy of the instruction, confirms the reliability of the source through instruction signature verification, and then locates the virtual cluster instance based on the tenant information in the instruction, checks whether the virtual components in the instance are operating normally, and ensures that the resource mapping layer is in an active state.
[0159] In step B20, the virtual control plane components are deployed in a dedicated namespace of the underlying physical Kubernetes cluster.
[0160] In this embodiment, the virtual control plane component is a collection of components in the virtual cluster instance that are responsible for control and management functions, including a virtual API server, a virtual controller manager, and a virtual scheduler; the underlying physical Kubernetes cluster is the physical cluster that actually runs the containers; and the dedicated namespace is an isolated space specifically allocated in the physical cluster for the virtual control plane component, which is used to limit the operating scope of the component.
[0161] As an optional implementation, the system uses a cluster deployment tool to point the deployment configuration file of the virtual control plane component to a dedicated namespace pre-created in the physical Kubernetes cluster. The configuration file contains information such as the component's image address and resource limits. The deployment tool automatically completes component deployment based on the configuration file.
[0162] As another optional implementation, the system uses a dynamic namespace creation mechanism. After receiving the deployment instructions, it automatically creates a dedicated namespace with the tenant identifier as a suffix in the physical cluster, then deploys the virtual control plane components to the namespace and configures the access permission policy of the namespace.
[0163] In response to a trigger operation of dynamic resource demand, the step of parsing the operation type and operation content of the trigger operation includes: Step B30, in response to a click operation on the virtual control plane component, determining an API interface; Step B40: If the API interface belongs to the virtual cluster instance associated with the tenant identifier of the tenant, the step of parsing the operation type and the operation content of the triggering operation is performed.
[0164] In this embodiment, the tap operation is an operation request initiated by the virtual control plane component; the API interface is an application programming interface used for interaction between different components or systems; and the tenant identifier is identification information used to uniquely identify the tenant.
[0165] As an optional implementation method, after the system detects a tap operation on the virtual control plane component, it captures the API interface call information corresponding to the operation, extracts the identification information of the API interface, and compares it with the API interface list contained in the virtual cluster instance associated with the tenant identifier. If there is a match, the parsing step is executed.
[0166] As another optional implementation, the system sets a dedicated access token for the API interface of each virtual cluster instance. When an API interface request corresponding to a tap operation is detected, the token carried in the request is verified. If the token matches the virtual cluster instance associated with the tenant identifier, it is confirmed that the API interface belongs to the instance, and then the parsing step is executed.
[0167] For example, a tenant triggers a virtual cluster isolation instruction. The system responds by determining the tenant's virtual cluster instance, which includes a virtual API server, a virtual controller manager, and a virtual scheduler, and is isolated from the physical cluster through a resource mapping layer. The system deploys these virtual control plane components into the "tenant-123-vcluster" dedicated namespace of the underlying physical Kubernetes cluster. When the tenant performs a click operation through the virtual controller manager, the system determines the corresponding API interface and checks that it belongs to the virtual cluster instance associated with the tenant identifier "123." The system then performs the steps of parsing the operation type and content of the triggering operation.
[0168] This embodiment implements tenant isolation based on virtual clusters by determining virtual cluster instances, deploying virtual control plane components to dedicated namespaces, and verifying API interface ownership, ensuring that tenant operations are only valid within their own virtual cluster instances, enhancing the security and isolation of the container management system, and ensuring the accuracy and pertinence of operation parsing triggered by dynamic resource demands.
[0169] The present application provides a control device for a container management system. The control device includes: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the container management system control method described in the first embodiment.
[0170] Reference below Figure 4 , which shows a schematic diagram of the hardware architecture of a control device suitable for implementing the container management system according to embodiments of the present application. The control device of the container management system according to embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, and in-vehicle terminals, as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The control device of the container management system shown is only an example and should not limit the functions and scope of use of the embodiments of the present application.
[0171] like Figure 4 As shown, the control device of the container management system may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the control device of the container management system. Processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems may be connected to I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, hard disk, etc.; and communication devices 1009. Communication devices 1009 may allow the control device of the container management system to communicate with other devices wirelessly or by wire to exchange data. While the diagram illustrates the control device of the container management system with various systems, it should be understood that implementation or presence of all of the illustrated systems is not required. More or fewer systems may alternatively be implemented or present.
[0172] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.
[0173] The container management system control device provided in this application, employing the container management system control method described in the aforementioned embodiments, can address the technical issue of some tenants excessively occupying resources and restricting the operation of other tenants' applications in scenarios with dynamic tenant business needs. Compared to existing technologies, the container management system control device provided in this application offers the same beneficial effects as those provided in the aforementioned embodiments. Other technical features of this container management system control device are the same as those disclosed in the aforementioned embodiments and are not further elaborated here.
[0174] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0175] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0176] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon. The computer-readable program instructions are used to execute the control method of the container management system in the above-mentioned embodiment.
[0177] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.
[0178] The computer-readable storage medium may be included in the control device of the container management system; or may exist independently without being assembled into the control device of the container management system.
[0179] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the control device of the container management system, the control device of the container management system: responds to the trigger operation of dynamic resource demand, parses the operation type of the trigger operation and the operation content of the trigger operation, wherein the operation type includes deploying new applications, performing scaling operations on existing applications, updating application image versions, and adjusting application runtime resource request parameters; obtains the initial resource quota allocated when the tenant corresponding to the trigger operation is registered, and the initial resource quota includes a first resource allocation amount, a second resource allocation amount, a third resource allocation amount, and a fourth resource allocation amount; determines the resource demand change corresponding to the trigger operation according to the operation type and the operation content; substitutes the used resource amount, the initial resource quota, and the resource demand change into the preset resource allocation formula to calculate the tenant's real-time resource allocation; adjusts the tenant's resource usage quota according to the real-time resource allocation, and synchronously updates the resource scheduling policy of the cluster node.
[0180] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0181] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0182] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0183] The computer-readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned container management system control method. This computer-readable storage medium can address the technical issue of some tenants excessively occupying resources, leading to operational limitations on other tenants' applications, in scenarios with dynamic tenant business needs. Compared to existing technologies, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the container management system control method provided in the aforementioned embodiments and are not further elaborated here.
[0184] An embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the computer program implements the steps of the above-mentioned method for controlling a container management system.
[0185] The computer program product provided in this application can address the technical issue of some tenants excessively occupying resources, leading to restricted operation of other tenants' applications, in scenarios with dynamic tenant business needs. Compared to the prior art, the beneficial effects of the computer program product provided in this embodiment are similar to those of the container management system control method provided in the aforementioned embodiment, and will not be elaborated upon here.
[0186] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent processing scope of the present application.
Claims
1. A method for controlling a container management system, characterized in that: Applied to a container management system, the control method of the container management system includes: In response to a trigger operation caused by dynamic resource demand, the operation type and operation content of the trigger operation are parsed, wherein the operation type includes deploying a new application, performing replica operations on an existing application, updating an application image version, and adjusting application runtime resource request parameters; Obtaining an initial resource quota allocated when the tenant corresponding to the trigger operation registers, the initial resource quota including a first resource allocation amount, a second resource allocation amount, a third resource allocation amount, and a fourth resource allocation amount; Determining a resource demand change corresponding to the triggering operation according to the operation type and the operation content; Substitute the used resources, initial resource quota, and resource demand change into the preset resource allocation formula to calculate the tenant's real-time resource allocation; The resource usage quota of the tenant is adjusted according to the real-time resource allocation amount, and the resource scheduling policy of the cluster node is synchronously updated.
2. The control method of the container management system according to claim 1, characterized in that: The step of parsing the operation type and operation content of the triggering operation in response to the dynamic resource demand includes: If the operation type is to deploy a new application, the operation content is parsed to obtain the preset number of replicas, the CPU request amount, memory request amount, GPU request amount, and storage request amount required for each replica; If the operation type is to scale replicas of an existing application, the operation content is parsed to obtain the current number of replicas, the target number of replicas, and the fixed resource configuration information of a single replica. The fixed resource configuration information includes the baseline amount of CPU, memory, GPU, and storage occupied by a single replica. If the operation type is to update the application image version, the operation content is parsed to obtain the resource requirement parameters of the current image version and the resource requirement parameters of the new version image. The resource requirement parameters include the CPU upper limit, memory upper limit, GPU usage threshold, and storage IOPS requirement; If the operation type is to adjust the resource request parameters of the application runtime, the operation content is parsed to obtain the resource type to be adjusted and the corresponding request amount adjustment value, where the resource type includes at least one of CPU, memory, GPU, and storage.
3. The control method of the container management system according to claim 1, wherein: The first resource is a CPU, the second resource is a memory, the third resource is a GPU, and the fourth resource is a storage. Before the step of obtaining the initial resource quota allocated when the tenant corresponding to the trigger operation is registered, the step includes: Receive a registration request from a tenant, parse the registration request to obtain tenant type information, business scenario description, and estimated resource requirements, where the tenant type information includes individual tenant, enterprise tenant, or priority tenant, and the business scenario description includes compute-intensive, storage-intensive, or network-intensive; The basic quota coefficient is determined based on the preset resource quota benchmark table and tenant type information. The basic quota coefficient for enterprise tenants is higher than that for individual tenants, and the basic quota coefficient for priority tenants is higher than that for ordinary tenants of the same type. Match the corresponding resource weight matrix based on the business scenario description. In compute-intensive scenarios, the weight of CPU and GPU is higher than that of memory and storage. In storage-intensive scenarios, the weight of storage is higher than that of CPU, GPU, and memory. Combining the basic quota coefficient, the resource weight matrix, and the estimated resource demand, the first resource allocation amount, the second resource allocation amount, the third resource allocation amount, and the fourth resource allocation amount are calculated to determine the tenant's initial resource quota; The initial resource quota is associated with the tenant identifier and stored, and a resource upper limit threshold for the tenant is set, where the resource upper limit threshold is a preset multiple of the initial resource quota and does not exceed a preset proportion of the total available resources of the cluster.
4. The control method of the container management system according to claim 1, wherein: The step of determining the resource demand change corresponding to the triggering operation according to the operation type and the operation content includes: If the operation type is to deploy a new application, the resource demand changes for CPU, memory, GPU, and storage are calculated based on the parsed CPU request, memory request, GPU request, and storage request for a single replica and the preset number of replicas. If the operation type is to scale replicas of an existing application, the resource demand change for each resource type is calculated based on the difference between the current number of replicas and the target number of replicas, combined with the fixed resource configuration information for a single replica. If the operation type is to update the application image version, calculate the difference in resource requirement parameters between the new version image and the current version image, and use the difference as the resource requirement change for each resource type; If the operation type is to adjust the resource request parameters when the application is running, the adjusted value of the request amount of each resource type obtained by parsing will be directly used as the corresponding resource demand change.
5. The control method of the container management system according to claim 1, wherein: The real-time resource allocation amount is obtained by the initial resource quota, the used resource amount, the resource type weight, the resource demand change amount and the tenant priority factor.
6. The method for controlling a container management system according to claim 1, wherein: After the step of adjusting the tenant's resource usage quota according to the real-time resource allocation amount and synchronously updating the resource scheduling policy of the cluster node, the method includes: For CPU, memory, GPU, and storage, calculate the ratio of used resources to the initial resource quota to obtain the utilization rate of each resource type; Based on the preset health assessment model, the utilization rate of each resource type is mapped to the corresponding health value; Determine the health weight of each resource type, with CPU and GPU having higher weights than memory and storage. The overall health of the tenant is calculated by calculating the health value of each resource type and configuring the health weight.
7. The control method of the container management system according to claim 1, wherein: After the step of adjusting the tenant's resource usage quota according to the real-time resource allocation amount and synchronously updating the resource scheduling policy of the cluster node, the method includes: Determine the initial GPU resource limits for the computing GPU, rendering GPU, and general-purpose GPU, respectively, based on the initial resource quotas corresponding to the GPUs. Get the tenant's initial GPU resource limit usage; Extract the tenant's historical GPU resource consumption data, including daily peak usage, average usage, and resource request frequency, as well as current GPU resource usage and request queue length. Generates dynamic correction factors based on daily peak usage, average usage, resource request frequency, current GPU resource usage, and request queue length; Adjust the GPU resource limit of tenants based on dynamic correction factors.
8. The control method of the container management system according to claim 1, wherein: Before the step of analyzing the operation type and the operation content of the triggering operation in response to the dynamic resource demand, the method includes: In response to a virtual cluster isolation instruction, determining a virtual cluster instance of the tenant from which the virtual cluster isolation instruction originated, the virtual cluster instance comprising a virtual API server, a virtual controller manager, and a virtual scheduler, wherein the virtual cluster instance and the physical cluster implement isolated scheduling of container resources through a resource mapping layer; Deploy the virtual control plane components in a dedicated namespace within the underlying physical Kubernetes cluster. The step of parsing the operation type and operation content of the triggering operation in response to the dynamic resource demand includes: In response to a click operation on the virtual control plane component, determining an API interface; If the API interface belongs to the virtual cluster instance associated with the tenant identifier of the tenant, the step of parsing the operation type of the triggering operation and the operation content of the triggering operation is performed.
9. A control device for a container management system, characterized in that: The control device of the container management system includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the control method of the container management system according to any one of claims 1 to 8.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method for controlling the container management system according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Resource scheduling method and device, electronic equipment and storage medium
CN115658304A
Resource quota adjustment method and device, electronic equipment and storage medium
CN116302472A
Cluster resource management system
CN118484297A
Multi-tenant management method and system based on container and Kubernetes
CN119668886A
Cited By
Method and system for maintaining consensus cluster majority in double-node cluster by utilizing migratable virtual consensus node, terminal and storage medium
CN122160242A
A method, system, terminal, and storage medium for maintaining a majority consensus in a two-node cluster using portable virtual consensus nodes.
CN122160242B