Multi-tenant management method and system based on container and Kubernetes
By adopting a multi-tenant management method based on containers and Kubernetes on the AI computing platform and using dynamic scheduling algorithms to allocate GPU resources, it solves the problem that traditional platforms are difficult to adapt to the high concurrency and dynamic computing power requirements of multiple users, and achieves efficient GPU resource management and improvement of computing efficiency.
Patent Information
- Application Number
- CN202510192108.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-21
AI Technical Summary
Traditional AI computing platforms are difficult to flexibly adapt to the high concurrency and dynamic computing power requirements of multiple users, resulting in low resource utilization and limited computing task scheduling efficiency. Especially when the scale of AI models continues to expand, fixed allocation methods are difficult to meet the computing power requirements in different scenarios.
The multi-tenant management method based on containers and Kubernetes is adopted, and computing power resource management is managed through distributed deployment architecture and containerization technology, and dynamic scheduling algorithms are used to allocate GPU resources to achieve efficient management of GPU resources.
It improves the management efficiency of GPU resources, ensures flexible adaptation and efficient utilization of resources, meets the computing power needs in different scenarios, and improves computing efficiency and resource utilization.
Smart Images

Figure CN119668886B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer resource management, and in particular to a multi-tenant management method and system based on containers and Kubernetes. Background Art
[0002] With the rapid development of AI and deep learning technologies, the demand for computing resources of AI algorithms continues to grow, especially in the training, data analysis and reasoning of large-scale models. The dependence on high-performance computing devices is increasing. Among them, GPU plays a key role in deep learning tasks due to its powerful parallel computing capabilities. Due to the continuous expansion of model scale and the increase in computational complexity, the required computing power is growing exponentially, which not only increases the demand for the number of GPUs, but also puts higher requirements on the architecture, storage bandwidth and energy efficiency of computing clusters.
[0003] The prior art has the following defects:
[0004] Traditional AI computing platforms usually adopt a centralized architecture or a fixed resource allocation model. When faced with high concurrency and dynamic computing power requirements of multiple users, they are often difficult to adapt flexibly, resulting in low resource utilization and limited computing task scheduling efficiency. As the scale of AI models continues to expand, the demand for computing resources for training and inference tasks has become more complex and changeable. A single fixed allocation method is difficult to meet the computing power requirements in different scenarios. Fixed allocation of GPU resources may cause some tenants to have excessive or insufficient resource allocation, thereby reducing the management efficiency of GPU resources.
[0005] Based on this, the present invention proposes a multi-tenant management method and system based on containers and Kubernetes, adopts a distributed deployment architecture, manages computing resources through containerization technology and Kubernetes (K8s) clusters, and uses a dynamic scheduling algorithm to allocate GPU resources, effectively improving the management efficiency of GPU resources. Summary of the invention
[0006] The purpose of the present invention is to provide a multi-tenant management method and system based on containers and Kubernetes to solve the deficiencies in the background technology.
[0007] In order to achieve the above object, the present invention provides the following technical solution: a multi-tenant management method based on containers and Kubernetes, the management method comprising the following steps:
[0008] The management system treats enterprises in the AI platform as independent tenants and uses UUID to generate encrypted directory names for each tenant for data isolation.
[0009] Encapsulate tenant tasks into several container instances and submit them to the Kubernetes cluster for computing. Each container instance performs Gang scheduling and Bin-Pack computing in the Kubernetes cluster.
[0010] Perform critical analysis on the GPU resources of the AI platform. When the critical edge is reached, classify all tenants according to their payment status, sort all categories by scheduling priority based on the cluster center value, and use a dynamic scheduling algorithm in each category to dynamically allocate GPU resources for each tenant and generate management strategies.
[0011] In a preferred embodiment, a dynamic scheduling algorithm is used in each category to dynamically allocate GPU resources to each tenant, including the following steps:
[0012] Get the GPU resource allocation of each category, and calculate the initial GPU resource limit usage rate of each tenant in the category based on the GPU resource allocation of the category;
[0013] After collecting the tenant's historical data and real-time data, a dynamic correction factor is generated for each tenant, and the initial GPU resource limit usage rate of each tenant is dynamically corrected by comparing the dynamic correction factor with the first correction threshold and the second correction threshold;
[0014] If the tenant's dynamic correction factor is greater than or equal to the first correction threshold, and the dynamic correction factor is less than or equal to the second correction threshold, the initial GPU resource limit usage rate is used as the tenant's current GPU resource limit usage rate;
[0015] If the tenant's dynamic correction factor is less than the first correction threshold, the tenant's initial GPU resource limit usage rate is reduced and then the current GPU resource limit usage rate is obtained;
[0016] If the tenant's dynamic correction factor is greater than the second correction threshold, the tenant's initial GPU resource limit usage rate is increased and then the current GPU resource limit usage rate is obtained.
[0017] In a preferred embodiment, a dynamic scheduling algorithm is used in each category to dynamically allocate GPU resources for each tenant and then generate a management strategy, including the following steps:
[0018] After dynamically correcting the initial GPU resource limit usage of all tenants to obtain the current GPU resource limit usage, obtain the real-time GPU resource usage of the tenants in the category, sum the real-time GPU resource usage of the tenants in the category to obtain the category real-time GPU resource usage, and finally sum the category real-time GPU resource usage of all categories to obtain the total real-time GPU resource usage of the AI platform;
[0019] If the total real-time GPU resource usage of the AI platform is greater than or equal to the usage threshold, the AI platform will replenish GPU resources. If the total real-time GPU resource usage of the AI platform is less than the usage threshold, the AI platform will allocate idle GPU resources to the categories sorted by scheduling priority.
[0020] In a preferred embodiment, after collecting the historical data and real-time data of the tenants, a dynamic correction factor is generated for each tenant. The historical data includes the energy efficiency ratio index, and the real-time data includes the usage rate fluctuation index and the usage rate growth index. The energy efficiency ratio index, the usage rate fluctuation index and the usage rate growth index are comprehensively calculated to obtain the dynamic correction factor, and the expression is: , where is the dynamic correction factor, is the energy efficiency ratio index, is the utilization growth index, is the utilization rate fluctuation index, , are the proportional coefficients of the energy efficiency ratio index and the utilization rate growth index, respectively, and , are greater than 0, is the monitoring period, Indicates the tenant's The GPU calculation output generated when Indicates the tenant's Energy consumption when Indicates the tenant's GPU usage at Represents a time increment.
[0021] In a preferred embodiment, the calculation expression of the usage rate fluctuation index is: , where To monitor the number of time points, For the GPU usage at each monitoring time point, It is the average GPU usage.
[0022] In a preferred embodiment, criticality analysis is performed on the GPU resources of the AI platform. When the critical edge is reached, all tenants are classified according to their payment status, and scheduling priorities are sorted for all categories in combination with the cluster center value, including the following steps:
[0023] Get the total GPU usage rate and usage rate threshold of the AI platform. The usage rate threshold is used to determine whether the GPU resources of the AI platform have reached the critical edge. If the total GPU usage rate is less than the usage rate threshold, it is determined that the GPU resources of the AI platform have not reached the critical edge. If the total GPU usage rate is greater than or equal to the usage rate threshold, it is determined that the GPU resources of the AI platform have reached the critical edge.
[0024] When it is determined that the GPU resources of the AI platform have reached the critical edge, the payment prices of all tenants are obtained, and all tenants are classified according to the preset price range. After all tenants are classified, the average payment price of each category is calculated and used as the cluster center value of the category;
[0025] After obtaining the cluster center values of all categories, all categories are sorted by scheduling priority from large to small according to the cluster center values.
[0026] In a preferred implementation, encapsulating the tenant's tasks into several container instances and submitting them to a Kubernetes cluster for computation includes the following steps:
[0027] According to the computing tasks submitted by the tenants, they are split into multiple computing subtasks, and the computing subtasks are packaged into container instance images using OCI-compatible tools. The container instance image contains the operating environment, dependent libraries, and execution scripts required for the subtasks. After the container instance image is built, it is uploaded to the container image repository supported by Kubernetes.
[0028] Generate a Kubernetes task definition file and submit it to the cluster, and tag the subtask with the tenant ID and job ID;
[0029] The Bin-Pack algorithm sorts all computing nodes and generates a node list. When a new task exists, an idle node is selected according to the positive order of the node list, and the new task is scheduled to be calculated on the node.
[0030] After all container instances obtain resources, the Gang scheduler starts the task, the task container runs, and the Kubernetes monitoring component is used to collect the GPU usage of the task. After the task is completed, the GPU resources are automatically released. If the task fails, it is automatically restarted according to the retry strategy of Kubernetes-Job.
[0031] In a preferred embodiment, the Bin-Pack algorithm sorts all computing nodes and generates a node list, including the following steps:
[0032] The Bin-Pack algorithm obtains the load rate, temperature fluctuation amplitude and energy efficiency of the node GPU, normalizes the load rate, temperature fluctuation amplitude and energy efficiency, maps the value range of the load rate, temperature fluctuation amplitude and energy efficiency to [0,1], obtains the normalized value of the load rate, the normalized value of the temperature fluctuation amplitude and the normalized value of the energy efficiency, adds the normalized value of the load rate to the normalized value of the temperature fluctuation amplitude and subtracts the normalized value of the energy efficiency to obtain the sorting coefficient, sorts all computing nodes from small to large according to the sorting coefficient, and generates a node list.
[0033] In a preferred embodiment, the management system treats the enterprises in the AI platform as independent tenants, uses UUID to generate an encrypted directory name for each tenant, and then performs data isolation, including the following steps:
[0034] The management system manages each enterprise in the AI platform as an independent tenant and generates a unique encrypted directory name for each tenant through UUID. When creating a tenant directory, it uses the access control mechanism to isolate each tenant's computing resources, storage space, and log files.
[0035] The data access scope is limited through a multi-layer permission management mechanism, and tenants' data access rights are managed by combining RBAC and ACL.
[0036] A multi-tenant management system based on containers and Kubernetes, including data isolation module, computing module, classification module, and scheduling module;
[0037] Data isolation module: treats enterprises in the AI platform as independent tenants, uses UUID to generate encrypted directory names for each tenant, and then performs data isolation. Tenant information and data isolation information are sent to the computing module;
[0038] Computing module: encapsulates tenant tasks into several container instances and submits them to the Kubernetes cluster for calculation. The calculated GPU resource usage results are sent to the classification module.
[0039] Classification module: When the critical edge is reached, all tenants are classified according to their payment status, and the classification results are sent to the scheduling module;
[0040] Scheduling module: After sorting the scheduling priorities of all categories based on the cluster center value, a dynamic scheduling algorithm is used in each category to dynamically allocate GPU resources for each tenant and generate management strategies.
[0041] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0042] 1. The present invention encapsulates tenants' tasks into several container instances and submits them to the Kubernetes cluster for calculation. Each container instance is Gang scheduled in the Kubernetes cluster, and the GPU resources of the AI platform are critically analyzed. When the critical edge is reached, all tenants are classified according to the tenant's payment situation, and all categories are prioritized by the cluster center value. Then, a dynamic scheduling algorithm is used in each category to dynamically allocate GPU resources for each tenant and generate management strategies. A distributed deployment architecture is adopted, computing power resources are managed through containerization technology and Kubernetes (K8s) cluster, and GPU resources are allocated using a dynamic scheduling algorithm, which effectively improves the management efficiency of GPU resources;
[0043] 2. In the present invention, when the total real-time GPU resource usage rate of the AI platform is greater than or equal to the usage rate threshold, the AI platform supplements GPU resources (the AI platform can dynamically expand the cloud GPU, for example, add 10 GPUs to ensure that all tenants can continue to run stably) to ensure stable use of all tenants. When the total real-time GPU resource usage rate of the AI platform is less than the usage rate threshold, the AI platform allocates idle GPU resources to the categories sorted by scheduling priority in sequence. When there are fewer idle GPU resources, they are preferentially allocated to the category with the first scheduling priority. When there are more idle GPU resources, GPU resources are allocated to the subsequent categories. In this way, the AI platform maximizes the use of GPU resources and improves computing efficiency while ensuring overall stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0045] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0047] Example 1: Please refer to Figure 1As shown, the multi-tenant management method based on containers and Kubernetes in this embodiment includes the following steps:
[0048] The management system treats the enterprises in the AI platform as independent tenants, uses UUID to generate encrypted directory names for each tenant, and then performs data isolation. During data isolation, each tenant can only access the data submitted by himself. The tenant's tasks are encapsulated into several container instances and submitted to the Kubernetes cluster for calculation. Each container instance is Gang scheduled in the Kubernetes cluster so that all computing tasks can obtain resources at the same time to avoid some tasks being blocked due to waiting for resources. The Bin-Pack algorithm is used to fill the existing GPU servers with computing tasks as much as possible, reducing the waste of GPU resources and improving the overall computing power utilization. The AI platform's GPU resources are critically analyzed. When the critical edge is reached, all tenants are classified according to the tenant's payment situation. After all categories are prioritized according to the cluster center value, a dynamic scheduling algorithm is used in each category to dynamically allocate each tenant's GPU resources and generate management strategies.
[0049] This application encapsulates tenants' tasks into several container instances and submits them to the Kubernetes cluster for calculation. Each container instance is Gang scheduled in the Kubernetes cluster, and critical analysis is performed on the GPU resources of the AI platform. When the critical edge is reached, all tenants are classified according to their payment status, and all categories are prioritized by the cluster center value. In each category, a dynamic scheduling algorithm is used to dynamically allocate GPU resources for each tenant and generate management strategies. A distributed deployment architecture is adopted, computing resources are managed through containerization technology and Kubernetes (K8s) clusters, and GPU resources are allocated using dynamic scheduling algorithms, effectively improving the management efficiency of GPU resources.
[0050] In this application, UUID is a universal unique identifier used to generate a unique encrypted directory name for each tenant to ensure data isolation. Even if the directory structures of different tenants are similar, the random string generated by UUID can prevent directory name conflicts and improve security. For example, the data directory of tenant A may be / data / 3f9a0c6e-12d4-4a8b-97f3-5bcd4567e89a; the data directory of tenant B may be / data / 8d7c3e12-5e9f-4b1a-bc2e-7fa8d234c6b2; in this way, even if the system paths of the two tenants are similar, there will be no data mixing or access conflicts.
[0051] Containers are a lightweight virtualization technology that can encapsulate applications and all their dependencies, enabling them to run in any environment. Each tenant's AI computing tasks will be encapsulated into multiple container instances and then submitted to Kubernetes for scheduling. For example: training task 1 → runs in container ai-train-job-001; training task 2 → runs in container ai-train-job-002; data processing task → runs in container data-process-job-003; these containers will run in the Kubernetes cluster, and Kubernetes will be responsible for managing the allocation of computing resources.
[0052] Kubernetes (K8s) - container orchestration system: Kubernetes is an open source container orchestration system for managing containerized applications. In this solution, the tenant's container instance is submitted to the Kubernetes cluster for computing, and Kubernetes is responsible for allocating GPU resources, task scheduling and optimization.
[0053] Gang scheduling is a task group scheduling strategy used for parallel computing tasks to ensure that all tasks obtain computing resources at the same time, otherwise no task will be started. If a computing task A depends on task B and task C to run together, Gang scheduling will ensure that tasks A, B, and C obtain GPU resources at the same time, otherwise task A will not be started to prevent resource waste. All computing nodes start synchronously to improve the efficiency of distributed training. Example: Suppose there are 3 computing tasks: task A (4 GPUs), task B (2 GPUs), and task C (1 GPU. If only task A obtains 2 GPU resources, Gang scheduling will not let it start, but wait until B and C also obtain the required GPUs before running together.
[0054] The Bin-Pack algorithm is an algorithm for optimizing resource allocation. Its goal is to fill existing GPU servers as efficiently as possible to reduce the waste of idle resources and prevent multiple tasks from occupying the GPU but not fully utilizing it, resulting in waste of computing resources. It allows tasks to fill up the GPU server as much as possible, reduce the number of idle GPUs, and improve AI computing efficiency.
[0055] Embodiment 2: The management system treats the enterprises in the AI platform as independent tenants, uses UUID to generate an encrypted directory name for each tenant, and then performs data isolation. During data isolation, each tenant can only access the data submitted by himself, including the following steps:
[0056] The management system manages each enterprise in the AI platform as an independent tenant and generates a unique encrypted directory name for each tenant through UUID (Universally Unique Identifier). This encrypted directory name not only prevents directory name conflicts, but also improves data security and makes unauthorized access more difficult. When creating a tenant directory, the management system will combine the access control mechanism to ensure that each tenant's computing resources, storage space, and log files are isolated to avoid cross-tenant data leakage or misoperation;
[0057] To ensure data isolation, each tenant can only access the data submitted by themselves, and the system limits the scope of data access through a multi-layer permission management mechanism. The tenant's data access rights are managed by combining RBAC (role-based access control) and ACL (access control list) to ensure that data between different tenants cannot be read, modified or deleted. In addition, the system supports fine-grained permission configuration, which can set different data access levels for different users within the tenant to ensure that team members can perform data operations within the appropriate permission range.
[0058] Generate a unique encrypted directory name for each tenant through UUID. When creating a tenant directory, combine the access permission control mechanism to isolate each tenant's computing resources, storage space, and log files, including the following steps:
[0059] The system generates a UUID (Universally Unique Identifier) for each tenant and creates a unique encrypted directory name based on the UUID to ensure the randomness and unpredictability of the directory name and prevent data conflicts or malicious access between tenants. The encrypted directory name is further processed by an encryption algorithm to enhance security and map the directory structure to a distributed storage system to improve data access efficiency.
[0060] When creating a tenant directory, the system categorizes them by computing resources, storage space, and log files, and allocates an independent storage path to each tenant. For example:
[0061] Computing resource directory ( / compute / ${UUID} / ): stores computing task instances such as AI training tasks and inference tasks submitted by tenants.
[0062] Storage space directory ( / storage / ${UUID} / ): used to store training data, model files, inference results, etc.
[0063] Log file directory ( / logs / ${UUID} / ): stores tenant task execution logs, error logs, and system access records for subsequent auditing and analysis.
[0064] The encrypted directory name is further processed by an encryption algorithm, which includes the following steps:
[0065] AES (Advanced Encryption Standard) is used to encrypt directory names. AES has efficient encryption performance and strong security, and supports 128, 192, and 256-bit key lengths. The encryption key (Secret-Key) is generated by the system Key Management Service (KMS) and stored in a secure manner. The randomly generated 16-byte data ensures that the same UUID will not produce the same result after encryption, thereby improving security. For example:
[0066] Key: F1A2B3C4D5E6F7G8H9I0J1K2L3M4N5O6;
[0067] IV (initialization vector): 7D8A9C0E1F23456789ABCDEF01234567;
[0068] Use AES-CBC (Cipher Block Chaining Mode) or AES-GCM (Galois / Counter-Mode) to encrypt the UUID to obtain the encrypted directory name. For example, use AES-256-CBC for encryption:
[0069] Original UUID: 550e8400-e29b-41d4-a716-446655440000;
[0070] After AES encryption (Base64 encoding): QmFzZTY0RW5jb2RlZERhdGE=;
[0071] The encrypted data can be Base64 encoded for easy storage and transmission. The encrypted UUID is combined with the tenant storage path to form the final directory name, for example:
[0072] / tenant_data / QmFzZTY0RW5jb2RlZERhdGE=;
[0073] In this way, the directory structure of each tenant can ensure uniqueness and prevent the tenant's UUID from being directly parsed. The encryption key and IV are stored by the Key Management Service (KMS) to prevent unauthorized access. Only system administrators and authorized processes can decrypt the directory name, and ordinary users cannot directly access or parse the original UUID.
[0074] The data access scope is limited by a multi-layer permission management mechanism. The tenant's data access rights are managed by combining RBAC and ACL, including the following steps:
[0075] By default, the data access rights of each tenant are subject to tenant-level permission control, that is, tenants can only access their own computing resources, storage space, and log files. The system binds the storage directory through the tenant ID to ensure that the resources of different tenants are completely isolated logically. All resource access requests must be authenticated, and unauthorized tenants cannot access the data or computing resources of other tenants.
[0076] Within a tenant, the system sets the permission levels for different users based on RBAC (Role-Based-Access-Control). Each user's permissions are determined by their role rather than directly assigned permissions, which simplifies management and improves security.
[0077] Super-Admin: can manage the computing resources, storage space, and user rights allocation of the entire tenant.
[0078] Resource Administrator (Resource-Admin): Responsible for managing tenants' computing tasks, storage files, logs and other resources.
[0079] Ordinary users (User): can only submit computing tasks and read the data uploaded by themselves, but cannot manage other users or tenant resources.
[0080] RBAC mainly manages user roles, while ACL (Access-Control-List) further provides fine-grained permission management, allowing users within a tenant to control access to specific resources. For example, ACL allows tenant administrators to define whether specific users can:
[0081] Read: View stored data, log files, or calculation results.
[0082] Write: upload files, submit computing tasks, or modify existing data.
[0083] Execute: Run specific computing tasks, call API interfaces, or manage computing instances.
[0084] ACL rules can be bound to specific data directories, computing tasks, or API requests to ensure precise control of access rights and prevent unauthorized users from tampering with critical data or running high-consumption computing tasks.
[0085] When users access tenant resources, the system uses multi-factor authentication (MFA) to improve data security. Common authentication methods include:
[0086] Password + One-Time-Password (OTP): After logging in, users need to enter a dynamic verification code obtained via SMS, email, or an identity verification application (such as Google Authenticator).
[0087] Token-based authentication (OAuth / JWT): Users access tenant resources through OAuth authorization or JWT tokens to ensure the security of requests.
[0088] IP Whitelisting: Tenant administrators can set the IP address range that is allowed to access resources to prevent unauthorized devices from accessing.
[0089] Encapsulate tenant tasks into several container instances and submit them to the Kubernetes cluster for computing. Each container instance is Gang scheduled in the Kubernetes cluster so that all computing tasks can obtain resources at the same time, avoiding some tasks from being blocked due to waiting for resources. The Bin-Pack algorithm is used to fill the existing GPU servers with computing tasks as much as possible, reducing the waste of GPU resources and improving the overall computing power utilization. The following steps are included:
[0090] According to the computing tasks submitted by the tenants, they are split into multiple independent computing subtasks, and the computing subtasks are packaged into container instance images using OCI-compatible tools. The container instance image contains the operating environment, dependent libraries, and execution scripts required for the subtasks. After the container instance image is built, it is uploaded to the container image repository supported by Kubernetes (such as Docker-Hub, Harbor, ECR) to ensure that the subtasks can be pulled and run on different nodes.
[0091] Generate a Kubernetes task definition file (YAML), specify computing resource requirements, environment variables, storage volumes, network policies, etc., and submit it to the cluster. Subtasks are marked with tenant IDs (Tenant-ID) and job IDs (Job-ID) to ensure that tasks of different tenants do not interfere with each other.
[0092] Gang scheduling mechanism: All container instances in a task group must obtain resources at the same time before they can be started, to prevent some tasks from being blocked due to insufficient resources, resulting in overall computing failure or delay.
[0093] All subtasks in a task group need to match the same resource constraints. If sufficient resources cannot be allocated at one time, the scheduler will wait for the resources to be released before allocating them as a whole.
[0094] The Bin-Pack algorithm sorts all computing nodes and generates a node list. When there is a new task, an idle node is selected according to the positive order of the node list, and the new task is scheduled to be calculated on the node.
[0095] After all container instances obtain resources, the Gang scheduler starts the task and the task container officially runs. Kubernetes monitoring components (such as Prometheus and Metrics-Server) are used to collect the GPU usage of the task. After the task is completed, the GPU resources are automatically released. If the task fails, the system can automatically restart according to the retry strategy of Kubernetes-Job, or manually intervene after analyzing the cause according to the fault log.
[0096] The Bin-Pack algorithm sorts all computing nodes and generates a node list. When a new task exists, an idle node is selected according to the node list in positive order, and the new task is scheduled to be calculated on the node, including the following steps:
[0097] The Bin-Pack algorithm obtains the load rate, temperature fluctuation amplitude and energy efficiency of the node GPU, normalizes the load rate, temperature fluctuation amplitude and energy efficiency, maps the value range of the load rate, temperature fluctuation amplitude and energy efficiency to [0,1], obtains the normalized value of the load rate, the normalized value of the temperature fluctuation amplitude and the normalized value of the energy efficiency, adds the normalized value of the load rate to the normalized value of the temperature fluctuation amplitude and subtracts the normalized value of the energy efficiency to obtain the sorting coefficient, sorts all computing nodes from small to large according to the sorting coefficient, and generates a node list.
[0098] The calculation logic of the temperature fluctuation amplitude is as follows: obtain the temperature values of the computing node GPU at multiple time points within the monitoring period, calculate the temperature standard deviation based on the temperature values at multiple time points, and use the temperature standard deviation as the temperature fluctuation amplitude. The larger the temperature fluctuation amplitude, the greater the temperature fluctuation of the computing node GPU within the monitoring period, that is, the worse the stability of the computing node.
[0099] The calculation logic of the load rate is as follows: obtain the total GPU running time of the computing node and the current computing time of the GPU of the computing node, and divide the current computing time of the GPU by the total GPU running time to obtain the load rate. The larger the load rate, the longer the current GPU usage time of the computing node is, and the more it needs to stop and rest.
[0100] The calculation logic of energy efficiency is: divide the load rate of the GPU of the computing node by the current power consumption to obtain the energy efficiency. The greater the energy efficiency, the lower the energy consumption of the computing node.
[0101] Perform critical analysis on the GPU resources of the AI platform. When the critical edge is reached, classify all tenants according to their payment status and prioritize all categories based on the cluster center value. This includes the following steps:
[0102] Get the total GPU usage and usage threshold of the AI platform. The usage threshold is used to determine whether the GPU resources of the AI platform have reached the critical edge. If the total GPU usage is less than the usage threshold, it is determined that the GPU resources of the AI platform have not reached the critical edge. If the total GPU usage is greater than or equal to the usage threshold, it is determined that the GPU resources of the AI platform have reached the critical edge.
[0103] When it is determined that the GPU resources of the AI platform have reached the critical edge, the payment prices of all tenants are obtained, and all tenants are classified according to the preset price range. After all tenants are classified, the average payment price of each category is calculated and used as the cluster center value of the category;
[0104] After obtaining the cluster center values of all categories, all categories are sorted by scheduling priority from large to small according to the cluster center values, so that the categories with higher rankings are given priority in resource scheduling analysis.
[0105] In the GPU resource scheduling process of the AI platform, when GPU resources are close to the critical edge, the scheduling priority is sorted by comprehensively considering the payment situation of the tenants to ensure that high-paying tenants can obtain resources first. The platform first obtains the current utilization rate of all GPU nodes and determines whether the GPU resources have reached the critical edge based on the set utilization threshold.
[0106] Assume that the current total GPU usage is 85%, and the preset usage threshold is 80%. Since 85%>80%, it can be determined that the GPU resources have reached the critical edge.
[0107] The platform obtains the payment information of all current tenants, and classifies all tenants according to the payment price based on the preset price range. For example, the following price range can be set:
[0108] High payment: greater than or equal to 100 yuan / hour;
[0109] Medium payment: greater than or equal to 50 yuan / hour, but less than 100 yuan / hour;
[0110] Low fee: less than 50 yuan / hour.
[0111] Tenants are classified according to payment prices. The cluster center value refers to the average payment price of each category, which is used to indicate the representative payment level of the category. The categories are sorted according to the cluster center value of each category, and the categories with larger cluster center values are given higher scheduling priorities. According to the scheduling priority, the platform will first allocate GPU resources to high-paying categories, followed by medium-paying categories, and finally low-paying categories. Within each category, tenants will further schedule resources according to their specific resource requirements. Scheduling priority sorting ensures that resources can be allocated to tenants with higher payments first, improves the profitability of the platform, and protects the interests of high-paying tenants when resources are tight. For example:
[0112] High-paying category (priority 1): The platform will prioritize allocating resources to tenants A, C, and E.
[0113] Medium payment category (priority 2): The platform allocates resources to tenant D.
[0114] Low-paying category (priority 3): Due to resource constraints, Tenant B may need to wait or delay resource allocation.
[0115] Through the above process, when GPU resources are tight, the AI platform can classify and prioritize tenants based on their payment status, ensuring that high-paying tenants get priority in resource allocation. This strategy helps improve the fairness of resource allocation and protect the commercial interests of the platform.
[0116] In each category, a dynamic scheduling algorithm is used to dynamically allocate GPU resources for each tenant and generate a management policy, including the following steps:
[0117] Get the GPU resource allocation of each category, and calculate the initial GPU resource limit usage rate of each tenant in the category based on the GPU resource allocation of the category;
[0118] After collecting the historical data and real-time data of tenants, a dynamic correction factor is generated for each tenant. The historical data includes the energy efficiency ratio index, and the real-time data includes the utilization rate fluctuation index and the utilization rate growth index. The energy efficiency ratio index, the utilization rate fluctuation index and the utilization rate growth index are comprehensively calculated to obtain the dynamic correction factor. The expression is: , where is the dynamic correction factor, is the energy efficiency ratio index, is the utilization growth index, is the utilization rate fluctuation index, , are the proportional coefficients of the energy efficiency ratio index and the utilization rate growth index, respectively, and , are greater than 0, is the monitoring period, Indicates the tenant's GPU computing output (such as AI reasoning frame rate) generated when Indicates the tenant's Energy consumption (e.g. kWh), Indicates the tenant's GPU usage at represents the time increment, It represents the limit operation, that is, when the time increment approaches zero infinitely, the instantaneous rate of change of GPU utilization is calculated;
[0119] The calculation expression of the utilization rate fluctuation index is: , where To monitor the number of time points, For the GPU usage at each monitoring time point, is the average GPU usage rate, when , it indicates that the tenant's GPU usage is on an upward trend. The larger the usage fluctuation index, the more unstable the tenant's GPU usage. Therefore, the usage growth index is corrected by the usage fluctuation index to make the calculation more accurate.
[0120] Dynamically correct the initial GPU resource limit usage rate of each tenant by comparing the dynamic correction factor with the first correction threshold and the second correction threshold;
[0121] If the tenant's dynamic correction factor is greater than or equal to the first correction threshold, and the dynamic correction factor is less than or equal to the second correction threshold, the initial GPU resource limit usage rate is used as the tenant's current GPU resource limit usage rate;
[0122] If the tenant's dynamic correction factor is less than the first correction threshold, the tenant's initial GPU resource limit usage rate is reduced to obtain the current GPU resource limit usage rate, that is, the tenant's initial GPU resource limit usage rate is reduced by 10% to obtain the current GPU resource limit usage rate;
[0123] If the tenant's dynamic correction factor is greater than the second correction threshold, the tenant's initial GPU resource limit usage rate is increased to obtain the current GPU resource limit usage rate, that is, the tenant's initial GPU resource limit usage rate is increased by 10% to obtain the current GPU resource limit usage rate.
[0124] After dynamically correcting the initial GPU resource limit usage of all tenants to obtain the current GPU resource limit usage, obtain the real-time GPU resource usage of the tenants in the category, sum the real-time GPU resource usage of the tenants in the category to obtain the category real-time GPU resource usage, and finally sum the category real-time GPU resource usage of all categories to obtain the total real-time GPU resource usage of the AI platform;
[0125] If the total real-time GPU resource utilization of the AI platform is greater than or equal to the utilization threshold, the AI platform will supplement GPU resources (the AI platform can dynamically expand the cloud GPU, for example, add 10 GPUs to ensure that all tenants can continue to run stably) to ensure stable use by all tenants. If the total real-time GPU resource utilization of the AI platform is less than the utilization threshold, the AI platform will allocate idle GPU resources to the categories sorted by scheduling priority in sequence. When there are fewer idle GPU resources, they will be allocated to the category with the first scheduling priority. When there are more idle GPU resources, GPU resources will be allocated to the subsequent categories. In this way, the AI platform can maximize the use of GPU resources and improve computing efficiency while ensuring overall stability.
[0126] Embodiment 3: The multi-tenant management system based on containers and Kubernetes described in this embodiment includes a data isolation module, a computing module, a classification module, and a scheduling module;
[0127] Data isolation module: treats enterprises in the AI platform as independent tenants, uses UUID to generate encrypted directory names for each tenant, and then performs data isolation. Tenant information and data isolation information are sent to the computing module;
[0128] Computing module: encapsulates tenant tasks into several container instances and submits them to the Kubernetes cluster for calculation. The calculated GPU resource usage results are sent to the classification module.
[0129] Classification module: When the critical edge is reached, all tenants are classified according to their payment status, and the classification results are sent to the scheduling module;
[0130] Scheduling module: After sorting the scheduling priorities of all categories based on the cluster center value, a dynamic scheduling algorithm is used in each category to dynamically allocate GPU resources for each tenant and generate management strategies.
[0131] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0132] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0133] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only specific implementation methods. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A multi-tenant management method based on containers and Kubernetes, characterized by: The management method comprises the following steps: The management system treats enterprises in the AI platform as independent tenants and uses UUID to generate encrypted directory names for each tenant for data isolation. Encapsulate tenant tasks into several container instances and submit them to the Kubernetes cluster for computing. Each container instance performs Gang scheduling and Bin-Pack computing in the Kubernetes cluster. Perform critical analysis on the GPU resources of the AI platform. When the critical edge is reached, classify all tenants according to their payment status, sort all categories by scheduling priority based on the cluster center value, and use the dynamic scheduling algorithm in each category to dynamically allocate GPU resources for each tenant and generate management strategies. Generate a dynamic correction factor for each tenant, the expression is: , where is the dynamic correction factor, is the energy efficiency ratio index, is the utilization growth index, is the utilization rate fluctuation index, and , where To monitor the number of time points, For the GPU usage at each monitoring time point, is the average GPU usage, , are the proportional coefficients of the energy efficiency ratio index and the utilization rate growth index, respectively, and , are greater than 0, is the monitoring period, Indicates the tenant's The GPU calculation output generated when Indicates the tenant's Energy consumption when Indicates the tenant's GPU usage at represents the time increment; The initial GPU resource limit usage rate of each tenant is dynamically corrected by comparing the dynamic correction factor with the first correction threshold and the second correction threshold.
2. The multi-tenant management method based on containers and Kubernetes according to claim 1 is characterized in that: In each category, a dynamic scheduling algorithm is used to dynamically allocate GPU resources to each tenant, including the following steps: Get the GPU resource allocation of each category, and calculate the initial GPU resource limit usage rate of each tenant in the category based on the GPU resource allocation of the category; After collecting the tenant's historical data and real-time data, a dynamic correction factor is generated for each tenant, and the initial GPU resource limit usage rate of each tenant is dynamically corrected by comparing the dynamic correction factor with the first correction threshold and the second correction threshold; If the tenant's dynamic correction factor is greater than or equal to the first correction threshold, and the dynamic correction factor is less than or equal to the second correction threshold, the initial GPU resource limit usage rate is used as the tenant's current GPU resource limit usage rate; If the tenant's dynamic correction factor is less than the first correction threshold, the tenant's initial GPU resource limit usage rate is reduced and then the current GPU resource limit usage rate is obtained; If the tenant's dynamic correction factor is greater than the second correction threshold, the tenant's initial GPU resource limit usage rate is increased and then the current GPU resource limit usage rate is obtained.
3. The multi-tenant management method based on containers and Kubernetes according to claim 2 is characterized in that: In each category, the dynamic scheduling algorithm is used to dynamically allocate GPU resources to each tenant and generate a management policy, which includes the following steps: After dynamically correcting the initial GPU resource limit usage of all tenants to obtain the current GPU resource limit usage, obtain the real-time GPU resource usage of the tenants in the category, sum the real-time GPU resource usage of the tenants in the category to obtain the category real-time GPU resource usage, and finally sum the category real-time GPU resource usage of all categories to obtain the total real-time GPU resource usage of the AI platform; If the total real-time GPU resource usage of the AI platform is greater than or equal to the usage threshold, the AI platform will replenish GPU resources. If the total real-time GPU resource usage of the AI platform is less than the usage threshold, the AI platform will allocate idle GPU resources to the categories sorted by scheduling priority.
4. The multi-tenant management method based on containers and Kubernetes according to claim 3 is characterized in that: After collecting the tenants' historical data and real-time data, a dynamic correction factor is generated for each tenant. The historical data includes the energy efficiency ratio index, and the real-time data includes the usage rate fluctuation index and the usage rate growth index.
5. The multi-tenant management method based on containers and Kubernetes according to claim 4 is characterized in that: Perform critical analysis on the GPU resources of the AI platform. When the critical edge is reached, classify all tenants according to their payment status and prioritize all categories based on the cluster center value. This includes the following steps: Get the total GPU usage rate and usage rate threshold of the AI platform. The usage rate threshold is used to determine whether the GPU resources of the AI platform have reached the critical edge. If the total GPU usage rate is less than the usage rate threshold, it is determined that the GPU resources of the AI platform have not reached the critical edge. If the total GPU usage rate is greater than or equal to the usage rate threshold, it is determined that the GPU resources of the AI platform have reached the critical edge. When it is determined that the GPU resources of the AI platform have reached the critical edge, the payment prices of all tenants are obtained, and all tenants are classified according to the preset price range. After all tenants are classified, the average payment price of each category is calculated and used as the cluster center value of the category; After obtaining the cluster center values of all categories, all categories are sorted by scheduling priority from large to small according to the cluster center values.
6. The multi-tenant management method based on containers and Kubernetes according to claim 5 is characterized in that: Encapsulate the tenant's tasks into several container instances and submit them to the Kubernetes cluster for computing, including the following steps: According to the computing tasks submitted by the tenants, they are split into multiple computing subtasks, and the computing subtasks are packaged into container instance images using OCI-compatible tools. The container instance image contains the operating environment, dependent libraries, and execution scripts required for the subtasks. After the container instance image is built, it is uploaded to the container image repository supported by Kubernetes. Generate a Kubernetes task definition file and submit it to the cluster, and tag the subtask with the tenant ID and job ID; The Bin-Pack algorithm sorts all computing nodes and generates a node list. When a new task exists, an idle node is selected according to the positive order of the node list, and the new task is scheduled to be calculated on the node. After all container instances obtain resources, the Gang scheduler starts the task, the task container runs, and the Kubernetes monitoring component is used to collect the GPU usage of the task. After the task is completed, the GPU resources are automatically released. If the task fails, it is automatically restarted according to the retry strategy of Kubernetes-Job.
7. The multi-tenant management method based on containers and Kubernetes according to claim 6 is characterized in that: The n-Pack algorithm sorts all computing nodes and generates a node list, including the following steps: The Bin-Pack algorithm obtains the load rate, temperature fluctuation amplitude and energy efficiency of the node GPU, normalizes the load rate, temperature fluctuation amplitude and energy efficiency, maps the value range of the load rate, temperature fluctuation amplitude and energy efficiency to [0,1], obtains the normalized value of the load rate, the normalized value of the temperature fluctuation amplitude and the normalized value of the energy efficiency, adds the normalized value of the load rate to the normalized value of the temperature fluctuation amplitude and subtracts the normalized value of the energy efficiency to obtain the sorting coefficient, sorts all computing nodes from small to large according to the sorting coefficient, and generates a node list.
8. The multi-tenant management method based on containers and Kubernetes according to claim 7 is characterized in that: The management system treats enterprises in the AI platform as independent tenants, uses UUID to generate encrypted directory names for each tenant, and then performs data isolation, including the following steps: The management system manages each enterprise in the AI platform as an independent tenant and generates a unique encrypted directory name for each tenant through UUID. When creating a tenant directory, it uses the access control mechanism to isolate each tenant's computing resources, storage space, and log files. The data access scope is limited through a multi-layer permission management mechanism, and tenants' data access rights are managed by combining RBAC and ACL.
9. A multi-tenant management system based on containers and Kubernetes, used to implement the management method according to any one of claims 1 to 8, characterized in that: Including data isolation module, calculation module, classification module and scheduling module; Data isolation module: treats enterprises in the AI platform as independent tenants, uses UUID to generate encrypted directory names for each tenant, and then performs data isolation. Tenant information and data isolation information are sent to the computing module; Computing module: encapsulates tenant tasks into several container instances and submits them to the Kubernetes cluster for calculation. The calculated GPU resource usage results are sent to the classification module. Classification module: When the critical edge is reached, all tenants are classified according to their payment status, and the classification results are sent to the scheduling module; Scheduling module: After sorting the scheduling priorities of all categories based on the cluster center value, a dynamic scheduling algorithm is used in each category to dynamically allocate GPU resources for each tenant and generate management strategies.
Citation Information
Patent Citations
Distributed processing method, device and equipment
CN118780342A