Computing power resource isolation method and device for intelligent computing center cloud platform
By creating a dedicated VKS for each tenant in the Kubernetes cluster and setting up controllers and synchronizers, the problems of uneven distribution of computing resources and insufficient security are solved, achieving efficient resource isolation and security management, and improving the utilization rate and isolation of computing resources of the intelligent computing center cloud platform.
Patent Information
- Application Number
- CN202510948342.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-10-31
AI Technical Summary
In Kubernetes clusters, the allocation and utilization of computing resources are low, and the lack of effective isolation mechanisms leads to resource waste and security risks. Traditional namespaces are insufficient in terms of security and permission management, affecting performance and reliability.
Create a dedicated VKS for each tenant on the Kubernetes physical cluster, achieve logical isolation of computing resources through independent namespaces, set up controllers and synchronizers, perform permission verification and forwarding of operation requests based on preset rules to ensure the execution of legitimate requests, and configure network policies and storage isolation mechanisms.
It achieves efficient isolation of computing resources, improves resource utilization and security, and ensures computing resource management capabilities and security in a multi-tenant environment.
Smart Images

Figure CN120872587A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of intelligent computing centers, smart computing centers, computing infrastructure, and smart cloud technologies, specifically to a method and apparatus for isolating computing resources on an intelligent computing center cloud platform. Background Technology
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "smart computing centers" have emerged.
[0003] An "intelligent computing center" refers to a facility that provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as deep learning model development, model fine-tuning, and model inference) by utilizing large-scale heterogeneous computing resources, including general-purpose and intelligent computing power. Intelligent computing centers encompass facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.
[0004] "Intelligent computing center" includes, but is not limited to, "intelligent computing center".
[0005] "Intelligent computing center" or artificial intelligence computing center is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting artificial intelligence computing architecture.
[0006] "Computing power" is the core of "intelligent computing center" and "smart computing center". It is the ability of computer equipment or computing / data center to process information. It is the ability of computer hardware and software to work together to perform a certain computing requirement. It is the computing power to achieve the target output by processing information data. It is a new type of productivity that integrates information computing power, network carrying capacity and data storage capacity. It mainly provides services to society through computing power infrastructure.
[0007] When intelligent computing centers undertake computing power tasks, an efficient computing resource management and scheduling platform becomes crucial. Kubernetes (K8s), as a popular container orchestration platform, has been widely adopted in enterprise production environments. Kubernetes' flexibility and scalability enable it to meet the needs of applications of different sizes and types. However, in traditional Kubernetes clusters, the allocation and utilization of computing resources are relatively inefficient. Users often request more computing resources than they actually need, leading to wasted resources. Furthermore, due to the lack of effective isolation mechanisms, applications from different tenants may interfere with each other due to competition for computing resources, causing performance degradation or unavailability. This deficiency in computing resource management not only increases the complexity of operation and maintenance but also limits the efficiency of the cluster. While traditional namespaces provide a certain degree of computing resource isolation, they are still insufficient in terms of security and access control, potentially leading to data leaks and abuse of computing resources. Simultaneously, the lack of fine-grained network control makes network isolation difficult, increasing security risks.
[0008] In summary, since the emergence of intelligent computing centers, how to achieve efficient isolation of computing resources in Kubernetes clusters and improve the utilization, security, and isolation of computing resources has become an urgent technical problem to be solved. Summary of the Invention
[0009] This invention provides a method and apparatus for isolating computing resources in a cloud platform for intelligent computing centers, in order to solve the problem of how to achieve efficient isolation of computing resources in Kubernetes clusters and improve the utilization, security and isolation of computing resources since the emergence of intelligent computing centers.
[0010] To solve the above-mentioned technical problems, the present invention is implemented as follows:
[0011] In a first aspect, the present invention provides a method for isolating computing resources in an intelligent computing center cloud platform, the method comprising:
[0012] Step S1: On the Kubernetes physical cluster, create a dedicated VKS for each tenant. Each VKS achieves logical isolation of computing resources through its own independent namespace, and each VKS is equivalent to an isolated, virtual Kubernetes cluster.
[0013] Step S2: Receive the user's operation request for the VKS control plane via kubectl;
[0014] Step S3: The operation request is judged based on the pre-set controller and preset rules to obtain the judgment result; wherein, the controller is used to receive the instruction issued by the user through the intelligent computing center cloud platform and determine the preset rules according to the instruction;
[0015] Step S4: Based on the judgment result, allow or block the operation request;
[0016] Step S4 includes at least one of the following:
[0017] Step S41: If the judgment result indicates that the operation request is an unauthorized operation request, then the controller intercepts the operation request and returns an error response to the control plane of the VKS;
[0018] Step S42: If the judgment result indicates that the operation request passes the verification, then based on the controller, the operation request that passes the verification is forwarded to the synchronizer for conversion, and then synchronized to the control plane of the Kubernetes physical cluster to perform computing resource orchestration operations, wherein the controller is set between the control plane of the VKS and the synchronizer.
[0019] Optionally, step S4 further includes:
[0020] Step S43: If the operation request is a persistent volume declaration, determine whether the tag of the persistent volume declaration matches the tag of the corresponding VKS.
[0021] If not, the controller intercepts the operation request and returns an error response to the control plane of the corresponding VKS;
[0022] If so, the operation request is allowed, and the operation request is synchronized to the corresponding VKS.
[0023] Optionally, after step S2, the method further includes at least one of the following:
[0024] Step S5: Based on the controller, repair the operation request and schedule the computing power resource load generated by the repaired operation request to one or a group of specified computing power nodes in the Kubernetes physical cluster; wherein, repairing the operation request includes at least: adding computing power resources, adjusting the Pod resources according to a preset container template, and tagging the operation request;
[0025] Step S6: Tag the namespace of a specific VKS to increase the computing power resources of the specific VKS;
[0026] Step S7: Tag the namespace of a specific VKS to limit the computing resources of that specific VKS.
[0027] Optionally, the computing resource orchestration operation includes at least one of the following: creating, deleting, or scheduling Kubernetes resources.
[0028] Optionally, prior to step S1, the method further includes:
[0029] Step Sa: Allocate independent storage resources for each VKS through persistent volumes and persistent volume declaration mechanisms;
[0030] Each VKS's persistent volume declaration can only access persistent volumes bound to the namespace of the corresponding VKS, in order to achieve tenant-level isolation of the storage resources.
[0031] Furthermore, VKS, persistent volume declarations, and persistent volumes correspond one-to-one through the same tags. Persistent volume declarations, persistent volumes, and storage resources between different VKS are invisible to each other and cannot be accessed by each other.
[0032] Optionally, prior to step S1, the method further includes:
[0033] Step Sb: Configure network policies in each VKS namespace, wherein the network policies are used to restrict the communication rules of Pods, and the communication rules include at least one of the following:
[0034] Allows communication between specified Pods within the same VKS;
[0035] Access to unauthorized networks across VKS instances is prohibited;
[0036] Assign a dedicated IP network segment to each tenant to achieve network isolation across tenants.
[0037] Secondly, the present invention provides a computing resource isolation device for an intelligent computing center cloud platform, the device comprising:
[0038] Create a module to perform step S1: On the Kubernetes physical cluster, create a dedicated VKS for each tenant. Each VKS achieves logical isolation of computing resources through its own independent namespace, and each VKS is equivalent to an isolated, virtual Kubernetes cluster.
[0039] The execution module is used to perform step S2: receiving the user's operation request to the control plane of the VKS via kubectl;
[0040] Step S3: The operation request is judged based on the pre-set controller and preset rules to obtain the judgment result; wherein, the controller is used to receive the instruction issued by the user through the intelligent computing center cloud platform and determine the preset rules according to the instruction;
[0041] Step S4: Based on the judgment result, allow or block the operation request;
[0042] Step S4 includes at least one of the following:
[0043] Step S41: If the judgment result indicates that the operation request is an unauthorized operation request, then the controller intercepts the operation request and returns an error response to the control plane of the VKS;
[0044] Step S42: If the judgment result indicates that the operation request passes the verification, then based on the controller, the operation request that passes the verification is forwarded to the synchronizer for conversion, and then synchronized to the control plane of the Kubernetes physical cluster to perform computing resource orchestration operations, wherein the controller is set between the control plane of the VKS and the synchronizer.
[0045] Thirdly, the present invention provides a server, comprising: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the computing resource isolation method for an intelligent computing center cloud platform as described in the first aspect above.
[0046] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a computing resource isolation method for an intelligent computing center cloud platform as described in the first aspect above.
[0047] Fifthly, the present invention provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of a computing resource isolation method for an intelligent computing center cloud platform as described in the first aspect above.
[0048] In this invention, a controller is positioned between the control plane and the synchronizer of the VKS. The controller can perform permission verification on user operation requests based on preset rules, ensuring that only users who pass the verification can access the corresponding computing resources on the intelligent computing center cloud platform. At the same time, it prevents unauthorized operations, thereby providing strong security and computing resource management capabilities for the intelligent computing center cloud platform in a multi-tenant environment. It can realize logical isolation of computing resources based on the user dimension, thus achieving efficient computing resource isolation and effectively improving the utilization rate, security, and isolation of the computing resources of the intelligent computing center cloud platform. Attached Figure Description
[0049] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0050] Figure 1 A flowchart illustrating a method for isolating computing resources in an intelligent computing center cloud platform, provided by the present invention;
[0051] Figure 2 This invention provides a structural block diagram for creating a VKS on Kubernetes;
[0052] Figure 3 A structural block diagram of a computing resource isolation device for an intelligent computing center cloud platform provided by the present invention;
[0053] Figure 4 This is a schematic diagram of the electronic device of the present invention. Detailed Implementation
[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0055] The technical terms involved in this invention will be briefly explained below.
[0056] The “computing power” mentioned in this invention refers to: the ability of computer equipment or computing / data center to process information; the ability of computer hardware and software to work together to perform a certain computing requirement; the computing power to achieve the target result output by processing information data; and a new type of productivity that integrates information computing power, network carrying capacity, and data storage capacity, mainly providing services to society through computing power infrastructure.
[0057] The "computational power" (CP) described in this invention refers to the ability of a data center server to process data and output results. It is a comprehensive indicator of a data center's computing power, encompassing general computing power, supercomputing power, and intelligent computing power. The commonly used unit of measurement is floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), with higher values indicating stronger overall computing power. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe-2A supercomputers, 500,000 mainstream server CPUs, or 2 million mainstream laptops. The calculation formula is: CP = CP 通用 +CP 智能 +CP 超级
[0058] The "Network Power" (NP) mentioned in this invention refers to the performance of data transmission capability of computing facilities, which includes comprehensive capabilities such as network architecture, network bandwidth, transmission latency, intelligent management and scheduling, and involves network transmission within and between data centers. It is a comprehensive indicator for measuring network transmission scheduling capability.
[0059] The "Storage Power" (SP) described in this invention refers to the comprehensive capabilities of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon operation. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and internal storage devices in servers. The commonly used unit of measurement for storage capacity is exabytes (EB, 1EB = 2^60 bytes), the commonly used unit of measurement for performance is the number of read / write operations per second per unit capacity (IOPS / TB), and the disaster recovery ratio is an important indicator of security and reliability.
[0060] The "computing infrastructure" mentioned in this invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage capacity, enabling centralized computing, storage, transmission, and application of information.
[0061] The "new information infrastructure" mentioned in this invention refers to network infrastructure such as 5G networks, fiber optic broadband networks, backbone networks, international communication networks, and satellite internet; computing infrastructure such as data centers, general computing centers, intelligent computing centers, and supercomputing centers; and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0062] The “computing power” mentioned in this invention includes: “general computing power”, “intelligent computing power” and “supercomputing power”.
[0063] The "general computing power" mentioned in this invention refers to the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0064] The "intelligent computing power" mentioned in this invention refers to: a computing platform deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various artificial intelligence innovative applications, such as natural language processing and machine vision.
[0065] The “supercomputing power” mentioned in this invention refers to the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and uses a dedicated operating system to handle extremely complex or data-intensive problems. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, and gene analysis.
[0066] The "intelligent computing center" described in this invention refers to a facility that, through the use of large-scale heterogeneous computing resources, including general-purpose computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), primarily provides the necessary computing power, data, and algorithms for artificial intelligence applications (such as deep learning model development, model fine-tuning, and model inference). The intelligent computing center encompasses facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enablement.
[0067] The "intelligent computing center cloud platform" mentioned in this invention, abbreviated as "intelligent computing cloud", refers to a cloud computing platform that integrates hardware and software resources based on an intelligent computing center.
[0068] The "intelligent computing center" mentioned in this invention includes, but is not limited to, "smart computing center".
[0069] The "intelligent computing center" mentioned in this invention, also known as an artificial intelligence computing center, is a type of computing infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications, based on artificial intelligence theory and adopting an artificial intelligence computing architecture.
[0070] The "computing center" mentioned in this invention refers to a facility that is mainly composed of infrastructure such as wind, thermal, hydro, and electricity, and IT hardware and software equipment, and has computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0071] The "supercomputing center" mentioned in this invention refers to a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters. It can provide large-scale computing, storage and network services and is widely used in aerospace, defense, oil exploration, climate modeling and genome sequencing and other application scenarios.
[0072] The “computing resources” mentioned in this invention refer to the technologies and facilities required for the development of the digital society that have the ability to compute, transmit, store and apply information, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and supporting and guaranteeing resources such as wind, fire, water and electricity.
[0073] The “computing power running task” mentioned in this invention refers to a specific workload or job executed on computing power resources that requires a certain amount of computing power support, usually involving complex data processing, numerical calculation, model training or simulation scenarios.
[0074] The "computing node" mentioned in this invention refers to the computing resources of a server / container capable of processing computing tasks.
[0075] The "VKS (virtual Kubernetes service)" described in this invention refers to a technical architecture that creates a lightweight, isolated virtual Kubernetes cluster within an existing Kubernetes host cluster, often referred to as Kubernetes-in-Kubernetes (K8s-in-K8s). Its core objective is to provide a highly isolated virtual control plane, offering a near-native K8s experience, to different teams or tenants on a shared underlying infrastructure.
[0076] In this invention, "Kubernetes" refers to an open-source system for automatically deploying, scaling, and managing containerized applications. It is widely used for container orchestration, helping developers and operations personnel efficiently manage large numbers of containers.
[0077] The "Kubernetes resources" mentioned in this invention refer to persistent entities created, managed, and configured through the Kubernetes API. These are the basic units representing the cluster state in the Kubernetes system, such as Pods, Deployments, and StatefulSets. A Pod is the smallest deployable and manageable computing unit in Kubernetes. A Pod contains one or more tightly coupled containers sharing a network and storage; it is the basic unit for scheduling, running, and scaling. A Deployment is a controller that manages Pod replica sets and provides declarative updates. It is used to deploy stateless applications and supports rolling updates, rollbacks, and scaling. A StatefulSet is used to manage stateful applications (such as databases), providing stable network identifiers (hostnames), ordered deployment / scaling / deletion, and persistent storage.
[0078] The "kubectl" mentioned in this invention refers to the command-line tool provided by Kubernetes for interacting with the Kubernetes API server. This command-line tool enables the creation, querying, updating, and deletion of Kubernetes cluster resources (such as containers, services, and storage volumes) through declarative or imperative commands, and is the core client tool for managing Kubernetes clusters.
[0079] Figure 1 This illustrates a method for isolating computing resources in an intelligent computing center cloud platform according to the present invention, such as... Figure 1 As shown, the method includes:
[0080] Step S1: On the Kubernetes physical cluster, create a dedicated VKS for each tenant;
[0081] Each VKS achieves logical isolation of computing resources through its own independent namespace, and each VKS is equivalent to an isolated, virtual Kubernetes cluster.
[0082] Step S2: Receive user operation requests to the VKS control plane via kubectl;
[0083] Step S3: Based on the pre-set controller and preset rules, judge the operation request and obtain the judgment result;
[0084] The controller is used to receive instructions from users through the intelligent computing center cloud platform and determine preset rules based on the instructions.
[0085] Step S4: Based on the judgment result, allow or block the operation request.
[0086] It should be noted that step S4 includes at least one of the following: Step S41: If the judgment result indicates that the operation request is an unauthorized operation request, then the controller intercepts the operation request and returns an error response to the VKS control plane; Step S42: If the judgment result indicates that the operation request passes the verification, then the controller forwards the verified operation request to the synchronizer for conversion, and then synchronizes it to the control plane of the Kubernetes physical cluster to perform computing resource orchestration operations. The controller is set between the VKS control plane and the synchronizer.
[0087] It's important to note that, firstly, a dedicated VKS needs to be created for each tenant on the Kubernetes physical cluster. Each VKS is associated with an independent namespace, thus achieving logical isolation of computing resources between tenants. Furthermore, each VKS itself is equivalent to a complete, isolated, and virtual Kubernetes cluster, providing users with an independent Kubernetes experience environment. Users use the standard command-line tool kubectl to initiate various operation requests to their VKS through the VKS control plane. Then, each operation request is carefully evaluated based on pre-configured controllers and preset rules to obtain a clear judgment result. Next, based on the controller and the judgment result, a decisive action is taken on the operation request: either allow it or block it. Specifically, in step S4: if the judgment result indicates that the operation request is an unauthorized operation request, the controller blocks the operation request and returns an error response to the VKS control plane to notify the user; if the judgment result indicates that the operation request passes the verification, the controller forwards the verified operation request to the synchronizer. The synchronizer is responsible for translating operation requests at the VKS level into instructions that the Kubernetes physical cluster can understand. Ultimately, the control plane of the Kubernetes physical cluster executes the specific computing resource orchestration operations. Figure 1 The key to the method shown lies in setting up the controller (such as...) Figure 2 As shown in the controller (in the diagram), the controller is located between the control plane and the synchronizer of VKS. It is responsible for receiving user operation requests, performing policy judgments, and controlling the flow of operation requests (intercepting or forwarding them to the synchronizer).
[0088] and Figure 2In this context, User refers to a user, kubectl refers to a command-line tool, ns-1namespace refers to the ns-1 namespace, ns-2namespace refers to the ns-2 namespace, high-level resources refers to high-level resources, low-level resources refers to low-level resources, deployment refers to deployment, svc-a refers to service 'a', pod1 refers to container group '1', statefulset refers to a stateful set, syncer refers to a synchronizer, kube-system namespace refers to the Kubernetes system namespace, VKS control-plane refers to the VKS control plane, controller refers to the controller (also known as rule reviewer), regularnetworking refers to regular networking, and host-namespace refers to the host namespace.
[0089] Therefore, by setting up a controller between the VKS control plane and the synchronizer, the controller can perform permission verification on user operation requests based on preset rules, ensuring that only users who pass the verification can access the corresponding computing resources on the intelligent computing center cloud platform, while preventing unauthorized operations. In a multi-tenant environment, this provides the intelligent computing center cloud platform with strong security and computing resource management capabilities, enabling logical isolation of computing resources based on the user dimension. This achieves efficient computing resource isolation and effectively improves the utilization, security, and isolation of the intelligent computing center cloud platform's computing resources.
[0090] In one possible implementation, the computing resource orchestration operation includes at least one of the following: creating, deleting, or scheduling Kubernetes resources.
[0091] It's important to note that in the VKS implementation mechanism, compute resource orchestration is a core function of the Kubernetes physical cluster control plane. Essentially, it involves the full lifecycle management and physical resource scheduling and allocation of Kubernetes resources defined by tenants in VKS. When a user submits an operation request (such as creating an application or deleting a service) to the VKS control plane via kubectl, and this request is approved by the controller, it is transformed by the synchronizer and passed to the Kubernetes physical cluster control plane. At this point, the Kubernetes physical cluster control plane performs the actual compute resource orchestration operations: generating new Kubernetes resource entities within the specified namespace of the Kubernetes physical cluster; destroying Kubernetes resource instances specified by the tenant (such as deleting no longer needed Pods or StatefulSets); and determining the running location of Kubernetes resources on compute nodes (such as assigning newly created Pods to specific compute nodes with sufficient compute resources).
[0092] Thus, through the coordinated execution of three operations—creating, deleting, and scheduling Kubernetes resources—the control plane of the Kubernetes physical cluster translates user requests for VKS into actual management and control of computing resources. This not only ensures the accurate implementation of tenant operation requests within the underlying Kubernetes physical cluster but also guarantees the efficient utilization of computing resources and the isolation of computing resources between tenants through a fine-grained scheduling mechanism.
[0093] In one possible implementation, step S4 further includes:
[0094] Step S43: If the operation request is a Persistent Volume Claim (PVC), determine whether the label of the persistent volume claim matches the label of the corresponding VKS.
[0095] If not, the controller intercepts the operation request and returns an error response to the corresponding VKS control plane;
[0096] If so, the operation request is allowed, and the operation request is synchronized to the corresponding VKS.
[0097] It's important to note that in the VKS operating mechanism, when a user initiates an operation request for a persistent volume declaration via the command-line tool kubectl, the system executes a security verification process. The controller verifies whether the tags carried in the persistent volume declaration strictly match the tags bound to the VKS being operated on. The controller first intercepts the user's operation request, extracts the explicitly defined tag information from the persistent volume declaration, and obtains the exclusive tags preset during VKS initialization. Then, the system compares the two item by item to determine if the tags of the persistent volume declaration are completely consistent with the tags of the corresponding VKS. If the tag comparison result is a mismatch, the controller immediately intercepts the operation request, preventing its transmission to the Kubernetes physical cluster, and returns a clear error response to the control plane of the VKS being operated on. If the tag comparison result is a match, the controller fully allows the operation request, synchronizes it to the corresponding VKS, transforms it via the synchronizer, and submits it to the Kubernetes physical cluster for execution.
[0098] This enables hard isolation and security control of storage resources in a multi-tenant environment, while maintaining the native experience of tenants operating storage resources within VKS.
[0099] In one possible implementation, after step S2, the method further includes at least one of the following:
[0100] Step S5: Based on the controller patching operation request, schedule the computing power resource load generated by the patched operation request to one or a group of specified computing power nodes in the Kubernetes physical cluster; wherein, the patching operation request includes at least: adding computing power resources, adjusting the resources of Pods according to the preset container template, and tagging the operation request;
[0101] Step S6: Tag the namespace of a specific VKS to increase the computing power resources of that specific VKS;
[0102] Step S7: Tag the namespace of a specific VKS to limit the computing resources of that specific VKS.
[0103] It should be noted that, optionally, after a user's operation request (such as creating or scaling a workload) is permitted by the controller, the controller can proactively patch the operation request. Patching the operation request includes at least: adding computing resources, adjusting Pod resources according to a preset container template, and tagging the operation request. The patched operation request is then submitted to the synchronizer for transformation and finally processed by the Kubernetes physical cluster's control plane. The control plane can schedule the workload (such as a Pod) associated with the operation request to one or a group of specified computing nodes in the Kubernetes physical cluster. Tagging specific VKS namespaces is a core method for controlling computing resources. Specifically, tagging specific VKS namespaces can increase or limit their computing resources.
[0104] Therefore, based on the repair of operation requests, the resulting computing resource load can be scheduled to a designated computing node or computing node group to ensure hardware compatibility and physical isolation; it can also increase the upper limit of the total amount of computing resources available to tenants to support business expansion; and it can also reduce the consumption threshold of tenant computing resources to prevent excessive occupation of computing resources.
[0105] In one possible implementation, prior to step S1, the method further includes:
[0106] Step Sa: Allocate independent storage resources for each VKS through persistent volumes and persistent volume declaration mechanisms;
[0107] Each VKS persistent volume declaration can only access persistent volumes bound to the namespace of the corresponding VKS, so as to achieve isolation of storage resources at the tenant level.
[0108] Furthermore, VKS, persistent volume declarations, and persistent volumes correspond one-to-one through the same tags. Persistent volume declarations, persistent volumes, and storage resources between different VKS are invisible to each other and cannot be accessed by each other.
[0109] It's important to note that before deploying tenant-specific VKS, the solution pre-emptively ensures that each virtual cluster has its own independent storage resources through Kubernetes' built-in Persistent Volumes (PVs) and persistent volume declaration mechanisms. Specifically, the platform configures a dedicated persistent volume resource pool for each tenant's VKS namespace. When a tenant deploys an application requiring persistent storage within its VKS, the application requests the necessary storage resources through persistent volume declarations. The Kubernetes system ensures that these persistent volume declarations can only allocate and bind storage resources from persistent volumes bound to the same tenant's VKS namespace. In this way, persistent volume declarations issued by different tenants' VKSs can only access persistent volume resources within their respective namespaces. Persistent volume declarations, persistent volumes, and storage resources between different VKSs are invisible and inaccessible to each other, thus achieving tenant-level isolation of storage resources. For example, if tenant A runs the `kubectl get pvc` command in its VKS, it can only query and declare persistent volumes belonging to its own namespace.
[0110] This ensures that in a multi-tenant environment sharing the same Kubernetes physical cluster, each tenant's VKS not only enjoys logical isolation of computing resources, but also strict isolation of the persistent storage resources required by its applications. Each tenant has its own dedicated and isolated storage space within its VKS namespace, which fully guarantees data storage security and the independence between tenants.
[0111] In one possible implementation, prior to step S1, the method further includes:
[0112] Step Sb: Configure network policies in each VKS namespace, where network policies are used to restrict the communication rules of Pods, and the communication rules include at least one of the following:
[0113] Allows communication between specified Pods within the same VKS;
[0114] Access to unauthorized networks across VKS instances is prohibited;
[0115] Assign a dedicated IP network segment to each tenant to achieve network isolation across tenants.
[0116] It's important to note that before creating a dedicated VKS for each tenant, network policies can be further configured within the namespace of each VKS. Different isolation policies and methods can be set through configuration. This network policy is the core network control mechanism of Kubernetes, used to finely restrict Pod communication rules and provide network-level security isolation for tenant applications. Communication rules include at least one of the following: allowing communication between specified Pods within the same VKS to ensure that application components within the cluster can interconnect on demand; prohibiting unauthorized network access across VKSs; and deeper isolation measures, such as assigning a dedicated IP address range to each tenant, ensuring that Pods from different tenants run within completely isolated IP address ranges. For example, using different IP address ranges to allocate access policies, achieving isolation at layers 4 and 7 (layer 4 is the transport layer, layer 7 is the application layer). Through the synergistic effect of these rules, communication rules for Pods can be restricted on top of the shared physical network infrastructure to achieve a mandatory network isolation effect, realizing network isolation between different tenant virtual environments and eliminating possible pathways for unauthorized network access.
[0117] Therefore, at the network layer, a strict security boundary can be established for each tenant's VKS. Network isolation ensures that tenant application Pods can only communicate within the authorized scope (within the same VKS), and by default, they cannot make unauthorized network connections with other tenants' VKS. In addition, combined with the use of dedicated IP network segments, cross-tenant network isolation is achieved at the network layer, making each tenant run as if in an independent network environment, which greatly improves the security and isolation of the multi-tenant environment.
[0118] In summary, multiple tenants can securely share the same Kubernetes physical cluster. Each tenant experiences what appears to be an independent, isolated, and virtual Kubernetes cluster, providing a dedicated and isolated experience. Furthermore, the controller intercepts any illegal access or operation requests, ensuring that only user requests with strict permission verification are ultimately executed, effectively guaranteeing the security and compliance of computing resource isolation. The underlying resource synchronization is efficiently completed by the Kubernetes physical cluster after conversion via a synchronizer. This provides robust security and computing resource management capabilities for the intelligent computing center cloud platform in a multi-tenant environment, enabling logical isolation of computing resources based on the user dimension. This achieves efficient computing resource isolation, effectively improving the utilization, security, and isolation of the intelligent computing center cloud platform's computing resources.
[0119] Figure 3 This invention illustrates a computing resource isolation device for an intelligent computing center cloud platform, such as... Figure 3 As shown, device 30 includes:
[0120] Create module 301 to execute step S1: On the Kubernetes physical cluster, create a dedicated VKS for each tenant. Each VKS achieves logical isolation of computing resources through its own independent namespace, and each VKS is equivalent to an isolated, virtual Kubernetes cluster.
[0121] Execution module 302 is used to execute step S2: receiving user operation requests to the VKS control plane via kubectl;
[0122] Step S3: Based on the pre-set controller and preset rules, the operation request is judged to obtain the judgment result; wherein, the controller is used to receive the instructions issued by the user through the intelligent computing center cloud platform and determine the preset rules according to the instructions;
[0123] Step S4: Based on the judgment result, allow or block the operation request;
[0124] Step S4 includes at least one of the following:
[0125] Step S41: If the judgment result indicates that the operation request is an unauthorized operation request, then the controller intercepts the operation request and returns an error response to the VKS control plane;
[0126] Step S42: If the judgment result indicates that the operation request passes the verification, then based on the controller, the operation request that passes the verification is forwarded to the synchronizer for conversion, and then synchronized to the control plane of the Kubernetes physical cluster to perform computing resource orchestration operations. The controller is set between the control plane of VKS and the synchronizer.
[0127] In one possible implementation, step S4 further includes:
[0128] Step S43: If the operation request is a persistent volume declaration, determine whether the label of the persistent volume declaration matches the label of the corresponding VKS.
[0129] If not, the controller intercepts the operation request and returns an error response to the corresponding VKS control plane;
[0130] If so, the operation request is allowed, and the operation request is synchronized to the corresponding VKS.
[0131] In one possible implementation, after step S2, the method further includes at least one of the following:
[0132] Step S5: Based on the controller patching operation request, schedule the computing power resource load generated by the patched operation request to one or a group of specified computing power nodes in the Kubernetes physical cluster; wherein, the patching operation request includes at least: adding computing power resources, adjusting the resources of Pods according to the preset container template, and tagging the operation request;
[0133] Step S6: Tag the namespace of a specific VKS to increase the computing power resources of that specific VKS;
[0134] Step S7: Tag the namespace of a specific VKS to limit the computing resources of that specific VKS.
[0135] In one possible implementation, the computing resource orchestration operation includes at least one of the following: creating, deleting, or scheduling Kubernetes resources.
[0136] In one possible implementation, prior to step S1, the method further includes:
[0137] Step Sa: Allocate independent storage resources for each VKS through persistent volumes and persistent volume declaration mechanisms;
[0138] Each VKS persistent volume declaration can only access persistent volumes bound to the namespace of the corresponding VKS, so as to achieve isolation of storage resources at the tenant level.
[0139] Furthermore, VKS, persistent volume declarations, and persistent volumes correspond one-to-one through the same tags. Persistent volume declarations, persistent volumes, and storage resources between different VKS are invisible to each other and cannot be accessed by each other.
[0140] In one possible implementation, prior to step S1, the method further includes:
[0141] Step Sb: Configure network policies in each VKS namespace, where network policies are used to restrict the communication rules of Pods, and the communication rules include at least one of the following:
[0142] Allows communication between specified Pods within the same VKS;
[0143] Access to unauthorized networks across VKS instances is prohibited;
[0144] Assign a dedicated IP network segment to each tenant to achieve network isolation across tenants.
[0145] In this invention, a controller is positioned between the control plane and the synchronizer of the VKS. The controller can perform permission verification on user operation requests based on preset rules, ensuring that only users who pass the verification can access the corresponding computing resources on the intelligent computing center cloud platform. At the same time, it prevents unauthorized operations, thereby providing strong security and computing resource management capabilities for the intelligent computing center cloud platform in a multi-tenant environment. It can realize logical isolation of computing resources based on the user dimension, thus achieving efficient computing resource isolation and effectively improving the utilization rate, security, and isolation of the computing resources of the intelligent computing center cloud platform.
[0146] Please refer to Figure 4 The present invention also provides an electronic device 40, including a processor 401, a memory 402, and a computer program stored in the memory 402 and executable on the processor 401. When the computer program is executed by the processor 401, it implements the steps of the above-mentioned intelligent computing center cloud platform computing resource isolation method and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0147] This invention also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the steps of the aforementioned method for isolating computing resources in an intelligent computing center cloud platform, achieving the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0148] The present invention also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the above-described method for isolating computing resources in an intelligent computing center cloud platform, and achieve the same technical effect. To avoid repetition, the details will not be repeated here.
[0149] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0150] Through the above description of the embodiments, those skilled in the art can clearly understand that the above methods can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the present invention.
[0151] The present invention has been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other modifications under the guidance of the present invention without departing from the spirit and scope of the claims, and all such modifications are within the protection scope of the present invention.
Claims
1. A method for isolating computing resources in an intelligent computing center cloud platform, characterized in that, The method includes: Step S1: On the Kubernetes physical cluster, create a dedicated VKS for each tenant. Each VKS achieves logical isolation of computing resources through its own independent namespace, and each VKS is equivalent to an isolated, virtual Kubernetes cluster. Step S2: Receive the user's operation request for the VKS control plane via kubectl; Step S3: The operation request is judged based on the pre-set controller and preset rules to obtain the judgment result; wherein, the controller is used to receive the instruction issued by the user through the intelligent computing center cloud platform and determine the preset rules according to the instruction; Step S4: Based on the judgment result, allow or block the operation request; Step S4 includes at least one of the following: Step S41: If the judgment result indicates that the operation request is an unauthorized operation request, then the controller intercepts the operation request and returns an error response to the control plane of the VKS; Step S42: If the judgment result indicates that the operation request passes the verification, then based on the controller, the operation request that passes the verification is forwarded to the synchronizer for conversion, and then synchronized to the control plane of the Kubernetes physical cluster to perform computing resource orchestration operations, wherein the controller is set between the control plane of the VKS and the synchronizer.
2. The method according to claim 1, characterized in that, Step S4 further includes: Step S43: If the operation request is a persistent volume declaration, determine whether the tag of the persistent volume declaration matches the tag of the corresponding VKS. If not, the controller intercepts the operation request and returns an error response to the control plane of the corresponding VKS; If so, the operation request is allowed, and the operation request is synchronized to the corresponding VKS.
3. The method according to claim 1, characterized in that, Following step S2, the method further includes at least one of the following: Step S5: Based on the controller, repair the operation request and schedule the computing power resource load generated by the repaired operation request to one or a group of specified computing power nodes in the Kubernetes physical cluster; wherein, repairing the operation request includes at least: adding computing power resources, adjusting the Pod resources according to a preset container template, and tagging the operation request; Step S6: Tag the namespace of a specific VKS to increase the computing power resources of the specific VKS; Step S7: Tag the namespace of a specific VKS to limit the computing resources of that specific VKS.
4. The method according to any one of claims 1-3, characterized in that, The computing resource orchestration operations include at least one of the following: creating, deleting, or scheduling Kubernetes resources.
5. The method according to claim 1, characterized in that, Prior to step S1, the method further includes: Step Sa: Allocate independent storage resources for each VKS through persistent volumes and persistent volume declaration mechanisms; Each VKS's persistent volume declaration can only access persistent volumes bound to the namespace of the corresponding VKS, in order to achieve tenant-level isolation of the storage resources. Furthermore, VKS, persistent volume declarations, and persistent volumes correspond one-to-one through the same tags. Persistent volume declarations, persistent volumes, and storage resources between different VKS are invisible to each other and cannot be accessed by each other.
6. The method according to claim 1, characterized in that, Prior to step S1, the method further includes: Step Sb: Configure network policies in each VKS namespace, wherein the network policies are used to restrict the communication rules of Pods, and the communication rules include at least one of the following: Allows communication between specified Pods within the same VKS; Access to unauthorized networks across VKS instances is prohibited; Assign a dedicated IP network segment to each tenant to achieve network isolation across tenants.
7. A computing resource isolation device for an intelligent computing center cloud platform, characterized in that, The device includes: Create a module to perform step S1: On the Kubernetes physical cluster, create a dedicated VKS for each tenant. Each VKS achieves logical isolation of computing resources through its own independent namespace, and each VKS is equivalent to an isolated, virtual Kubernetes cluster. The execution module is used to perform step S2: receiving the user's operation request to the control plane of the VKS via kubectl; Step S3: The operation request is judged based on the pre-set controller and preset rules to obtain the judgment result; wherein, the controller is used to receive the instruction issued by the user through the intelligent computing center cloud platform and determine the preset rules according to the instruction; Step S4: Based on the judgment result, allow or block the operation request; Step S4 includes at least one of the following: Step S41: If the judgment result indicates that the operation request is an unauthorized operation request, then the controller intercepts the operation request and returns an error response to the control plane of the VKS; Step S42: If the judgment result indicates that the operation request passes the verification, then based on the controller, the operation request that passes the verification is forwarded to the synchronizer for conversion, and then synchronized to the control plane of the Kubernetes physical cluster to perform computing resource orchestration operations, wherein the controller is set between the control plane of the VKS and the synchronizer.
8. A server, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of a computing resource isolation method for an intelligent computing center cloud platform as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a computing resource isolation method for an intelligent computing center cloud platform as described in any one of claims 1-6.
10. A computer program product, characterized in that, It includes computer instructions, which, when executed by a processor, implement the steps of a computing resource isolation method for an intelligent computing center cloud platform as described in any one of claims 1-6.