Switching scheduling method and system for containerized computing environment
By using containerization technology, persistent storage and CRIU technology in heterogeneous computing power environments, the efficiency and consistency of heterogeneous computing power resource switching is solved, efficient resource allocation and task migration is achieved, and the flexibility and resource utilization of the system are improved.
Patent Information
- Application Number
- CN202510277093.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-10
AI Technical Summary
In the heterogeneous computing power environment, it is difficult for the prior art to switch between different computing power resources quickly and efficiently, resulting in computing interruptions, delays and data loss, and resource allocation lacks real-time and dynamic prediction capabilities.
Containerization technology is used to design standardized mirror templates, combining persistent storage and CRIU technology to achieve consistency and availability of model data, task migration and state pre-allocation and scheduling computing resources.
It improves the flexible allocation of resources, reduces computational interrupts and delays, ensures data consistency and availability, and improves resource utilization and task execution efficiency.
Smart Images

Figure CN119759599B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a switching scheduling method and system for a containerized computing power environment, in particular to a method and system for resource switching and scheduling between CPUs and GPUs in a containerized heterogeneous computing power environment, which is suitable for efficient dynamic scheduling and operation of artificial intelligence tasks and belongs to the field of cloud computing and containerization technology. Background Art
[0002] With the rapid development of artificial intelligence and deep learning, the computing demand for model training, reasoning, and development is growing, especially in large-scale model training and real-time reasoning tasks, where the demand for computing resources is particularly prominent. Traditional computing resources usually rely on homogeneous hardware environments (such as pure GPU clusters), but with the diversification of computing tasks and the popularization of cloud native technologies, heterogeneous computing environments (such as the mixed use of CPU, GPU, TPU, FPGA, etc.) have gradually become mainstream. This heterogeneous computing environment can not only provide more diverse computing resources, but also improve resource utilization and save hardware costs.
[0003] However, the switching and scheduling of heterogeneous computing resources has always been a difficult point in system design. How to quickly and efficiently switch between different computing resources during task execution or model training to reduce computing interruptions and delays has become a key issue. In a heterogeneous computing environment, technologies such as image switching, environment switching, and resource switching need further optimization. Current research has not fully solved the problem of fast switching between different computing power specifications. For example, switching to a new computing environment requires reconfiguring the operating environment and loading data, resulting in significant delays; secondly, the model state cannot be saved and migrated during the switching process, which may cause computing interruptions or even data loss; and resource allocation lacks real-time and dynamic prediction capabilities, resulting in slow resource switching response. Existing technologies still have a lot of room for improvement in resource switching efficiency and consistency. Summary of the invention
[0004] In view of the deficiencies in the prior art, the present invention provides a switching scheduling method and system for a containerized computing environment, which can effectively enhance the flexible allocation capability of resources, adapt to different task requirements of instances, and improve switching efficiency.
[0005] The present invention adopts the following technical solution:
[0006] On the one hand, the present invention provides a method for switching and scheduling a containerized computing environment, comprising the following steps:
[0007] S1: Before switching resources in a heterogeneous computing environment, we use containerization technology to design and create standardized image templates that adapt to different resources to speed up container startup and configuration.
[0008] S2: During the resource switching process in a heterogeneous computing environment, the data generated during the entire process of model design, model training, and model reasoning is ensured to be consistent and available during the switching process by using persistent storage technology.
[0009] S3: In the process of switching resources in heterogeneous computing environments, the process state preservation and recovery mechanism of CRIU technology is combined to achieve task migration and reduce computing interruptions caused by environment switching;
[0010] S4: In the process of resource scheduling, by designing a resource pre-allocation strategy, it is responsible for dynamically predicting and reasonably scheduling the computing resources required for the task according to the priority of the task, and finally completing the resource switching and scheduling in a heterogeneous computing environment.
[0011] Preferably, in step S1, the standardized image template includes a basic CPU image, a Pytorch CPU image, and a basic GPU image, wherein the basic CPU image contains the most basic Python development environment, as well as data analysis and machine learning toolkits, which are suitable for applications that require less computing resources;
[0012] The PytorchCPU image includes the Python development environment and toolkit in the base CPU image, as well as the CPU-based Pytorch image, which includes the Pytorch framework for deep learning development.
[0013] The basic GPU image includes the basic environment with GPU acceleration support (such as CUDA, cuDNN, etc.), which is suitable for tasks that require high-performance computing;
[0014] The integrated development environment is built into the standardized image templates to support seamless code writing and debugging.
[0015] Preferably, the integrated development environment is Jupyter Notebook.
[0016] Preferably, step S2 comprises:
[0017] S21: Designing a persistent storage architecture
[0018] In heterogeneous computing clusters, we build a persistent storage system based on Kubernetes. We design various persistent volume claim (PVC) templates for different types of computing tasks, such as different types of models. The persistent volume claim template predefines the storage capacity, access mode, and storage category of the persistent volume claim. Users can select the appropriate persistent volume claim template to create a persistent volume claim based on task requirements.
[0019] The cluster administrator configures various types of persistent storage volumes. Persistent storage volumes are based on different storage backend technologies. Through the dynamic provisioning mechanism, the persistent storage volume claim (PVC) is automatically bound to the appropriate persistent storage volume (PV) to ensure that the task has storage resources that meet its needs.
[0020] S22: Data Mounting and Container Integration
[0021] In the configuration file of the task container, explicitly specify to mount the corresponding persistent storage volume declaration (PVC). When the container is started on a heterogeneous node, whether it is a CPU node or a GPU node, Kubernetes automatically mounts the storage volume corresponding to the persistent storage volume declaration (PVC) to the specified directory in the container according to the configuration of the task container;
[0022] S23: The container can store the model's training data, checkpoints, and parameters in persistent storage. Through the persistent storage volume declaration (PVC) technology provided by Kubernetes, the data will be stored on the persistent volume and will not be affected by the container lifecycle.
[0023] Preferably, in step S23, storing the data on a persistent volume means that: during the model training process, the model checkpoints are periodically saved to the mounted / data path (i.e., persistent storage) to ensure that the data will not be lost due to container migration or resource switching;
[0024] When resources are switched, containers can seamlessly access mounted storage. When the container switches from a CPU node to a GPU node, the mounted persistent storage remains unchanged, and the container continues to access the same data storage. Even if computing resources are switched, the data is always in the same storage location and does not need to be reloaded.
[0025] Preferably, in step S3, during the resource switching process, the state of the running model task can be saved (Checkpoint) through CRIU technology, including memory status, file descriptors, register contents, etc., and restored (Restore) when switching to the target device. This method can avoid computing interruptions caused by task restarts and achieve seamless switching from the original container to the new container, specifically:
[0026] For each task T, the current execution state S(T) contains information such as memory snapshot, process status, network connection, etc. Through CRIU technology, the task state is serialized into a checkpoint file C(T):
[0027] C(T) = checkpoint(T);
[0028] Where checkpoint() represents the serialization process of the checkpoint file;
[0029] On the target device, the task is continued through the recovery mechanism:
[0030] T′=restore(C(T));
[0031] Here, restore ( ) indicates the operation process of restoring the task status from the checkpoint file.
[0032] Preferably, the implementation process of step S4 is:
[0033] S41: Define system resource status as input to resource demand assessment, including current CPU utilization and current memory utilization ;
[0034] S42: Periodically collect the current CPU utilization and current memory utilization of system resources through the resource monitoring module;
[0035] S43: define a pre-allocation strategy, i.e., a decision function based on resource utilization, for determining whether switching between heterogeneous resources (especially switching from CPU to GPU) is required and whether resource pre-allocation is required;
[0036] S44: Sort the task priorities and allocate resources.
[0037] Preferably, in step S43, the decision function is set to ,when =1, indicating that resource switching or pre-allocation is required; =0, indicating that no resource switching or pre-allocation is required, the expression is:
[0038] ;
[0039] in Indicates A task is 1 or 0, , are weight coefficients, which respectively represent the influence of CPU utilization and memory utilization on the switching decision, satisfying ,These weight coefficients can be adjusted according to the system’s reliance on different resources and the ,application scenarios; , They are the thresholds of CPU utilization and memory utilization, respectively.
[0040] Preferably, in step S44, when multiple tasks request resources at the same time, the task requesting resource switching is first stored in a queue, and then the tasks in the queue are prioritized to determine the order of resource allocation; when the priorities of the tasks are the same, they are prioritized according to the order in which they enter the queue.
[0041] The task priority index P is designed to measure the priority of the task. Its calculation is based on the current node resource utilization obtained by the resource monitoring module. The formula is:
[0042] ;
[0043] Indicates The priority index of a task. The higher the priority index, the more urgent the task's demand for resources is, and resources should be allocated first.
[0044] On the other hand, the present invention provides a switching scheduling system for a containerized computing environment, which is used to implement the switching scheduling method for the containerized computing environment, including:
[0045] The containerized environment construction and initialization module is configured to: before switching resources in a heterogeneous computing environment, accelerate the startup and configuration of containers by designing and producing standardized image templates that adapt to different resources in combination with containerization technology;
[0046] The persistent storage module is configured to: during the switching of resources in a heterogeneous computing environment, ensure the consistency and availability of data generated during the entire process of model design, model training, and model reasoning by using persistent storage technology;
[0047] The task state preservation and interruption recovery module is configured to: in the process of switching resources in a heterogeneous computing environment, combine the process state preservation and recovery mechanism of CRIU technology to achieve task migration and reduce computing interruptions caused by environment switching;
[0048] The resource pre-allocation and scheduling decision module is configured as follows: in the process of resource scheduling, by designing a resource pre-allocation strategy, it is responsible for dynamically predicting and reasonably scheduling the computing resources required for the task according to the priority of the task, and finally completing resource switching and scheduling in a heterogeneous computing environment.
[0049] For any details not provided in the present invention, please refer to the prior art.
[0050] The beneficial effects of the present invention are:
[0051] In terms of container management, standardized image templates greatly shorten the time it takes to start and configure containers, greatly improving efficiency compared to the traditional method of building an environment from scratch; containerized isolation and portability reduce task conflicts, simplify operations and maintenance, enhance system flexibility and scalability, and effectively solve problems such as poor compatibility of traditional environments.
[0052] In terms of data and task management, the reliable system built with persistent storage technology ensures the consistency and availability of data during heterogeneous switching, overcoming the drawbacks of traditional data storage; CRIU technology enables efficient task migration, completely saves and restores key process states, avoids restart and recalculation losses after task interruption, and ensures task continuity.
[0053] In terms of resource utilization, the heterogeneous resource pre-allocation strategy uses precise resource monitoring and intelligent decision-making functions to achieve early dynamic prediction of the computing power requirements of model tasks. When faced with a large number of switching demands, it rationally schedules and allocates resources based on the task priority index to avoid idle resources and improve overall utilization. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The drawings in the specification, which constitute a part of the present application, are used to provide further understanding of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.
[0055] Figure 1 This is a flow chart of the switching scheduling method of the containerized computing environment of the present invention;
[0056] Figure 2 It is a schematic diagram of the overall framework of switching scheduling of the containerized computing environment of the present invention;
[0057] Figure 3 This is the flowchart of task migration and recovery based on CRIU;
[0058] Figure 4 Decision flow chart for dynamic resource switching. DETAILED DESCRIPTION
[0059] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of the present invention are clearly and completely described below in conjunction with the drawings in the implementation of this specification, but are not limited to this. Anything not fully described in the present invention shall be based on the conventional technology in the art.
[0060] Terminology explanation:
[0061] 1. Heterogeneous computing environment: refers to a computing system that contains multiple different types of computing resources. These resources differ in architecture, performance characteristics, instruction sets, etc., and work together to process computing tasks. In this embodiment, it specifically refers to the CPU-GPU heterogeneous computing environment.
[0062] 2. Containerization technology: Containerization technology is a lightweight operating system-level virtualization method that packages software and its dependencies into independent and portable containers. Container images are built in a layered storage manner and contain all the elements for running applications. The container runtime is responsible for creating and managing the container lifecycle. Orchestration tools such as Kubernetes can automatically deploy and schedule containers on multiple hosts. This technology has the advantages of efficient resource utilization, strong application portability, and good environmental isolation. It can greatly improve the efficiency of software deployment and operation and maintenance and promote the development of fields such as cloud computing.
[0063] 3. Persistent storage: Persistent storage is a storage method used in computer systems to preserve data for a long time. In cloud computing and container orchestration (such as Kubernetes), persistent storage is implemented through the mechanisms of volume and volume claim. Volumes are configured by cluster administrators and provide physical storage resources based on different storage backend technologies. Volume claims are users requesting appropriate storage resources based on task requirements. The two are automatically bound, allowing containers to mount these storage resources to store important information, ensuring that data is not affected by the container life cycle and ensuring data consistency and availability between different nodes.
[0064] 4. CRIU technology: Checkpoint / Restore In Userspace is a user space process checkpoint preservation and recovery technology that can capture and serialize the memory, file descriptors, registers and other states of the running process. In scenarios such as heterogeneous resource switching, the process can be restored from the checkpoint to continue execution to avoid computing interruptions. It can be combined with containers and orchestration systems to enhance the flexibility and stability of task scheduling in complex computing environments.
[0065] Example 1
[0066] A switching scheduling method for a containerized computing environment, such as Figure 1 As shown, the following steps are included:
[0067] S1: Before switching resources in a heterogeneous computing environment, we use containerization technology to design and create standardized image templates that adapt to different resources to speed up container startup and configuration.
[0068] S2: During the resource switching process in a heterogeneous computing environment, the data generated during the entire process of model design, model training, and model reasoning is ensured to be consistent and available during the switching process by using persistent storage technology.
[0069] S3: In the process of switching resources in heterogeneous computing environments, the process state preservation and recovery mechanism of CRIU technology is combined to achieve task migration and reduce computing interruptions caused by environment switching;
[0070] S4: In the process of resource scheduling, by designing the resource pre-allocation strategy, it is responsible for dynamically predicting and reasonably scheduling the computing resources required for the task according to the priority of the task, and finally completing the resource switching and scheduling in the heterogeneous computing environment.
[0071] Example 2
[0072] A method for switching and scheduling a containerized computing environment, as shown in Example 1, except that, in step S1, the Docker technology is used to encapsulate the commonly used data analysis and machine learning toolkit operating environment required for processes such as model design, model training and model reasoning into a portable container image, and the standardized image template includes a basic CPU image, a Pytorch CPU image and a basic GPU image. Exemplarily, different basic CPU images include the most basic Python development environment (Python 2.x, Python 3.x), as well as some commonly used data analysis and machine learning toolkits, such as NumPy, Pandas, Scikit-learn, etc., which are suitable for applications that require less computing resources, such as model design, data preprocessing and preliminary model training;
[0073] The PytorchCPU image in different standardized image templates includes the Python development environment and toolkit in the basic CPU image, as well as the CPU-based Pytorch image, which includes the Pytorch framework for deep learning development (optimized for the CPU version);
[0074] The basic GPU image in different standardized image templates includes a basic environment with GPU acceleration support (such as CUDA, cuDNN, etc.), which is suitable for tasks that require high-performance computing, such as large-scale deep learning training and model reasoning. Configure deep learning frameworks compatible with GPU hardware, such as Pytorch (GPU version) and TensorFlow (GPU version), and support distributed training (for example, multi-card parallel training through tools such as Horovod). Install drivers, libraries, and tools related to GPU hardware (such as NVIDIA GPU drivers, CUDA Toolkit, cuDNN, TensorRT, etc.) to maximize the utilization of GPU computing resources.
[0075] Integrated development environments (IDEs), such as Jupyter Notebook, are built into standardized image templates to support seamless code writing and debugging. By running code in Jupyter Notebook, developers can efficiently design, train, and reason about models, and track and tune experiments directly in the container.
[0076] Jupyter Notebook not only supports Python, but also supports other languages such as R, Julia, etc. through the kernel, enhancing multi-language support and flexibility; it provides an interactive data analysis environment, allowing developers to view and modify training data, model parameters, and model prediction results in real time within the container, thereby accelerating the development and experimental process.
[0077] Example 3
[0078] A method for switching and scheduling a containerized computing environment, as shown in Example 2, except that step S2 includes:
[0079] S21: Designing a persistent storage architecture
[0080] In heterogeneous computing clusters, we build a persistent storage system based on Kubernetes. We design various persistent volume claim (PVC) templates for different types of computing tasks, such as different types of models. The persistent volume claim template predefines the storage capacity, access mode (such as read-only, read-write, multi-read multi-write) and storage category (such as general-purpose SSD storage and capacity storage properties) of the persistent volume claim. Users can select the appropriate persistent volume claim template to create a persistent volume claim based on task requirements.
[0081] Cluster administrators configure various types of persistent volumes (PVs). Based on different storage backend technologies, the persistent volume declaration (PVC) is automatically bound to the appropriate persistent volume (PV) through the dynamic provisioning mechanism to ensure that the task has storage resources that meet its needs. For example, for large-scale deep learning model training tasks, you can bind to a persistent volume (PV) stored in a general-purpose SSD to meet high read and write speeds and large-capacity storage requirements; for small inference tasks, you can use the PV of a capacity storage device to reduce costs and increase data access speed.
[0082] S22: Data Mounting and Container Integration
[0083] In the configuration file of the task container, explicitly specify to mount the corresponding persistent storage volume claim (PVC). When specifying to mount the corresponding PVC in Kubernetes, it is necessary to define the volumeMounts and volumes fields in the Pod configuration file, and specify the name of the PVC to be used through claimName. Finally, a variety of PVC templates are formed. Users can flexibly select the appropriate PVC to mount to the specified directory in the container according to the specific needs of the task.
[0084] When a container is started on a heterogeneous node, whether it is a CPU node or a GPU node, Kubernetes automatically mounts the storage volume corresponding to the persistent storage volume declaration (PVC) to the specified directory in the container according to the configuration of the task container. This process is transparent to the user. The task process does not need to care about the actual storage location of the data, and can access the data simply by mounting the directory. For example, in a model training container based on TensorFlow, the PVC is mounted to the / data directory. The training script can directly read the training data set, load model parameters, and save checkpoints from the / data directory to ensure consistency in data access when the task runs on different nodes.
[0085] S23: During the task switching process, it is crucial to ensure data consistency. To this end, the container can store the model's training data, checkpoints, and parameters in persistent storage. Through the persistent storage volume declaration (PVC) technology provided by Kubernetes, the data will be stored on the persistent volume and will not be affected by the container lifecycle.
[0086] In step S23, data will be stored on the persistent volume, which means that during the model training process, the model checkpoints are regularly saved to the mounted / data path (i.e., persistent storage) to ensure that the data will not be lost due to container migration or resource switching;
[0087] When resources are switched, containers can seamlessly access mounted storage. When the container switches from a CPU node to a GPU node, the mounted persistent storage remains unchanged, and the container continues to access the same data storage. Even if computing resources are switched, the data is always in the same storage location and does not need to be reloaded.
[0088] Example 4
[0089] A method for switching and scheduling a containerized computing environment, as shown in Example 3, is different in that in step S3, during the resource switching process, the state of the running model task can be saved (Checkpoint) through the CRIU technology, including memory status, file descriptors, register contents, etc., and restored (Restore) when switching to the target device. This method can avoid computing interruptions caused by task restarts and achieve seamless switching from the original container to the new container. The task migration and recovery flowchart based on CRIU is shown in the figure. Figure 3 As shown, specifically:
[0090] For each task T, the current execution state S(T) contains information such as memory snapshot, process status, network connection, etc. Through CRIU technology, the task state is serialized into a checkpoint file C(T):
[0091] C(T) = checkpoint(T);
[0092] Where checkpoint() represents the serialization process of the checkpoint file;
[0093] On the target device, the task is continued through the recovery mechanism:
[0094] T′=restore(C(T));
[0095] Here, restore ( ) indicates the operation process of restoring the task status from the checkpoint file.
[0096] Through the checkpoint operation, the status information of task T can be processed and converted into a storable checkpoint file C(T), that is, C(T) = checkpoint(T). This process is like recording the "appearance" of the task at a certain moment and saving it to a file, so that the task can be restored to its previous state based on this file at an appropriate time.
[0097] The core function of restore is to reload and apply the task status information previously saved to the checkpoint file through the checkpoint operation to the target environment, so that the task can continue to execute from the previous paused position. This process can effectively avoid computing interruptions caused by task restarts and achieve seamless switching from the original container to the new container and from one computing resource (such as CPU) to another computing resource (such as GPU).
[0098] In the process of implementing the above technologies, it is necessary to identify the core states that really need to be saved and restored in the task, such as model parameters and optimizer states in deep learning tasks, avoid saving unnecessary system-level states, and reduce the workload of saving and restoring. For example, when performing a checkpoint operation, only the memory data directly related to model training is saved, while ignoring some temporary file descriptors and other information. For some difficult-to-handle resources, such as the complex state of specific hardware (such as the fine state of GPU video memory), they can be excluded when saving the state and processed by reinitialization when restoring. In some appropriate cases, the checkpoint saving and recovery functions provided by PyTorch or TensorFlow are used to save core information such as model parameters and optimizer states. In this way, the core idea of CRIU technology can be reasonably utilized instead of using CRIU technology itself, which can better improve the switching efficiency of models in containerized computing environment switching.
[0099] How to integrate CRIU (Checkpoint / Restore In Userspace) technology to implement the checkpoint saving and recovery mechanism of process status in the containerized runtime environment, so as to avoid task interruption during resource switching. Specifically, the integration of CRIU technology mainly involves the following steps:
[0100] Process status saving (Checkpoint): Save the memory, registers, file descriptors and other states of the current container process to disk.
[0101] Process Restore: In the target resource environment, use CRIU to restore the saved process state so that the task can continue from the point where it was suspended.
[0102] This method combines CRIU with Kubernetes to ensure that containers maintain computing consistency and continuity during resource switching.
[0103] Furthermore, combined with the PVC technology mentioned above, PVC persistent storage is used as the disk of the job instance to store all file data, including software packages, configuration files, data sets, etc., which can make data storage management more centralized and organized. This is just like the hard disk used for long-term data storage in traditional computer systems. PVC provides a reliable repository to ensure that data remains available during the switching of heterogeneous resources.
[0104] For example, when switching from a CPU cloud host to a GPU cloud host, the data in the PVC, such as the model training dataset and pre-trained model files, can be directly accessed on the new host, and CRIU is responsible for restoring key status information in the memory, such as the current progress of model training and temporary variables in the memory, so that the task can continue to execute seamlessly.
[0105] The overall architecture diagram of switching scheduling in containerized heterogeneous computing environment is as follows: Figure 2 As shown. Create a container from the standardized image templates of CPU and GPU. When this container is started, specify the PVC to be mounted in the volumeMounts and volumes fields through the Kubernetes Pod configuration file according to the task configuration. For example, if the task needs to access some basic data sets, mount the corresponding PVC to the specified directory in the container, such as the / data directory. In this way, the container can access the data in the persistent storage through this directory during operation. In the stage where the user requests CPU resources, that is, the stage where resources are in low demand, CPU resources are allocated to the container because it is in the stage where resources are in low demand. When the container needs to be paused or resources need to be adjusted, the CRIU technology is used to save the state of the container process (Checkpoint). The container's memory status, file descriptors, register contents and other information are saved to the disk to form a checkpoint file. At the same time, the data generated by the container during operation, such as intermediate calculation results and temporary files, are also stored in the mounted PVC persistent storage to ensure that the data will not be lost due to container suspension.
[0106] As the user's task requirements change, the user wants to switch from CPU resources to higher-demand GPU resources to meet the needs of model training. The resource demand phase has entered a high stage, and it is necessary to switch from CPU resources to GPU resources. At this time, since PVC provides persistent storage, the mounted PVC can continue to be used by the new container as the container migrates. When the new container is started, the previous PVC is also specified in the Pod configuration file to ensure that the data generated by the previous container and the basic data set can be accessed. After the new container is started, the previously saved checkpoint file is restored (Restore) using CRIU technology. The previously saved process status information is reloaded into the new container so that the task can continue to execute from the previously paused location. Combined with the data stored in the PVC, the new container can seamlessly continue to complete complex tasks because it has both restored the process status and accessed the required persistent data.
[0107] Figure 2K8S scheduling refers to the process in which the scheduler in Kubernetes allocates Pods to suitable nodes to run according to a series of rules and policies, that is, the scheduling policies designed in the following. In the above scenario, K8S scheduling is responsible for resource allocation and PVC association. During the resource allocation and scheduling process, the container is scheduled to the appropriate CPU or GPU node according to the resource requirements of the container (such as CPU, memory, GPU, etc.) and the available resources of the node. For example, when the user selects CPU resources, the container is scheduled to the appropriate CPU node; when the user chooses to switch to GPU resources, the container that requires GPU resources is scheduled to the node with the appropriate GPU device. In PVC association, ensure that the container can correctly mount the required PVC during the scheduling process. When the container is scheduled to a new node, K8S will ensure that the corresponding PVC storage volume can be correctly mounted to the specified directory in the container.
[0108] The low resource demand stage refers to the stage where the user selects CPU resources at the beginning. The tasks in this stage have relatively low requirements for computing resources, which may be some simple data processing, preliminary preparations for model reasoning, or lightweight monitoring tasks. For example, preliminary cleaning and preprocessing of small-scale data sets can meet the computing needs using CPU resources without the need for additional GPU acceleration. At this stage, in order to efficiently utilize resources, the system will allocate CPU resources to the corresponding containers, and can use CRIU technology to flexibly save the process status of the container so that it can be adjusted according to demand later.
[0109] The high resource demand phase refers to the phase where users need to switch to GPU resources. When the complexity of the task increases, such as large-scale deep learning model training, complex image or video processing, etc., it enters the high resource demand phase. These tasks require a lot of computing resources, and the performance requirements cannot be met by relying solely on the CPU, so GPU resources are needed. At this stage, the system will switch the container environment through technologies such as the K8S scheduler, schedule containers that require GPU acceleration to nodes with suitable GPU devices, and ensure data continuity by mounting PVC persistent storage. At the same time, it uses CRIU technology to restore the container status so that the task can continue to execute seamlessly.
[0110] Example 5
[0111] A method for switching and scheduling a containerized computing environment is shown in Example 4, except that the implementation process of step S4 is as follows:
[0112] S41: Each task model has different computing requirements, especially in the process of model training and inference, the resource consumption may vary greatly due to the complexity of the model. Therefore, we can use a series of parameterized formulas to map the computing resource requirements of the model. Define the system resource state (SystemResource Utilization) as the input of resource demand assessment, including the current CPU utilization and current memory utilization , the unit is percentage (%).
[0113] S42: The resource monitoring module is built on a heterogeneous computing environment. It aims to accurately and in real time collect the utilization of various key resources in the system, and provide accurate data support for subsequent resource prediction and pre-allocation strategies. This module is tightly integrated with the underlying hardware resources, container orchestration platforms (such as Kubernetes), and various dedicated monitoring tools (such as Prometheus and NVIDIA DCGM) to form a comprehensive resource monitoring system. The resource monitoring module regularly collects the utilization of system resources.
[0114] The resource monitoring module regularly collects the current CPU utilization and current memory utilization of system resources;
[0115] Deploy the Prometheus node exporter on each node of the Kubernetes cluster. Configure the Prometheus server, set the CPU utilization and memory utilization collection interval to 15 seconds, and obtain the total memory, used memory, and available memory of the node. Calculate the memory utilization accurately according to the calculation formula (used memory / total memory*100%) and store it in the time series database.
[0116] The Prometheus server obtains and stores CPU utilization and memory utilization data from related components at set intervals to facilitate subsequent query and analysis.
[0117] S43: defining a pre-allocation strategy, i.e., a decision function based on resource utilization, for determining whether it is necessary to switch between heterogeneous resources (especially from CPU to GPU) and whether it is necessary to pre-allocate resources. This function makes decisions based on the current resource load of the system (CPU utilization and memory utilization data provided by the resource monitoring module) and pre-defined thresholds.
[0118] Let the decision function be ,when =1, indicating that resource switching or pre-allocation is required; =0, indicating that no resource switching or pre-allocation is required, the expression is:
[0119] ;
[0120] in Indicates A task is 1 or 0, , are weight coefficients, which respectively represent the influence of CPU utilization and memory utilization on the switching decision, satisfying These weight coefficients can be adjusted according to the system's dependence on different resources and application scenarios. For example, in a system where GPU accelerated computing tasks account for a large proportion, the weight coefficients can be appropriately increased. The value of , They are the thresholds of CPU utilization and memory utilization. When the utilization of the corresponding resources exceeds the threshold, it indicates that the resources are under high load or are about to face resource shortage. At this time, it is more likely to need resource switching or pre-allocation. For example, , It can be set to 80%, which means that when the CPU utilization and memory utilization exceed 80%, you may need to consider switching or pre-allocating resources.
[0121] S44: sorting task priorities and allocating resources;
[0122] When multiple tasks request resources at the same time, the tasks requesting resource switching are first placed in the queue, and then the tasks in the queue are prioritized to determine the order of resource allocation; when the priorities of the tasks are the same, they are sorted in the order in which they enter the queue.
[0123] The task priority index P is designed to measure the priority of the task. Its calculation is based on the current node resource utilization obtained by the resource monitoring module. The formula is:
[0124] ;
[0125] Indicates The priority index of each task. The higher the priority index, the more urgent the task's demand for resources, and resources should be allocated first. For example, in a multi-task concurrent scenario, by calculating the priority index of each task, the tasks are sorted from high to low according to priority, and resources are allocated first to tasks with high priority indexes to ensure that system resources can give priority to tasks that need resources most, thereby improving the overall resource utilization efficiency and task execution efficiency of the system. When allocating resources, computing power resources are allocated according to the priority index of the task to achieve reasonable allocation and efficient utilization of resources. The dynamic resource switching decision flow chart is as follows: Figure 4 shown.
[0126] For example, in a scenario where only CPU resources are used, the CPU resources are 2 cores and 4G. Set to 0.5. Set to 0.5; , It can be set to 80%, assuming that the CPU utilization of user 1 is The memory utilization is 85%. The CPU utilization of users 2 to 7 is assumed to be 90%. and memory utilization The priority index is obtained through a decision function based on resource utilization for , calculated is 0.075, and we can find arrive , as shown in Table 1 below:
[0127] Table 1 Priority index and decision function table of users 1 to 7
[0128]
[0129] At this time, users 1, 4, 5, and 7 are stored in the queue, and the GPU resources are marked according to the user serial number to ensure that the GPU scheduling strategy of the user demand of R=0 will not be affected. At this time, the scheduling is executed according to the default scheduling strategy in the cluster. Further, when these users stored in the queue choose to switch from CPU resources to the same GPU resources at the same time, the next step is to determine the priority index of users 1, 4, 5, and 7 and sort them. The priority of user 4 = the priority of user 7 > the priority of user 5 > the priority of user 1. If the user priorities are the same, the GPU resources are allocated according to the time order of which user meets R=1 in advance, that is, the earlier the user triggers the switching condition, the higher the priority of the user, that is, the priority of user 4 > the priority of user 7 > the priority of user 5 > the priority of user 1. The GPU resources are allocated according to the size of the user priority.
[0130] Example 6
[0131] A switching scheduling system for a containerized computing environment, used to implement the switching scheduling method for a containerized computing environment of the above-mentioned embodiment 5, comprises:
[0132] The containerized environment construction and initialization module is configured to: before switching resources in a heterogeneous computing environment, accelerate the startup and configuration of containers by designing and producing standardized image templates that adapt to different resources in combination with containerization technology;
[0133] The persistent storage module is configured to: during the switching of resources in a heterogeneous computing environment, ensure the consistency and availability of data generated during the entire process of model design, model training, and model reasoning by using persistent storage technology;
[0134] The task state preservation and interruption recovery module is configured to: in the process of switching resources in a heterogeneous computing environment, combine the process state preservation and recovery mechanism of CRIU technology to achieve task migration and reduce computing interruptions caused by environment switching;
[0135] The resource pre-allocation and scheduling decision module is configured as follows: in the process of resource scheduling, by designing a resource pre-allocation strategy, it is responsible for dynamically predicting and reasonably scheduling the computing resources required for the task according to the priority of the task, and finally completing resource switching and scheduling in a heterogeneous computing environment.
[0136] The above is a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for switching and scheduling a containerized computing environment, characterized in that: The steps include: S1: Before switching resources in a heterogeneous computing environment, we use containerization technology to design and create standardized image templates that adapt to different resources to speed up container startup and configuration. S2: During the resource switching process of heterogeneous computing environment, the data generated during the entire process of model design, model training, and model reasoning is ensured to be consistent and available during the switching process by using persistent storage technology. S3: In the process of switching resources in heterogeneous computing environments, the process state preservation and recovery mechanism of CRIU technology is combined to achieve task migration and reduce computing interruptions caused by environment switching; S4: In the process of resource scheduling, the resource pre-allocation strategy is designed to dynamically predict and reasonably schedule the computing resources required for the task according to the priority of the task, and finally complete the resource switching and scheduling in the heterogeneous computing environment; The implementation process of step S4 is: S41: Define the system resource status as input to resource demand assessment, including current CPU utilization CPU util and current memory utilization MEM util ; S42: Periodically collect the current CPU utilization and current memory utilization of system resources through the resource monitoring module; S43: defining a pre-allocation strategy, i.e., a decision function based on resource utilization, for determining whether switching between heterogeneous resources is required and whether resource pre-allocation is required; S44: sorting task priorities and allocating resources; In step S43, let the decision function be R. When R=1, it means that resource switching or pre-allocation is required; when R=0, it means that resource switching or pre-allocation is not required. The expression is: Where R i Indicates whether the i-th task is 1 or 0, ω1 and ω2 are weight coefficients, which respectively represent the influence of CPU utilization and memory utilization on the switching decision, satisfying ω1+ω2=1, and θ1 and θ2 are the thresholds of CPU utilization and memory utilization, respectively.
2. The method for switching and scheduling a containerized computing environment according to claim 1, characterized in that: In step S1, the standardized image template includes a basic CPU image, a Pytorch CPU image, and a basic GPU image, wherein the basic CPU image includes a Python development environment, as well as a data analysis and machine learning toolkit, which is suitable for applications that require less computing resources; The PytorchCPU image includes the Python development environment and toolkit in the base CPU image, as well as the CPU-based Pytorch image, which includes the Pytorch framework for deep learning development. The basic GPU image includes a basic environment with GPU acceleration support, which is suitable for tasks that require high-performance computing; The integrated development environment is built into the standardized image templates to support seamless code writing and debugging.
3. The method for switching and scheduling a containerized computing environment according to claim 2, characterized in that: The integrated development environment is Jupyter Notebook.
4. The method for switching and scheduling a containerized computing environment according to claim 2, characterized in that: Step S2 includes: S21: Designing a persistent storage architecture In heterogeneous computing clusters, a persistent storage system based on Kubernetes is built. A variety of persistent storage volume declaration templates are designed for different types of computing tasks. The persistent storage volume declaration template predefines the storage capacity, access mode, and storage category of the persistent storage volume declaration. Users select the appropriate persistent storage volume declaration template to create a persistent storage volume declaration based on task requirements. The cluster administrator configures various types of persistent storage volumes. Persistent storage volumes are based on different storage backend technologies. Through the dynamic provisioning mechanism, persistent storage volume declarations are automatically bound to appropriate persistent storage volumes to ensure that tasks have storage resources that meet their needs. S22: Data Mounting and Container Integration In the configuration file of the task container, explicitly specify the corresponding persistent storage volume declaration to be mounted. When the container is started on a heterogeneous node, whether it is a CPU node or a GPU node, Kubernetes automatically mounts the storage volume corresponding to the persistent storage volume declaration to the specified directory in the container according to the configuration of the task container; S23: The container stores the model's training data, checkpoints, and parameters in persistent storage. Through the persistent storage volume declaration technology provided by Kubernetes, the data will be stored on the persistent volume and will not be affected by the container lifecycle.
5. The method for switching and scheduling a containerized computing environment according to claim 4, characterized in that: In step S23, data will be stored on the persistent volume, which means that during the model training process, model checkpoints are regularly saved to the mounted / data path to ensure that data will not be lost due to container migration or resource switching; When resources are switched, containers can seamlessly access the mounted storage. When the container switches from a CPU node to a GPU node, the mounted persistent storage remains unchanged and the container continues to access the same data storage.
6. The method for switching and scheduling a containerized computing environment according to claim 5, characterized in that: In step S3, during the resource switching process, the state of the running model task is saved through the CRIU technology, including the memory state, file descriptor, and register content, and restored when switching to the target device, specifically: For each task T, the current execution state S(T) contains memory snapshot, process state, and network connection information. Through CRIU technology, the task state is serialized into a checkpoint file C(T): C(T) = checkpoint(T); Where checkpoint() represents the serialization process of the checkpoint file; On the target device, the task is continued through the recovery mechanism: T′=restore(C(T)); Where restore() represents the operation process of restoring the task status from the checkpoint file.
7. The method for switching and scheduling a containerized computing environment according to claim 1, characterized in that: In step S44, when multiple tasks request resources at the same time, the tasks requesting resource switching are first stored in the queue, and then the tasks in the queue are prioritized to determine the order of resource allocation; when the priorities of the tasks are the same, they are sorted in the order of entering the queue. The priority of the task is measured by designing the task priority index P, the formula is: P i =ω1×(CPU util -θ1)+ω2×(MEM util -θ2); P i Represents the priority index of the i-th task. The higher the priority index, the more urgent the task's demand for resources is, and resources should be allocated first.
8. A switching scheduling system for a containerized computing environment, characterized in that: The method for implementing the switching scheduling of the containerized computing environment according to claim 1 comprises: The containerized environment construction and initialization module is configured to: before switching resources in a heterogeneous computing environment, accelerate the startup and configuration of containers by designing and producing standardized image templates that adapt to different resources in combination with containerization technology; The persistent storage module is configured to: during the switching of resources in a heterogeneous computing environment, ensure the consistency and availability of data generated during the entire process of model design, model training, and model reasoning by using persistent storage technology; The task state preservation and interruption recovery module is configured to: in the process of switching resources in a heterogeneous computing environment, combine the process state preservation and recovery mechanism of CRIU technology to achieve task migration and reduce computing interruptions caused by environment switching; The resource pre-allocation and scheduling decision module is configured as follows: in the process of resource scheduling, by designing a resource pre-allocation strategy, it is responsible for dynamically predicting and reasonably scheduling the computing resources required for the task according to the priority of the task, and finally completing resource switching and scheduling in a heterogeneous computing environment.
Citation Information
Patent Citations
Dynamic load balancing resource scheduling method based on Kubernetes
CN110780998A
Container cross-heterogeneous cluster reconstruction method for domestic platform
CN110851237A