Resource scheduling method and device, electronic equipment and computer program product

By combining the distributed execution engine and container orchestration engine with the Ray and Kubernetes frameworks, the problem of inflexible resource scheduling in cloud-native environments is solved, efficient and flexible resource management and task scheduling for AI computing are achieved, and distributed computing deployment is simplified.

CN120832235APending Publication Date: 2025-10-24SUPCON TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510863290.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

When existing technologies perform AI computing in a cloud-native environment, resource scheduling is inflexible, computing efficiency is low, task management is complex, and it is difficult to dynamically adjust computing resources, resulting in complex and inefficient AI application deployment.

Method used

By using a distributed execution engine and a container orchestration engine, we can obtain task resource demand information and container information of the target resource cluster, identify non-running containers, schedule containers to execute tasks based on demand, and use the Ray and Kubernetes frameworks to achieve flexible resource scheduling and management.

Benefits of technology

It realizes flexible resource scheduling for machine learning tasks, improves computing efficiency and resource utilization, simplifies the complexity of distributed computing, and supports automated management and task scheduling of heterogeneous computing nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120832235A_ABST
    Figure CN120832235A_ABST
Patent Text Reader

Abstract

The invention discloses a resource scheduling method and device, electronic equipment and a computer program product. The method comprises the steps that task resource demand information of a to-be-scheduled task is obtained, and the task resource demand information at least comprises the resource demand quantity needed by the to-be-scheduled task; resource node information of a target resource cluster is obtained, the target resource cluster is created by using a distributed execution engine in advance, the target resource cluster comprises a plurality of containers created on resource nodes by using a container arrangement engine in advance, and the resource node information is at least used for representing the running condition of each container; determining non-running containers in a target resource cluster according to the resource node information; and under the condition that the resources of the non-running containers meet the resource demand quantity, scheduling the non-running containers to execute the to-be-scheduled task according to the resource demand quantity. According to the invention, the technical problem of inflexible resource scheduling under the condition of machine learning calculation in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of computers, and in particular, to a resource scheduling method and device, electronic equipment and computer program product. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, the demand for computing resources is increasing. The computing tasks in the AI field usually require a large amount of computation. According to the type of the trained model, whether it is gradient boosting or neural network, some problems may be encountered when training the model: too long time is needed to complete the training; the data is too large to be accommodated by a machine; the model is too large to be placed in a machine, etc. The traditional single-machine computing mode has been difficult to meet the needs of large-scale AI training and reasoning. The emergence of cloud-native technology provides elastic and scalable computing resources for AI computing. The technical personnel in this field begin to consider using the computing resources of multiple machines for computing tasks in the AI field, but how to efficiently use these resources to realize distributed AI computing is still a technical challenge.

[0003] In the prior art, professional basic setting knowledge is needed to manage and schedule computing resources. If the computing resources used by enterprises or technical personnel have heterogeneous nodes, this undoubtedly increases the complexity of deployment and use. In order to realize distributed computing, developers often need to make a lot of modifications to the existing AI code to adapt to the distributed architecture, which is not only time-consuming but also prone to errors; after modifying the code to realize the use of multi-node computing resources, the communication between nodes needs to be configured separately between each two nodes to ensure that the nodes can communicate with each other, and the start command needs to be executed on each node, which is tedious and complex; different users or teams need independent running and scheduling environments when training AI models, and the inconsistency of environment configuration also increases the difficulty of development and cooperation.

[0004] The traditional framework needs to separately set the communication between the master node and other nodes when expanding the computing nodes, which is tedious and costly; it does not support a dynamic task allocation mechanism, making it difficult to dynamically adjust the computing resources according to real-time needs, and the computing resources are not fully utilized; and the system maintenance complexity is high, requiring high technical requirements for AI field personnel.

[0005] The prior art faces problems such as inflexible resource scheduling, low computing efficiency, and complex task management when performing AI computing in a cloud-native environment. These problems limit the rapid development and deployment of AI applications.

[0006] In view of the problem of inflexible resource scheduling in the prior art when performing machine learning computing, no effective solution has been proposed so far. SUMMARY

[0007] Embodiments of the present application provide a resource scheduling method and device, electronic equipment and computer program product to at least solve the technical problem of inflexible resource scheduling in the prior art when machine learning is performed.

[0008] According to an aspect of embodiments of the present application, a resource scheduling method is provided, comprising: obtaining task resource requirement information of a to-be-scheduled task, wherein the to-be-scheduled task is used to indicate calling a resource node to perform machine learning, and the task resource requirement information at least includes a resource requirement amount required by the to-be-scheduled task; obtaining resource node information of a target resource cluster, wherein the target resource cluster is created in advance using a distributed execution engine, the target resource cluster includes a plurality of containers created in advance on a resource node using a container orchestration engine, and the resource node information at least indicates a running state of each container; determining a container that is not running in the target resource cluster according to the resource node information; and scheduling the container that is not running to perform the to-be-scheduled task according to the resource requirement amount if the resource of the container that is not running meets the resource requirement amount.

[0009] Optionally, after obtaining the task resource requirement information of the to-be-scheduled task, the method further comprises: identifying a task running environment required by the to-be-scheduled task from the task resource requirement information; and determining the target resource cluster corresponding to the task running environment from a plurality of preset resource clusters created in advance using the distributed execution engine, wherein the distributed execution engine creates a plurality of preset resource clusters corresponding to a plurality of preset running environments in advance.

[0010] Optionally, after obtaining the task resource requirement information of the to-be-scheduled task, the method further comprises: identifying a task running environment required by the to-be-scheduled task from the task resource requirement information; and using the container orchestration engine to orchestrate the resource node conforming to the resource type to obtain the target resource cluster including a plurality of containers.

[0011] Optionally, before obtaining the resource node information of the target resource cluster, the method further comprises: obtaining a computer cluster for performing the to-be-scheduled task, wherein the computer cluster includes at least one computing node for performing the to-be-scheduled task; importing the distributed execution engine into each computing node to obtain a plurality of resource nodes, wherein the distributed execution engine is used to create at least one resource node on each computing node; and generating at least one preset resource cluster according to the plurality of resource nodes, wherein the preset resource cluster at least includes the target resource cluster.

[0012] Optionally, generating the at least one preset resource cluster according to the plurality of resource nodes comprises: using the distributed execution engine to containerize each of the resource nodes to obtain a plurality of containers; and dividing the plurality of containers into each of the preset resource clusters, wherein each of the preset resource clusters comprises a plurality of the containers.

[0013] Optionally, generating the at least one preset resource cluster according to the plurality of resource nodes comprises: identifying a resource type of the resource nodes, wherein the resource type at least comprises a first type and a second type; using the distributed execution engine to containerize the resource nodes of the first type to obtain first containers; using the distributed execution engine to containerize the resource nodes of the second type to obtain second containers; and generating a plurality of the preset resource clusters according to the plurality of first containers and the plurality of second containers, wherein the plurality of containers in the preset resource cluster are all the first containers, or the plurality of containers in the preset resource cluster are all the second containers, or the plurality of containers in the preset resource cluster are the first containers and the second containers.

[0014] Optionally, in a case where resources of the non-running containers meet the resource demand amount, scheduling the non-running containers to execute the to-be-scheduled task according to the resource demand amount comprises: selecting, in the target resource cluster, a plurality of non-running containers that meet the resource demand amount to obtain a to-be-scheduled resource set; selecting an arbitrary container as a head node and selecting other containers as worker nodes in the to-be-scheduled resource set; scheduling the head node to receive the to-be-scheduled task; and scheduling, by the head node, the worker nodes in the same to-be-scheduled resource set to execute the to-be-scheduled task.

[0015] According to another aspect of the embodiment of the present application, a resource scheduling apparatus is further provided, comprising: a first obtaining module configured to obtain task resource demand information of a to-be-scheduled task, wherein the to-be-scheduled task is used to instruct to call a resource node to perform machine learning, and the task resource demand information at least comprises a resource demand amount required by the to-be-scheduled task; a second obtaining module configured to obtain resource node information of a target resource cluster, wherein the target resource cluster is created in advance using a distributed execution engine, and the target resource cluster comprises a plurality of containers created on resource nodes in advance using a container orchestration engine, and the resource node information at least indicates a running state of each of the containers; a determining module configured to determine, according to the resource node information, non-running containers in the target resource cluster; and a scheduling module configured to, in a case where resources of the non-running containers meet the resource demand amount, schedule the non-running containers to execute the to-be-scheduled task according to the resource demand amount.

[0016] According to another aspect of the embodiments of the present application, an electronic device is provided, comprising a memory and a processor, the memory storing a computer program, and the processor being configured to execute the above-mentioned resource scheduling method by the computer program.

[0017] According to another aspect of the embodiments of the present application, a computer program product is provided, comprising computer instructions, which, when executed by a processor, implement the steps of the above-mentioned resource scheduling method.

[0018] In the embodiments of the present application, the task resource requirement information of the to-be-scheduled task is acquired, wherein the to-be-scheduled task is used to indicate calling the resource node to perform machine learning, and the task resource requirement information at least includes a resource requirement amount required by the to-be-scheduled task; the resource node information of the target resource cluster is acquired, wherein the target resource cluster is created in advance using a distributed execution engine, and the target resource cluster includes a plurality of containers created on the resource node in advance using a container orchestration engine, and the resource node information at least indicates the running state of each container; according to the resource node information, the container not running in the target resource cluster is determined; in the case that the resource of the container not running meets the resource requirement amount, the container not running is scheduled to perform the to-be-scheduled task according to the resource requirement amount; thereby, by monitoring the container not running in the target resource cluster, the container not running can be flexibly scheduled to perform the to-be-scheduled task of machine learning in a distributed manner, and the technical effect of flexibly scheduling the resource node required by machine learning is achieved, thereby solving the technical problem of inflexible resource scheduling in the prior art in the case of machine learning calculation. BRIEF DESCRIPTION OF DRAWINGS

[0019] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the application. In the drawings:

[0020] Figure 1 is a flowchart of a resource scheduling method according to an embodiment of the present application;

[0021] Figure 2 is a schematic diagram of a Ray technical framework according to an embodiment of the present application;

[0022] Figure 3 is a schematic diagram of a Ray cluster distributed computing framework according to an embodiment of the present application;

[0023] Figure 4 is a schematic diagram of a Ray-based distributed computing platform architecture according to an embodiment of the present application;

[0024] Figure 5is a schematic diagram of a cluster AI task scheduling process according to an embodiment of the application;

[0025] Figure 6 is a schematic diagram of a resource scheduling device according to an embodiment of the application;

[0026] Figure 7 is a structural block diagram of a computer terminal according to an embodiment of the application. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0029] First, some of the nouns or terms appearing in the description of the embodiments of the present application are applicable to the following explanations:

[0030] Ray: an open source framework for extending AI and Python applications such as machine learning, which provides a computing layer for parallel processing, greatly reducing the complexity of running distributed computing.

[0031] AI distributed computing: using the computing power of multiple virtual or physical computers, working cooperatively through network connection, decomposing large-scale data and computing tasks to multiple nodes, each node computing simultaneously, efficiently completing complex AI computing tasks.

[0032] K8S: the abbreviation of Kubernetes, is an open source container orchestration system that automates the deployment, expansion and management of containerized applications, which can be used to manage distributed computing resources and support automatic expansion and resource scheduling of Ray clusters, so that applications can run reliably in different environments.

[0033] According to an embodiment of the present application, a resource scheduling method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0034] Figure 1 is a flowchart of a resource scheduling method according to an embodiment of the present application, as shown in Figure 1 The method comprises the following steps:

[0035] Step S102, obtaining task resource demand information of a to-be-scheduled task, wherein the to-be-scheduled task is used to indicate calling a resource node for machine learning, and the task resource demand information at least includes a resource demand amount required by the to-be-scheduled task;

[0036] Step S104, obtaining resource node information of a target resource cluster, wherein the target resource cluster is created in advance using a distributed execution engine, and the target resource cluster includes a plurality of containers created in advance on the resource node using a container orchestration engine, and the resource node information at least indicates the running status of each container;

[0037] Step S106, determining a non-running container in the target resource cluster according to the resource node information;

[0038] Step S108, in the case that the resource of the non-running container meets the resource demand amount, scheduling the non-running container to execute the to-be-scheduled task according to the resource demand amount.

[0039] In the embodiment of the present application, the task resource requirement information of the task to be scheduled is obtained, wherein the task to be scheduled is used to indicate calling the resource node to perform machine learning, and the task resource requirement information at least includes the resource requirement amount required by the task to be scheduled; the resource node information of the target resource cluster is obtained, wherein the target resource cluster is created in advance using the distributed execution engine, and the target resource cluster includes: a plurality of containers created on the resource node in advance using the container orchestration engine, and the resource node information at least indicates the running condition of each container; according to the resource node information, the container not running in the target resource cluster is determined; in the case that the resource of the container not running meets the resource requirement amount, the container not running is scheduled to perform the task to be scheduled according to the resource requirement amount; thereby, by monitoring the container not running in the target resource cluster, the container not running can be flexibly scheduled to perform the machine learning task to be scheduled in a distributed manner, and the technical effect of flexibly scheduling the resource node required by machine learning is achieved, thereby solving the technical problem of inflexible resource scheduling in the prior art in the case of machine learning calculation.

[0040] In the step S102, the task to be scheduled can be obtained from the cluster task queue, which can be a first-in first-out queue, that is, the scheduling task added first to the cluster task queue will be output from the cluster task queue first, and the task to be scheduled is the scheduling task output from the cluster task queue.

[0041] In the step S102, the task to be scheduled can be a training task of machine learning, which needs to use at least one resource node for execution.

[0042] Optionally, the cluster task queue can use the cluster task queue information to describe the total number of tasks in the cluster task queue and the task resource requirement (i.e., the task resource requirement information) of each task in the cluster task queue.

[0043] In the step S104, the distributed execution engine uses the Ray framework.

[0044] In the step S104, the target resource cluster can be a Ray cluster created in advance using the distributed execution engine, which includes a plurality of resource nodes capable of providing distributed computing.

[0045] Figure 2 is a schematic diagram of a Ray technical framework according to the embodiment of the present application, as Figure 2 shown, Ray is an open-source unified framework for expanding AI and Python application programs such as machine learning, which is composed of three layers:

[0046] S1, Ray's high-level libraries - Ray AI Libraries, provide scalable and unified ML application toolkits for AI domain computing tasks, which can simplify the development and deployment of AI applications and enable developers to focus more on algorithm and model design.

[0047] S2, Ray Core, an open-source Python general-purpose distributed computing library provided by Ray, enables ML developers to scale Python applications and accelerate computing task workloads. Core functions of distributed computing are provided, including task scheduling, resource management, and data parallel processing, etc.

[0048] S3, Ray Cluster, an ever-growing integration ecosystem, can automatically scale and reduce computing resources according to the resources requested by the applications running on the cluster, improving resource utilization.

[0049] Figure 3 is a schematic diagram of a Ray cluster distributed computing framework according to an embodiment of the present application, as Figure 3 shown, the training function is a user-defined python function training code containing end-to-end model training loop logic; Trainer is a Trainer class designed by Ray Train for different frameworks in the AI field, which can start distributed training by starting Ray Cluster according to the configuration of scalingconfiguration, and distribute the user-defined trainingfunction to multiple workers in the cluster, each worker will execute this training function to realize parallel processing of AI domain computing tasks among multiple computing resources.

[0050] It should be noted that Ray Train is a scalable machine learning library for distributed training and fine-tuning, which supports many commonly used frameworks in the AI field, such as PyTorch, Hugging Face, TensorFlow, etc. It can extend AI domain computing tasks such as model training code from a single computer to a computer cluster and abstract the complexity of distributed computing. Whether it is a large model or a large data set, workloads can be easily scaled and distributed among multiple GPU computing resources without the need to change the code.

[0051] It should be noted that Ray Cluster includes a single Head node (i.e., head node) and any number of worker nodes (i.e., worker nodes), which can be automatically expanded and reduced according to the resources requested by the application running on the cluster. Each Ray cluster has a resource node designated as the Head node of the cluster. The Head node is the same as other Worker nodes and can be scheduled for tasks on Ray. The difference is that it also runs the Driver process responsible for cluster management, GCS, etc. The Head node is responsible for allocating resources, creating worker nodes and ensuring their normal operation.

[0052] In the above step S104, the resource nodes in the target resource cluster can be containerized using a container orchestration engine. Each resource node can generate at least one container, and containerizing multiple resource nodes using a container orchestration engine can generate multiple containers.

[0053] In the above step S106, resource node information can be obtained by monitoring the running status of each container in the target resource cluster in real time. According to the resource node information, the number of non-running containers in the target resource cluster can be counted in real time, and it can be determined whether the non-running containers in the target resource cluster meet the resource requirement of the task to be scheduled.

[0054] In the above step S108, each container in the same target resource cluster can have the same data processing capability, i.e., each container can provide the same amount of resource processing. Therefore, according to the resource requirement of the task to be scheduled, at least one container that meets the resource requirement can be scheduled to execute the task to be scheduled.

[0055] Optionally, the product of the number of scheduled containers and the resource processing amount of each container is not less than the resource requirement of the task to be scheduled.

[0056] As an optional embodiment, after obtaining the task resource requirement information of the task to be scheduled, the method further includes: identifying the task running environment required by the task to be scheduled from the task resource requirement information; determining the target resource cluster corresponding to the task running environment from a plurality of preset resource clusters created in advance using a distributed execution engine, wherein the distributed execution engine creates a plurality of preset resource clusters corresponding to a plurality of preset running environments in advance.

[0057] In the above embodiments of the present application, the target resource cluster can be created in advance using resource nodes, wherein the resource nodes can be various resource types such as CPU resources or GPU resources. Different resource types of resource nodes or combinations of resource nodes can support different task running environments. Therefore, using the container orchestration engine to orchestrate the selected resource nodes can obtain a plurality of preset resource clusters corresponding to a plurality of preset running environments, and then after obtaining the resource node information of the target resource cluster, the target resource cluster that meets the required task running environment of the task to be scheduled can be selected from the plurality of preset resource clusters, and then all containers in the target resource cluster can be called to process the task to be scheduled, or part of the containers in the target resource cluster can be called to process the task to be scheduled according to the resource requirement of the task to be scheduled.

[0058] As an optional embodiment, after obtaining the task resource requirement information of the task to be scheduled, the method further includes: identifying the required task running environment of the task to be scheduled from the task resource requirement information; and using the container orchestration engine to orchestrate the resource nodes that meet the resource type to obtain the target resource cluster including a plurality of containers.

[0059] In the above embodiments of the present application, different tasks to be scheduled require different task running environments, and different resource types of resource nodes need to be configured for different task running environments. Therefore, when the task to be scheduled needs to be processed, the task running environment required by the task to be scheduled is identified through the task resource requirement information, and the resource nodes that meet the resource type are selected for orchestration by the container orchestration engine, so that a plurality of containers that meet the required task running environment of the task to be scheduled can be obtained, and then a plurality of containers that meet the required resource requirement of the task to be scheduled are selected to generate a target resource cluster that meets the required task running environment and resource requirement of the task to be scheduled.

[0060] Optionally, the container orchestration engine can use Kubernetes (k8s) as a container orchestration platform to realize unified management and scheduling of GPU resources in the entire cluster by using the powerful resource scheduling capability of Kubernetes (k8s). On this basis, the Ray cluster is deployed, the head node of Ray intelligently distributes tasks to each worker node, and parallel processing of distributed computing tasks is realized.

[0061] Optionally, the task running environment can indicate the resource type of the resource node required by the task to be scheduled, and then the resource nodes that meet the resource type can be selected for containerization processing according to the task running environment to realize the creation of the preset resource cluster.

[0062] Optionally, the resource type indicated resource node includes: a CPU node and a GPU node.

[0063] As an optional embodiment, before obtaining the resource node information of the target resource cluster, the method further comprises: obtaining a computer cluster for executing the to-be-scheduled task, wherein the computer cluster comprises: at least one computing node for executing the to-be-scheduled task; importing a distributed execution engine into each computing node to obtain a plurality of resource nodes, wherein the distributed execution engine is used to create at least one resource node on each computing node; and generating at least one preset resource cluster according to the plurality of resource nodes, wherein the preset resource cluster at least comprises the target resource cluster.

[0064] In the above embodiments of the present application, the computer cluster can include at least one computing node with computing capability, such as a CPU or a GPU. The distributed execution engine is imported into each computing node, which can enable each computing node to have distributed computing capability, and at least one resource node is created for each computing node. For example, for each CPU, a plurality of CPU nodes can be obtained, and for each GPU, a plurality of GPU nodes can be obtained. Then, by integrating all the computing nodes in the computer cluster, a plurality of resource nodes can be obtained. Based on the plurality of resource nodes in the computer cluster, a preset resource cluster can be generated, which realizes the creation of the preset resource cluster. Further, the target resource cluster capable of processing the to-be-scheduled task can be selected from the plurality of preset resource clusters.

[0065] It should be noted that if you want to simply use these AI distributed computing related libraries, that is, to enable each computing node in the computer cluster to provide distributed computing capability, you need to install the distributed execution engine Ray on the operating system of each computing node, that is, you can seamlessly use the above open source libraries. For example, using the pip method:

[0066] pip install“ray[train,tune,...]”==2.x.x.

[0067] Next, start a Python session, in which Ray can be easily imported and initialized:

[0068] “import ray

[0069] ray.init()”。

[0070] Using the above two lines of code, a Ray cluster can be started on the local computer, which can use all available kernels in the computing resource as a worker. Users do not have to worry about how to parallelize the code, but only need to “connect” the large dataset to Ray Train.

[0071] Here, taking the commonly used PyTorch framework pseudo code as an example, the specific implementation is as follows:

[0072] step1: Import the packages or libraries required by the distributed computing platform.

[0073] step2: Define the train_func function, which wraps the code in the train_func function, that is, the user-defined code training logic is placed here.

[0074] step3: Define the number of distributed training worker threads and whether to use GPU resources, where num_workers: the number of worker nodes to which the distributed training work is assigned to the training task, that is, the user's training task is assigned to several workers; use_gpu: whether each worker node uses GPU resources.

[0075] step4: Create a run_config object to specify the path to save the results (model parameters, model structure, etc.); storage_path here can specify a shared storage location (such as cloud storage or NFS), which is optional for single-node clusters, but is mandatory and effective for multi-node clusters.

[0076] Step5: Use the integrated TorchTrainer class to start the distributed training job according to the set scaling_config.

[0077] Optionally, using the commonly used PyTorch framework pseudo code as an example, the implementation code is as follows:

[0078] “step1:

[0079] from ray.train.torch import TorchTrainer

[0080] from ray.train import ScalingConfig

[0081] step2:

[0082] def train_func():

[0083] #Your PyTorch training code here. ...

[0085] step3:

[0086] scaling_config=ScalingConfig(num_workers=2,use_gpu=True)

[0087] step4:

[0088] run_config = RunConfig(storage_path=" / some / local / path", name="run_name")

[0089] step5:

[0090] trainer = TorchTrainer(train_func, scaling_config=scaling_config)

[0091] result = trainer.fit()”.

[0092] As an optional embodiment, generating the at least one preset resource cluster according to the plurality of resource nodes comprises: using a distributed execution engine to containerize each resource node to obtain a plurality of containers; and dividing the plurality of containers into each preset resource cluster, wherein each preset resource cluster comprises: a plurality of containers.

[0093] In the above embodiments of the present application, each resource node is containerized using a distributed execution engine to obtain a container that can be independently deployed on a resource node, and the plurality of containers are divided into clusters to obtain a plurality of preset resource clusters, wherein at least one container can be allocated to each preset resource cluster, thereby achieving the creation of a preset resource cluster.

[0094] Optionally, the number of containers for each preset resource cluster is set in advance, and the containers are allocated to each preset resource cluster according to the set number of resources, thereby achieving the cluster division of the containers obtained from the plurality of resource nodes.

[0095] Optionally, dividing the plurality of containers into each preset resource cluster comprises: dividing the plurality of containers obtained from the same resource node into the same preset resource cluster.

[0096] Optionally, dividing the plurality of containers into each preset resource cluster comprises: identifying the container types of the containers, and then dividing the plurality of containers into the preset resource clusters to which the containers belong according to the container types.

[0097] Optionally, the containers in each preset resource cluster can have the same container type.

[0098] Optionally, each preset resource cluster can also be a combination of containers of different container types.

[0099] Optionally, the container type is determined according to the resource type of the resource node that generates the container.

[0100] As an optional embodiment, generating the at least one preset resource cluster according to the plurality of resource nodes comprises: identifying a resource type of the resource nodes, wherein the resource type at least comprises: a first type and a second type; containerizing the resource nodes of the first type using the distributed execution engine to obtain first containers; containerizing the resource nodes of the second type using the distributed execution engine to obtain second containers; and generating a plurality of preset resource clusters according to the plurality of first containers and the plurality of second containers, wherein the plurality of containers in a preset resource cluster are all first containers, or the plurality of containers in a preset resource cluster are all second containers, or the plurality of containers in a preset resource cluster are a combination of first containers and second containers.

[0101] In the above embodiments of the present application, the required resources of the to-be-scheduled task can be a single type of resource or a combination of multiple types of resources. For example, the to-be-scheduled task can be executed by a CPU-based container alone, or can be executed by a GPU-based container alone, or can be executed by a combination of a CPU-based container and a GPU-based container. Therefore, in the case of generating a preset resource cluster, different first containers and second containers of different container types can be generated based on resource nodes of different resource types. Then, based on the plurality of first containers and the plurality of second containers, a preset resource cluster of all first containers, a preset resource cluster of all second containers, and a preset resource cluster of a combination of first containers and second containers can be generated, thereby achieving the construction of the preset resource cluster.

[0102] Optionally, the resource nodes of the first type can be resource nodes obtained based on a CPU, and the resource nodes of the second type can be resource nodes obtained based on a GPU.

[0103] As an optional embodiment, in the case where the resources of the non-running containers meet the resource demand amount, scheduling the non-running containers to execute the to-be-scheduled task according to the resource demand amount comprises: in the target resource cluster, selecting a plurality of non-running containers that meet the resource demand amount to obtain a to-be-scheduled resource set; in the to-be-scheduled resource set, selecting an arbitrary container as a head node and selecting other containers as worker nodes; scheduling the head node to receive the to-be-scheduled task; and scheduling the worker nodes in the same to-be-scheduled resource set by the head node to execute the to-be-scheduled task.

[0104] In the above embodiments of the present application, the containers in the target resource cluster can be used as head nodes (such as Head nodes) or worker nodes (such as Work nodes). Therefore, in the case of resource scheduling, the head node can receive the to-be-scheduled task, and then the to-be-scheduled task can be distributed to the worker nodes, so that the worker nodes and the head node can jointly execute the to-be-scheduled task in a distributed manner, thereby achieving flexible scheduling of the required resources of the to-be-scheduled task and improving the execution efficiency of the to-be-scheduled task.

[0105] Optionally, both the head node and the worker node are schedulable resources, but the head node can also be responsible for the driver process Driver, GCS, etc. of cluster management, for example, the head node is responsible for allocating resources, creating worker nodes, and ensuring that they are running normally.

[0106] The application also provides a preferred embodiment, which provides a Ray-based cloud-native AI distributed computing scheme, which combines the open source framework Ray with the Kubernetes cluster management capability to simplify the deployment and communication setting of AI distributed computing tasks, dynamically adjusts the required computing resources according to the computing tasks, improves the resource utilization and computing efficiency, and significantly improves the computing efficiency and system scalability by optimizing the resource scheduling, task allocation and execution process.

[0107] As an optional embodiment, the Ray cluster allows the expansion of AI field application tasks without the need for professional infrastructure knowledge, and almost no code modification is needed to easily parallelize and distribute machine learning ML workloads among multiple GPUs.

[0108] In the cloud-native environment, the KubeRay Operator can be used to realize k8s management of the Ray cluster, build an AI distributed computing platform, and schedule tasks according to the needs and resource conditions of different users to deploy large-scale workloads.

[0109] Figure 4 is a schematic diagram of a Ray-based distributed computing platform architecture according to an embodiment of the application, as Figure 4 shown, the KubeRay Operator is used to realize k8s native Ray cluster management, and the Operator provides a k8s native way to manage the Ray cluster, each Ray cluster includes a Head node pod (such as a container serving as a head node) and a group of worker node pods (such as containers serving as worker nodes), and the Ray Autoscaler extension supports the Operator to adjust the size of the Ray cluster (i.e. adjust the number of required workers, computing resources GPU, storage resources, etc.) according to the requirements of different RayApplications, and add and delete pods as needed, this way also supports heterogeneous computing nodes (including GPU), and running runtime envs with different Ray Applications required in the same k8s cluster.

[0110] Optionally, as Figure 4As shown, Scalable Ray Applications, fetch different user-defined respective AI domain application codes, which can be Tensorflow architecture using GPU computing resources for text classification tasks, or pandas program code using CPU resources for custom models.

[0111] Optionally, as shown, Ray Core APIs, are the core application programming interfaces of the Ray framework, used to build distributed applications. Figure 4

[0112] Optionally, as shown, KubeRay Operator, is responsible for managing the life cycle of Ray clusters, including creation, update and deletion operations. Figure 4

[0113] Optionally, as shown, MinIO, is a distributed object storage service deployed through k8s, commonly used to store model files after model training. Figure 4

[0114] Optionally, as shown, Services, are third-party services integrated with Ray clusters, used to expose Ray cluster services. Figure 4

[0115] The above embodiments of the present application can achieve k8s management of the computing cluster, such as resource management (GPU, CPU, memory, storage, etc.), by deploying k8s applications in existing computing resources, such as multiple GPU computing nodes. Based on the Ray base image, different application running environment containers can be created and packaged according to different application programs or user needs. Ray Cluster custom resources can be created using k8s configuration scripts, and multiple Ray Cluster clusters can be created in k8s to support multiple environments. The above containers can be deployed in different Ray clusters to isolate the environments required for running and debugging different applications or users, such as Figure 5 “Text classifier runtime env1” in “Custom deeplearning runtime env2” and “Custom model runtime env3”, at this time each cluster has a Head node; in addition, through the Services service of k8s, each Ray cluster service can be exposed and its respective monitoring page port can be accessed for application connection.

[0116] ​​​​As an optional example, the current computing cluster is two computing nodes across physical nodes, each computing node has 2 GPU resources, and there are currently three application programs 1, 2 and 3.

[0117] Optionally, the running environment required by the program 1 is runtime_env1, which requires 4 GPUs (i.e. num_workers = 4, at this time it is assumed that each pod resource in the cluster is set to 1 GPU and 4 CPUs through k8s) for processing. After adjusting the code of the application program 1 into a distributed scalable program, the address address of the Ray Cluster1 in which the first running environment runtime_env1 is deployed is specified in ray.init(address), and the cluster RayCluster1 is connected according to the address address; according to the num_workers = 4 set in the application program, the entire computing cluster resource needs to be called, at this time the application program is started, and the service discovery mechanism of k8s will schedule the computing resources on two physical computing nodes, which is 4 GPUs in total, at this time there is only one Head node in Ray Cluster1, and Ray Autoscaler will create a group of 3 worker nodes according to num_workers = 4, each node has 1 GPU resource and corresponding memory resource, the application program 1 realizes distributed parallel computing tasks, and if there are a large number of model training files to be stored, the model files can be stored through the storage service such as MinIO deployed on k8s.

[0118] Optionally, after the application program 1 is started, the running environment required by the application program 2 is runtime_env2, which requires 2 GPUs (i.e. num_workers = 2), which is similar to the above method, but at this time k8s detects that the current computing cluster does not have available computing resources to distribute the expanded program 2.

[0119] Figure 5 is a schematic diagram of a cluster AI task scheduling process according to an embodiment of the application, as shown in Figure 6 , the steps are as follows:

[0120] Step S51, obtaining a cluster task queue.

[0121] Step S52, obtaining resource node information, wherein the resource node information includes: used resource nodes and unused resource nodes in each resource node, and the number and type of each resource node.

[0122] Step S53, obtaining resource node cluster task queue information, wherein the resource node cluster task queue information includes: the total number of tasks in the cluster task queue, and the task resource demand (such as resource demand) of each task in the cluster task queue.

[0123] Step S54, scheduling tasks based on the cluster task queue and allocating the required computing resources (such as containers) for each task; wherein the tasks to be scheduled are queued in the cluster task queue and wait until the required task resources are sufficient, and then scheduled to the corresponding allocated resource nodes.

[0124] In the above embodiments of the present application, program 2 is queued in the cluster task queue in a pending state, waiting for other tasks to complete and release resources, and then scheduled to Ray Cluster2 and allocated corresponding worker resource nodes.

[0125] The above is a description of the Ray-based AI distributed computing scheme under cloud-native. The embodiments in the present application are described in a progressive manner, and the same or similar parts between embodiments can be referred to. Each embodiment focuses on the differences from other embodiments. Through the description of the above embodiments, those skilled in the art can understand that, in addition, those skilled in the art can transform the same type of technology in the present application, such as container orchestration technology k8s or file storage service MinIO, or add other additional services for use by the skilled person.

[0126] The above embodiments of the present application use the Ray distributed framework to extend AI field computing tasks. The framework provides a variety of AI libraries and its own ecological environment, simplifies the complexity of distributed computing deployment, and supports a variety of commonly used AI field frameworks such as PyTorch, Hugging Face, TensorFlow, etc., providing extensive applicability.

[0127] The above embodiments of the present application use Ray clusters in combination with Kubernetes, which can support heterogeneous computing resources, achieve automated resource management and task scheduling in an isolated environment, and seamlessly expand and allocate workloads across nodes; and can carry a variety of third-party services for additional AI distributed computing service needs.

[0128] The Ray-based AI distributed computing scheme provided by the present application combines Kubernetes (k8s) technology, which can simply and efficiently utilize computing resources, seamlessly expand and allocate machine learning workloads, and has the following effects compared to the prior art:

[0129] 1. Simplified distributed task deployment: allows users to deploy distributed tasks almost without modifying existing code, simplifying the complexity of distributed computing.

[0130] 2. Computing capability across physical nodes: capable of large-scale workload deployment and computing in a cross-physical-node environment without complex and error-prone node communication settings, supporting multi-node distributed computing.

[0131] 3. Automated resource management and scheduling: the combination of Ray cluster and KubeRay Operator realizes automated resource management and task scheduling, dynamically adjusts computing resources according to application requirements, and improves resource utilization.

[0132] It should be noted that in specific scenarios, such as distributed computing tasks that require support for more languages, Ray can be combined with other orchestration tools (such as Apache Mesos) to provide extended support.

[0133] According to an embodiment of the present application, a resource scheduling device is also provided. It should be noted that the resource scheduling device can be used to execute the resource scheduling method of the present application. The resource scheduling method of the present application can be executed in the resource scheduling device.

[0134] Figure 6 is a schematic diagram of a resource scheduling device according to an embodiment of the present application, as shown in Figure 7 The device can include: a first acquisition module 62 for acquiring task resource demand information of a to-be-scheduled task, wherein the to-be-scheduled task is used to indicate calling a resource node for machine learning, and the task resource demand information at least includes a resource demand amount required by the to-be-scheduled task; a second acquisition module 64 for acquiring resource node information of a target resource cluster, wherein the target resource cluster is created in advance using a distributed execution engine, and the target resource cluster includes: a plurality of containers created on resource nodes in advance using a container orchestration engine, and the resource node information at least indicates the running status of each container; a determination module 66 for determining a container that is not running in the target resource cluster according to the resource node information; and a scheduling module 68 for scheduling the container that is not running to execute the to-be-scheduled task according to the resource demand amount if the resource of the container that is not running meets the resource demand amount.

[0135] It should be noted that the first acquisition module 62 in this embodiment can be used to execute step S102 in the present application, the second acquisition module 64 in this embodiment can be used to execute step S104 in the present application, the determination module 66 in this embodiment can be used to execute step S106 in the present application, and the scheduling module 68 in this embodiment can be used to execute step S108 in the present application. The above-mentioned modules and the corresponding steps realize the same examples and application scenarios, but are not limited to the contents disclosed in the above-mentioned embodiments.

[0136] In the embodiment of the present application, task resource requirement information of a to-be-scheduled task is acquired, wherein the to-be-scheduled task is used to indicate calling a resource node to perform machine learning, and the task resource requirement information at least includes a resource requirement amount required by the to-be-scheduled task; resource node information of a target resource cluster is acquired, wherein the target resource cluster is created in advance using a distributed execution engine, and the target resource cluster includes a plurality of containers created in advance on a resource node using a container orchestration engine, and the resource node information at least indicates a running condition of each container; based on the resource node information, a container not running in the target resource cluster is determined; in a case where a resource of the container not running satisfies the resource requirement amount, the container not running is scheduled to perform the to-be-scheduled task according to the resource requirement amount; thereby, by monitoring the container not running in the target resource cluster, the container not running can be flexibly scheduled to perform the to-be-scheduled task of machine learning in a distributed manner, and a technical effect of flexibly scheduling the resource node required by machine learning is achieved, thereby solving the technical problem of inflexible resource scheduling in the prior art in the case of machine learning calculation.

[0137] As an optional embodiment, the apparatus further includes a first identification submodule configured to identify a task running environment required by the to-be-scheduled task from the task resource requirement information after acquiring the resource node information of the target resource cluster; and a determination submodule configured to determine, from a plurality of preset resource clusters created in advance using the distributed execution engine, the target resource cluster corresponding to the task running environment, wherein the distributed execution engine creates a plurality of preset resource clusters corresponding to a plurality of preset running environments in advance.

[0138] As an optional embodiment, the apparatus further includes a second identification submodule configured to identify a task running environment required by the to-be-scheduled task from the task resource requirement information after acquiring the task resource requirement information of the to-be-scheduled task; and an orchestration submodule configured to use the container orchestration engine to orchestrate resource nodes of a resource type to obtain the target resource cluster including a plurality of containers.

[0139] As an optional embodiment, the apparatus further includes an acquisition submodule configured to acquire a computer cluster for performing the to-be-scheduled task before acquiring the task resource requirement information of the to-be-scheduled task, wherein the computer cluster includes at least one computing node for performing the to-be-scheduled task; an import submodule configured to import the distributed execution engine into each computing node to obtain a plurality of resource nodes, wherein the distributed execution engine is used to create at least one resource node on each computing node; and a generation submodule configured to generate at least one preset resource cluster based on the plurality of resource nodes, wherein the preset resource cluster at least includes the target resource cluster.

[0140] As an optional embodiment, the generating submodule comprises: a first containerization processing unit configured to containerize each resource node using a distributed execution engine to obtain a plurality of containers; and a dividing unit configured to divide the plurality of containers into each preset resource cluster, wherein each preset resource cluster comprises the plurality of containers.

[0141] As an optional embodiment, the generating submodule comprises: an identifying unit configured to identify a resource type of the resource node, wherein the resource type comprises at least a first type and a second type; a second containerization processing unit configured to containerize the resource node of the first type using the distributed execution engine to obtain a first container; a second containerization processing unit configured to containerize the resource node of the second type using the distributed execution engine to obtain a second container; and a generating unit configured to generate a plurality of preset resource clusters according to the plurality of first containers and the second containers, wherein the plurality of containers in each preset resource cluster are all the first containers, or the plurality of containers in each preset resource cluster are all the second containers, or the plurality of containers in each preset resource cluster are the first containers and the second containers.

[0142] As an optional embodiment, the scheduling module comprises: a first selecting unit configured to select a plurality of non-running containers satisfying the resource requirement amount in the target resource cluster to obtain a to-be-scheduled resource set; a second selecting unit configured to select any one container as a head node and select other containers as worker nodes in the to-be-scheduled resource set; a first scheduling unit configured to schedule the head node to receive a to-be-scheduled task; and a second scheduling unit configured to schedule the worker nodes in the same to-be-scheduled resource set to execute the to-be-scheduled task by the head node.

[0143] Embodiments of the present application can provide an electronic device, which can be a computer terminal, and the computer terminal can be any one of computer terminal devices in a computer terminal group. Alternatively, in the present embodiment, the computer terminal can be replaced by a mobile terminal or other terminal device.

[0144] Alternatively, in the present embodiment, the computer terminal can be located in at least one network device of a plurality of network devices of a computer network.

[0145] In the embodiment, the computer terminal can execute program codes of the following steps in the resource scheduling method: obtaining task resource demand information of a to-be-scheduled task, wherein the to-be-scheduled task is used to indicate calling a resource node to perform machine learning, and the task resource demand information at least includes a resource demand amount required by the to-be-scheduled task; obtaining resource node information of a target resource cluster, wherein the target resource cluster is created in advance using a distributed execution engine, and the target resource cluster includes a plurality of containers created in advance on the resource node using a container orchestration engine, and the resource node information at least indicates a running condition of each container; determining a non-running container in the target resource cluster according to the resource node information; and scheduling the non-running container to perform the to-be-scheduled task according to the resource demand amount in a case that a resource of the non-running container meets the resource demand amount.

[0146] Figure 7 is a structural block diagram of a computer terminal according to an embodiment of the present application, as shown in the figure, the computer terminal 70 can include one or more (only one is shown in the figure) processors 72 and a memory 74. Figure 7

[0147] The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the resource scheduling method and device in the embodiments of the present application, and the processor executes various functions and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned resource scheduling method. The memory can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, and these remote memories can be connected to the terminal 70 through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0148] The processor can call information and application programs stored in the memory through a transmission device to execute the following steps: obtaining task resource demand information of a to-be-scheduled task, wherein the to-be-scheduled task is used to indicate calling a resource node to perform machine learning, and the task resource demand information at least includes a resource demand amount required by the to-be-scheduled task; obtaining resource node information of a target resource cluster, wherein the target resource cluster is created in advance using a distributed execution engine, and the target resource cluster includes a plurality of containers created in advance on the resource node using a container orchestration engine, and the resource node information at least indicates a running condition of each container; determining a non-running container in the target resource cluster according to the resource node information; and scheduling the non-running container to perform the to-be-scheduled task according to the resource demand amount in a case that a resource of the non-running container meets the resource demand amount.

[0149] ​Optionally, the above-mentioned processor can also execute the program code of the following steps: identifying the task running environment required for the task to be scheduled from the task resource requirement information; determining the target resource cluster corresponding to the task running environment among multiple preset resource clusters created in advance using the distributed execution engine, wherein the distributed execution engine pre-creates preset resource clusters corresponding to multiple preset running environments.

[0150] Optionally, the processor may also execute the program code of the following steps: identifying the task running environment required for the task to be scheduled from the task resource requirement information; and using the container orchestration engine to orchestrate resource nodes that meet the resource type to obtain a target resource cluster including multiple containers.

[0151] Optionally, the processor may also execute the program code of the following steps: obtaining a computer cluster for executing the task to be scheduled, wherein the computer cluster includes: at least one computing node for executing the task to be scheduled; importing a distributed execution engine into each computing node to obtain multiple resource nodes, wherein the distributed execution engine is used to create at least one resource node on each computing node; generating at least one preset resource cluster based on the multiple resource nodes, wherein the preset resource cluster includes at least: a target resource cluster.

[0152] Optionally, the processor may also execute the program code of the following steps: using a distributed execution engine to containerize each resource node to obtain multiple containers; dividing the multiple containers into each preset resource cluster, wherein each preset resource cluster includes: multiple containers.

[0153] Optionally, the processor may also execute program code for the following steps: identifying the resource type of the resource node, wherein the resource type includes at least: a first type and a second type; containerizing the resource node of the first type using a distributed execution engine to obtain a first container; containerizing the resource node of the second type using a distributed execution engine to obtain a second container; generating a plurality of preset resource clusters based on the plurality of first containers and the second container, wherein the plurality of containers in the preset resource cluster are all first containers, or the plurality of containers in the preset resource cluster are all second containers, or the plurality of containers in the preset resource cluster are all first containers and the second container.

[0154] Optionally, the above-mentioned processor can also execute the program code of the following steps: in the target resource cluster, select multiple non-running containers that meet the resource requirements to obtain a set of resources to be scheduled; in the set of resources to be scheduled, select any one container as the head node, and other containers as working nodes; schedule the head node to receive the tasks to be scheduled; and the head node schedules the working nodes in the same set of resources to be scheduled to execute the tasks to be scheduled.

[0155] In an embodiment of the present invention, task resource requirement information of a task to be scheduled is obtained, wherein the task to be scheduled is used to indicate the call of a resource node for machine learning, and the task resource requirement information includes at least: the resource requirement required for the task to be scheduled; resource node information of a target resource cluster is obtained, wherein the target resource cluster is pre-created using a distributed execution engine, and the target resource cluster includes: multiple containers pre-created on the resource node using a container orchestration engine, and the resource node information is at least used to represent the running status of each container; based on the resource node information, the non-running containers in the target resource cluster are determined; when the resources of the non-running containers meet the resource requirement, the non-running containers are scheduled to execute the task to be scheduled according to the resource requirement; thereby, by monitoring the non-running containers in the target resource cluster, the non-running containers can be flexibly scheduled to distribute and execute the task to be scheduled for machine learning, thereby achieving the technical effect of flexibly scheduling the resource nodes required for machine learning, thereby solving the technical problem of inflexible resource scheduling in the prior art when performing machine learning calculations.

[0156] It can be understood by those skilled in the art that Figure 7 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 7 It does not limit the structure of the above electronic device. For example, the computer terminal 70 may also include Figure 7 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with ​ Different configurations shown.

[0157] A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a computer program. The computer program can be stored in a non-volatile medium. The non-volatile storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0158] The embodiment of the present invention further provides a non-volatile storage medium. Optionally, in this embodiment, the non-volatile storage medium can be used to store the program code executed by the resource scheduling method provided in the embodiment.

[0159] Optionally, in the embodiment, the non-volatile storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0160] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining task resource requirement information of a to-be-scheduled task, wherein the to-be-scheduled task is used to indicate calling a resource node to perform machine learning, and the task resource requirement information at least includes a resource requirement amount required by the to-be-scheduled task; obtaining resource node information of a target resource cluster, wherein the target resource cluster is created in advance using a distributed execution engine, and the target resource cluster includes a plurality of containers created on resource nodes in advance using a container orchestration engine, and the resource node information at least indicates a running state of each container; determining a non-running container in the target resource cluster according to the resource node information; and scheduling the non-running container to perform the to-be-scheduled task according to the resource requirement amount in a case where a resource of the non-running container meets the resource requirement amount.

[0161] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: identifying a task running environment required by the to-be-scheduled task from the task resource requirement information; and determining a target resource cluster corresponding to the task running environment from a plurality of preset resource clusters created in advance using the distributed execution engine, wherein the distributed execution engine creates a plurality of preset resource clusters corresponding to a plurality of preset running environments in advance.

[0162] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: identifying a task running environment required by the to-be-scheduled task from the task resource requirement information; and using the container orchestration engine to orchestrate resource nodes of a resource type to obtain the target resource cluster including a plurality of containers.

[0163] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: obtaining a computer cluster for performing the to-be-scheduled task, wherein the computer cluster includes at least one computing node for performing the to-be-scheduled task; importing the distributed execution engine into each computing node to obtain a plurality of resource nodes, wherein the distributed execution engine is used to create at least one resource node on each computing node; and generating at least one preset resource cluster according to the plurality of resource nodes, wherein the preset resource cluster at least includes the target resource cluster.

[0164] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: using the distributed execution engine, containerizing each resource node to obtain a plurality of containers; and dividing the plurality of containers into each preset resource cluster, wherein each preset resource cluster comprises the plurality of containers.

[0165] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: identifying a resource type of the resource node, wherein the resource type at least comprises a first type and a second type; using the distributed execution engine to containerize the resource node of the first type to obtain a first container; using the distributed execution engine to containerize the resource node of the second type to obtain a second container; and generating a plurality of preset resource clusters according to the plurality of first containers and the plurality of second containers, wherein the plurality of containers in the preset resource cluster are all the first containers, or the plurality of containers in the preset resource cluster are all the second containers, or the plurality of containers in the preset resource cluster are the first containers and the second containers.

[0166] Optionally, in the embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: in the target resource cluster, selecting a plurality of non-running containers satisfying the resource demand amount to obtain a to-be-scheduled resource set; in the to-be-scheduled resource set, selecting an arbitrary container as a head node and selecting other containers as worker nodes; scheduling the head node to receive a to-be-scheduled task; and scheduling, by the head node, the worker nodes in the same to-be-scheduled resource set to execute the to-be-scheduled task.

[0167] The embodiment of the present application further provides a computer program product comprising a computer program. Optionally, in the embodiment, the computer program is executed by a processor to realize the steps of the resource scheduling method provided in the above-described embodiments.

[0168] The above-described serial numbers of the embodiments of the present application only serve for description, and do not represent the advantages or disadvantages of the embodiments.

[0169] In the above-described embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0170] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented in other manners. For example, the described unit embodiments can be divided into other ways, and the functions of the units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be implemented by some interfaces, and the indirect couplings or communication connections can be implemented in some other manners, which are not shown or discussed.

[0171] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments.

[0172] In addition, each functional unit in the various embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0173] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a non-volatile storage medium. Based on this understanding, the technical solutions of the present application, essentially or the part that contributes to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a non-volatile storage medium, including a number of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned non-volatile storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0174] The above description is only the preferred embodiments of the present application, and it should be pointed out that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which should be considered as the protection scope of the present application.

Claims

1. A resource scheduling method, characterized in that: The method comprises: obtaining task resource requirement information of a task to be scheduled, wherein the task to be scheduled is used to indicate calling a resource node to perform machine learning, and the task resource requirement information at least includes a resource requirement amount required by the task to be scheduled; obtaining resource node information of a target resource cluster, wherein the target resource cluster is created in advance using a distributed execution engine, and the target resource cluster includes a plurality of containers created in advance on a resource node using a container orchestration engine, and the resource node information at least indicates a running state of each container; determining a container that is not running in the target resource cluster according to the resource node information; in a case where a resource of the container that is not running satisfies the resource requirement amount, scheduling the container that is not running to perform the task to be scheduled according to the resource requirement amount.

2. The method of claim 1, wherein, After obtaining the task resource requirement information of the task to be scheduled, the method further comprises: identifying a task running environment required by the task to be scheduled from the task resource requirement information; determining the target resource cluster corresponding to the task running environment from a plurality of preset resource clusters created in advance using the distributed execution engine, wherein the distributed execution engine creates a plurality of preset resource clusters corresponding to a plurality of preset running environments in advance.

3. The method of claim 1, wherein, After obtaining the task resource requirement information of the task to be scheduled, the method further comprises: identifying a task running environment required by the task to be scheduled from the task resource requirement information; orchestrating the resource node conforming to the resource type using the container orchestration engine to obtain the target resource cluster including a plurality of containers.

4. The method of claim 1, wherein, Before obtaining the resource node information of the target resource cluster, the method further comprises: obtaining a computer cluster for performing the task to be scheduled, wherein the computer cluster includes at least one computing node for performing the task to be scheduled; importing the distributed execution engine into each computing node to obtain a plurality of resource nodes, wherein the distributed execution engine is used to create at least one resource node on each computing node; generating at least one preset resource cluster according to a plurality of resource nodes, wherein the preset resource cluster at least includes the target resource cluster.

5. The method of claim 4, wherein, Generating at least one preset resource cluster according to a plurality of resource nodes comprises: containerizing each resource node using the distributed execution engine to obtain a plurality of containers; dividing a plurality of containers into each preset resource cluster, wherein each preset resource cluster includes a plurality of containers.

6. The method of claim 4, wherein, Generating at least one preset resource cluster according to a plurality of resource nodes comprises: identifying a resource type of the resource node, wherein the resource type at least includes a first type and a second type; containerizing the resource node of the first type using the distributed execution engine to obtain a first container; containerizing the resource node of the second type using the distributed execution engine to obtain a second container; According to the first container and the second container, a plurality of preset resource clusters are generated, wherein the containers in the preset resource cluster are all the first containers, or the containers in the preset resource cluster are all the second containers, or the containers in the preset resource cluster are the first containers and the second containers.

7. The method of claim 1, wherein, In a case where the resources of the non-operating container meet the resource demand amount, the non-operating container is scheduled to execute the to-be-scheduled task according to the resource demand amount. In the target resource cluster, a plurality of non-operating containers meeting the resource demand amount are selected to obtain a to-be-scheduled resource set. In the to-be-scheduled resource set, any one of the containers is selected as a head node, and the other containers are selected as worker nodes. The head node is scheduled to receive the to-be-scheduled task. The worker nodes in the same to-be-scheduled resource set are scheduled by the head node to execute the to-be-scheduled task.

8. A resource scheduling apparatus, characterized by comprising: Comprise: The first acquisition module is configured to acquire task resource demand information of a to-be-scheduled task, wherein the to-be-scheduled task is used to instruct a resource node to perform machine learning, and the task resource demand information at least includes a resource demand amount required by the to-be-scheduled task. The second acquisition module is configured to acquire resource node information of a target resource cluster, wherein the target resource cluster is created in advance using a distributed execution engine, and the target resource cluster includes a plurality of containers created on resource nodes in advance using a container orchestration engine, and the resource node information at least indicates the running state of each container. The determination module is configured to determine, according to the resource node information, a non-operating container in the target resource cluster. The scheduling module is configured to, in a case where the resources of the non-operating container meet the resource demand amount, schedule the non-operating container to execute the to-be-scheduled task according to the resource demand amount.

9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the resource scheduling method of any one of claims 1 to 7 through the computer program.

10. A computer program product comprising computer instructions, characterized in that, The computer instructions are executed by the processor to implement the steps of the resource scheduling method of any one of claims 1 to 7.

Citation Information

Cited By

  • Multi-modal data processing method and device, equipment and medium

    CN121255403A