Distributed scheduling method, device and system based on deterministic tasks
By scheduling tasks based on task configuration and computing node resources in a cloud environment, the deterministic problem of task scheduling is solved, and deterministic scheduling and efficient utilization of tasks in various scenarios are achieved.
Patent Information
- Application Number
- CN202410301246.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-15
- Publication Date
- 2025-09-26
AI Technical Summary
The task scheduling framework in the existing cloud environment cannot achieve deterministic execution, cannot guarantee the determinism of task execution, and only supports scheduling in container scenarios.
The management node schedules tasks to computing nodes and running capsules that meet deterministic constraints based on the task configuration and computing node resources. It supports scheduling of various scenarios such as virtual machines, containers, processes, and threads, and performs lifecycle management through the container management and scheduling system.
It realizes deterministic scheduling of tasks on the cloud, improves the certainty of task execution and resource utilization of computing nodes, and supports scheduling requirements in various scenarios.
Smart Images

Figure CN120704808A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer software technology, and in particular to a distributed deterministic task-based scheduling method, device, and system. Background Art
[0002] Currently, the "task" scheduling frameworks available in the market for cloud environments cannot achieve deterministic execution of "tasks". They can only achieve batch allocation and scheduling of "tasks" based on priority order, resource occupancy, etc., and cannot guarantee the determinism of task execution. In addition, they only support scheduling in container scenarios. Summary of the Invention
[0003] In view of this, the embodiments of the present application provide a distributed deterministic task-based scheduling method, device and system for scheduling tasks on the cloud. The technical solution of the embodiments of the present application schedules computing nodes and computing resources for the task according to the configuration of the received task and the resources of each computing node, thereby realizing the scheduling of deterministic tasks on the cloud.
[0004] In the first aspect, an embodiment of the present application provides a distributed scheduling method based on deterministic tasks for scheduling tasks on a cloud, wherein the cloud includes a management node and several computing nodes, including: the management node schedules the task to a computing node that meets the configuration according to the configuration of the received task and the resources of each computing node, wherein the configuration includes resource requirements that meet the deterministic constraints of the task, and the resource requirements include computing resource requirements and transmission resource requirements; the management node schedules the task to a corresponding running capsule on the scheduled computing node according to the configuration, and the resources of the running capsule meet the configuration. The running capsule of each computing node is deployed on the elastic microkernel of the computing node, and the elastic microkernel of each computing node allocates resources to its running capsule. In some embodiments, the resources of the running capsule meet the configuration, and the running capsule supports one of the following scenarios: virtual machine, container, process, thread.
[0005] As described above, the configuration of the received task includes the resource requirements that meet the deterministic constraints of the task. According to the configuration of the received task and the resources of each computing node, the computing nodes, computing resources and transmission resources are scheduled for the task to realize the scheduling of deterministic tasks on the cloud.
[0006] In a possible implementation of the first aspect, it also includes: the management node decomposing the task into several subtasks, and decomposing the configuration into several subconfigurations, each subconfiguration meeting the deterministic resource requirements of a subtask; scheduling the task to a computing node that meets the configuration according to the received task configuration and the resources of each computing node, specifically including: scheduling each subtask to a computing node that meets the subconfiguration according to the subconfiguration of the subtask and the resources of each computing node, and the transmission delay between the scheduled computing nodes meets the deterministic requirements of data flow transmission between subtasks; and / or scheduling the task to a corresponding running capsule on the scheduled computing node according to the configuration, specifically including: scheduling the corresponding subtask to a corresponding running capsule on the scheduled computing node according to the subconfiguration, and the resources of the corresponding running capsule meet the subconfiguration of the corresponding subtask.
[0007] As shown above, when a task is scheduled to multiple computing nodes, the transmission delay between the scheduled computing nodes satisfies the deterministic requirements of data flow transmission between subtasks to improve the determinism of task implementation. The task is scheduled to multiple running capsules to increase parallelism, thereby further improving the determinism of task implementation.
[0008] In a possible implementation of the first aspect, when there are multiple computing nodes that meet the configuration, the computing nodes are scheduled for the task based on load balancing and / or shortest delay.
[0009] As described above, the computing nodes are scheduled for the received tasks among a plurality of computing nodes that meet resource requirements based on load balancing and / or delay, so as to improve the time certainty of the received tasks.
[0010] In a possible implementation of the first aspect, the configuration further includes a priority of the task, and the priority of the running capsule matches the priority of the task.
[0011] As described above, by matching the priority of the running capsule with the priority of the received task, the time certainty of the received task is improved.
[0012] In a possible implementation of the first aspect, the method further includes: the management node obtaining the configuration according to the deterministic constraints of the received task.
[0013] As described above, by obtaining the task configuration according to the deterministic constraint of the received task, the running capsule scheduled based on the configuration runs the task, thereby realizing the deterministic constraint of the task.
[0014] In a possible implementation of the first aspect, the management node further includes: deploying the task's operating environment and installing the task on the scheduled computing node, and performing lifecycle management on the task, including at least monitoring the resource load of the scheduled computing node and the resource load of the scheduled running capsule.
[0015] As described above, by performing lifecycle management on the received tasks on the scheduled computing nodes, not only the highly reliable operation of the tasks is achieved, but also the resource utilization of the computing nodes is improved.
[0016] In a possible implementation of the first aspect, the further step includes: when the resource load of the scheduled computing node is higher than a first set threshold, the management node reducing the proportion of subsequent scheduling of the computing node; when the resource load of the scheduled running capsule is higher than a second set threshold, the management node reducing the proportion of subsequent scheduling of the running capsule.
[0017] As described above, by reducing the subsequent scheduling ratio of the scheduled computing nodes and running capsules, the resources of the scheduled computing nodes and running capsules can be more used for the received tasks, thereby improving their determinism.
[0018] In a possible implementation of the first aspect, the further step includes: the elastic microkernel of the scheduled computing node dynamically adjusting resources for the running capsule according to the resource load of the running capsule when the task is running.
[0019] As described above, the elastic microkernel of the scheduled computing node dynamically adjusts resources for the running capsule according to the resource load of the scheduled running capsule when the received task is running, thereby improving the certainty of the received task.
[0020] In a possible implementation of the first aspect, the running capsule of each computing node is located in the corresponding adaptive partition, and further includes: the elastic microkernel of each computing node schedules the CPU running time of the corresponding adaptive partition budget for the task in the running capsule of each adaptive partition of the computing node, wherein, when the CPU running time of an adaptive partition budget of the computing node is remaining, the elastic microkernel of the computing node schedules the remaining CPU running time to the tasks in the running capsules of other application partitions, and the maximum CPU running time used by the task in a scheduling main frame is the sum of the CPU running time budgeted by the adaptive partition where the task is located and the remaining CPU running time of other adaptive partitions occupied by it.
[0021] As described above, when there is surplus CPU runtime in the budget of an adaptive partition, the remaining CPU runtime is scheduled to tasks in the running capsules of other application partitions to improve the scheduling of high-priority tasks.
[0022] In one possible implementation of the first aspect, the task and its configuration are encapsulated as an image, and the container management and scheduling system utilizes the image to schedule the compute nodes and run capsules for the task according to a customized scheduling rule for the task. The scheduling rule is customized with reference to the various possible implementations of the first aspect described above.
[0023] As described above, by encapsulating the received task and its configuration into an image, the container management and scheduling system schedules the compute nodes and runs the capsule for the task, fully utilizing the task capabilities of the container management and scheduling system. The container management and scheduling system is shown as Kubernetes.
[0024] In one possible implementation of the first aspect, the container management and scheduling system deploys and runs the task on the scheduled computing node and performs lifecycle management of the task via the container engine of the scheduled computing node. The container engine implements lifecycle management of the task within the scheduled running capsule. The management rules are customized based on the management method of deterministic tasks.
[0025] As shown above, the container management and scheduling system implements task lifecycle management through the container engine on the compute node, fully leveraging the capabilities of the container management and scheduling system to also implement lifecycle management for embedded deterministic tasks. This container management and scheduling system is illustrated as Kubernetes.
[0026] In a second aspect, an embodiment of the present application provides a distributed deterministic task-based scheduling device for scheduling tasks on a cloud, wherein the cloud includes a management node and a computing node, and includes: a task receiving module for the management node to receive a task and its configuration from the cloud, wherein the configuration includes resource requirements that meet the determinism of the task; a task scheduling module for the management node to schedule the task to a computing node that meets the configuration based on the received task configuration and the resources of each computing node, wherein the configuration includes resource requirements that meet the deterministic constraints of the task, wherein the resource requirements include computing resource requirements and transmission resource requirements; a task scheduling module for the management node to schedule the task to a corresponding running capsule on the scheduled computing node based on the configuration, wherein the running capsule of each computing node is deployed on the elastic microkernel of the computing node, and the elastic microkernel of each computing node allocates resources to its running capsule. In some embodiments, the resources of the running capsule meet the configuration, and the running capsule supports one of the following scenarios: virtual machine, container, process, thread.
[0027] As described above, the configuration of the received task includes the resource requirements that meet the deterministic constraints of the task. According to the configuration of the received task and the resources of each computing node, the computing nodes, computing resources and transmission resources are scheduled for the task to realize the scheduling of deterministic tasks on the cloud.
[0028] In a possible implementation of the second aspect, the task is decomposed into several subtasks, and the configuration is decomposed into several subconfigurations, each subconfiguration meeting the deterministic resource requirements of a subtask; when the task scheduling module schedules the task to the computing node that meets the configuration according to the configuration of the received task and the resources of each computing node, it is specifically used to schedule each subtask to the computing node that meets the subconfiguration according to the subconfiguration of the subtask and the resources of each computing node, and the transmission delay between the scheduled computing nodes meets the deterministic requirements of data flow transmission between subtasks.
[0029] From the above, when a task is scheduled to multiple computing nodes, the transmission delay between the scheduled computing nodes satisfies the deterministic requirements of data flow transmission between subtasks, thereby improving the determinism of task implementation.
[0030] In a possible implementation of the second aspect, when there are multiple computing nodes that meet the configuration, the task scheduling module schedules computing nodes for the task based on load balancing and / or shortest delay.
[0031] As described above, the computing nodes are scheduled for the received tasks among a plurality of computing nodes that meet resource requirements based on load balancing and / or delay, so as to improve the time certainty of the received tasks.
[0032] In a possible implementation of the second aspect, the configuration further includes a priority of the task, and the priority of the running capsule matches the priority of the task.
[0033] As described above, by matching the priority of the running capsule with the priority of the received task, the time certainty of the received task is improved.
[0034] In a possible implementation of the second aspect, the task scheduling module obtains the configuration according to a deterministic constraint of the received task.
[0035] As described above, by obtaining the task configuration according to the deterministic constraint of the received task, the running capsule scheduled based on the configuration runs the task, thereby realizing the deterministic constraint of the task.
[0036] In a possible implementation of the second aspect, the system further includes: a management module, configured for the management node to deploy the operating environment of the task and install the task on the scheduled computing node, and perform lifecycle management of the task, including at least monitoring the resource load of the scheduled computing node and the resource load of the scheduled running capsule.
[0037] As described above, by performing lifecycle management on the received tasks on the scheduled computing nodes, not only the highly reliable operation of the tasks is achieved, but also the resource utilization of the computing nodes is improved.
[0038] In a possible implementation of the second aspect, the task scheduling module is further used to reduce the proportion of subsequent scheduling of the computing node when the resource load of the scheduled computing node is higher than the first set threshold; and reduce the proportion of subsequent scheduling of the running capsule when the resource load of the scheduled running capsule is higher than the second set threshold.
[0039] As described above, by reducing the subsequent scheduling ratio of the scheduled computing nodes and running capsules, the resources of the scheduled computing nodes and running capsules can be more used for the received tasks, thereby improving their determinism. In one possible implementation of the second aspect, the elastic microkernel of the scheduled computing nodes dynamically adjusts resources for the running capsules based on the resource load of the running capsules when the tasks are executed.
[0040] As described above, the elastic microkernel of the scheduled computing node dynamically adjusts resources for the running capsule according to the resource load of the scheduled running capsule when the received task is running, thereby improving the certainty of the received task.
[0041] In a possible implementation of the second aspect, the running capsule of each computing node is located in the corresponding adaptive partition, and further includes: a resource scheduling module, which is used for the elastic microkernel of each computing node to schedule the CPU running time of the corresponding adaptive partition budget for the tasks in the running capsules of each adaptive partition of the computing node, wherein, when the CPU running time budgeted by an adaptive partition of the computing node is surplus, the elastic microkernel of the computing node schedules the remaining CPU running time to the tasks in the running capsules of other application partitions, and the maximum CPU running time used by the task in a scheduling main frame is the sum of the CPU running time budgeted by the adaptive partition where the task is located and the remaining CPU running time of other adaptive partitions occupied by it.
[0042] As described above, when there is surplus CPU runtime in the budget of an adaptive partition, the remaining CPU runtime is scheduled to tasks in the running capsules of other application partitions to improve the scheduling of high-priority tasks.
[0043] In one possible implementation of the second aspect, the task scheduling module encapsulates the task and its configuration into an image. The container management and scheduling system schedules compute nodes and runtime capsules for the task based on the image. The container management and scheduling system uses the image to perform lifecycle management of the task based on customized scheduling rules for the task. The scheduling rules are customized with reference to the various possible implementations of the first aspect. Kubernetes is illustratively used as the container management and scheduling system.
[0044] As described above, by encapsulating the received task and its configuration into an image, the container management and scheduling system schedules the compute node and runtime capsule for the task, fully leveraging the capabilities of the container management and scheduling system.
[0045] In one possible implementation of the second aspect, the container management and scheduling system manages the lifecycle of the task on the scheduled compute node and runs the task. The management module uses the container engine of the scheduled compute node to manage the lifecycle of the task within the scheduled running capsule. The container engine implements lifecycle management of the task within the scheduled running capsule. The management rules are customized based on the management method of deterministic tasks.
[0046] As described above, the container management and scheduling system implements task lifecycle management through the container engine on the computing node, fully utilizing the capabilities of the container management and scheduling system, and also implements lifecycle management of embedded deterministic tasks.
[0047] In a third aspect, an embodiment of the present application provides a distributed deterministic task-based scheduling system for scheduling tasks on a cloud, wherein the cloud includes a management node and several computing nodes, and includes: a deterministic task scheduling framework for scheduling received deterministic tasks according to custom rules based on a task scheduling and execution framework of a container management and scheduling system, wherein the custom rules include any of the implementations described in the first aspect of the present application. The container management and scheduling system is exemplified as Kubernetes.
[0048] In a possible implementation of the second aspect, it also includes: a deterministic resource management scheduling management framework, which is used to deploy and run the task on the scheduled computing node based on the resource management framework of the container management and scheduling system, and perform lifecycle management of the task.
[0049] In a possible implementation of the second aspect, the deterministic task scheduling framework includes: a task interface, a scheduler and a resource controller; the task interface is used by the management node to receive tasks from the cloud; the scheduler is used by the management node to perform the scheduling; the resource controller is used by the management node to check whether the configuration of the received task is correct and whether the resources of each computing node meet the configuration.
[0050] In a fourth aspect, an embodiment of the present application provides a computing device, comprising:
[0051] bus;
[0052] a communication interface connected to the bus;
[0053] at least one processor connected to the bus; and
[0054] At least one memory is connected to the bus and stores program instructions, and when the program instructions are executed by the at least one processor, the at least one processor executes any implementation method of the first aspect of the present application.
[0055] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a computer, causes the computer to execute any of the implementations described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 This is a schematic diagram of the structure of a new industrial operating system for this application;
[0057] Figure 2 This is a schematic diagram of a first embodiment of a distributed deterministic task-based scheduling method of the present application;
[0058] Figure 3 This is a schematic diagram of a second embodiment of a distributed deterministic task-based scheduling method of the present application;
[0059] Figure 4 This is a structural diagram of an embodiment of a distributed deterministic task-based scheduling device of the present application;
[0060] Figure 5 This is a structural diagram of an embodiment of a distributed deterministic task-based scheduling system of the present application;
[0061] Figure 6 A schematic diagram of the structure of a computing device according to various embodiments of the present application. DETAILED DESCRIPTION
[0062] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0063] In the following description, the terms "first\second\third, etc." or module A, module B, module C, etc. are only used to distinguish similar objects, or to distinguish different embodiments, and do not represent a specific ordering of the objects. It can be understood that the specific order or sequence can be interchanged where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0064] In the following description, the numbers representing the steps, such as S110, S120, etc., do not necessarily mean that the steps must be executed in this manner. If permitted, the order of the steps can be interchanged or they can be executed simultaneously.
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0066] Embodiments of the present application provide a distributed deterministic task-based scheduling method, device, and system for scheduling tasks on a cloud, wherein the cloud includes a management node and a computing node, including: the management node schedules the task to a computing node that meets the configuration of the received task and the resources of each computing node according to the configuration, wherein the configuration includes resource requirements that meet the deterministic constraints of the task, and the resource requirements include computing resource requirements and transmission resource requirements; the management node schedules the task to a corresponding running capsule on the scheduled computing node according to the configuration, wherein the resources of the running capsule meet the configuration, and the running capsule supports one of the following scenarios: virtual machines, containers, processes, and threads, and the running capsule of each computing node is deployed on the elastic microkernel of the computing node, and the elastic microkernel of each computing node allocates resources to its running capsule.
[0067] The technical solution of the embodiment of the present application schedules a computing node for the task based on the configuration of the received task and the resources of each computing node, and the computing node schedules computing resources for the task, thereby realizing the scheduling of deterministic tasks on the cloud and supporting deterministic scheduling of tasks in various scenarios such as virtual machines, containers, processes, and threads.
[0068] The following describes various embodiments of the present application in conjunction with the accompanying drawings. First, the scenarios in which the embodiments of the present application are used are described.
[0069] The embodiments of this application are used for task scheduling in industrial cloud scenarios. The industrial cloud includes several regional sub-clouds (also called edge clouds). Each sub-cloud includes several physical nodes. The physical nodes include management nodes and several computing nodes. The industrial cloud of this application is managed by a new industrial operating system. Figure 1 Introducing a new industrial operating system. Figure 1 The structure of a novel industrial operating system of the present application is shown, which includes, from bottom to top: a base layer, a platform layer, and a service layer. The embodiments of the present application are used to run on the platform layer.
[0070] The base layer is deployed on each physical node in the industrial cloud and includes the node's elastic microkernel and several runtime capsules. The elastic microkernel allocates hardware resources to the runtime capsules. One possible implementation of the base layer is in the new Rust language, leveraging its features to provide memory safety. The system utilizes a standardized functional component design approach, enabling modular assembly. This allows for flexible, on-demand integration of advanced features to support virtualization and provide enhanced security isolation.
[0071] Among them, the elastic microkernel includes componentized hardware resources, which are used to allocate componentized hardware resources to each running capsule. The elastic microkernel manages the hardware resources of the physical node in a componentized manner, and each componentized hardware resource is a standardized resource component. The elastic microkernel elastically loads and / or deletes each resource component according to demand to achieve elastic management of the hardware resources on the physical node. The resources include the CPU core on the chip, the motherboard and / or chip memory, the physical node peripherals, etc. The elastic microkernel is a super-elastic microkernel that can randomly and quickly combine the size of chip resources according to demand. The elastic microkernel is also used to statically and / or dynamically allocate componentized hardware resources to the running capsules on the physical node, realizing elastic allocation of resource components.
[0072] Among them, the running capsule supports one of the following scenarios: thread, process, real-time container, non-real-time container, real-time virtual machine, non-real-time virtual machine. The running capsule includes the corresponding running environment for running tasks in the service components of the new industrial operating system.
[0073] For example, Figure 1 Two partitioned virtual machines and one non-partitioned container are shown in the figure. The real-time running environment of one partitioned virtual machine supports real-time application tasks, the high-security running environment of one partitioned virtual machine supports high-security application tasks, and the non-real-time running environment in the non-partitioned container supports the application tasks of the non-real-time container.
[0074] Run capsules can also be divided into real-time run capsules and non-real-time run capsules, which are used to run real-time tasks and non-real-time tasks respectively. Real-time run capsules are time-critical run capsules.
[0075] The platform layer is used to schedule running capsules with matching capabilities on several physical nodes for each task in the service component of the new industrial operating system from the industrial cloud, so that at least the predicted delay of the service component obtained based on the computational delay of the scheduled running capsules and the transmission delay between the scheduled physical nodes meets the deterministic constraints of the service component.
[0076] In a possible implementation of the platform layer, the platform layer includes an orchestrator, a scheduler, and middleware.
[0077] Among them, the coordinator is deployed on the physical node used for management of the industrial cloud in each region, and is used to schedule several physical nodes from the industrial cloud for each task in the service components of the new industrial operating system.
[0078] The orchestrator is deployed regionally, on physical nodes managed within the central and / or edge clouds of each region's industrial cloud. It schedules multiple physical nodes from the industrial cloud for tasks within the new industrial operating system's service components. The orchestrator operates in a hierarchical manner, providing cross-system and cross-regional device coordination and centralized control, real-time monitoring, and remote collaboration.
[0079] For example, the industrial cloud includes a sub-cloud in region A and a sub-cloud in region B, and each sub-cloud includes a central cloud and an edge cloud. A coordinator is deployed on the management node of the central cloud of the sub-cloud in region A and the sub-cloud in region B respectively. When the running capsules of the physical nodes in the sub-cloud in region A and the sub-cloud in region B can satisfy the sub-tasks in the scheduled service components, the collaborative scheduling of these coordinators is based on load balancing to schedule several physical nodes from a regional sub-cloud for each task in the service component; when the sub-task in the scheduled service component is used to control an industrial actuator or sensor in an edge cloud in a region, a physical node is scheduled from the edge cloud in the region based on the shortest predicted delay. The scheduling method distinguishing these two scenarios realizes the cross-regional cloud-edge collaborative scheduling.
[0080] The scheduler is deployed on the physical node for management of the edge cloud in each region. Figure 6 The central distributed deterministic scheduler is the scheduler, and the edge distributed platform is the physical node used for management.
[0081] The scheduler uses an end-to-end task delay analysis algorithm to calculate the worst-case delay constraints of tasks in a multi-level dynamic scheduling framework. It adopts real-time scheduling analysis technology for multi-core systems, establishes a time-predictable basic structure model of multi-core processors, and performs scheduling analysis based on it, thereby fully utilizing the powerful parallel computing capabilities provided by multi-core processors.
[0082] The scheduler schedules running capsules that match the capacity of each task in each service component from the physical nodes scheduled by the orchestrator, and predicts the predicted delay of running the service component based on the computational delay of the scheduled running capsules and the transmission delay between the scheduled physical nodes. The predicted delay at least meets the deterministic constraints of the scheduled service component.
[0083] For example, the scheduler schedules three serial tasks A1, B1, and C1 of a service component to run capsules A2, B2, and C2, respectively. The physical nodes where run capsules A2, B2, and C2 reside are connected via a TSN network. Based on the computing power of run capsules A2, B2, and C2, the scheduler uses a preset model to predict the completion times t1, t2, and t3, respectively. It also predicts the transmission delay p1 from run capsule A2 to run capsule B2, and the transmission delay p2 from run capsule B2 to run capsule C2. The estimated delay for completing the service component is (t1+t2+t3+p1+p2). If this estimated delay is less than the deterministic constraint of the service component, the service component can complete the deterministic computation.
[0084] The scheduler also ensures that the running capsules scheduled by the scheduler for each service component meet the security constraints of the service component, thereby making the new industrial operating system a secure system. The scheduler also ensures that the running capsules scheduled by the scheduler for each service component meet the trust constraints of the service component, thereby making the new industrial operating system a trustworthy system.
[0085] When the scheduler cannot schedule a running capsule that meets the requirements of the task in the service component from the existing running capsules on each physical node, the scheduler is also used to schedule a physical node to dynamically create a running capsule that meets the requirements; the scheduler also schedules a running capsule with the closest capability to a physical node, and then notifies the elastic microkernel of the physical node to dynamically increase hardware resources for the running capsule, so that it becomes a running capsule that meets the requirements of the task in the scheduled service component.
[0086] Middleware manages communication between physical nodes, stores data, and supports artificial intelligence engines. These middleware runs on runtime capsules on physical nodes dedicated to their specific functions and can be scheduled statically or dynamically. Together with the runtime capsules dispatched by the scheduler, the middleware helps the service components of the new industrial operating system complete their functions.
[0087] In another possible implementation of the platform layer, the platform layer further includes a manager.
[0088] The manager manages the lifecycle of tasks in the running capsules in the related physical nodes, including image management, storage management, network management (including TSN network management), event management, and interface management of the applications in the running capsules. This is similar to the abstract management of container applications, making the deployment of each running capsule independent of the hardware of the specific physical node, thus achieving rapid deployment and highly reliable operation and maintenance management of the running capsules.
[0089] The platform layer provides distributed collaboration and deterministic control capabilities based on a distributed coordination framework and deterministic scheduling. This includes: elastic resource allocation and scheduling based on the base layer, distributed deterministic communication capabilities, and standardized consistency protocols and mechanisms. This addresses issues such as real-time determinism and data consistency in distributed control and computing. Furthermore, the platform layer provides security and privacy protection mechanisms that are applied to the collaboration process to ensure data security and confidentiality. Finally, protocol definition and standardization are key to ensuring collaboration and control across different layers. Through unified protocol specifications, seamless integration and communication between different devices, edge platforms, and cloud platforms can be achieved.
[0090] The platform layer also provides ubiquitous industrial connectivity, addressing interoperability issues across diverse industrial devices, and software-defined control, addressing the need for flexible, on-demand deployment of control systems. The platform layer also rapidly detects when industrial control terminals connect to the network. Combining platform-layer collaborative control with deterministic scheduling and communication, it enables rapid switching of device control across regions.
[0091] The service layer includes the Industrial Application Cloud Development Kit, deployed on physical nodes within the Industrial Cloud with an integrated development environment. This kit is used to develop service components that form various service suites. The Industrial Application Cloud Development Kit decomposes each service suite into several service components. Each service component is open to users and broken down into tasks that can be subscribed to. Service component tasks are typically dispatched by the platform layer to runtime capsules for execution.
[0092] The service layer also deploys industrial control suites and industrial simulation cloud platforms. The industrial control suites include time-critical industrial control components. The tasks of the industrial control components are scheduled by the platform layer to real-time running capsules (such as partition-based real-time containers or real-time virtual machines). The service layer is also used to start related service components in the industrial control suite to control and access industrial actuators or sensors connected to the edge cloud in the industrial cloud through the platform layer and the base layer to complete industrial control.
[0093] The following combination Figure 2 The present invention introduces a first embodiment of a distributed deterministic task-based scheduling method.
[0094] Figure 2 The flowchart of a first embodiment of a distributed scheduling method based on deterministic tasks is shown, including steps S210 to S230.
[0095] S210: The management node receives tasks from the service layer of the industrial cloud.
[0096] Among them, the management node and the computing node are both physical nodes, and the management node implements the functions of the platform layer.
[0097] Among them, each service component of the service layer generates tasks, including industrial control components, industrial artificial intelligence components, etc.
[0098] In some embodiments, the received task includes a configuration of the task, and the configuration includes hardware resource requirements for running the task, so that the management node performs deterministic scheduling for the received task.
[0099] In some embodiments, a received task includes deterministic constraints for the task, including latency constraints. The management node converts the deterministic constraints into the configuration of the task, i.e., into the resource requirements of the task, to facilitate deterministic scheduling of the received task by the management node. The resource requirements include computing resource requirements and transmission resource requirements. The received task is an actual task in a new industrial operating system, and its task determinism is planned in advance, including its latency constraints. The computing power and transmission requirements of the latency constraints are also planned in advance. The computing resource requirements and transmission resource requirements are obtained based on the computing power and transmission requirements. The transmission resource requirements are the transmission resources of the deterministic network used, such as the bandwidth and time slice of the TSN network.
[0100] In some embodiments, tasks from the service layer are decomposed into finer-grained tasks to be scheduled to running capsules of different types and / or different scenarios at the base layer.
[0101] S220: The management node schedules the received task to the corresponding computing node according to the configuration of the received task and the resources of each computing node.
[0102] The management node checks the configuration of the received tasks and schedules the tasks that pass the check.
[0103] The management node also checks whether the resources of each computing node on the cloud meet the configuration of the received task and the task's deterministic requirements. It also checks whether the transmission resources of each computing node meet the requirements. It schedules computing nodes from those that meet the requirements to achieve deterministic computing.
[0104] The scheduling is performed by defining a scheduling rule, which includes: the resources of the scheduled computing nodes and the transmission resources meet the requirements.
[0105] In some embodiments, the scheduling rules also include: when there are multiple computing nodes that meet the configuration of the received task, the management node schedules the computing nodes for the received task based on load balancing and / or shortest latency, which not only meets the determinism of the task but also achieves the utilization of the computing nodes.
[0106] In some embodiments, the received task is decomposed into several subtasks, and its configuration is decomposed into several subconfigurations, each subconfiguration meets the deterministic resource requirements of a subtask; the scheduling rules also include: scheduling each subtask to a computing node that meets the corresponding subconfiguration based on the subconfiguration of the subtask and the resources of each computing node, and the transmission delay between the scheduled computing nodes meets the deterministic requirements of data flow transmission between subtasks.
[0107] In some embodiments, when a management node manages compute nodes via Kubernetes, it represents the received task configuration as a container dependency and composes the container image with the task. The node then schedules the compute nodes using the Volcano component scheduling framework in a containerized manner according to customized rules. This leverages Kubernetes' Volcano functionality to implement the parallel computing scheduling framework's capabilities in non-containerized scenarios. Customized rules include the methods described in this step.
[0108] In some embodiments, the scheduling rules further include: scheduling different types of computing nodes based on the type of received tasks, with each type of computing node deploying an execution capsule with related components. For example, AI tasks are scheduled to computing nodes with deployed AI operators, which are pre-packaged with AI components.
[0109] In some embodiments, this step is run in the corresponding task of the management node on the orchestrator of the platform layer.
[0110] S230: The management node schedules a corresponding running capsule for the task on the scheduled computing node according to the configuration of the received task.
[0111] The run capsule includes resources for the following scenarios: virtual machines, containers, processes, and threads. It supports running tasks in virtual machines, containers, processes, and threads.
[0112] The scheduling of this step continues according to the defined scheduling rules. The scheduling rules also include: the resources of the scheduled computing nodes and transmission resources meet the requirements. The computing power resources and transmission resources used by the running capsule meet the configuration requirements of the received task. In some embodiments, this step is executed in the corresponding task of the management node's scheduler at the platform layer.
[0113] In some embodiments, the received task configuration further includes a priority of the task; and the scheduling rule further includes: when the priority of the scheduled running capsule meets the task configuration, the certainty of the task is improved through the priority.
[0114] In some embodiments, the configuration of the received task also includes a time determinism requirement for the task. When the received subtask is scheduled into multiple running capsules, the delay determined by the running time of each running capsule for the subtask and the transmission delay between each running capsule must meet the time determinism requirement of the task.
[0115] In some embodiments, the management node deploys the task's runtime environment and installs the task's application on the scheduled computing node, starts the task, and performs lifecycle management on the task. This lifecycle management runs on the scheduling manager at the platform layer.
[0116] In some embodiments, when the management node manages each computing node through Kubernetes, the management node represents the received tasks and configurations as containers. The management node deploys the task's operating environment and installation tasks on the scheduled computing nodes in a container image format through the Kubernetes management framework. Its management rules are customized rules based on the management of the determined tasks, and include: performing lifecycle management on the task through the container engine of the scheduled computing node, and the container engine implements the lifecycle management of the task in the scheduled running capsule, thereby using Kubernetes to implement lifecycle management in non-container scenarios. Lifecycle management includes at least monitoring the resource load of the scheduled computing node and the resource load of the scheduled running capsule.
[0117] In some embodiments, the scheduling rules also include: when the resource load of the scheduled computing node is higher than a first set threshold, reducing the proportion of subsequent scheduling of the computing node; when the resource load of the scheduled running capsule is higher than a second set threshold, reducing the proportion of subsequent scheduling of the running capsule to reduce the resource load of the scheduled computing node or the resource load of the scheduled running capsule.
[0118] In some embodiments, the elastic microkernel of the scheduled computing node dynamically adjusts resources for the running capsule according to the resource load of the scheduled running capsule when the received task is executed.
[0119] In some embodiments, when an adaptive partition of a scheduled compute node has excess CPU runtime budget, the elastic microkernel of the compute node schedules the remaining CPU runtime to the task received for this embodiment. Each compute node's runtime capsule is located in a corresponding adaptive partition, each adaptive partition being a resource combination. The elastic microkernel of each compute node schedules the CPU runtime budget of the corresponding adaptive partition for the task in the runtime capsule of each adaptive partition of the compute node.
[0120] In the prior art, the microkernel schedules resources for execution capsules and resources for tasks within the execution capsules. After the microkernel allocates a budgeted CPU runtime to the execution capsule, the tasks within the execution capsule can only use that budgeted CPU runtime. In this embodiment, the elastic microkernel of each compute node schedules the adaptive partition CPU runtime for the tasks within the execution capsules of each adaptive partition of the compute node. If the budgeted CPU runtime of an adaptive partition of the compute node is surplus, the elastic microkernel of the compute node schedules the surplus CPU runtime to high-priority tasks within execution capsules of other application partitions. The maximum CPU runtime used by the task in a scheduled main frame is the sum of the budgeted CPU runtime of the adaptive partition in which the task resides and the remaining CPU runtime of the other adaptive partitions occupied by the task, thereby improving the determinism of the task. In this embodiment, after the priority of a task received is set to high and all other tasks with higher priority than the task on the scheduled compute node have been scheduled, if the remaining CPU runtime of an adaptive partition of the scheduled compute node is surplus, the surplus CPU runtime can be scheduled to the task received in this embodiment, thereby improving its determinism.
[0121] In summary, embodiment 1 of a distributed scheduling method based on deterministic tasks schedules a computing node for the task according to the configuration of the received task and the resources of each computing node, and the computing node schedules computing resources for the task, thereby realizing the scheduling of deterministic tasks on the cloud and supporting deterministic scheduling of tasks in various scenarios such as virtual machines, containers, processes, and threads.
[0122] A second embodiment of a distributed deterministic task scheduling method is a specific implementation of the first embodiment, offering all of its advantages. Specifically, at the platform layer of a novel industrial operating system, this method implements unified management and scheduling of deterministic resources through a container management and scheduling system, and implements deterministic task allocation and execution through the container management and scheduling system's batch task scheduling and execution framework. This embodiment is described below using Kubernetes as the container management and scheduling system.
[0123] In the second embodiment of a distributed scheduling method based on deterministic tasks, the following execution framework is first defined in Kubernetes.
[0124] The deterministic-resource dispatch management framework (DDMF) corresponds to the platform-layer manager and is used to manage the lifecycle of deterministic tasks in various scenarios. DDMF calls the Kubernetes resource management and dispatch management framework, and its management rules are customized based on deterministic resource management, including load monitoring of each running capsule.
[0125] Define a deterministic job dispatch management framework (JDMF) for scheduling deterministic tasks in various scenarios. This framework implements platform-level coordination and scheduling for new industrial operating systems, corresponding to the platform-level orchestrator and scheduler. JDMF leverages the Kubernetes Volcano component's task scheduling framework, and its scheduling rules are based on deterministic resource management scheduling definitions.
[0126] The scheduler in the deterministic task scheduling framework is defined as JDMF-SC (scheduler), which is used to schedule deterministic tasks in various scenarios. JDMF-SC calls the task scheduler framework of the Kubernetes Volcano component. Its scheduling rules are customized based on deterministic resource management scheduling, referring to the various scheduling methods in steps S220 and S230 of the first embodiment of a distributed scheduling method based on deterministic tasks.
[0127] The task interface in the deterministic task scheduling framework is defined as JDMF-CLI (Command Line Interface), which is used to receive deterministic tasks in various scenarios.
[0128] Defines the JDMF-CM (controller manager) resource management controller in the deterministic task scheduling framework, which is used to manage resources for deterministic tasks in various scenarios. JDMF-CM calls the Kubernetes resource management controller, and its control rules are customized based on deterministic resource management.
[0129] Figure 3 The flowchart of a second embodiment of a distributed scheduling method based on deterministic tasks is shown, including steps S310 to S380.
[0130] S310: Receive tasks from the service layer through JDMF CLI.
[0131] For the definition of the received task, please refer to the description of the first embodiment of a distributed scheduling method based on deterministic tasks.
[0132] Among them, JDMFCLI decomposes the received tasks into volcano jobs. One task is decomposed into multiple jobs, and the task configuration is decomposed into the configuration of multiple jobs, including resource requirements and priorities. Resource requirements include computing power resource requirements and transmission resource requirements.
[0133] S320: Send the job to JDMF-CM through JDMF CLI.
[0134] The JDMF CLI releases multiple jobs of a receiving task to the JDMF-CM in parallel.
[0135] S330: Check the job configuration and computing node cluster resources through JDMF-CM.
[0136] Among them, check whether the job configuration is correct and whether it meets the scheduling parameter requirements of Volcano.
[0137] Among them, JDMF-CM's control rules are customized based on deterministic resource management control, checking whether the computing resources and transmission resources of the computing node group meet the requirements, including computing power type (CPU and / or GPU, etc.), computing power size, transmission resource type and latency.
[0138] When the configuration of the job is correct and the computing node group resources meet the requirements, step S340 is executed.
[0139] S340: Send a job scheduling request to JDMF-SC through JDMF-CM.
[0140] Among them, JDMF-CM sends multiple job scheduling requests of a receiving task to JDMF-SC in parallel, so that JDMF-SC can perform parallel scheduling.
[0141] S350: Execute scheduling through JDMF-SC.
[0142] JDMF-SC's scheduling rules are based on deterministic resource management scheduling customization. Jobs and their configurations are encapsulated as images. JDMF-SC leverages Volcano's parallel scheduling architecture to schedule each job to the corresponding compute node for execution based on the compute node cluster's resource availability and job configuration. The scheduling rules include:
[0143] JDMF-SC selects the appropriate compute node to execute the job based on the resource usage of the compute node cluster.
[0144] Among them, when there are multiple computing nodes that meet the configuration of a job, JDMF-SC schedules the computing nodes for the job based on load balancing and / or minimum latency.
[0145] When JDMF-SC schedules the computing nodes, it also schedules a computing node from the multiple computing nodes whose overall priority of running jobs is lower than the priority of the received job.
[0146] When JDMF-SC schedules compute nodes, it assigns them to different types of compute nodes based on the type of job received. Each type of compute node deploys a runtime capsule containing relevant components. For example, AI tasks are scheduled to compute nodes with deployed AI operators, which are pre-packaged with AI components.
[0147] Among them, when each job is scheduled to different computing nodes, the transmission delay between computing nodes meets the data transmission requirements of the job.
[0148] JDMF-SC allocates resources to each job on the scheduled compute nodes based on the job configuration.
[0149] The resource is a running capsule on a computing node, which can include resources in scenarios such as virtual machines, containers, processes, and threads. Each job can be a task in a virtual machine, container, process, or thread.
[0150] Among them, the scheduling rules also include the resources for running capsules and the transmission resources they use to meet the resource requirements of the job.
[0151] S360: Execute the job on the scheduled computing resources through DDMF and return the execution result to JDMF-CM.
[0152] Among them, each business and configuration is represented as a container, and the operating environment and installation operations of the job are deployed on the scheduled computing node in the form of container images. The lifecycle management of the job is performed based on the container management method, thereby using Kubernetes to implement lifecycle management in non-container scenarios.
[0153] The elastic microkernel of the scheduled computing node dynamically adjusts resources for the running capsule according to the resource load of the scheduled running capsule when the received task is running.
[0154] DDMF also monitors the load of scheduled computing nodes and running capsules. The scheduling rules also include: when the resource load of a scheduled computing node is higher than a first set threshold, the proportion of subsequent scheduling of the computing node is reduced; when the resource load of a scheduled running capsule is higher than a second set threshold, the proportion of subsequent scheduling of the running capsule is reduced.
[0155] Each compute node's runtime capsule is located in a corresponding adaptive partition, each of which is a resource combination. The elastic microkernel of each compute node schedules the CPU runtime budgeted by the corresponding adaptive partition for the tasks in the runtime capsules of each adaptive partition of the compute node. In this embodiment, the priority of the tasks received is set to the highest priority. When the CPU runtime budget of an adaptive partition of the scheduled compute node remains, the elastic microkernel of the compute node schedules the remaining CPU runtime to the task received in this embodiment. The maximum CPU runtime used by the task in a scheduling main frame is the sum of the CPU runtime budget of the adaptive partition in which the task resides and the remaining CPU runtime of other adaptive partitions occupied by the task, thereby improving the determinism of the task.
[0156] S370: Feedback the execution result to the JDMF CLI through the JDMF-CM.
[0157] JDMF-CM feeds back the execution results of each job to JDMF CLI in parallel.
[0158] S380: Feedback the task execution results to the application layer through the JDMF CLI.
[0159] The JDMF CLI combines the execution results of each job into the execution result of the received task.
[0160] In summary, embodiment 2 of a distributed deterministic task-based scheduling method schedules and executes tasks that require deterministic execution through an execution framework defined in Kubernetes. For example, in the field of flight control, when an attitude change task is submitted, its execution result should be deterministic through the execution framework of this embodiment.
[0161] The following combination Figure 4 An embodiment of a distributed scheduling device based on deterministic tasks is introduced.
[0162] A distributed deterministic task-based scheduling device embodiment is used to run a distributed deterministic task-based scheduling method embodiment 1, with all its advantages.
[0163] Figure 4 The structure of an embodiment of a distributed deterministic task-based scheduling device is shown, which is located on a management node and runs in a corresponding platform layer module deployed on the management node, including: a task receiving module 410 and a task scheduling module 420.
[0164] The task receiving module 410 is used to manage nodes receiving tasks from the service layer of the industrial cloud. For its working principle and advantages, please refer to step S210 of the first embodiment of a distributed scheduling method based on deterministic tasks.
[0165] Task scheduling module 420 is used by the management node to schedule received tasks to corresponding compute nodes based on the received task configuration and the resources of each compute node. It also schedules the corresponding running capsule for the task on the scheduled compute node based on the received task configuration. For details on its operating principles and advantages, please refer to steps S220 and S230 of Example 1 of a distributed deterministic task scheduling method.
[0166] In some embodiments, a resource scheduling module is also included, which is used to schedule the remaining CPU running time of an adaptive partition budget of the scheduled computing node to the task received for this embodiment when the remaining CPU running time is left.
[0167] The following combination Figure 5 An embodiment of a distributed deterministic task-based scheduling system is introduced.
[0168] A distributed deterministic task-based scheduling system for scheduling tasks on a cloud, wherein the cloud includes a management node and several compute nodes. The system includes a deterministic task scheduling framework (JDMF) for scheduling received deterministic tasks according to custom rules based on the scheduling and execution framework of a container management and scheduling system. The custom rules include the method described in Example 1 of a distributed deterministic task-based scheduling method of the present application. This example is described below using Kubernetes as the container management and scheduling system.
[0169] In some embodiments, a distributed deterministic task-based scheduling system also includes: a deterministic-resource dispatch management framework DDMF (deterministic-resource dispatch management framework), which is used to deploy and run the task on the scheduled computing node based on the resource management framework of the container management and scheduling system, and perform lifecycle management of the task.
[0170] In some embodiments, the deterministic task scheduling framework JDMF includes: a task interface JDMF-CLI (Command Line Interface), a scheduler JDMF-SC (scheduler) and a resource controller JDMF-CM (controller manager); the task interface is used by the management node to receive tasks from the cloud; the scheduler is used by the management node to perform the scheduling; the resource controller is used by the management node to check whether the configuration of the received task is correct and whether the resources of each computing node meet the configuration.
[0171] A detailed embodiment of a distributed deterministic task-based scheduling system is used to run the method described in embodiment 2 of a distributed deterministic task-based scheduling method, with all its advantages. Figure 5 The structure of a distributed deterministic task-based scheduling system embodiment in a detailed manner is shown, including: DDMF 510, JDMF 520, wherein JDMF 520 includes JDMF-SC 521, JDMF-CLI 522 and JDMF-CM 523.
[0172] DDMF 510 is a deterministic resource management scheduling management framework DDMF, which is used for lifecycle management of deterministic tasks in this scenario and runs step S360 of the second embodiment of a distributed scheduling method based on deterministic tasks.
[0173] JDMF 520 is a deterministic task scheduling framework (JDMF), which is used to schedule deterministic tasks in various scenarios, including:
[0174] JDMF-SC 521 determines the scheduler in the task scheduling framework, which is used for scheduling deterministic tasks in various scenarios (including platform layer coordinator and scheduler functions), and runs step S350 of the second embodiment of a distributed scheduling method based on deterministic tasks;
[0175] JDMF-CLI 522 is a task interface in the deterministic task scheduling framework, used for receiving deterministic tasks in various scenarios and running steps S310, S320 and S380 of Example 2 of a distributed scheduling method based on deterministic tasks;
[0176] JDMF-CM 523 is a resource management controller in a deterministic task scheduling framework, used for resource management of deterministic tasks in various scenarios, and executes steps S330, S340, and S370 of a second embodiment of a distributed scheduling method based on deterministic tasks.
[0177] The present application embodiment also provides a computing device, Figure 6Detailed introduction.
[0178] The computing device 600 includes a processor 610 , a memory 620 , a communication interface 630 , and a bus 640 .
[0179] It should be understood that the communication interface 630 in the computing device 600 shown in this figure can be used to communicate with other devices.
[0180] The processor 610 may be connected to a memory 620. The memory 620 may be used to store the program code and data. Therefore, the memory 620 may be a storage unit within the processor 610, an external storage unit independent of the processor 610, or a component including both a storage unit within the processor 610 and an external storage unit independent of the processor 610.
[0181] Optionally, computing device 600 may further include a bus 640. Memory 620 and communication interface 630 may be connected to processor 610 via bus 640. Bus 640 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. Bus 640 may be classified as an address bus, a data bus, a control bus, and the like. For ease of illustration, the figure uses only one line, but this does not imply that there is only one bus or only one type of bus.
[0182] It should be understood that in the embodiment of the present application, the processor 610 can adopt a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Alternatively, the processor 610 adopts one or more integrated circuits to execute relevant programs to implement the technical solutions provided in the embodiment of the present application.
[0183] The memory 620 may include a read-only memory and a random access memory, and provides instructions and data to the processor 610. A portion of the processor 610 may also include a non-volatile random access memory. For example, the processor 610 may also store information about the device type.
[0184] When the computing device 600 is running, the processor 610 executes the computer-executable instructions in the memory 620 to perform the operating steps of each method embodiment.
[0185] It should be understood that the computing device 600 according to the embodiment of the present application can correspond to the corresponding subject in executing the method according to each embodiment of the present application, and the above-mentioned and other operations and / or functions of each module in the computing device 600 are respectively for implementing the corresponding processes of each method of the present embodiment. For the sake of brevity, they will not be repeated here.
[0186] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0187] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0188] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0189] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0190] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0191] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0192] An embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, it is used to perform the operating steps of each method embodiment.
[0193] The computer storage medium of the embodiment of the present application can adopt any combination of one or more computer-readable media.Computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium.Computer-readable storage medium can be, for example, but not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof.More specific examples (non-exhaustive list) of computer-readable storage medium include, an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.In this document, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, a device or a device or used in combination with it.
[0194] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0195] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0196] The computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0197] Note that the above are only preferred embodiments of the present application and the technical principles employed. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of protection of the present application, all of which fall within the scope of protection of the present application.
Claims
1. A distributed deterministic task-based scheduling method, characterized in that: Used to schedule tasks on the cloud, which includes a management node and several computing nodes, including: The management node schedules the task to a computing node that meets the configuration according to the received task configuration and the resources of each computing node, wherein the configuration includes resource requirements that meet the deterministic constraints of the task, and the resource requirements include computing resource requirements and transmission resource requirements; The management node schedules the task to the corresponding running capsule on the scheduled computing node according to the configuration, and the resources of the running capsule meet the configuration. The running capsule of each computing node is deployed on the elastic microkernel of the computing node, and the elastic microkernel of each computing node allocates resources to its running capsule.
2. The method according to claim 1, characterized in that Also includes: The management node decomposes the task into a number of subtasks and decomposes the configuration into a number of subconfigurations, each subconfiguration meeting the deterministic resource requirements of a subtask; The step of scheduling the received task to a computing node that satisfies the configuration based on the configuration of the task and the resources of each computing node specifically includes: scheduling each subtask to a computing node that satisfies the subconfiguration based on the subconfiguration of the subtask and the resources of each computing node, wherein the transmission delay between the scheduled computing nodes satisfies the deterministic requirements of data flow transmission between the subtasks; and / or Scheduling the task into a corresponding running capsule on the scheduled computing node according to the configuration specifically includes: scheduling the corresponding subtask into a corresponding running capsule on the scheduled computing node according to the subconfiguration, wherein resources of the corresponding running capsule satisfy the subconfiguration of the corresponding subtask.
3. The method according to claim 1, characterized in that When there are multiple computing nodes that meet the configuration, the management node schedules computing nodes for the task based on load balancing and / or shortest latency.
4. The method according to claim 1, characterized in that The configuration further includes a priority of the task, and the priority of the running capsule matches the priority of the task.
5. The method according to claim 1, characterized in that: Also includes: The management node obtains the configuration according to the deterministic constraints of the received task.
6. The method according to claim 1, characterized in that Also includes: The management node deploys and runs the task on the scheduled computing node and performs life cycle management on the task.
7. The method according to claim 1, characterized in that: Also includes: When the resource load of the scheduled computing node is higher than a first set threshold, the management node reduces the proportion of subsequent scheduling of the computing node; When the resource load of the scheduled running capsule is higher than the second set threshold, the management node reduces the proportion of subsequent scheduling of the running capsule.
8. The method according to claim 1, characterized in that: The running capsule of each computing node is located in a corresponding adaptive partition, and the method further includes: The elastic microkernel of each computing node schedules the CPU runtime budget of the corresponding adaptive partition for the tasks in the running capsules of each adaptive partition of the computing node. When the CPU runtime budget of an adaptive partition of the computing node has surplus, the elastic microkernel of the computing node schedules the surplus CPU runtime to the tasks in the running capsules of other application partitions. The maximum CPU runtime used by the task in a scheduling main frame is the sum of the CPU runtime budget of the adaptive partition where the task is located and the remaining CPU runtime of other adaptive partitions occupied by it.
9. A distributed deterministic task-based scheduling device, characterized in that: Used to schedule tasks on the cloud, which includes management nodes and computing nodes, including: A task receiving module, configured for the management node to receive a task and its configuration from the cloud, wherein the configuration includes deterministic resource requirements to meet the task, and the resource requirements include computing resource requirements and transmission resource requirements; A task scheduling module is used for the management node to schedule the task to a computing node that meets the configuration of the received task and the resources of each computing node according to the configuration of the received task and the resources of each computing node, and to schedule the task to a corresponding running capsule on the scheduled computing node. The resources of the running capsule meet the configuration, and the running capsule supports one of the following scenarios: virtual machine, container, process, thread. The running capsule of each computing node is deployed on the elastic microkernel of the computing node, and the elastic microkernel of each computing node allocates resources to its running capsule.
10. A distributed deterministic task-based scheduling system, characterized in that: Used for task scheduling on the cloud, which includes a management node and several computing nodes, including: A deterministic task scheduling framework is used to schedule received deterministic tasks according to custom rules based on a task scheduling and execution framework of a container management and scheduling system, wherein the custom rules include the method described in any one of claims 1 to 8.
11. The system according to claim 10, characterized in that: Also includes: A deterministic resource management scheduling management framework is used to deploy and run the task on the scheduled computing node based on the resource management framework of the container management and scheduling system, and to perform life cycle management on the task.
12. The system according to claim 10, characterized in that: The deterministic task scheduling framework includes: a task interface, a scheduler, and a resource controller; the task interface is used for the management node to receive tasks from the cloud; The scheduler is used by the management node to perform the scheduling; The resource controller is used by the management node to check whether the configuration of the received task is correct and whether the resources of each computing node meet the configuration.
13. A computing device, characterized in that include, bus; a communication interface connected to the bus; at least one processor connected to the bus; as well as at least one memory connected to the bus and storing program instructions, wherein when the program instructions are executed by the at least one processor, the at least one processor executes the method according to any one of claims 1 to 8.
14. A computer-readable storage medium, characterized in that Program instructions are stored thereon, and when the program instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1 to 8.