Novel industrial operating system, device, and storage medium
By deploying elastic microkernels and operating capsules on the industrial cloud, combining platform layer scheduling and service layer development, the real-time and data sharing problems in cloud edge-end collaborative control technology are solved, real-time and certainty of the industrial control system are achieved, and high-end manufacturing needs are met.
Patent Information
- Application Number
- PCT/CN2024/143579
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-15
- Filing Date
- 2024-12-30
- Publication Date
- 2025-08-28
AI Technical Summary
The existing cloud-edge collaborative control technology has shortcomings in real-time, data sharing and consistency, and cannot meet the needs of future high-end manufacturing.
By deploying elastic microkernels and running capsules on the industrial cloud, the flexible use of hardware resources of each physical node is achieved, combined with the platform layer scheduling and service layer component development, it meets the time-critical calculation and transmission delay constraints of industrial control components, and realizes real-time and deterministic control.
It improves the real-time and certainty of industrial control systems, ensures that the calculation and transmission delay of time-critical industrial control components meets constraints, and achieves cross-regional coordination and data consistency.
Smart Images

Figure CN2024143579_28082025_PF_FP_ABST
Abstract
Description
A new industrial operating system, equipment and storage medium
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on February 19, 2024, application number 202410185400.8, and application name "A New Industrial Operating System, Equipment and Storage Medium".
[0002] This application also claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on March 15, 2024, with application number 2024103012076 and application name “A New Industrial Operating System, Equipment and Storage Medium”. Technical Field
[0003] The present application relates to the field of industrial control, and in particular to a new industrial operating system, equipment, and storage medium. Background Art
[0004] With the continuous advancement of technologies such as cloud computing, big data analysis and artificial intelligence, the industrial production system has become further flattened, evolving from traditional vertical architecture to cloud-edge architecture.
[0005] Current cloud computing control technologies for cloud-edge-end collaboration suffer from issues such as poor real-time control and are unable to meet the demands of future high-end manufacturing. To address the shortcomings of cloud-edge-end industrial control systems, a new industrial operating system (IOS) is needed to address these issues from an architectural perspective. This new IOS, developed based on industrial digitization, integrated network control, and cross-regional cloud-edge-end collaboration, aligns with industrial development trends, meets the connectivity and management needs of diverse devices, and ensures system determinism and real-time performance.
[0006] Furthermore, the current cross-regional cloud collaboration control technology of cloud computing still has problems such as poor real-time control, poor data sharing and poor consistency. Summary of the Invention
[0007] In view of this, the embodiments of the present application provide a new industrial operating system, equipment and storage medium, which realizes the elastic use of chip resources of each physical node through dynamic scheduling of running capsules of physical nodes on the industrial cloud, so that the calculation and transmission delays of time-critical industrial control components meet their delay constraints, and completes real-time and deterministic control in the industrial field.
[0008] In the first aspect, an embodiment of the present application provides a new industrial operating system, including: a base layer, a platform layer and a service layer; the base layer is deployed on each physical node on the industrial cloud, including an elastic microkernel and several running capsules, and the elastic microkernel is used to adaptively allocate hardware resources to the running capsules; each of the running capsules supports one or more of multiple running scenarios; the platform layer is used to schedule the running capsules with matching capabilities on several physical nodes for each task in each service component of the system from the industrial cloud, wherein the predicted delay of each service component satisfies the deterministic constraints of the service component, and the predicted delay is obtained based on the computational delay of the scheduled running capsules and the transmission delay between the scheduled physical nodes; the service layer is used to develop, deploy and start the service components, and the service components include time-critical industrial control components.
[0009] As described above, by scheduling tasks in the service components of the new industrial operating system to run in the running capsules of physical nodes on the industrial cloud, the hardware resources of each physical node on the industrial cloud can be elastically used, so that the calculation and transmission delays of time-critical industrial control components meet their delay constraints, and the real-time and deterministic control of the underlying industrial control field is completed.
[0010] In a possible implementation of the first aspect, the running capsules of each physical node are located in an adaptive partition of the physical node, and each adaptive partition is configured with a budgeted CPU running time. When the elastic microkernel allocates hardware resources to the running capsules on the physical node, it is specifically configured to schedule the budgeted CPU running time of each adaptive partition for tasks in the running capsules of the adaptive partition. When the CPU running time actually used by any adaptive partition is less than its budgeted CPU running time, the remaining CPU running time of the adaptive partition is allocated to other adaptive partitions with tasks waiting to be scheduled.
[0011] As described above, by scheduling tasks in capsules running on adaptive partitions on the elastic microkernel, the architecture supports allocating the remaining CPU runtime of an adaptive partition to other adaptive partitions with tasks waiting to be scheduled, thereby achieving dynamic scheduling of the CPU runtime of running capsules.
[0012] In a possible implementation of the first aspect, the elastic microkernel is specifically used to statically and / or dynamically allocate componentized hardware resources to the running capsules on the physical node; the elastic microkernel is also specifically used to elastically load and / or delete each componentized hardware resource according to demand.
[0013] As mentioned above, by componentizing the hardware resources of each physical node, the elastic microkernel achieves elastic and scalable loading of hardware resources, and dynamically schedules componentized chip resources for each task in the service component, fully utilizing the computing power of multi-core chips, which not only improves resource efficiency but also increases determinism.
[0014] In one possible implementation of the first aspect, the physical node is connected to a TSN network, and a CPU time slice scheduled by an elastic microkernel of the physical node for running a capsule is aligned in timing with a TSN time slice used by the running capsule. A TSN time slice is a time slice for scheduling a TSN data flow in a TSN switch.
[0015] As described above, the CPU time slices scheduled by the microkernel for the real-time running capsule are aligned in timing with the TSN time slices used by the running capsule in the TSN data stream, so that the data of each task in the real-time running capsule can be transmitted deterministically, realizing the real-time and deterministic performance of time-critical industrial control components.
[0016] In one possible implementation of the first aspect, an elastic microkernel of a physical node connected to a TSN network schedules each task running in a capsule running on the physical node within the CPU time slice based on a task priority of each task running in the capsule, where the task priority matches a service priority of a TSN data stream of the task. The TSN service priority is a priority defined by a Pri field in a VLAN tag in the TSN data stream.
[0017] As described above, by scheduling the task priority of each task running in the running capsule with TSN, the task priority is matched with the service priority of TSN, so that the time when each task running in the running capsule is scheduled and its data is transmitted is matched, further realizing the real-time and deterministic performance of time-critical industrial control components.
[0018] In a possible implementation of the first aspect, when the elastic microkernel allocates hardware resources to a running capsule on a physical node, it is specifically used to adaptively allocate hardware resources to the running capsule based on the latency requirement of the task running in the running capsule, so that the computing latency of the running environment provided by the running capsule is better than or equal to the latency requirement of the task, and the relationship between the computing latency and the allocated hardware resources is obtained based on a preset prediction model.
[0019] As described above, by allocating hardware resources to running capsules based on the latency requirements of their tasks, tasks in time-critical industrial control components can be calculated in real time, thus achieving determinism in industrial control component calculations.
[0020] In a possible implementation of the first aspect, the elastic microkernel manages the hardware resources of the physical node based on at least one of the following methods: accessing the shared area of the running capsule based on messages and service interruptions, accessing the critical area of the CPU running the capsule based on the mapping of the CPU critical area to the address on the memory, and accessing the cache of the CPU running the capsule based on cache partitioning.
[0021] As described above, through the above resource management method, lock-free access to shared memory of tasks in time-critical industrial control components, lock-free access to CPU critical areas, and fast access to CPU cache areas are achieved, thereby achieving deterministic calculation of industrial control components.
[0022] In a possible implementation of the first aspect, the running capsule of each physical node is encapsulated based on its capability, and the base layer is further configured to register the current running capsule with the platform layer, where the registration information includes at least the capability of the running capsule.
[0023] From the above, by encapsulating the running capsule based on its capabilities, the platform layer is isolated from the devices of the physical nodes when scheduling the running capsule for the task in the service component, so that the platform layer can schedule the running capsule for various types of physical nodes, even including the running capsules encapsulated by other systems according to this application.
[0024] In a possible implementation of the first aspect, the execution capsule supports one of the following execution scenarios: thread, process, real-time container, non-real-time container, real-time virtual machine, and non-real-time virtual machine.
[0025] As shown above, running capsules can be compatible with various tasks, including tasks that require threads, processes, real-time containers, non-real-time containers, real-time virtual machines, or non-real-time virtual machines, to meet the needs of various tasks and improve the efficiency of physical node resources.
[0026] In a possible implementation of the first aspect, when a physical node runs the task of the industrial control component to control a device connected to the industrial cloud via wireless, each node determines its deterministic control range of the device based on its control delay of the device; the platform layer is also used for switching the control function of the device to a neighboring node when the device moves to the edge of the deterministic control range of the node and is within the deterministic control range of a neighboring control node of the node.
[0027] From the above, by sensing the delay of the node to the controlled device, it is determined whether the controlled device has left the deterministic control range and switched to the deterministic control range of the neighboring node, thus realizing the switching of the deterministic control function of the controlled device.
[0028] In a possible implementation of the first aspect, when a physical node runs the task of the industrial control component to control a device connected to the industrial cloud wirelessly, when each node determines that the control delay of its controlled device is greater than the delay constraint of the device minus a first set value, the device is outside the deterministic control range of the node; when each node determines that the control delay of its controlled device is less than or equal to the delay constraint minus a second set value, the device is within the deterministic control range of the node, and the second set value is greater than or equal to the first set value.
[0029] Based on the above, the control delay of each node on the controlled device and the delay constraint of the device minus the set value can be used to accurately detect whether the controlled device has left the deterministic control range.
[0030] In one possible implementation of the first aspect, the platform layer includes a scheduler and a coordinator, and the industrial cloud includes a central cloud and an edge cloud. The coordinator is configured to schedule multiple physical nodes for each task in the service component based on the requirement of minimizing the predicted latency from the industrial cloud. The scheduler is configured to schedule the running capsules that match the capabilities of the service component from the multiple physical nodes. The industrial cloud includes a central cloud and an edge cloud.
[0031] From the above, based on the requirement of shortest predicted latency, several physical nodes are scheduled for each task in the service component to achieve collaboration among the central cloud, edge cloud and terminal devices in a region, and improve the real-time and deterministic performance of the service component.
[0032] In a possible implementation of the first aspect, when the industrial cloud includes sub-clouds in multiple regions, the orchestrator is further configured to allocate corresponding regions to corresponding service components based on load balancing, so that corresponding running capsules are scheduled from corresponding sub-clouds in the corresponding regions.
[0033] From the above, based on the load balancing requirements, cross-regional collaboration of central cloud, edge cloud and terminal devices is achieved, and the real-time and deterministic performance of service components is improved.
[0034] In a possible implementation of the first aspect, the platform layer further includes deterministic communication middleware, which is used to generate TSN communication strategy parameters for the real-time capsule and predict the transmission delay of the real-time capsule based on the generated TSN communication strategy parameters.
[0035] As described above, the deterministic data transmission of the real-time running capsule is achieved through the deterministic communication middleware.
[0036] In a possible implementation of the first aspect, when the physical nodes are connected through a TSN network, the deterministic communication middleware is further used to obtain the network topology and the actual delay between the physical nodes through a DDS+TSN method.
[0037] As described above, the network topology and the actual delay between physical nodes are obtained through the DDS+TSN method, and the transmission delay is accurately predicted, which improves the certainty of real-time capsule transmission.
[0038] In a possible implementation of the present application, the running capsule scheduled by the platform layer for each of the service components satisfies the security constraints and / or trust constraints of the service component.
[0039] From the above, the security and trustworthiness of the service components are achieved through the security and trustworthiness of the running capsules at the base layer, thereby achieving the security and trustworthiness of the entire industrial control system.
[0040] In a possible implementation of the first aspect, when a running capsule that meets the requirements of the task in the service component cannot be scheduled from the existing running capsules on each physical node, the platform layer is further used to schedule a physical node so that the physical node dynamically creates a running capsule that meets the requirements.
[0041] As described above, by dynamically creating a running capsule when the existing running capsule on the physical node does not meet the requirements of the task in the service component, the dynamic use of hardware resources on the physical node is achieved, which improves the efficiency of resources and the real-time performance of the service component.
[0042] In a possible implementation of the first aspect, the platform layer further includes a manager configured to perform lifecycle management on the running capsule, wherein the lifecycle management includes performing lifecycle management on the running tasks therein.
[0043] As described above, by managing the lifecycle of running capsules, efficient creation, deletion, and maintenance of running capsules is achieved, improving the efficiency of physical node resource utilization. For example, by abstracting running capsules based on containers, Kubernetes can be used to manage the lifecycle of running capsules across devices based on containers.
[0044] In one possible implementation of the first aspect, the platform layer further includes artificial intelligence middleware, which is used to implement an artificial intelligence computing engine; the scheduler is further used to schedule the corresponding running capsules for the artificial intelligence middleware; and the service layer further includes an artificial intelligence component, which is used to call the artificial intelligence middleware to perform industrial data analysis, and the analysis results are used by other service components. The analysis is performed based on the input and output of the service component.
[0045] From the above, by setting up artificial intelligence components to perform various artificial intelligence data analyses, the analysis results can be shared across regions and devices. By setting up artificial intelligence middleware, artificial intelligence components can share the artificial intelligence engine to improve the efficiency of artificial intelligence data analysis.
[0046] In a possible implementation of the first aspect, the platform layer further includes data storage middleware, which is used to enable the service components to share industrial data; and the scheduler is further used to schedule the corresponding running capsule for the data storage middleware.
[0047] From the above, by setting up data sharing middleware, data sharing of each service component of the application layer is realized, and data consistency across regions and devices is achieved.
[0048] In a possible implementation of the first aspect, the platform layer is further configured to implement a consistency protocol-based session between the scheduled running capsules.
[0049] From the above, the platform layer sets a consistency protocol to realize collaborative conversations between running capsules, thereby achieving control collaboration across regions and devices.
[0050] In the second aspect, an embodiment of the present application provides an industrial control method for performing industrial control on the system described in any implementation mode of the first aspect, including: calling the platform layer through the service layer to schedule each task in the service component to the running capsule of the corresponding physical node; starting the service component through the service layer, and each task of the service component runs in the scheduled running capsule to realize industrial control.
[0051] In a third aspect, an embodiment of the present application provides a computing device, including:
[0052] bus;
[0053] a communication interface connected to the bus;
[0054] at least one processor connected to the bus; and
[0055] At least one memory is connected to the bus and stores program instructions, and when the program instructions are executed by the at least one processor, the computing device is configured as the system described in any embodiment of the first aspect of the present invention.
[0056] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having program instructions stored thereon, which, when executed by a computer, causes the computer to be configured as the system described in any embodiment of the first aspect of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] FIG1 is a schematic diagram of the structure of the industrial cloud according to various embodiments of the present application;
[0058] FIG2 is a schematic structural diagram of an embodiment of a novel industrial operating system of the present application;
[0059] FIG3 is a schematic diagram showing the alignment of the time slices of the real-time running capsule scheduling of the present application with the TSN time slices used by the running capsule;
[0060] FIG4 is a schematic diagram showing the alignment of time slices of several task schedules in the real-time running capsule of the present application with the TSN time slices used by the several tasks;
[0061] FIG5 is a schematic structural diagram of a platform layer embodiment of a novel industrial operating system of the present application;
[0062] FIG6 is a schematic diagram of the structure of a cross-region industrial cloud in a platform layer embodiment of a novel industrial operating system of the present application;
[0063] FIG7 is a schematic structural diagram of a service layer embodiment of a novel industrial operating system of the present application;
[0064] FIG8 is a schematic structural diagram of an embodiment of a device for dynamically allocating CPU runtime at the base layer of a novel industrial operating system of the present application;
[0065] FIG9 is a schematic structural diagram of an embodiment of a control function switching device for an edge device in a novel industrial operating system of the present application;
[0066] FIG10 is a flow chart of an industrial control method of the present application;
[0067] FIG11 is a schematic structural diagram of a computing device according to various embodiments of the present application. DETAILED DESCRIPTION
[0068] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0069] In the following description, the terms "first\second\third, etc." or module A, module B, module C, etc. are only used to distinguish similar objects, or to distinguish different embodiments, and do not represent a specific ordering of the objects. It can be understood that the specific order or sequence can be interchanged where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0070] In the following description, the numbers representing the steps, such as S110, S120, etc., do not necessarily mean that the steps must be executed in this manner. If permitted, the order of the steps can be interchanged or they can be executed simultaneously.
[0071] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0072] An embodiment of the present application provides a new industrial operating system, device and storage medium, which includes: a base layer, a platform layer and a service layer; the base layer is deployed on each physical node on the industrial cloud, including an elastic microkernel and several running capsules, and the elastic microkernel is used to adaptively allocate hardware resources to the running capsules; each of the running capsules supports one or more of multiple operating scenarios; the platform layer is used to schedule the running capsules with matching capabilities on several physical nodes for each task in each service component of the system from the industrial cloud, wherein the predicted delay of each service component satisfies the deterministic constraints of the service component, and the predicted delay is obtained based on the computational delay of the scheduled running capsules and the transmission delay between the scheduled physical nodes; the service layer is used to develop, deploy and start the service components, and the service components include time-critical industrial control components.
[0073] The technical solution of the embodiment of the present application realizes the elastic use of chip resources of each physical node through dynamic scheduling of the running capsules of the physical nodes on the industrial cloud, so that the calculation and transmission delays of time-critical industrial control components meet their delay constraints, and completes the real-time and deterministic control in the industrial field.
[0074] The following first introduces the various embodiments of the present application in conjunction with the accompanying drawings. First, the application scenarios of the various embodiments of the present application are introduced in conjunction with FIG1 .
[0075] FIG1 shows the structure of the industrial cloud of each embodiment of the present application, which includes an OT network and an IT network in the industrial field.
[0076] The computing nodes in the OT network are located in the physical nodes of this application and are connected through TSN devices to realize the functions of the industrial controller and SCADA. The industrial controller can be a PLC controller or motion controller, etc. It controls robots in the industrial control field in real time through I / O modules and runs in the real-time operation capsule of the physical node. The real-time operation capsule has a real-time operation environment. SCADA displays and analyzes field data in the non-real-time operation capsule running in the real-time operation capsule of the physical node. The non-real-time operation capsule has a non-real-time operation environment. The industrial controller and SCADA can connect actuators and sensors through the AUTBUS network.
[0077] Among them, the IT network includes SCM servers, ERP servers, PLM servers and MES servers, which are also located in the physical nodes in this application and are also connected to the office network. Users in the office network access the SCM servers, ERP servers, PLM servers and MES servers through the IT network to perform related SCM, ERP, PLM and MES tasks.
[0078] Among them, the IT network is also connected to the OT network through TSN devices and / or 5G networks, and users in the office network access the controller and SCADA through the IT network and OT network.
[0079] Among them, the IT network is also connected to other industrial clouds through the Detnet network (deterministic network) to achieve cross-regional industrial cloud connection. The IT network is also connected to the external network through the Detnet network, and the Detnet network can be a TSN network.
[0080] An embodiment of a novel industrial operating system will be described below with reference to FIG. 2 and FIG. 3 .
[0081] FIG2 shows the structure of an embodiment of a novel industrial operating system of the present application, which includes, from bottom to top: a base layer 100 , a platform layer 200 , and a service layer 300 .
[0082] The base layer 100 is deployed on each physical node of the industrial cloud and includes an elastic microkernel 110 of the physical node and several running capsules 120. The elastic microkernel 110 is used to allocate hardware resources to the running capsules.
[0083] One possible implementation of the base layer 100 is using the new Rust language, leveraging its features to provide memory safety. The system utilizes standardized functional component design and offers component-based assembly, enabling flexible, on-demand integration of advanced features to support virtualization and provide enhanced security isolation.
[0084] The elastic microkernel 110 includes componentized hardware resources for dynamically allocating the componentized hardware resources to each running capsule 120 .
[0085] Among them, the running capsule 120 supports one of the following scenarios: thread, process, real-time container, non-real-time container, real-time virtual machine, non-real-time virtual machine, and the running capsule 120 includes the running environment of the corresponding scenario, which is used for tasks in the service component of the new industrial operating system. The running environment includes component resources that support the running capsule 120 of the corresponding scenario.
[0086] For example, Figure 2 shows two partitioned virtual machines and one non-partitioned container. The real-time operating environment of one partitioned virtual machine includes real-time functional components and supports real-time application tasks. The high-security operating environment of one partitioned virtual machine includes high-security functional components and supports high-security application tasks. The non-real-time operating environment in the non-partitioned container supports the application tasks of the non-real-time container.
[0087] The running capsule 120 can also be divided into a real-time running capsule and a non-real-time running capsule, which are used to run real-time tasks and non-real-time tasks respectively. The real-time running capsule is a time-critical running capsule 120.
[0088] In one possible implementation of this embodiment, the running capsule 120 of each physical node is located in the adaptive partition of the physical node, and each adaptive partition is configured with a budgeted CPU running time. When the elastic microkernel 100 of the physical node allocates hardware resources to the running capsule 120, it is specifically used to schedule the budgeted CPU running time of each adaptive partition on the physical node for the tasks in the running capsule 120 of the adaptive partition. When the CPU running time actually used by any adaptive partition on the physical node is lower than its budgeted CPU running time, the remaining CPU running time of the adaptive partition is allocated to the running capsules 120 of other adaptive partitions with tasks waiting to be scheduled on the physical node.
[0089] For a detailed description of the base layer 100 , please refer to an embodiment of a base layer of a new industrial operating system.
[0090] The platform layer 200 is used to schedule running capsules 120 with matching capabilities on several physical nodes for each task in the service component of the new industrial operating system from the industrial cloud, so that at least the predicted delay of the service component obtained based on the computing delay of the scheduled running capsules 120 and the transmission delay between the scheduled physical nodes meets the deterministic constraints of the service component.
[0091] In a possible implementation of this embodiment, the platform layer 200 includes an orchestrator 210 , a scheduler 220 , and middleware 230 .
[0092] The coordinator 210 is deployed on the physical node for management of the industrial cloud in each region, and is used to schedule several physical nodes for each task in the service component of the new industrial operating system from the industrial cloud.
[0093] Scheduler 220 is deployed on the physical nodes used for management in each region's industrial cloud, or on each physical node. It schedules a runtime capsule that matches the capabilities of each service component in service layer 300 from the physical nodes scheduled by orchestrator 210. It also predicts the predicted latency for running the service component based on the computational latency of the scheduled runtime capsule 120 and the transmission latency between the scheduled physical nodes. Real-time tasks are scheduled to run in real-time runtime capsules, which are time-critical tasks. Non-real-time tasks are scheduled to run in non-real-time runtime capsules.
[0094] Middleware 230 manages communication between physical nodes, stores data, and supports an artificial intelligence engine. These run on runtime capsules 120 on physical nodes dedicated to their respective functions and can be scheduled statically or dynamically. Together with the runtime capsules 120 scheduled by scheduler 220, middleware 230 helps the service components of the new industrial operating system complete their functions.
[0095] The platform layer 200 provides distributed collaboration and deterministic control capabilities based on a distributed coordination framework and deterministic scheduling. This includes: providing elastic resource allocation and scheduling based on the base layer 100, providing distributed deterministic communication capabilities based on the TSN time-sensitive network, and standardized consistency protocols and mechanisms, thereby solving problems such as real-time determinism and data consistency in distributed control and computing. At the same time, the platform layer 200 provides security and privacy protection mechanisms that are applied in the collaboration process to ensure data security and confidentiality. Finally, protocol definition and standardization are key to ensuring collaboration and control between different layers. Through unified protocol specifications, seamless integration and communication between different devices, edge platforms, and cloud platforms can be achieved.
[0096] Platform layer 200 also provides ubiquitous industrial connectivity, addressing interoperability issues across diverse industrial devices, and software-defined control, addressing the need for flexible, on-demand deployment of control systems. Platform layer 200 also rapidly detects the entry of industrial control terminals into the network. Combining platform layer 200 collaborative control with deterministic scheduling and communication, it enables rapid handover of device control across regions.
[0097] In one possible implementation of this embodiment, when a physical node runs a task of an industrial control component to control a device connected to an industrial cloud via a wireless connection, each node determines its deterministic control range of the device based on its control delay of the device; the platform layer 200 is also used to switch the control function of the device to a neighboring control node when the device leaves the deterministic control range of the node and is within the deterministic control range of the device of the node. When the control delay of each node for its controlled device is greater than the delay constraint of the device minus a first set value, it is determined that the device is outside the deterministic control range of the node; when the control delay of each node for its controlled device is less than or equal to the delay constraint minus a second set value, it is determined that the device is within the deterministic control range of the node, and the second set value is greater than or equal to the first set value, to avoid ping-pong switching.
[0098] For a detailed description of the platform layer 200 , please refer to an embodiment of a platform layer of a new industrial operating system.
[0099] The service layer 300 includes an industrial application cloud development kit 310, also known as an industrial application cloud development platform. Deployed on physical nodes within the industrial cloud with an integrated development environment, it is used to develop service components that form various service suites. The industrial application cloud development kit 310 decomposes each service suite into several service components. Each service component is open to users and decomposed into several tasks that can be subscribed to. The service component tasks are essentially dispatched by the platform layer 200 to the execution capsule 120 for execution.
[0100] The Industrial Application Cloud Development Kit 310 is based on cloud computing and microservice architecture to build an open, one-stop service platform for users. It enables the rapid development of automation systems for industrial equipment and provides component-based and template-based services, enabling enterprises to quickly build and deploy customized industrial applications.
[0101] The Industrial Application Cloud Development Kit 310 is a WEB-based integrated development environment that supports multiple programming languages and application development frameworks (such as artificial intelligence application development framework, data analysis application development framework, etc.), and integrates one or more artificial intelligence-assisted development tools to help developers improve work efficiency.
[0102] The Industrial Application Cloud Development Kit 310 also provides automated testing tools: it provides an automated testing framework and tools, including unit testing, integration testing, and end-to-end testing, to ensure the quality and stability of industrial applications.
[0103] The Industrial Application Cloud Development Kit 310 also provides CI / CD tools and platforms to automate the building, testing, packaging, and deployment of industrial applications to achieve fast and reliable application delivery.
[0104] The Industrial Application Cloud Development Kit 310 also provides security and vulnerability scanning tools: integrated security scanning tools and vulnerability detection tools to detect and fix security vulnerabilities and potential risks in industrial applications.
[0105] The Industrial Application Cloud Development Kit 310 also provides operation and monitoring tools: It provides operation and monitoring tools for industrial applications, including application performance monitoring, error log tracking and real-time monitoring, to ensure the high availability and performance of industrial applications.
[0106] The Industrial Application Cloud Development Kit 310 also provides collaboration and teamwork tools: It provides collaboration and teamwork tools for communication, collaboration and version control among team members to support multiple people to jointly develop industrial applications.
[0107] The Industrial Application Cloud Development Kit 310 deploys and containerizes service suite applications and supports tools like container orchestration and management to quickly deploy and manage service suites.
[0108] The service layer 300 also deploys an industrial control suite 320 and an industrial simulation cloud platform 330. The industrial control suite 320 includes time-critical industrial control components. The tasks of the industrial control components are scheduled by the platform layer 200 to real-time running capsules (such as partition-based real-time containers or real-time virtual machines). The service layer 300 is also used to start related service components in the industrial control suite to control and access industrial actuators or sensors connected to the edge cloud in the industrial cloud through the platform layer 200 and the base layer 100 to complete industrial control.
[0109] The industrial simulation cloud platform 330 includes simulation service components and provides a cloud-based simulation platform for simulating and testing various industrial application scenarios. Currently, there is a wide variety of industrial simulation software. Foreign companies have closed industrial chains, resulting in significant "lock-in" barriers. Domestic ecosystem deployment is limited, and unified software standards are lacking. The new industrial operating system simulation cloud platform integrates the capabilities of domestic CAD / CAE vendors and leverages the simulation capabilities of the new industrial operating system to transform from soft simulation to hard simulation, improving system reliability and transitioning from semi-physical simulation to full-physical simulation. It provides a unified development platform, integrated simulation environment, and solutions for multidisciplinary and multi-domain system analysis. Users can model, expand, or modify models through a graphical interface; it offers multiple simulation operation modes and optimization algorithms, as well as interfaces compatible with other software and unified interface standards. The simulation service components are dispatched through the platform layer 200 to non-real-time runtime capsules with matching computing power.
[0110] For a detailed description of the service layer 300, please refer to a service layer embodiment of a new industrial operating system.
[0111] 2 to 4 , an embodiment of a base layer of a novel industrial operating system will be described below.
[0112] The base layer 100 in FIG2 is a structure of a base layer embodiment of a novel industrial operating system.
[0113] The elastic microkernel 110 manages the hardware resources of the physical node in a componentized manner. Each componentized hardware resource is a standardized resource component. The elastic microkernel 110 elastically loads and / or deletes each resource component as needed, enabling flexible management of the hardware resources on the physical node. These hardware resources include the CPU core on the chip, the motherboard and / or chip memory, and physical node peripherals. Therefore, the elastic microkernel of this embodiment of the present invention is a highly elastic microkernel that can allocate chip resources in random and rapid combinations based on demand.
[0114] The elastic microkernel 110 is also used to statically and / or dynamically allocate componentized hardware resources to the runtime capsules 120 on the physical node, achieving flexible allocation of resource components. The elastic microkernel 110 dynamically allocates CPU runtime, as described in the embodiment of a dynamic CPU runtime allocation device for the base layer of a novel industrial operating system.
[0115] The elastic microkernel 110 accesses hardware resources allocated to the partition-based time-critical execution capsules based on at least one of the following methods:
[0116] Access to the shared area of each CPU running the capsule is based on inter-CPU messages and service interrupts, without the need for locks to synchronize the shared area, thereby improving the real-time performance of shared memory access based on the IPC of the time-critical running capsule of the partition;
[0117] Map the critical area of the CPU running the capsule to the memory, and access the memory corresponding to the critical area of the CPU running the capsule based on the mapped address, so as to achieve direct access to the critical area of the CPU running the capsule based on the variable name, with efficient and fast access speed;
[0118] Based on the cache partition of the CPU running the capsule, different running capsules have different base addresses. The cache of the CPU running the capsule is accessed according to the base address, realizing the shared cache isolation of the CPU running the capsule and improving the real-time performance of the program.
[0119] When a physical node has a time-critical runtime capsule 120, the elastic microkernel 110 of the physical node is also used to connect to the TSN network via the TSN network card to perform TSN communication for the runtime capsule 120 of the physical node. If a physical node does not have a time-critical runtime capsule 120, the elastic microkernel 110 does not need to load the TSN resource component for TSN communication. If a physical node has a time-critical runtime capsule 120, it needs to load the TSN resource component, and the elastic microkernel 110 performs TSN communication using the loaded TSN resource component.
[0120] The CPU time slices scheduled by the elastic microkernel 110 for the time-critical execution capsule 120 are aligned in timing with the TSN time slices used by the execution capsule 120. The TSN time slices are the time slices for scheduling TSN data streams in the TSN switch. That is, the tasks corresponding to the TSN data streams using a TSN time slice are scheduled within the CPU time slice of the execution capsule 120 corresponding to the TSN time slice. The tasks corresponding to these TSN data streams form a task group.
[0121] The elastic microkernel 110 schedules each task running in the time-critical running capsule 120 based on the task priority, which matches the service priority of TSN. That is, the tasks in a task group are scheduled based on the task priority that matches the service priority of TSN. The service priority of TSN is the priority defined by the Pri field in the VLANTag in the TSN data stream.
[0122] Specifically, after the TSN network-wide time is synchronized, the partition window time and Qbv gate interval time are dynamically adjusted according to the task priority of the service component; based on the service priority and the above-mentioned QoS mapping, the queue priority 0 to 7 corresponding to the TSN chip output port is generated and transmitted to the network protocol stack; when adding the VLANTag, the protocol stack directly maps the above-mentioned priority to the 3-bit pri field in the VLANTag; the message enters different port queues according to the pri, and finally the time determinism is guaranteed by the TSN network; for burst services, the Qbu function is enabled on the low-priority output port, that is, the fast frame (E frame) can preempt the preemptible frame (P frame) to ensure that the fast service passes in time.
[0123] Figure 3 shows a schematic diagram of the alignment of the time slices scheduled by a real-time run capsule with the TSN time slices used by the run capsule. The real-time run capsule in Figure 3 uses a partitioned run capsule 120 (partition-based run capsule) as an example. The number of physical nodes and the number of partitioned run capsules 120 are examples. Each partitioned run capsule 120 runs tasks on an industrial controller, connected to the TSN network via the TSN network card of each physical node and connected to actuators or sensors via remote unit I / O. The TSN network card, TSN gateway, and remote unit I / O are time-aligned based on TSN's IEEE 802.1AS standard. The tasks scheduled by the elastic microkernel 110 in each partitioned run capsule 120 are aligned with the time slices of their TSN data streams scheduled in the corresponding TSN gateway, thereby improving the real-time performance of data transmission within the real-time run capsule.
[0124] Figure 4 shows a schematic diagram of the alignment of the time slices used by several tasks in a real-time capsule (a time-critical capsule) with the TSN time slices used by these tasks. In Figure 4, RTOS represents the real-time capsule, thread represents the task in the real-time capsule, the table on the lower left shows the time slices used by the microkernel to schedule each task, and the table shows the time slices used by the microkernel to schedule each task. The number of real-time capsules and the number of tasks in the figure are examples, as are the number of TSN time slices and the length of each time slice. The first real-time capsule contains three tasks, thread1, thread2, and thread3, forming the first task group. The second real-time capsule contains one task, thread4, forming the second task group. The third real-time capsule contains two tasks, thread5 and thread6, forming the second task group.
[0125] Among them, the TSN time slices allocated to each task in the first real-time running capsule based on the IEEE802.1QBV protocol of TSN are t0 to t1. The microkernel also schedules the running time slice of the first real-time running capsule from t0 to t1. In the first real-time running capsule, that is, the first task group, thread 3 has the highest TSN service priority, so thread 3 also has the highest task priority. From t0 to t1, the microkernel gives priority to allocating CPU time slices, and from t0 to t1, the TSN switch gives priority to scheduling TSN time slices.
[0126] The second real-time running capsule, thread 4 in the second task group, is assigned a TSN time slice from t1 to t2 based on the IEEE802.1QBV protocol of TSN. Because this running capsule only has one task, thread 4, from t1 to t2, thread 4 is allocated CPU time slices by the microkernel and scheduled by the TSN switch.
[0127] Among them, the third real-time running capsule, that is, each task in the third task group is allocated a TSN time slice of t2 to t3 based on the IEEE802.1QBV protocol of TSN. The microkernel also schedules the running time slice of the third real-time running capsule from t2 to t3. In the third real-time running capsule, thread 6 has the highest TSN service priority, so thread 6 also has the highest task priority. From t2 to t3, the microkernel gives priority to allocating CPU time slices to thread 6, and from t2 to t3, the TSN switch gives priority to scheduling TSN time slices.
[0128] Each physical node's runtime capsule 120 is encapsulated based on its capabilities. The base layer 100 is further configured to register the current runtime capsule 120 with the platform layer 200. This registration information includes at least the capabilities of the registered runtime capsule. These capabilities include at least one of the following: computing power, security, and trust. Computational latency can be predicted based on the computing power and latency model.
[0129] The probability of the running capsule 120 obtaining the CPU core and memory in different running scenarios is different, and the computing power is different. For example, the running capsule 120 with a high probability of obtaining the CPU core and memory has a strong computing power; the degree of isolation of the CPU core and memory of the running capsule 120 in different running scenarios from other running capsules 120 is different, and the security is different. For example, the running capsule 120 based on the partitioned virtual machine has the highest security with an independent CPU core and independent memory; the trust capabilities of the running capsule 120 in different running scenarios are different, and the trust capabilities of the running capsule 120 are evaluated based on the trust evaluation method.
[0130] At the industrial cloud site, each physical node of the base layer 100 is connected to the actuators and controllers at the site via AUTBUS to connect to more site nodes.
[0131] 5 and 6 , a platform layer embodiment of a novel industrial operating system is introduced below.
[0132] Figure 5 shows the structure of a platform layer embodiment of a new industrial operating system, which inherits the collaborator 210 and scheduler 220 in the platform layer 200 of Figure 1, adds a manager 240, and divides the middleware 230 into deterministic communication middleware 231, pan-industrial communication middleware 232, data storage middleware 233 and artificial intelligence engine middleware 234.
[0133] Coordinators 210 are deployed regionally, on physical nodes managed by the central cloud within each region's industrial cloud. They are used to schedule multiple physical nodes from the industrial cloud for tasks within the new industrial operating system's service components. Coordinators 210 operate in a hierarchical manner, providing cross-system and cross-regional device coordination and centralized control, real-time monitoring, and remote collaboration.
[0134] Figure 6 shows the structure of a cross-region industrial cloud in this example. This cross-region industrial cloud, for example, includes a sub-cloud in region A and a sub-cloud in region B. Each sub-cloud includes a central cloud and an edge cloud. An orchestrator 210 is deployed on the management node of the central cloud in each of the sub-clouds in region A and region B.
[0135] Among them, when the running capsules 120 of the physical nodes in the sub-cloud of area A and the sub-cloud of area B can both meet the sub-tasks in the scheduled service components, the collaborative scheduling of these coordinators 210 is based on load balancing to schedule several physical nodes from a regional sub-cloud for each task in the service component; when the sub-task in the scheduled service component is used to control an industrial actuator or sensor in an edge cloud in a region, a physical node is scheduled from the edge cloud in the region based on the shortest predicted delay. The scheduling method distinguishing these two scenarios is used to realize cross-regional cloud-edge collaborative scheduling.
[0136] Collaborator 210 uses a standardized and consistent protocol to ensure data consistency across all parties in a collaborative session, and to achieve consistent on-site perception and control across systems and devices. For example, in the field of industrial control, including component manufacturing robot control, assembly robot control, and transport robot control, the standardized and consistent protocol of Collaborator 210 can be used to achieve collaborative control among various robots in this field. The standardized and consistent protocol of Collaborator 210 can be converted to a standard by middleware at the platform layer, or a converted driver can be generated and distributed to the runtime capsule where the corresponding task resides for standard conversion.
[0137] The collaborator 210 also uses security and privacy protection mechanisms applied in the collaboration process to ensure the security and confidentiality of data.
[0138] The scheduler 220 is deployed on a physical node for management of the edge cloud in each region. In FIG6 , the distributed deterministic scheduler is the scheduler 220 , and the edge distributed platform is the physical node for management.
[0139] Scheduler 220 adopts the task end-to-end delay analysis algorithm to calculate the worst delay constraint of the task in the multi-level dynamic scheduling framework, adopts the real-time scheduling analysis technology for multi-core systems, establishes the basic structure model of the time-predictable multi-core processor, and performs scheduling analysis based on it, so as to fully utilize the powerful parallel computing capabilities provided by the multi-core processor.
[0140] The scheduler 220 schedules running capsules that match the capacity of each task in each service component from the physical nodes scheduled by the orchestrator 210, and predicts the predicted delay of running the service component based on the computational delay of the scheduled running capsule 120 and the transmission delay between the scheduled physical nodes. The predicted delay at least meets the deterministic constraint of the scheduled service component (essentially a delay constraint).
[0141] For example, scheduler 220 schedules three serial tasks A1, B1, and C1 of a service component to execution capsules A2, B2, and C2, respectively. The physical nodes where execution capsules A2, B2, and C2 reside are connected via a TSN network. Based on the computing power of execution capsules A2, B2, and C2, scheduler 220 uses a preset model to predict the completion times t1, t2, and t3, respectively. It also predicts the transmission delay p1 from execution capsule A2 to execution capsule B2, and the transmission delay p2 from execution capsule B2 to execution capsule C2. The estimated completion delay for the service component is (t1+t2+t3+p1+p2). If this estimated delay is less than the deterministic constraint of the service component, the service component can complete the deterministic computation.
[0142] The scheduler 220 also schedules the running capsules for each service component to meet the security constraints of the service component, thereby making the new industrial operating system a secure system. The scheduler 220 schedules the running capsules for each service component to meet the trust constraints of the service component, thereby making the new industrial operating system a trustworthy system.
[0143] When the scheduler 220 cannot schedule a running capsule 120 that meets the requirements of the task in the service component from the existing running capsules 120 on each physical node, the scheduler 220 is also used to schedule a physical node to enable the physical node to dynamically create a running capsule 120 that meets the requirements; the scheduler 220 also schedules a running capsule 120 with the closest capability to a physical node, and then notifies the elastic microkernel 110 of the physical node to dynamically increase hardware resources for the running capsule 120, so that it becomes a running capsule 120 that meets the requirements of the task in the scheduled service component.
[0144] The deterministic communication middleware 231, the pan-industrial communication middleware 232, the data storage middleware 233 and the artificial intelligence engine middleware 234 are implemented based on software configuration, that is, through SDN.
[0145] The deterministic communication middleware 231 is deployed on the running capsule 120 on the physical node for managing deterministic communication through static scheduling of the scheduler 220. It is used to generate TSN communication strategy parameters for the real-time running capsule 120 and predict the transmission delay of the real-time running capsule based on this. The deterministic communication middleware 231 is also used to obtain the network topology and automatically sense the actual delay between physical nodes through the DDS+TSN method. Obtaining the network topology includes sensing the access of TSN terminals, including industrial control actuators and sensors to the pan-industrial communication network. For example, it can quickly sense the cross-regional movement of drones and air traffic control equipment. Specifically, the working methods of the deterministic communication middleware 231 include:
[0146] 1) Deterministic communication middleware 231 adopts a publish / subscribe architecture, emphasizes data-centricity, and provides a rich set of QoS service quality policies based on time-sensitive networks to ensure real-time, efficient, and flexible data distribution. This meets the deterministic requirements of distributed communication applications and enables reliable message transmission, data exchange, and collaborative operations between different nodes. By adhering to the DDS protocol standard and based on the subscription-publish model and topic approach, it achieves node decoupling and simplifies the application call interface.
[0147] 2) Deterministic communication middleware 231 uses IEEE 802.1AS for precise network-wide timing, with an accuracy of ±10ns between hops;
[0148] 3) The flexible QoS policies provided by the DDS of the deterministic communication middleware 231, such as RELIABILITY, which can ensure service reliability from the application layer, DEADLINE, which constrains the data transmission cycle, and TRANSPORT PRIORITY, which sets service priority;
[0149] 4) The deterministic communication middleware 231 uses time-sharing and partitioning IEEE 802.1Qbv gate intervals and IEEE 802.1Qbu frame preemption. The pan-industrial communication middleware 232 is used to generate TSN communication policy parameters for the pan-industrial communication running capsule and predict the transmission latency of the running capsule based on these parameters. Pan-industrial communication protocols include AUTBUS, MODTCPBUS, and E-CAT. The pan-industrial communication middleware 232 is deployed on the running capsule 120 on the physical node responsible for managing pan-industrial communication through static scheduling by the scheduler 220. The pan-industrial communication middleware 232 is also used to detect the entry of terminal devices into the pan-industrial communication network, including industrial control actuators and sensors. For example, it can quickly detect the cross-regional movement of drones and air traffic control equipment. The data storage middleware 233 is used to store industrial big data and enable sharing of industrial data among service components of various service suites. The data storage middleware 233 is deployed on the running capsule 120 on the physical node responsible for storage management through static scheduling by the scheduler 220.
[0150] The artificial intelligence engine middleware 234 is used to implement an engine for artificial intelligence computing, for calling artificial intelligence-related service components, and for artificial intelligence analysis of industrial big data. Among them, the artificial intelligence engine middleware 234 is used to run capsules 120 on physical nodes with artificial intelligence computing power (such as physical nodes with GPUs) through static scheduling deployment of the scheduler 220. In one possible implementation of this embodiment, the artificial intelligence engine middleware 234 is an AI training framework for distributed elastic resource scheduling for industrial sites, such as a training framework based on pytorch; an AI reasoning framework with real-time and deterministic performance, such as an onnxruntime reasoning framework; and an autonomous lifelong learning AI algorithm that can interact with data collected through Autbus. The artificial intelligence engine middleware 234 can also use artificial intelligence to provide solutions for adaptive training of industrial large models and general industrial scenarios.
[0151] The manager 240 is deployed on the management node of the edge cloud and is used to perform lifecycle management on the running capsules 120 in each physical node in the edge cloud. For example, the SDN in Figure 6 is the manager 240 of the edge cloud. The manager 240 is also deployed on the management node of the central cloud and is used to perform lifecycle management on the running capsules 120 in each physical node in the central cloud.
[0152] The manager 240 manages the lifecycle of the runtime capsules 120 in its associated physical nodes, including image management, storage management, network management (including TSN network management), event management, and interface management for the applications in the runtime capsules 120. The runtime capsules 120 employ an abstraction similar to that of container applications, making the deployment of each runtime capsule 120 independent of the hardware of the specific physical node, enabling rapid deployment and highly reliable operation and maintenance management of the runtime capsules 120.
[0153] For example, the OSL abstraction method is used for each running capsule to shield the implementation differences of running capsules in different systems. The management of the central running capsule cluster is achieved by modifying the Kubernetes method. The volcano method is used to implement Kubernetes-based "job" batch processing. KubeEdge is used to achieve deterministic management of edge clouds and terminal devices.
[0154] As described above, through the collaborative session of consistency protocol and mechanism provided by the orchestrator 210, the scheduling of deterministic computing for running capsules provided by the scheduler 220, the TSN network communication capability provided by the deterministic communication middleware 231, and the containerized abstract lifecycle management of the manager 240, the cloud-edge collaboration and rapid switching across systems and devices are achieved at the platform layer 200, thereby achieving deterministic control.
[0155] The above deterministic control is combined with the deterministic communication middleware 231 and / or the pan-industrial communication middleware 232 to quickly sense the access of industrial control terminals to the network, thereby achieving cross-regional coordination and rapid switching of industrial control. For details, please refer to the embodiment of a control function switching device for edge devices in a new industrial operating system.
[0156] For example, when a drone connected to the TSN network wirelessly moves from one industrial control area to another, the deterministic communication middleware 231 quickly senses that the drone has joined the network in the new area, and schedules the physical node on the edge cloud of the new area and the running capsule 120 on the physical node for the drone through the coordinator 210 and the scheduler 220. The control task for deterministic control of the drone is run on the running capsule 120. The coordinator 210 enables the control task to conduct a control session based on a protocol consistent with the entire network. The control task then controls the drone through the TSN network, enabling the drone to quickly switch between regions.
[0157] The following describes a service layer embodiment of a new industrial operating system with reference to FIG7 .
[0158] Figure 7 shows the structure of a service layer embodiment of a new industrial operating system, which inherits the industrial application cloud development kit 310, industrial control kit 320, and industrial simulation cloud platform 330 of the service layer 300 of Figure 1, and also deploys an industrial artificial intelligence kit 340 and an industrial application store 350.
[0159] The industrial application cloud development kit 310 is described in a novel industrial operating system embodiment.
[0160] The tasks of each service component in the industrial control suite 320, the industrial simulation cloud platform 330 and the industrial artificial intelligence suite 340 are all dispatched through the platform layer 200 to the corresponding running capsules on the corresponding physical nodes in the base layer 100 for execution.
[0161] The industrial control suite 320 includes time-critical industrial control components. The platform layer 200 schedules tasks of these components into corresponding real-time runtime capsules. The platform layer 200's deterministic communication middleware 231 manages the connection between the physical nodes where these scheduled runtime capsules 120 reside and the TSN network. It predicts the service component's latency based on the computational latency of the scheduled runtime capsules 120 and the transmission latency of the physical node through the TSN network, ensuring that the latency meets the industrial control component's latency constraints. The service layer 300 activates the relevant components of the industrial control suite 320 through user-friendly interfaces to perform industrial control.
[0162] The industrial simulation cloud platform 330 is described in detail in the following examples of a novel industrial operating system. It can leverage the analysis results of artificial intelligence components and / or manual configuration experience to obtain device-isolated simulation results, or it can be combined with specific devices to obtain device-specific simulation results. The industrial simulation cloud platform 330 also provides service components for the digital world of the industrial process metaverse, enabling the creation of an industrial digital twin. This platform can leverage the analysis results of artificial intelligence components and / or manual configuration experience to digitally iterate the industrial process metaverse.
[0163] The industrial artificial intelligence suite 340 includes artificial intelligence components, which are used to call the artificial intelligence engine middleware 234 of the platform layer to perform industrial data analysis based on the input and output of the service components, and obtain analysis results that are isolated from the physical hardware of the base layer 100. The analysis results are experience isolated from specific devices and specific areas, and can be used by other service components to realize experience sharing across devices and regions.
[0164] The artificial intelligence component provides data processing and analysis tools, tools and libraries for processing and analyzing industrial data, including tools for big data processing, data mining, machine learning and deep learning. The analysis results are an abstract shared knowledge base, which forms an interconnection standard through open interfaces, effectively promoting the sharing of knowledge and experience among various industrial sectors, and forming the service layer 300 into a knowledge sharing platform for industrial production control.
[0165] Industrial Application Store 350 is an open, one-stop service component store for users. Users simply select various service components, such as industrial apps, process algorithms, and industry know-how, pay by license, and purchase physical components at a low cost to build high-quality industrial equipment on demand. This allows the vast majority of people to focus their creativity on industry know-how and software algorithm optimization, creating greater social and economic value.
[0166] In summary, the technical solution of the embodiment of the present application provides a new type of industrial operating system, which creatively proposes key innovative technologies such as elastic microkernel, deterministic scheduling, deterministic computing, cloud-edge cross-regional collaboration and control, and industrial artificial intelligence. It realizes dynamic elastic allocation of chip resources, combines TSN network technology to support deterministic computing and deterministic control, and realizes collaboration between heterogeneous cores, collaboration between different chips, collaboration between different servers, and collaboration between different cloud centers based on cloud-edge and cross-regional cloud collaboration and control, ensuring collaboration time at different levels, thereby ensuring real-time and determinism, and can achieve collaborative control of hundreds of millions of nodes, realize the generalization of the control bottom layer in the industrial field, and become an application platform for industrial artificial intelligence. The new industrial operating system of the embodiment of the present application also provides a high-bandwidth, low-latency, and highly reliable industrial network to adapt to the trend of full connection of elements, full flow of data, and intelligent system in the era of industrial Internet, solve the extension of industrial network determinism from local to wide area, the sinking of new network technology from park to production line, and the support of the network for vertical industries from network interconnection to data interoperability. It needs to have the ability to access massive devices, interconnect heterogeneous systems, end-to-end deterministic transmission, and intelligent scheduling of network resources.
[0167] The following, in conjunction with Figure 8, introduces an embodiment of a dynamic CPU runtime allocation device for the base layer of a new industrial operating system. The device configures the data structure of the partition scheduler of each adaptive partition on the physical node and the budgeted CPU runtime of each adaptive partition; schedules the corresponding budgeted CPU runtime for each adaptive partition's task based on the data structure of each adaptive partition; when the CPU runtime actually used by any adaptive partition is lower than its budgeted CPU runtime, the remaining CPU runtime of the adaptive partition is allocated to other adaptive partitions with tasks waiting to be scheduled. The technical solution of the embodiment of the present application dynamically and elastically allocates chip resources on physical nodes through an elastic microkernel, dynamically allocates and adjusts computing resources according to changes in the demand for computing tasks of each adaptive partition, and ensures sudden CPU demand to meet the computing power requirements of different tasks.
[0168] Figure 8 shows the structure of an embodiment of a CPU runtime dynamic allocation device of the base layer of a new industrial operating system, including: a resource configuration module 410, a resource scheduling module 420, a budget judgment module 430, a task scheduling module 440, a remaining judgment module 450 and a remaining scheduling module 460.
[0169] The resource configuration module 410 is used to abstract the resources on the physical node into component resources and configure them into each adaptive partition.
[0170] The resources (including CPU cores, memory, storage space, devices, etc.) are abstracted as virtual component resources. Each adaptive partition is a combination of several component resources, ensuring that each adaptive partition has a set of engineered resources. Each adaptive partition can run multiple runtime capsules, which support one of the following scenarios: threads, processes, containers, or operating systems. Based on the CPU architecture (ARM, x86, etc.), the number of CPU cores, and the computing power of each CPU core, a single CPU runtime is abstracted to configure and schedule the CPU core resources of any chip in subsequent steps.
[0171] The data structure of the partition scheduler for each adaptive partition and the CPU runtime budget for each adaptive partition are configured. Unlike conventional partitions, adaptive partitions do not necessarily require a partition operating system to run. For example, when running threads, processes, and containers, a partition operating system is not required. The lack of a partition operating system results in the absence of a partition scheduler for the partition operating system in the adaptive partition. In this embodiment, the partition schedulers are all mounted in the elastic microkernel. The data structure of the partition scheduler for each adaptive partition defines the scheduling algorithm for that adaptive partition.
[0172] The data structure of each adaptive partition's partition scheduler is attached to the elastic microkernel's kernel scheduler, and the CPU runtime of each adaptive partition is allocated according to the budget ratio. The data structure of each adaptive partition's partition scheduler is configured based on one of the following methods: RMS monotonic rate, priority, or time schedule.
[0173] The resource scheduling module 420 is used to schedule the corresponding budgeted CPU running time for the tasks of each adaptive partition according to the data structure of the partition scheduler of each adaptive partition.
[0174] The budgeted CPU runtime for each adaptive partition is converted into a CPU runtime time slice within a scheduling master frame. When the elastic microkernel schedules a task for an adaptive partition, it assigns the time slice to the task. The task can run in one of the following scenarios: thread, process, container, or virtual machine.
[0175] Budget determination module 430 is used to check whether the CPU runtime of the adaptive partition exceeds the budget of the current adaptive partition before selecting a task to run. If so, resource scheduling module 420 is executed for the next adaptive partition. If not, task scheduling module 440 is executed.
[0176] The check is performed only within the budgeted CPU running time of the current adaptive partition. If the check is performed within the remaining time of other adaptive partitions, the check in this step is not performed.
[0177] The task scheduling module 440 is used to select tasks to be executed according to the data structure of the current partition scheduler, wherein the selection is based on one of the following methods: RMS monotonic rate, priority, and time schedule.
[0178] The remaining judgment module 450 is used to judge whether the actual CPU running time used by any adaptive partition is less than its budgeted CPU running time. If so, the remaining scheduling module 660 is executed, otherwise the resource scheduling module 420 is executed for the next adaptive partition.
[0179] The remaining scheduling module 460 is used to schedule the remaining CPU running time to the adaptive partition with the highest priority among other adaptive partitions having tasks waiting to be scheduled, and schedule the task with the highest priority from the adaptive partition with the highest priority to run.
[0180] Wherein, no matter whether the adaptive partition with the highest priority has been scheduled before in the current scheduling main frame, it can occupy the remaining CPU running time.
[0181] Among them, if the adaptive partition with the highest priority has not been scheduled before in this scheduling main frame, then after the remaining CPU running time is used up, the adaptive partition with the highest priority will continue to be scheduled until its scheduling time reaches the sum of its own budgeted CPU running time and the remaining CPU running time.
[0182] If the adaptive partition with the highest priority has been scheduled before in this scheduling main frame, then after the remaining CPU running time is used up, the adaptive partition with the highest priority will occupy the remaining CPU running time.
[0183] The CPU running time that a task can occupy in a scheduling main frame includes the sum of the CPU running time budgeted by the adaptive partition where the task is located and the CPU running time occupied by the task in other adaptive partitions.
[0184] 9 , an embodiment of a control function switching device for an edge device in a novel industrial operating system is described below.
[0185] In an embodiment of a control function switching device for edge devices in a novel industrial operating system, when a physical node runs the task of an industrial control component to control a device connected to an industrial cloud via a wireless connection, each node determines its deterministic control range of the device based on its control delay of the device; the platform layer 200 is also used for switching the control function of the device to a neighboring node when the device moves to the edge of the deterministic control range of the node and is within the deterministic control range of a neighboring control node of the node.
[0186] Figure 9 shows the structure of an embodiment of a control function switching device for an edge device in a new industrial operating system, including: a period measurement module 510, a measurement trigger module 520, a switching trigger module 530, a target measurement module 540, a switching response module 550 and a switching confirmation module 560.
[0187] This embodiment is introduced by taking an industrial cloud wirelessly connected to an edge device through a TSN network as an example. For ease of description, the edge device is referred to as device A, and its current control node is referred to as a source control node.
[0188] The period measurement module 510 is located at the source control node and is used for the source control node to periodically measure the first control delay between the source control node and the device A.
[0189] To achieve deterministic control of device A, device A must have a reliable connection to the source control node and be within the source control node's control range, ensuring that the latency meets deterministic requirements. The first control latency reflects the connection status between device A and the source control node. This first latency is periodically measured to accurately and in real time determine the connection status between device A and the source control node.
[0190] The control delay between the source control node and the device A is measured at the base layer 100 of the new industrial operating system and is periodically sensed by the deterministic communication middleware 231 of the platform layer 200 of the new industrial operating system.
[0191] The control function of device A is scheduled into a matching runtime capsule on the source control node. The elastic microkernel 110 of the source control node dynamically allocates hardware resources for the tasks corresponding to the control function of device A. When each node in the cloud environment is managed through Kubernetes, the control function of device A and its required resources are encapsulated as an image. Kubernetes then schedules the control function of device A into a matching runtime capsule on the source control node. When the control function is deployed on the source control node, its runtime scenario is restored from the containerized state.
[0192] The source control node measures the first control delay for device A. Specifically, the source control node sends a delay measurement command to device A and receives a corresponding measurement report from device A. The source control node then obtains the first control delay between itself and device A based on the measurement report. The measurement report includes not only the transmission delay but also the interface processing delay between device A and the source control node. To accurately measure the first control delay, it is typically obtained based on multiple measurement reports. The two-way delays from sending the measurement command to receiving the measurement command and from sending the measurement report to receiving the measurement report are also calculated.
[0193] Among them, the management tasks corresponding to the source control node in the platform layer 200 of the new industrial operating system also perform lifecycle management on the application tasks of the control function of device A on the source control node. Among them, each node in the cloud environment is managed by kubernetes, and the application of device A is encapsulated as a container and deployed in its control node; when the cloud environment includes sub-clouds in several regions, the nodes of each sub-cloud are managed by kubeEdge, thereby making use of the management capabilities of kubernetes.
[0194] The measurement trigger module 520 is located in the collaborator 210 of the platform layer 200 of the new industrial operating system, and is used to select a target node from other nodes in the cloud when the first control delay is greater than or equal to the delay constraint of device A minus the first set value, and enable each target node to measure the second control delay of device A.
[0195] Among them, when the first control delay is greater than or equal to the difference between the delay constraint of device A and the first set value, the connection between device A and the source control node will not meet the delay constraint requirements. It can be considered that the device in the cloud environment has deviated from the acceptable deterministic control range or the edge of deterministic control. At this time, the target node must be triggered in time to measure its connection status with device A.
[0196] Each target node is a potential handover target and can sense the presence of device A. The second control delay of each target node indicates its connection status with device A and is used to select the final handover target. Each target node can sense the presence of device A.
[0197] When the cloud includes sub-clouds in several regions, the target node may be another node in the sub-cloud where the source control node is located, or a node in a sub-cloud in an adjacent region.
[0198] Among them, the measurement trigger module 520 can be divided into a first trigger module and a second trigger module. The first trigger module is located on the source control node, and is used for when there are other nodes in the sub-cloud where the source control node is located, the source control node first selects the target node from other nodes in the sub-cloud where it is located; the second trigger module is located on the central management node of the sub-cloud where the source control node is located. When the second control delay of any target node in the sub-cloud where the source control node is located is not less than the delay constraint of device A minus the second set value, the central management node of the sub-cloud where the source control node is located selects the node of the sub-cloud in the adjacent area as the target node.
[0199] The switching trigger module 530 is located in the coordinator 210 of the platform layer 200 of the new industrial operating system, and is used to switch the control function of device A from the source control node to the target node where the second control delay is less than the delay constraint of device A minus the second set value.
[0200] If the second control delay of the selected target node is less than the delay constraint of device A minus the second set value, the connection between the target node and device A is considered reliable, supporting deterministic control of device A. Device A is within the deterministic control range of the target node, and the target node can be a handover target. To avoid ping-pong handovers, the second set value is greater than the first set value, meaning that the delay requirement between the target node and device A is higher.
[0201] When the target node is a node in a sub-cloud of an adjacent region, seamless switching of the control function of device A across regions is achieved.
[0202] When the second control delays of multiple target nodes meet the requirement, the source control node selects the target node with the smallest second control delay as the node to be switched, so as to achieve a more reliable connection.
[0203] Among them, when the second control delay of the corresponding task on the source control node at a target node is less than the difference between the delay constraint of device A and the second set value, the platform layer 200 of the new industrial operating system initiates the switching of the control function of device A to the target node, and the scheduler of the platform layer of the new industrial operating system schedules the corresponding running capsule at the target control node to realize the task corresponding to the control function.
[0204] The source control node switches the control function of device A to the target node, specifically by synchronizing the control function's relevant data with the target node and replicating the control function to the target node. When each node in the cloud environment is managed via Kubernetes, the control function of device A and its required resources are encapsulated as a mirror. Kubernetes then schedules the control function of device A into a runtime capsule with matching capabilities on the target node. When the control function is deployed on the target node, its runtime is restored from the container.
[0205] The target measurement module 540 is located in the target node and is configured to measure the second control delay between the target node and device A.
[0206] The target node measures the second control delay with device A. Specifically, the target node sends a delay measurement command to device A and receives a corresponding measurement report from device A. Based on the measurement report, the target node obtains the second control delay between itself and device A. This measurement report includes not only the transmission delay but also the interface processing delay between device A and each target node. To accurately measure the second control delay, it is typically obtained based on multiple measurement reports. The two-way delays from sending the measurement command to receiving the measurement command and from sending the measurement report to receiving the measurement report are also calculated.
[0207] In some of these scenarios, when the source control node controls device A, it always maintains a wireless connection with the wireless point corresponding to the source control node in the industrial cloud during the movement of device A. At this time, a target node measures the second control delay of device A by sending the identification code of the wireless point corresponding to the target node in the industrial cloud to the source control node (such as the SSID of WIFI). The source control node allows device A to connect to the wireless access point corresponding to the identification code and connect to the wireless access point to measure the second control delay with the target node. If device A has two wireless network cards, one maintains a wireless connection with the wireless point corresponding to the source control node in the industrial cloud, and the other is used to measure the second control delay. If device A only has one wireless network card, after measuring the second control delay of the target node to it, it returns to the wireless point corresponding to the source control node in the industrial cloud.
[0208] In some scenarios, as device A moves, it automatically selects the wireless access point with the best signal in the industrial cloud based on wireless technology (such as 5G automatic switching based on signal quality, or Wi-Fi switching based on the best signal for the same SSID). When a target node measures the second control delay to device A, device A directly measures the delay through the path to the wireless access point with the best signal, without transmitting the target node's corresponding wireless point identifier in the industrial cloud. In this case, the second control delay includes the delay between device A and the wireless access point with the best signal and the delay between the target node and the wireless access point with the best signal. The first control delay includes the delay between device A and the wireless access point with the best signal and the delay between the source control node and the wireless access point with the best signal. It is important to emphasize that each control node in the industrial cloud has a wireless access point with the shortest transmission delay. When device A leaves the wireless access point with the shortest transmission delay to the source control node, its first control delay with the source control node increases. When device A connects to the wireless access point with the shortest transmission delay to the target node, its second control delay with the target node decreases. This solution is generally used in device A with only one network card.
[0209] Among them, the process of the target node measuring the second control delay of device A runs on the running capsule of the target control node running the deterministic communication task, and sends the second control delay to the middleware 231 for deterministic communication of the platform layer 200 of the new industrial operating system.
[0210] The switching response module 550 is located in the coordinator 210 of the platform layer 200 of the new industrial operating system, and is used to select a target node from the target nodes whose second control delay is less than the difference between the delay constraint of device A and the second set value. Before the source control node switches the control function of the device to the target node, the target node decides whether to accept the switching application based on its own resources, and after deciding to accept the switch, sends an acceptance message to the source control node.
[0211] The switching confirmation module 560 is used to measure the third control delay of the control node over device A through the control function when the target node has the control function of device A; when the difference in the third control delay is less than the difference between the delay constraint of device A and the third set value, device A is within the deterministic control range of the target node, the target node becomes the new control node of the device, and notifies the source control node to release the control function over device A; otherwise, the target node deletes the control function over device A, and the third set value is equal to the second set value.
[0212] Among them, after the source control node switches the control function of device A to the target node, it also includes: when the target node has the control function of device A, measuring the third control delay of the control node to device A through the control function; when the difference in the third control delay is less than the difference between the delay constraint of device A and the third set value, the target node becomes the new control node of the device and notifies the source control node to release the control function of device A; otherwise, the target node deletes the control function of device A, and the third set value is equal to the second set value.
[0213] The target node measures the third control delay for device A, specifically by sending a delay measurement command to device A and receiving a corresponding measurement report from device A. The target node then obtains the third control delay between itself and device A based on the measurement report. The measurement report includes not only the transmission delay but also the interface processing delay between device A and the selected target node. To accurately measure the third control delay, it is typically obtained based on multiple measurement reports, and the two-way delays from sending the measurement command to receiving the measurement command and from sending the measurement report to receiving the measurement report are also calculated.
[0214] Among them, the process of the final target node measuring the third control delay of device A runs on the running capsule of the final target control node running the deterministic communication task, and sends the third control delay to the middleware 231 for deterministic communication of the platform layer 200 of the new industrial operating system.
[0215] The management tasks corresponding to the final target node in the platform layer 200 of the new industrial operating system also perform lifecycle management on the application tasks of the control function of device A on the final target node. Each node in the cloud environment is managed through Kubernetes, and the application tasks and resource requirements of the control function of device A are encapsulated as container images, deployed in the final target node, and managed through Kubernetes for lifecycle management. When the cloud environment includes sub-clouds in several regions, the nodes of each sub-cloud are managed through kubeEdge, thereby leveraging the management capabilities of Kubernetes. Thus, the control functions of edge devices and their required resources are encapsulated as container images to fully utilize the container management capabilities of Kubernetes for scheduling and lifecycle management of control functions.
[0216] An embodiment of an industrial control method of the present application is described below with reference to FIG10 .
[0217] This industrial control method embodiment performs industrial control based on the system described in a novel industrial operating system embodiment. FIG10 introduces a process of the industrial control method embodiment, including steps S810 to S830.
[0218] S810 : The service layer 300 calls the platform layer 200 to dispatch the service component of the industrial control to the running capsule 120 of the corresponding physical node of the base layer 100 .
[0219] The service component may be developed through the industrial application cloud development kit 310 of the service layer 300 or may be obtained through other means.
[0220] The platform layer 200 schedules the corresponding physical node through its orchestrator 210, and the orchestrator 210 schedules the corresponding running capsule 120 on the corresponding physical node and obtains the expected delay of the service component, and the expected delay satisfies the deterministic constraint of the service component.
[0221] The deployment in this step can be dynamic or static. In dynamic deployment, the platform layer 200 is used to quickly sense the network access of the industrial control terminal and automatically coordinate, schedule and deploy through the platform layer.
[0222] S820 : Start the service component through the service layer 300 , and enable the service component to run in the scheduled running capsule 120 .
[0223] When the scheduled running capsule 120 is running, the scheduler 220 of the platform layer 200 performs lifecycle management on the running capsule.
[0224] S830: The running tasks in the scheduled running capsule 120 are industrially controlled through the TSN network or the pan-industrial communication network.
[0225] Among them, the industrial control is based on the coordination of standardized protocols through the collaborator 210 of the platform layer 200, performs deterministic calculations based on the computing power of the scheduled running capsule 120, and performs deterministic communication based on the TSN network, so as to meet the deterministic constraints of its service components and be safe and reliable.
[0226] The embodiment of the present application also provides a computing device, which is described in detail below in conjunction with Figure 11.
[0227] The computing device 900 includes a processor 910 , a memory 920 , a communication interface 930 , and a bus 940 .
[0228] It should be understood that the communication interface 930 in the computing device 900 shown in this figure can be used to communicate with other devices.
[0229] The processor 910 may be connected to a memory 920. The memory 920 may be used to store the program code and data. Therefore, the memory 920 may be a storage unit within the processor 910, an external storage unit independent of the processor 910, or a component including both a storage unit within the processor 910 and an external storage unit independent of the processor 910.
[0230] Optionally, the computing device 900 may further include a bus 940. The memory 920 and the communication interface 930 may be connected to the processor 910 via the bus 940. The bus 940 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus 940 may be divided into an address bus, a data bus, a control bus, and the like. For ease of illustration, the figure shows only one line, but this does not mean that there is only one bus or only one type of bus.
[0231] It should be understood that in the embodiment of the present application, the processor 910 can adopt a central processing unit (CPU). The processor can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. Alternatively, the processor 910 uses one or more integrated circuits to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0232] The memory 920 may include a read-only memory and a random access memory, and provides instructions and data to the processor 910. A portion of the processor 910 may also include a non-volatile random access memory. For example, the processor 910 may also store information about the device type.
[0233] When the computing device 900 is running, the processor 910 executes the computer execution instructions in the memory 920 to configure the computer into the system described in the embodiment of a novel industrial operating system of the present application.
[0234] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0235] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0236] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0237] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0238] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the program is used to configure the computer into the system described in an embodiment of a novel industrial operating system of the present application.
[0239] The computer storage medium of the embodiment of the present application can adopt any combination of one or more computer-readable media.Computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium.Computer-readable storage medium can be, for example, but not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof.More specific examples (non-exhaustive list) of computer-readable storage medium include, an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.In this document, a computer-readable storage medium can be any tangible medium containing or storing a program, which can be used by an instruction execution system, a device or a device or used in combination with it.
[0240] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0241] Program code embodied on a computer readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0242] The computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0243] Note that the above are only preferred embodiments of the present application and the technical principles employed. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, readjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments and may include many other equivalent embodiments without departing from the scope of protection of the present application, all of which fall within the scope of protection of the present application.
Claims
1. A new industrial operating system, characterized in that: include: Base layer, platform layer and service layer; The base layer is deployed on each physical node on the industrial cloud and includes an elastic microkernel and a plurality of running capsules. The elastic microkernel is used to adaptively allocate hardware resources to the running capsules. Each of the operation capsules supports one or more of a plurality of operation scenarios; The platform layer is used to schedule the running capsules with matching capabilities on several physical nodes for each task in each service component of the system from the industrial cloud, wherein the predicted latency of each service component satisfies the deterministic constraints of the service component, and the predicted latency is obtained based on the computation latency of the scheduled running capsules and the transmission latency between the scheduled physical nodes; The service layer is used to develop, deploy and start the service components, which include time-critical industrial control components.
2. The system according to claim 1, characterized in that The running capsule of each physical node is located in the adaptive partition of the physical node, and each adaptive partition is configured with its budgeted CPU running time; When the elastic microkernel allocates hardware resources to the running capsules on the physical node, it is specifically used to schedule the budgeted CPU runtime of each adaptive partition for tasks in the running capsules of the adaptive partition. When the CPU runtime actually used by any adaptive partition is lower than its budgeted CPU runtime, the remaining CPU runtime of the adaptive partition is allocated to other adaptive partitions with tasks waiting to be scheduled.
3. The system according to claim 1, characterized in that The elastic microkernel is specifically used to statically and / or dynamically allocate componentized hardware resources to the running capsules on the physical node; The elastic microkernel is further specifically configured to elastically load and / or delete componentized hardware resources according to demand.
4. The system according to claim 1, characterized in that The physical node is connected to the TSN network, and the CPU time slice for scheduling the capsule running on the physical node by the elastic microkernel is aligned in timing with the TSN time slice used by the running capsule.
5. The system according to claim 4, characterized in that: The elastic microkernel of the physical node connected to the TSN network is scheduled within the CPU time slice based on the task priority of each task running in the capsule running thereon, and the task priority matches the service priority of the TSN data flow of the task.
6. The system according to claim 1, characterized in that: When the elastic microkernel allocates hardware resources to a running capsule on a physical node, it is specifically used to adaptively allocate hardware resources to the running capsule based on the latency requirement of the task running in the running capsule, so that the computing latency of the running environment provided by the running capsule is better than or equal to the latency requirement of the task. The relationship between the computing latency and the allocated hardware resources is obtained based on a preset prediction model.
7. The system according to claim 1, characterized in that The elastic microkernel manages the hardware resources of the physical node based on at least one of the following methods: Accessing the shared area of the running capsule based on messages and service interruptions; Accessing the critical region of the CPU running the capsule at an address in the memory based on the CPU critical region mapping; The cache of the CPU running the capsule is accessed based on a cache partitioning method.
8. The system according to claim 1, characterized in that The running capsule supports one of the following running scenarios: thread, process, real-time container, non-real-time container, real-time virtual machine, and non-real-time virtual machine.
9. The system according to claim 1, characterized in that: When a physical node runs the task of the industrial control component to control a device connected to the industrial cloud via wireless, each node determines a deterministic control range for the device based on its control delay for the device; The platform layer is also used for, when the device moves out of the deterministic control range of the node over the device and is within the deterministic control range of a neighboring control node of the node over the device, the node switches the control function of the device to the neighboring node.
10. The system according to claim 9, characterized in that: When a physical node runs a task of the industrial control component to control a device connected to the industrial cloud via wireless, and each node determines that the control delay of its controlled device is equal to the delay constraint of the device minus a first set value, the device is outside the deterministic control range of the node; When each node determines that the control delay of its controlled device is less than the delay constraint minus the second set value, the device is within the deterministic control range of the node, and the second set value is greater than or equal to the first set value.
11. The system according to claim 1, characterized in that: The platform layer includes a scheduler and a coordinator; The coordinator is used to schedule a number of physical nodes for each task in the service component based on the requirement of the shortest predicted delay from the industrial cloud; The scheduler is used to schedule the running capsules matching the capabilities of the service components from the multiple physical nodes.
12. The system according to claim 11, characterized in that: When the industrial cloud includes sub-clouds in multiple regions, the orchestrator is further configured to allocate corresponding regions to corresponding service components based on load balancing, so that corresponding running capsules are dispatched from corresponding sub-clouds in the corresponding regions.
13. The system according to claim 11, characterized in that The platform layer also includes deterministic communication middleware for generating TSN communication strategy parameters for real-time capsules and predicting the transmission delay of the real-time capsules based on the generated TSN communication strategy parameters.
14. The system according to claim 1, wherein: The running capsule scheduled by the platform layer for each of the service components satisfies the security constraints and / or trust constraints of the service component.
15. The system according to claim 1, wherein: When a running capsule that meets the requirements of the task in the service component cannot be scheduled from the existing running capsules on each physical node, the platform layer is further used to schedule a physical node to enable the physical node to dynamically create a running capsule that meets the requirements.
16. The system according to claim 11, characterized in that The platform layer also includes a manager for performing life cycle management on the running capsule.
17. The system according to claim 11, characterized in that The platform layer also includes artificial intelligence middleware, which is used to implement the engine of artificial intelligence computing; The scheduler is further configured to schedule the corresponding running capsule for the artificial intelligence middleware; The service layer also includes an artificial intelligence component for calling the artificial intelligence middleware to perform industrial data analysis, and the analysis results are used by other service components.
18. The system according to claim 11, characterized in that The platform layer also includes data storage middleware for enabling the service components to share industrial data; The scheduler is further configured to schedule the corresponding running capsule for the data storage middleware.
19. A computing device, characterized in that include: bus; a communication interface connected to the bus; at least one processor connected to the bus; as well as At least one memory is connected to the bus and stores program instructions, and when the program instructions are executed by the at least one processor, the computing device is configured as the system according to any one of claims 1 to 18.
20. A storage medium, characterized in that Program instructions are stored thereon, and when the program instructions are executed by a computer, the computer is configured to be configured as the system according to any one of claims 1 to 18.
Citation Information
Patent Citations
Task scheduling method based on cloud platform, cloud platform and computer storage medium
CN108563500A
Decentralized mobile edge computing resource discovery and selection method and system
CN111556514A
Industrial intelligent control system based on software definition
CN112181382A
Distributed unit cloud deployment method and system, equipment and storage medium
CN113992688A
Security service scheduling method and system, electronic equipment and storage medium
CN117376032A