SLA driven orchestration of software containers

By monitoring and predicting SLA violations in real time in an Industry 4.0 environment, and utilizing a combination of network orchestration and migration processing units, adaptive service migration is achieved, solving downtime issues caused by SLA violations and improving system flexibility and security.

CN120936987APending Publication Date: 2025-11-11NOKIA NETWORKS OY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380097124.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-02-21
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing orchestration systems are unable to effectively manage network and computing resources in Industry 4.0 environments, leading to service level agreements (SLAs) violations. In particular, they are unable to migrate services in a timely manner to avoid downtime in the event of a failure, which affects the safety and efficiency of the production line.

Method used

The system employs a combination of network orchestration units, detection and prediction units, and migration processing units to monitor and predict SLA violations in real time. If the violation time is less than the migration window, the system triggers the service migration to a suitable destination node and selects a migration strategy based on the service's criticality level, including generating replicas and reserving resources.

Benefits of technology

It effectively reduces downtime due to SLA violations, improves the flexibility and reliability of service migration in industrial environments, and ensures the continuity and security of critical services in the event of failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120936987A_ABST
    Figure CN120936987A_ABST
Patent Text Reader

Abstract

Described herein is an apparatus for controlling a plurality of working nodes on which a service-providing application runs, the apparatus comprising a network orchestration unit, a detection and prediction unit, and a migration processing unit, wherein the network arrangement unit is configured to monitor the plurality of working nodes and configure network resources for the plurality of working nodes; the detection and prediction unit is configured to detect or predict a service level agreement (SLA) violation at a target node included in the plurality of working nodes; and the migration processing unit is configured to trigger migration if a violation time related to the predicted or detected SLA violation is less than a duration of a migration sensitivity window related to migration of a target service provided by a target application running on the target node to a destination node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the orchestration of network resources in a variable network environment, and particularly to adaptive orchestration for preventing violations of service level agreements. Background Technology

[0002] Any discussion of the background art throughout the specification should not be considered as an admission that such art is widely known or forms part of common knowledge in the art.

[0003] Continuous reconfiguration of the factory floor is one of the key prerequisites for Industry 4.0 (I4.0). To handle flexible market demands and changing business objectives, the progressive software-based transformation of industrial components is required; however, this introduces new challenges due to the inherent tendency of software to fail and stringent regulations.

[0004] However, in order to meet the functional and non-functional requirements specified by the factory's objectives, network and computing resources must be managed to minimize service level agreement (SLA) violations.

[0005] Orchestration systems manage the lifecycle of software containers to deploy them precisely where and when SLAs need to be met. Currently available orchestration systems are tailored for containers in cloud environments and require complete redesign to meet I4.0 requirements, featuring intelligent placement that takes into account network resources and SLAs.

[0006] Specifically, it's crucial to avoid placing currently non-critical perceptions (potential consequences of perception failures and SLA violations) and ignoring network-related SLA parameters. For example, if the fleet manager in the production line fails to deliver its commands to the mobile robot in a timely manner, and its services cannot be reconfigured correctly and promptly by the orchestrator, the robot may lose control and injure nearby workers. The fleet manager is responsible for determining the mobile robot's path to avoid collisions and transmitting points along the path to the robot's software components. A robot losing control would require: 1) an automatic restart of its onboard computer, which takes approximately 5 minutes, or 2) manual intervention, which could take up to 20 minutes. These downtimes are unacceptable because they significantly impact factory objectives.

[0007] When an SLA violation is predicted, migrating the service to another location that complies with the SLA is a promising option due to the high number of nodes in the environment.

[0008] Regardless, migrating critical services such as industrial control loops requires additional measures: a suitable destination node must be selected, and network and computing resources must be reconfigured to minimize the probability of SLA violations and reduce downtime. In industrial verticals, even sub-second downtime can lead to severe system outages and serious consequences.

[0009] Returning to the fleet manager's example, once a waiting time fault is detected, it can be moved as close to the production line as possible to meet its expected response time and avoid robot malfunctions.

[0010] In addition, each service may have a wide range of varying connectivity requirements, expressed in terms of a specific SLA (Service Level Agreement), which is represented using a mathematical relationship between defined Key Performance Indicators (KPIs) and mathematical thresholds.

[0011] The key is to ensure that the migration process does not lead to unacceptable SLA violations (regarding time or value) and that SLA compliance is maintained after the migration. The problem is that the current migration process does not consider industry-specific SLA-related KPIs during node selection, and SLA compliance cannot be guaranteed if network conditions change. The main problem can be broken down into two sub-problems: when to migrate services, given that the current approach involves recovery after a failure, i.e., after a few minutes of downtime; and where to migrate services, selecting the optimal destination node.

[0012] Therefore, the following issue needs to be addressed: In the event of changes in a given environment (e.g., failure of connected devices, network degradation, and moving to a shadow area), the best strategy for migrating a service to another location is to take into account its SLA requirements (e.g., bandwidth, latency, jitter, resilience, availability, reliability, and security, as well as any other task-specific KPIs) to keep up with the application's requirements.

[0013] In "State machine replication in containers managed by Kubemetes" (Netto, Lung, Correia, Luiz, Sá de Souza, JSA 2016), state machine replication is implemented in Kubemetes through consensus of stateful replicas to improve the fault tolerance of stateful services. In "Automatic Integration of BFT State-Machine Replication into IoT Systems" (Berger, Reiser, Hauck, Held, Domaschka, EDCC 2022), a framework for integrating state machine replication into k3s is proposed, emphasizing an event-driven interaction model on a client-server model, build-block principles, and automatic replica deployment. However, neither paper addresses criticality, migration, or factory environments.

[0014] US10776244B2 provides a method for modeling anticipated system migration between server systems. The method includes: a remote system detecting that a first server system does not have a gateway installed, and, following the detection, activating the gateway on the first server system to operate at least partially as a control point on the first server system, which provides access and access security to a second server system and provides services for migrating applications from the first server system to the second server system. However, the method does not include explicit strategies regarding the timing and location of the migration decision.

[0015] The paper "Proactive Virtual Machine Migration in Fog Environments" (Goncalves, Velasquez, Curado, Bittencourt and Madeira, ISCC 2018) proposes a VM (virtual machine) migration method based on mobility prediction. It minimizes a general cost function using an ILP model with predicted future data for the VM, taking the communication latency with a user in a fog cloud as an example. However, no decision-making process is provided for migration time, and it is unrelated to SLA level, criticality, and fault tolerance.

[0016] Given the above, there is a need to provide pre-emptive mitigation of SLA violations with optimal strategies in terms of timing and cost (hardware and time). In particular, in addition to in-situ mitigation as a countermeasure, for example, after an SLA violation is detected, it is also necessary to consider the service's criticality level to provide adaptive migration of the relevant service to the most appropriate destination node. Summary of the Invention

[0017] According to an aspect of this disclosure, an apparatus is provided for controlling multiple worker nodes of an application running thereon to provide services, the apparatus comprising a network orchestration unit, a detection and prediction unit, and a migration processing unit, wherein: The network orchestration unit is configured to monitor the plurality of working nodes and configure network resources for the plurality of working nodes; The detection and prediction unit is configured to detect or predict service level agreement (SLA) violations at a target node included in the plurality of working nodes; and The migration processing unit is configured to trigger the migration if the violation time associated with the predicted or detected SLA violation is less than the duration of the migration sensitivity window associated with the migration of the target service provided by the target application running on the target node to the destination node.

[0018] In some examples, the migration processing unit is configured to trigger the migration if the violation time is greater than the duration of the in-situ mitigation window associated with the in-situ mitigation of the detected or predicted SLA violation.

[0019] In some examples, the migration processing unit is configured to trigger on-site mitigation of the detected or predicted SLA violation if the violation time is less than the duration of the on-site mitigation window associated with the on-site mitigation.

[0020] In some examples, the duration of the migration sensitivity window is equal to the estimated migration time multiplied by a criticality factor, and the criticality factor is related to the criticality level of the target service.

[0021] In some examples, the values ​​of the criticality factors increase with respect to criticality levels in the order of first criticality level, second criticality level, third criticality level, and fourth criticality level, and the apparatus includes a criticality management unit configured to record and output the criticality levels.

[0022] In some examples, for the first critical level: The network orchestration unit is configured to preferably generate a copy of the target service at the destination node when the SLA violation is detected or predicted; and The migration processing unit is configured to migrate the target service to the destination node.

[0023] In some examples, for the second critical level: The network orchestration unit is configured to generate a cold standby copy of the target service at the destination node; and The migration processing unit is configured to preferably activate the cold standby copy at the destination node when an SLA violation is detected or predicted, and is configured to migrate the target service from the target node to a node different from the destination node among the plurality of nodes.

[0024] In some examples, the network orchestration unit is configured to send a delay instruction to the target node to delay the target service if the destination node is unable to mitigate the SLA violation within the estimated migration time.

[0025] In some examples, the network orchestration unit is configured to: target the third critical level: If no SLA violation has been detected or predicted at the target node, the target node is configured as the primary node for the target service, and a hot standby replica for the target service is generated at the destination node; and If an SLA violation is detected or predicted, the destination node is configured as the primary node for the target service, and the target node is configured as the hot standby replica.

[0026] In some examples, the network orchestration unit is configured to: target the fourth critical level: If the detected SLA violation is a violation of a security state, the application running on the target node is replaced with a security container used to return to the security state.

[0027] In some examples, the SLA includes a mathematical and statistical relationship between key performance indicators (KPIs) specified for the plurality of nodes and numerical KPI thresholds.

[0028] In some examples, the migration processing unit is configured to obtain the violation time for the target node based on a regression model built for sampled KPIs associated with the target node, wherein the violation time is calculated as the time after the regression model intersects with one or more KPI thresholds specified for the target node.

[0029] In some examples, the prediction and detection unit is configured to predict and / or detect the SLA violation based on sampled values ​​of the KPIs of the plurality of nodes.

[0030] In some examples, the apparatus further includes a score ranking unit configured to calculate a score for each of the plurality of nodes and an SLA associated with each node; rank the plurality of nodes based on the calculated scores; and the criticality management unit is configured to examine the plurality of nodes with scores, starting from the highest-ranked node, in descending order of score, determine whether a node has a score that is at least a predetermined threshold higher than the score of the target node, and select the node as the destination node.

[0031] In some examples, the device is configured to manage factory-related network and computing resources, and / or to manage production lines and deliver commands to industrial components, such as mobile, static, or semi-static components.

[0032] According to another aspect of this disclosure, a method is provided performed by an apparatus for controlling multiple worker nodes of an application running thereon to provide services, the apparatus including a network orchestration unit, a detection and prediction unit, and a migration processing unit, wherein the method includes: The network orchestration unit monitors the multiple working nodes and configures network resources for the multiple working nodes; The detection and prediction unit detects and predicts service level agreement (SLA) violations at the target node included in the plurality of working nodes; and The migration is triggered by the migration processing unit if the violation time associated with the predicted or detected SLA violation is less than the duration of the migration sensitivity window associated with the migration of the target service provided by the target application running on the target node to the destination node.

[0033] According to another aspect of this disclosure, a system including the device and a plurality of working nodes controlled by the device are provided, wherein each node includes a KPI sampling unit configured to collect sampled KPI values ​​of KPIs specified in an SLA for each node, and configured to provide the collected values ​​to the device.

[0034] In some examples, the device and the plurality of working nodes are configured in the cloud.

[0035] According to some example embodiments, a computer program is also provided, which includes instructions for causing a device to perform the methods disclosed herein.

[0036] According to some example embodiments, a memory is also provided that stores computer-readable instructions for causing the apparatus to perform methods as disclosed in this disclosure.

[0037] Furthermore, according to some other example embodiments, a computer program product is provided, for example, for a wireless communication device including at least one processor, including software code portions for performing corresponding steps disclosed herein when the product is run on the device. The computer program product may include a computer-readable medium on which the software code portions are stored. Furthermore, the computer program product may be directly loaded into the internal memory of a computer and / or may be transmitted via a network through at least one of the steps of uploading, downloading, and pushing.

[0038] While this document will describe some example embodiments with specific reference to the above applications, it should be understood that this disclosure is not limited to this field of use, but can be applied in a broader context.

[0039] It is worth noting that the methods according to this disclosure relate to methods of operating apparatus according to the above-described exemplary embodiments and variations thereof, and the corresponding statements made regarding the apparatus also apply to the corresponding methods, and vice versa; thus, similar descriptions may be omitted for brevity. Furthermore, even without explicit disclosure, the aspects described above can be combined in many ways. Those skilled in the art will understand that combinations of these aspects and features / steps are possible unless they create an explicitly excluded contradiction.

[0040] Implementations of the disclosed apparatus may include, but are not limited to, the use of one or more processors, one or more application-specific integrated circuits (ASICs) and / or one or more field-programmable gate arrays (FPGAs). Implementations of the apparatus may also include the use of other conventional and / or custom hardware such as software-programmable processors (such as graphics processing unit (GPU) processors).

[0041] Other and further exemplary embodiments of this disclosure will become apparent during the following discussion and by referring to the accompanying drawings. Attached Figure Description

[0042] Exemplary embodiments of the present disclosure will now be described by way of example only with reference to the accompanying drawings, in which:

[0043] Figure 1 An example of a mitigation process is illustrated schematically;

[0044] Figure 2 An example of estimating mitigation time is illustrated schematically;

[0045] Figure 3 An example of an SLA violation prediction algorithm is illustrated schematically;

[0046] Figure 4 and 5 Examples of migration time estimates for SLA criticality at different levels are illustrated schematically.

[0047] Figure 6 An example architecture of an orchestration system is illustrated schematically;

[0048] Figure 7 The illustration shows an example of usage for a performance SLA level;

[0049] Figure 8 The illustration shows an example of a use case for a high availability SLA level;

[0050] Figure 9 The illustrations depict examples of usage scenarios for resilient SLA levels; and

[0051] Figure 10 The illustration shows an example of how to use a safety SLA level. Description of Example Implementations

[0052] In the following description, examples of communication networks that can be used as examples of embodiments to which they can be applied will be used to describe different exemplary embodiments based on the communication network architecture of 3GPP standards for communication networks such as 5G / NR, without limiting the embodiments to such architecture. It will be apparent to those skilled in the art that the embodiments can also be applied to other types of communication networks where mobile communication principles are integrated with D2D (device-to-device) or V2X (vehicle-to-everything) configurations, such as SL (sidelink), for example, Wi-Fi, Global Microwave Interconnection Access (WiMAX), Bluetooth®, Personal Communication Services (PCS), ZigBee®, Wideband Code Division Multiple Access (WCDMA), systems using Ultra Wideband (UWB) technology, Mobile Ad Hoc Networks (MANET), wired access, etc. Furthermore, without loss of generality, while some examples of the embodiments are described in relation to mobile communication networks, the principles of this disclosure can be extended and applied to any other type of communication network, such as wired communication networks.

[0053] The following examples and embodiments should be understood as illustrative examples only. Although the specification may refer to "a," "an," or "some" (or more) examples or embodiments in several places, this does not necessarily mean that each such reference relates to the same (or more) examples or embodiments, or that a feature applies only to a single example or embodiment. Individual features of different embodiments may also be combined to provide other embodiments. Furthermore, terms such as "comprising" and "including" should be understood not to limit the described embodiments to consisting only of those features already mentioned; such examples and embodiments may also include features, structures, units, modules, etc., not specifically mentioned.

[0054] The infrastructure of a mobile communication system's (remote) communication network, including examples of applicable embodiments, may include an architecture comprising one or more communication networks, including (multiple) radio access network subsystems and (multiple) core networks. Such an architecture may include one or more communication network control elements or functions, access network elements, radio access network elements, access service network gateways, or base transceivers (such as base stations (BS), access points (AP), node Bs (NBs), eNBs or gNBs, distributed units (DUs), or centralized / central units (CUs)) controlling a corresponding coverage area or (multiple) cells, and having one or more communication stations, such as communication elements or functions, like user equipment or terminal equipment, like user equipment (UE), or another device with similar functionality, such as modem chipsets, chips, modules, etc., which may also be part of a communication-enabled station, element, function, or application, such as a UE, element, or function that can be used in a machine-to-machine communication architecture, or attached as a separate element to such a communication-enabled element, function, or application, capable of communicating via one or more channels, via one or more communication bundles, for transmitting several types of data in multiple access areas. In addition, it may include core network elements or network functions (such as gateway network elements / functions, mobility management entities, mobile switching centers, servers, databases, etc.).

[0055] The following description provides further details on alternatives, modifications, and changes: gNB includes, for example, nodes that provide NR user plane and control plane protocol termination to the UE, and are connected to the 5GC via the NG interface, for example, according to Section 3.2 of 3GPP TS 38.300 V16.6.0 (2021-06) which has been incorporated.

[0056] The gNB Central Unit (gNB-CU) includes, for example, a logical node that hosts, for example, the gNB's RRC, SDAP, and PDCP protocols and the en-gNB's RRC and PDCP protocols that control the operation of one or more gNB-DUs. The gNB-CU terminates the FI interface connected to the gNB-DU.

[0057] A gNB Distributed Unit (gNB-DU) includes, for example, a logical node that hosts the RLC, MAC, and PHY layers of, for example, a gNB or en-gNB, and its operation is partially controlled by the gNB-CU. A gNB-DU supports one or more cells. A cell is supported by only one gNB-DU. The gNB-DU terminates the F1 interface connected to the gNB-CU.

[0058] The gNB-CU control plane (gNB-CU-CP) includes, for example, a logical node that hosts the control plane portions of the gNB-CU's RRC and PDCP protocols, for example, for the en-gNB or gNB. The gNB-CU-CP terminates the E1 interface connected to the gNB-CU-UP and the F1-C interface connected to the gNB-DU.

[0059] The gNB-CU user plane (gNB-CU-UP) includes, for example, a logical node that hosts, for example, the user plane portion of the PDCP protocol for the gNB-CU for the en-gNB, and the user plane portion of the PDCP protocol for the gNB-CU for the gNB and the SDAP protocol. The gNB-CU-UP terminates the E1 interface connected to the gNB-CU-CP and the F1-U interface connected to the gNB-DU, for example, according to Section 3.1 of 3GPP TS 38.401 V16.6.0 (2021-07) which was incorporated.

[0060] Different functional divisions between central and distributed units are possible, for example, referred to as options: Option 1 (Class 1A split): The functional partitioning in this option is similar to the 1A architecture in the DC. The RRC is in the central unit. The PDCP, RLC, MAC, physical layer, and RF are in the distributed unit. Option 2 (3C Class Segmentation): The functional partitioning in this option is similar to the 3C architecture in a DC (Distributed Control) unit. RRC and PDCP are in the central unit. RLC, MAC, physical layer, and RF are in the distributed unit. Option 3 (Internal RLC Segmentation): Low RLC (partial RLC functionality), MAC, physical layer, and RF are located in the distributed unit. PDCP and high RLC (other RLC functionality) are located in the central unit. Option 4 (RLC-MAC splitting): The MAC, physical layer, and RF are located in the distributed unit. PDCP and RLC are located in the central unit. Alternatively, for example, according to Section 11 of 3GPP TR 38.801 V14.0.0 (2017-03) incorporated by reference.

[0061] gNB supports different protocol layers, such as Layer 1 (L1) - the physical layer.

[0062] NR's Layer 2 (L2) is divided into the following sublayers: Media Access Control (MAC), Radio Link Control (RLC), Packet Data Convergence Protocol (PDCP), and Service Data Adaptation Protocol (SDAP), among which, for example: o The physical layer provides a transmission channel to the MAC sublayer; o The MAC sublayer provides logical channels to the RLC sublayer; o The RLC sublayer provides RLC channels to the PDCP sublayer; o The PDCP sublayer provides radio bearers to the SDAP sublayer; o The SDAP sublayer provides 5GC QoS flows; o Comp refers to header compression and Segm refers to segmentation; o Control channels include (BCCH, PCCH).

[0063] Layer 3 (L3) includes, for example, Radio Resource Control (RRC), as per Section 6 of 3GPP TS38.300 V16.6.0 (2021-06) incorporated by reference.

[0064] RAN (Radio Access Network) nodes or network nodes, such as gNBs, base stations, gNB CUs, gNB DUs, or portions thereof, may be implemented using means, for example, having at least one processor and / or at least one memory (with computer-readable instructions (computer program)), which is configured to support and / or provide and / or process functions and / or features associated with the CU and / or DU, and / or at least one protocol (sub) layer of the RAN (Radio Access Network), such as layer 2 and / or layer 3.

[0065] The gNB CU and gNB DU portions may, for example, be located in the same location or physically separated. The gNB DU may even be further divided into, for example, two portions, one including processing equipment and the other including an antenna. The Central Unit (CU) may also be referred to as BBU / REC / RCC / C-RAN / V-RAN, O-RAN, or a portion thereof. The Distributed Unit (DU) may also be referred to as RRH / RRU / RE / RU, or a portion thereof. In the various exemplary embodiments of this disclosure herein, the CU-DP (or more generally, the CU) may also be referred to as a (first) network node supporting at least one of the Central Unit Control Plane Functions or Layer 3 Protocols of the Radio Access Network; and similarly, the DU may be referred to as a (second) network node supporting at least one of the Distributed Unit Functions or Layer 2 Protocols of the Radio Access Network.

[0066] gNB-DU supports one or more cells and can therefore be used as, for example, a serving cell for a user equipment (UE).

[0067] User equipment (UE) may include wireless or mobile devices, devices having a radio interface for interacting with a RAN (Radio Access Network), smartphones, in-vehicle devices, IoT devices, M2M devices, etc. Such a UE or device may include: at least one processor; and at least one memory including computer program code; wherein the at least one memory and the computer program code are configured, together with the at least one processor, to cause the device to perform at least certain operations, such as, for example, an RRC connection to the RAN. The UE may be configured, for example, to generate messages (e.g., including a cell ID) to be transmitted via radio to the RAN (e.g., to arrive at and communicate with the serving cell). The UE may generate and send and receive RRC messages containing one or more RRC PDUs (Packet Data Units).

[0068] The UE can have different states (e.g., according to sections 42.1 and 4.4 of 3GPP TS 38.331 V16.5.0 (2021-06) which have been incorporated).

[0069] When an RRC connection has been established, the UE is in, for example, the RRC_CONNECTED state or the RRC_INACTIVE state.

[0070] A UE in the RRC_CONNECTED state can: o Store AS background; o Transmit unicast data to / from the UE; o Monitor the control channels associated with the shared data channels to determine whether data should be scheduled for the data channels; o Provides channel quality and feedback information; o Perform neighboring cell measurements and measurement reports.

[0071] The RRC protocol includes, for example, the following main functions: o RRC connection control; o Measurement Configuration and Reporting o Create / change / release measurement configurations (e.g., intra-frequency, inter-frequency, and inter-RAT measurements); o Set and release the measurement gap; o Measurement report.

[0072] The general functions and interconnections described for the components and functions also depend on the actual network type, as is known to those skilled in the art and described in the corresponding descriptions; therefore, for the sake of brevity, their detailed descriptions may be omitted herein. However, it should be noted that, in addition to those described in detail below, several additional network components and signaling links may be employed for communication to or from components, functions, or applications, such as communication endpoints, communication network control elements (e.g., servers, gateways, wireless network controllers), and other components of the same or other communication networks.

[0073] The communication network architecture considered in the examples of the embodiments can also communicate with other networks, such as the public switched telephone network or the Internet. The communication network can also support the use of cloud services for virtual network elements or their functions. It should be noted that the virtual network portion of a telecommunications network can also be provided by non-cloud resources (e.g., internal networks). It should be understood that network elements and / or corresponding functions of access systems, core networks, etc., can be implemented using any nodes, hosts, servers, access nodes, or entities suitable for such use. Generally, network functions can be implemented as network elements on dedicated hardware, as software instances running on dedicated hardware, or as virtualized functions instantiated on a suitable platform (e.g., cloud infrastructure).

[0074] Furthermore, network elements (such as communication elements (like UE, terminal equipment)), control elements or functions (such as access network elements (like base stations / BS, gNB, radio network controllers)), core network control elements or functions (such as gateway elements), or other elements or functions, as well as other elements, functions, or applications, as described herein, may be implemented by software (e.g., by computer program products for computers) and / or by hardware. The devices, nodes, functions, or network elements used to perform their respective processes and are used may include several parts, modules, units, components, etc. (not shown) required for control, processing, and / or communication / signaling functions. Such components, modules, units, and parts may include, for example: one or more processors or processor units including one or more processing sections for executing instructions and / or programs and / or processing data; storage and memory units or components (e.g., ROM, RAM, EEPROM, etc.) for storing instructions, programs, and / or data and for serving as working areas for processors or processing sections; input or interface components (e.g., floppy disks, CD-ROMs, EEPROMs, etc.) for inputting data and instructions via software; user interfaces (e.g., screens, keyboards, etc.) for providing users with the possibility of monitoring and manipulation; and other interfaces or components for establishing links and / or connections under the control of processor units or sections (e.g., wired or wireless interface components, radio interface components including, for example, antenna units, components for forming radio communication sections, etc.), wherein the corresponding components forming the interface (such as radio communication sections) may also be located at remote sites (e.g., radio heads or radio base stations, etc.). It should be noted that in this specification, "processing section" should not be considered merely as a physical portion of one or more processors, but can be considered as a logical portion involving processing tasks executed by one or more processors. It should be understood that, based on some examples, a so-called "fluid" or flexible network concept can be adopted, in which the operation and function of network elements, network functions, or another entity of the network can be performed in a flexible manner across different entities or functions (such as nodes, hosts, or servers). In other words, the "division of labor" between the network elements, functions, or entities involved can vary depending on the circumstances.

[0075] Referring now to the accompanying drawings. In particular, it should be noted that, unless otherwise indicated, the same or similar reference numerals used in the drawings of this disclosure may indicate the same or similar elements, and thus, for the sake of brevity, their repetitive descriptions may be omitted. It should also be noted that, as those skilled in the art may understand or realize, even though the drawings may appear to refer to some specific / definite message name / type, these messages may, of course, have different names and / or be communicated / exchanged in different forms / formats depending on various implementations (e.g., underlined techniques).

[0076] An orchestration system is a distributed system responsible for automatically placing, deploying, monitoring, and migrating packaged software (e.g., containers) on computing infrastructure, acting as a cloud operating system. An orchestration system consists of a control plane and a cluster of computers. The control plane receives requests for deployment, monitors the application's status, and manages the lifecycle of the packaged software. The cluster of computers consists of worker nodes on which the packaged applications are deployed (according to, for example, MA Rodriguez and R. Buyya, “Container-Based Cluster Orchestration Systems: A Taxonomy and Future Directions,”). Wiley Software: Practice and Experience, (2019). On deployment requests, the control plane places the application on worker nodes via a scheduling process. The orchestration system reacts accordingly when the current state deviates from the desired stable state. For example, the application's desired stable state requires three load-balanced service replicas, but one is unresponsive. A replacement is performed on a different worker node, and the resources of the failed replica are released. This process, referred to herein as migration, involves additional maintenance for stateful service requirements.

[0077] In this disclosure, the migration of a service includes generating a copy at a worker node (or worker node, used interchangeably) selected as the destination node for (re)generating the service, which was previously provided at another node where an SLA violation was detected or predicted. This process may involve reserving relevant resources at the destination node by a network orchestration system or platform.

[0078] Preferably, service migration refers to the migration of the software application(s) providing the service, provided that the service (or, alternatively, the software application providing the service) is stateless (i.e., does not maintain any state). In the case of stateful services, migration involves recreating the service, and more preferably, physically replicating the storage running the service from the previous node to the destination node.

[0079] The purpose of this disclosure is to ensure SLA-aware migration of services, such as those in a variable factory floor, based on digital KPIs.

[0080] The decision-making process determines when to trigger service migration to reduce downtime.

[0081] This disclosure includes elements as described below. 1. A differentiated mitigation method for SLA violations based on SLA criticality levels; 2. An algorithm / method for selecting a time window during which migration is convenient and causes little or no service interruption. The algorithm is based on migration time estimates and can be fine-tuned based on several parameters: a. Ensure sufficient margin to avoid time deviations in downtime. b. Lag parameters that ensure algorithm stability c. Avoid gain thresholds for migrations between comparable nodes. d. Key mission requirements; 3. An algorithm / method for selecting the most suitable destination node for migrating one or more service replicas, wherein the service replicas are characterized by a scoring function with configurable weights. The destination node is selected based on the following: a. The cost of migration (in terms of time). b. The occupancy status of the destination node. c. The number and location of customers served. d. Availability of access technologies (e.g., Wi-Fi, private / public LTE, Multifire, and 5G) at the destination node. e. The overall state of the network; 4. An architecture for implementing prior methods and algorithms for automated resource management;

[0082] This disclosure relies on the following assumptions: i) Multiple nodes exist in different locations, and each has a different number and type of resources controlled by an orchestration system (e.g., Kubernetes, Docker, etc.). ii) Complex changes exist in the environment, such as robot movement, newly deployed services (e.g., telemetry streaming, camera feeds, process control), and resource reallocation. iii) A programmable framework is available that enables KPI-driven operations (e.g., latency toward a given destination, bandwidth, jitter, RSSI, etc.), enabling network services to be provisioned across multiple access interfaces. iv) Each service has network requirements in terms of KPIs, expressed as an SLA to be met at runtime, and v) SLAs are divided into four critical levels: performance, high availability, resilience, and security.

[0083] Figure 1 The activity diagram depicts the mitigation differentiation methods provided in this disclosure. Unless otherwise specified, activities also belong to higher-level flows for high availability and high SLA levels. Figure 1 As shown, different SLA levels are depicted with different patterns of text boxes. See details. Figure 1 The different text boxes are shown at the bottom.

[0084] An I4.0 application typically consists of multiple services (or, in general, software wrapped in virtualization technology) running in one or more containers that implement business logic for use cases. The constituent services belong to four levels of SLAs characterized by increasing criticality levels as described below.

[0085] Performance SLA This refers to non-critical services that require guaranteed performance at all times, where temporary breaches of these requirements will not cause any critical failures. In this disclosure, the performance SLA level is also referred to as the first critical level.

[0086] High Availability SLA High Availability SLA Levels are required for services that can experience SLA breaches but only last for a negligible duration (e.g., 1 minute) within a given time window (e.g., 1 month). Downtime must be minimized through redundancy techniques such as cold standby replicas. In this disclosure, the High Availability SLA Level is also referred to as the Second Critical Level.

[0087] Flexible SLA This is required by services that cannot be subject to any form of SLA breach. These services may reserve resources for multiple copies (hot or cold) to avoid interruptions regardless of the failure of any subsystem or component (according to, for example, J.-C. Laprie, “From Dependability to Resilience”). IEEE / IFIP Inti. Conf. Dependable Systems and Networks, (2008). In this disclosure, the Resilient SLA level is also referred to as the Third Critical Level.

[0088] Safety SLA Safety-critical services are those for which failure could result in hazardous conditions for people and / or the environment. These services must be deployed on certified hardware and must be reconfigured to provide safety checks and safety-stopping behaviors depending on the nature of the use case. For example, in the case of a malfunctioning robot, safety stopping involves stopping any movement and releasing the clutches in its joints. In this disclosure, the safety SLA level is also referred to as Level 4 Critical.

[0089] The criticality level (or interchangeable criticality levels) in this disclosure is "a designation of the level of fault tolerance guarantee required for a system component," as exemplified by, for example, A. Bums and RI Davis, "Mixed Criticality Systems - A Review: (Feb. 2022)," York, 2022.

[0090] The higher the criticality level of a service, the more severe the consequences of a failure related to that service will be, and therefore the corresponding mitigation of SLA violations should, for example, be earlier in time than the timing of detected or predicted SLA violations. Criticality level is related to risk, i.e., the product of the probability of a failure and the potential losses that the failure may cause.

[0091] In this disclosure, SLA violation mitigation includes in-situ mitigation, i.e., countermeasures performed locally at the node where the SLA violation is detected or predicted. SLA violation mitigation also includes migrating the service violating the SLA to a destination node that complies with the SLA, wherein the migration is preferably performed before the SLA violation is detected or occurs, and more preferably when the SLA violation has been predicted.

[0092] The core idea of ​​the method proposed in this disclosure is to mitigate SLA violations locally when possible, trigger migration when local countermeasures are insufficient to restore the SLA, and employ adaptive strategies for each SLA criticality level. Services are migrated and generated elsewhere, for example, in factories that comply with the SLA.

[0093] In this context, the SLA can be specified as a general logical function of conditions evaluated between the measured value of the defined KPI and the specified threshold using statistical and data operators (e.g., (95th percentile of network latency < 15ms and average bandwidth in the last second > 10mbit / s) or (packet delivery delay == 0 for the last 50 packets)).

[0094] Continuous monitoring can detect different types of anomalies. For example, a brief outage may cause network anomalies. In this context, the most important thing is to make quick decisions (in...) Figure 1 In the "Migration Time Estimation and Determination" step, it is necessary to determine whether to apply appropriate in-situ mitigation, i.e., to evaluate KPI mitigation without a migration container.

[0095] Examples of mitigation include interface switching, such as switching from 5G to Wi-Fi. On the other hand, examples of network-level mitigation include slicing rearrangement, reallocation of network resources, and antenna tilting.

[0096] These recovery actions can be sufficient to restore the SLA. On the other hand, in the case of permanent failures such as hardware failures to the network interface card (NIC), these recovery actions will be useless and migration will be required. Figure 1 The "Trigger migration for backup" step in the process.

[0097] In summary, migrating critical services like those in an industrial control loop is a significant task and requires additional safeguards. In an industrial context, there are several services with deadlines on the order of milliseconds, and anomaly detection and mitigation strategies can provide more robust safeguards. Furthermore, migrations take several seconds, and deadlines for critical services must be met under all circumstances. Mitigation methods depend on the criticality level because, for example, if a service is critical and has copies of other services, an expensive migration can be triggered immediately without impacting service availability, while non-replicated services should be restored as appropriately as possible, as the migration will severely impact their availability.

[0098] In the following, combined Figure 1 A differentiated approach for migrating SLA violations is provided, highlighting specific aspects of each criticality level. Performance SLA

[0099] Suppose there is an edge service performing some non-critical tasks, and an SLA violation is predicted or detected, where the response time is higher than a threshold.

[0100] Along with the network, node resources must be rearranged to relocate services to optimal locations to comply with the response times specified in the SLA, ensuring low latency.

[0101] Therefore, the orchestration system reserves resources for a replica to start service, migrating it when the replica starts. The replica can be a hot (standby) replica or a cold (standby) replica. At this critical level, a few seconds of service downtime will not cause catastrophic failure, so reserving resources in advance is not strictly necessary, leading to underutilization. Therefore, they must be reserved as late as possible, i.e., directly when an SLA violation is detected or predicted. Figure 1 The "Find and Reserve Destination Node" step in the process.

[0102] In cases with numerous false alarms, more migrations than needed will be triggered, degrading QoS. On the other hand, false alarms will cause brief outages. In both cases, there will be no serious consequences. SLA violation predictors can rely on more complex models than SLA violation detectors, which should be easier to run in real-time with low overhead. High Availability SLA

[0103] Migrating services with high availability requirements necessitates a different approach to the previous situation. In industrial environments, high availability requirements can also be interpreted as low mean time to repair (MTTR) versus downtime requirements, as high-availability services cannot tolerate even a few seconds of downtime. The orchestration system finds a new location for the service, rescheduling compute and communication resources with negligible or no application interruption.

[0104] Therefore, the main difference in the approach lies in the use of fault-tolerance techniques, such as redundancy, to maintain the availability requested by the service. In practice, relying solely on prediction, for example, could lead to significant downtime. The orchestration system extracts resources reserved for cold standby copies. Due to the constantly evolving environment, cold copies are periodically moved to ensure the lowest possible downtime. Copy placement can utilize additional information, such as industrial processing information.

[0105] In the absence of additional information, the orchestration system can provide feedback to the factory floor to delay or avoid SLA violations. Figure 1 The "feedback to the shop floor" step in this context reduces operational technology (OT) capabilities and eliminates compromises, allowing for proper placement of backups. For example, OT entities can receive feedback: a robot can slow down to save time understanding what's happening, or a robot arm can stop, and actuators can adaptively send commands at a lower sampling frequency to avoid complete failure, etc. In summary, each service interacting with the real-world environment can have a degraded mode where it still functions to prevent complete failure / SLA violations.

[0106] After the migration, the network and node resources previously allocated to the failed replica are rearranged to create another cold standby replica, which may replace the failed replica. Flexible SLA

[0107] In addition to technologies used for high availability SLAs, hot standby replicas or other advanced technologies, such as dual redundancy or triple modular redundancy (TMR) models implemented in critical contexts, may also be involved. These schemes guarantee the lowest possible probability of service interruption, potentially avoiding violations entirely. Service must still be guaranteed despite component failures.

[0108] For example, in TMR, the output of a service is provided by a majority vote among three replicas. If one of them fails, the service is still guaranteed.

[0109] Therefore, for this critical level, the method anticipates that after SLA violation detection / prediction, the replica manager must consider the failure and may relinquish the role of the primary replica with a hot standby replica scheme. Figure 1 The "transfer of primary role" step in the process. Next, network and computing resources are rearranged to recreate a hot spare copy (or peer copy) elsewhere, where the SLA is adhered to. Figure 1 The phrase "rearrange the network and nodes used for the backup replica migration" appears in the text.

[0110] However, redundancy is only useful when replicas do not have common-cause failures. Therefore, the node selection algorithm must also evaluate common-cause failures during the rescheduling of resources used to generate replicas. Safety SLA

[0111] This criticality level is assigned to services where misconduct would harm the environment and people. All technologies used in resilient SLAs remain effective. Furthermore, the concept of safety reconfiguration exists to prevent hazards. Safety status can be identified through fault data analysis (FDA). Safety-critical events are detected in real time through environmentally aware monitoring, and the container implementing the equipment's behavior must be replaced with a safety container to allow only the previously identified safety status to be considered. Figure 1 (Refer to "Reconfiguration of safety state" in the document). If identifying a safety state is not feasible, then a safety state or a safety stop is a valid option.

[0112] Therefore, considering the four severity levels with progressively increasing criticality, a combined adaptive and automatic mitigation approach is implemented through the mitigation process proposed above. This approach uses in-situ mitigation when possible and service migration when in-situ mitigation is insufficient to address violations in a timely manner. This proposed approach enables the selection of the optimal mitigation method for SLA violations, and it is also suitable for the service's severity level. Therefore, mitigation is service-oriented and efficient in terms of time and resources.

[0113] Below, a method for migration assessment and decision-making is provided, explaining the corresponding time required for local / in-situ mitigation and migration in the event of an SLA violation being detected or predicted. An example algorithm is also shown.

[0114] The goal is to have a "connect first, disconnect later" approach, meaning to prevent SLA violations from occurring before they happen, whenever possible. The following inequality must be satisfied: t_monitoring+t_mitigation <t_viol。

[0115] Therefore, t_monitoring is the time required for sampling, computation, and running the anomaly detection / prediction algorithm. t_mitigation is the duration of the migration process until the SLA is restored, determining the Mean Time to Recovery (MTTR). On the right-hand side (RHS), t_viol is the time for SLA violations due to fault / error propagation or due to unexpected scenarios. t_monitoring is included on the left-hand side (LHS) because, in the worst case, the fault is detected at the end of the monitoring window, after it has already propagated.

[0116] Decomposing LHS: t_sampling+t_detection+t_decision+t_mit+t_node_sel+t_exc+t_download+t_act <t_viol

[0117] Combining the LHS inequality above, Figure 2 Example time estimates for mitigation are shown in the worst-case scenario where an SLA violation is detected and both local mitigation and migration are used to mitigate the SLA violation after detection.

[0118] Specifically, t_sampling is needed to collect n samples, t_detection is needed to run anomaly detection / prediction algorithms, t_decision is needed to determine the mitigation of the application, t_mitigation is needed to determine the duration of local mitigation, t_node_sel is needed to run destination node selection algorithms, t_exc is needed to exchange information between the control plane and the compute cluster, t_download is needed to download containers on the destination node, and t_act is needed to reserve resources and start containers.

[0119] The decision-making algorithm determines when to implement local mitigation or trigger a migration. `t_migration` must be estimated based on the service's criticality level. This decision-making process begins with SLA violation prediction, which is typically derived from error detection or trend estimation. In the former case, an error is detected, such as a hardware / software component malfunction that could lead to an SLA violation. In the latter case, mutation mitigation causes a decrease in KPIs even in the absence of a fault.

[0120] The regression coefficients for each KPI are estimated on the n most recent samples of the KPI values ​​for each node. The number of samples n to be considered is calculated as a function of the statistical exponent of the KPIs considered, i.e., between 10 and 60 samples, with more samples if the R^2 exponent is far from 1. The regression model used is a linear model because it is simple enough to be run regularly at the high frequency required in industrial scenarios.

[0121] If the regression line intersects with one of the KPI thresholds specified in the SLA, causing the entire SLA's logistic function to no longer hold (i.e., a SAT solver solution), then an SLA violation can be predicted, and the evaluation phase can begin. The time of violation, t_viol, can be calculated as the time following the intersection.

[0122] During this phase, local mitigation should be implemented, if possible, based on non-compliant KPIs.

[0123] The estimate of t_migration is calculated, and if t_viol is greater than the duration s The sensitivity window for t_migration indicates that the violation is still far off in time and migration must be implemented. Otherwise, migration is triggered. The s value depends on the criticality level and must be defined in the SLA.

[0124] If the regression coefficients increase and there are no longer any violations of the prediction, the evaluation phase stops.

[0125] Below, example algorithms for migration evaluation and decision-making are provided. Do_mitigation (t_viol) / / Start the evaluation phase, triggered when estimate_ttv returns a non-inf value. { Start_inplace_mitigation (); t_migration = fetch t_migration / / t_migration is updated periodically by another task; the most recent value is retrieved here. While Detection or (Prediction != inf) { if t_viol t_migration / / Conditions for immediate migration { trigger_migration(); } Else / / Periodically sample to check for migration or application of in-situ mitigation. { Wait for the end of t_sampling; / / In principle, this can be a monitoring cycle phase different from the mitigation function. Prediction = estimate_ttv (KPIs, SLA); Detection = SLA_violation_detection (KPIs, SLA); } } Stop_inplace_mitigation(); } Update_migration() { t_migration = estimate_migration (); Store t_migration Wait for a period of time; } estimate_ttv (KPIs, SLA) {​ For each KPI in KPIs { n = compute_number_samples (KPI); RC = regression_on_the_KPI (KPI, n); If RC<0 / / For GT KPI operator <0, for LT KPI operator >0. { ttv_vec.append (RC / (Current-SLA_thresh)); violated_KPIs.append (KPI) } } Violation_predicted = SAT_solver(SLA, violated_KPIs) / / Check if any violated KPIs could lead to a violation of the entire SLA. If violation_predicted { ttv = min {ttv_vec | ttv belongs to violated_KPIs} / / Find the set of KPIs that caused the violation. return ttv } Else return inf }

[0126] Using the migration assessment and decision-making method proposed above, SLA violations can be predicted in advance, and the time required for mitigation (including in-situ mitigation and migration) can be estimated. Furthermore, the configuration of the migration sensitivity (time) window provides specific conditions for deciding whether to perform a migration, i.e., when the predicted time of the violation is less than the duration of the migration sensitivity (time) window. This ensures sufficient migration time for mitigating the violation, particularly since the configuration of the migration sensitivity (time) window equals the estimated migration time multiplied by the criticality level perception factor. Therefore, the adaptive migration assessment and decision-making method can be applied to specific criticality levels of services: for example, for services with higher criticality levels, a larger factor can be configured to ensure sufficient advance execution of migration. Preferably, the criticality level perception factor is, for example, an integer in the range of 1 to 10. Thus, the most appropriate migration decision can be provided for specific services with corresponding criticality levels.

[0127] When an SLA violation is predicted to occur in the future, the algorithm for evaluating migration or in-situ mitigation begins. Specifically, if the time of the violation is longer than the duration of the migration sensitivity (time) window, in-situ mitigation can be performed based on the violation's KPIs; if the time of the violation is shorter than the duration of the migration sensitivity (time) window, preemptive migration can be performed.

[0128] like Figure 2 The estimated migration time shown is preferably performed before any in-situ mitigation or migration is initiated, allowing for an advance decision on whether to perform in-situ mitigation or migration before a predicted violation occurs. This improves the efficiency of mitigation and enables responses to violations even before they occur.

[0129] Figure 3 An example implementation of an SLA violation prediction method / algorithm is shown: t_s represents the sampling time, and the circles are samples of one of the KPIs specified in the SLA. The sample sequence shows a decreasing trend. The horizontal dashed line represents the threshold considered acceptable for a given SLA. RC is the regression coefficient. Thanks to this, the intersection between the threshold and the actual sampled KPI trend can be calculated. Therefore, t_viol, the time interval from the most recent KPI sample to the predicted violation, can be derived.

[0130] Figure 4 and 5 Example estimates of migration time are shown for different levels of SLA criticality. In real-world scenarios, time can vary considerably. Figure 4 An estimate of the time required to migrate the service to a hot standby copy is shown, and Figure 5 The estimated time required to migrate a service to a cold standby copy is shown. The basic understanding is that a hot copy is active, receiving input and computing output without sending them, while a cold copy is inactive and must start with reserved resources.

[0131] The methods for destination node selection are provided below. Example algorithms are also shown.

[0132] The service experience is characterized by a complete mitigation workflow with an SLA of at least C criticality level. The factory floor consists of N nodes, of which K nodes are suitable for hosting workloads with at least C criticality. Each node is represented by a vector of resources. and S features The vector description is given by S, where S is the number of active SLAs, i.e., the SLAs for system admission. S distinct vectors are needed because each SLA can specify a particular request, such as latency from a specific node, making it impossible to have a unique set of KPIs that represents the capabilities of nodes adhering to each SLA in the system. The number of R free hardware resources for a node, such as free memory, CPU, disk, available sensors, accelerometers, actuators, etc. It includes F KPIs sampled periodically, such as network latency; these represent the guarantees that the SLA can be met. This information is aggregated into a single score for each SLA, representing the ability to comply with the SLA.

[0133] Nodes must be filtered based on their available resources and capabilities to comply with the KPIs specified in the SLA, i.e., thresholding the R and F vectors based on requirements. The remaining nodes suitable for hosting services must be ranked based on their highest SLA compliance score and lowest migration cost. The node with the highest score is selected for migration.

[0134] The scores are updated regularly by the KPI monitoring system.

[0135] Periodically, and after migration, the network re-optimization process is simulated to discover the optimal resource allocation. This process can involve different techniques: network slicing rearrangement, antenna tilting or repositioning, etc.

[0136] First, based on The process filters out k suitable nodes with sufficient resources from a pool of k nodes. Nodes with hardware utilization exceeding a specified threshold are also filtered out to prevent resource waste. The remaining nodes are ranked by score, which is updated periodically, and the ranking for each SLA is tracked by the control plane. The final ranking also includes penalties for resource utilization and migration costs. The former prevents congestion of the best nodes, as a node with a high score for its SLA may have a high score for other SLAs that attract service. The latter takes into account the migration cost between two nodes and is given by the following formula. cost = t_exc+t_download+t_issue.

[0137] In cases of elastic SLA or higher, the penalty for nodes suffering common-cause failures is also subtracted.

[0138] If the first node in the ranking has a score that is at least as high as the threshold parameter as the current node, then it is selected as the target node for migration.

[0139] The score is calculated using the following equation: Where i is the feature vector of SLA s KPI index in j It is one of the K nodes, and des This is the expected SLA requirement. w_i It is the weight of KPI and Furthermore, the specific requirements differ for each SLA.

[0140] The score follows the rule that the higher the better. The idea is that a score that is slightly worse or slightly better than the current score makes a big difference, while a very high or very low score has an impact, but it is unrelated to how high or low the score is.

[0141] Next, the orchestration system can attempt to schedule deployment requests from the queue, potentially preempting lower-priority deployment requests.

[0142] The relocation of computing and network resources must be joint because the two are functionally interdependent: network relocation is used to create dedicated channels to accelerate the migration process, and node resources may be necessary for hosting network functions.

[0143] Below, example algorithms for node scoring and destination node selection are shown. Node_selection() { Filtered0 = Filter_out_by_criticality (All_nodes); Filtered1=Filter_out_by_KPIs (Filtered0); Filtered2=Filter_out_by_available_resources (Filtered1); For all nodes in filtered2 { Migration_cost = (t_exc+t_download+t_issue ) weight? If SLA_level>=resiliency CCF_penalty = compute_common_cause_failures (node); Else CCF_penalty = 0 Score = compute_score (KPIs, requirements); Overcrowding_penalty = compute_ overcrowding (node,score); overall_score =score - migration_cost - overcrowding_penalty - CCF_penalty; } Pick node with highest score in filtered2 } / / Note: The filtering function can be implemented in one step, but for clarity, it is implemented in three steps in the algorithm. Compute_overcrowding (score) { Penalty = 0 For all resources in { if occupancy > threshold Penalty += Occupancy - threshold } Return penalty score } Compute_score() { } KPIs update

[0144] Utilizing the node scoring and destination node selection methods proposed above, this paper ranks the node list for each SLA based on node and network environment mutations, continuously updating node-based resources and capabilities to meet the KPIs specified in the SLA. Thus, at each time step, and particularly when predicting an SLA violation, the top-ranked node with the highest score can be selected as the destination node for migration, conditioned on the score of this top-ranked node being higher than a predetermined threshold by the current node providing the service to be migrated. This ensures that the most suitable node can be selected immediately upon making a migration decision. Therefore, the efficiency of preempting SLA violations is further improved.

[0145] Figure 6 A system diagram of the main components of the framework is provided. Items enclosed in bold text boxes indicate elements introduced into the architecture by this disclosure, while other elements represent basic Kubernetes components, i.e., with the orchestrator as a reference.

[0146] The architecture is divided into a control plane, which consists of services deployed on one or more master nodes of the control cluster, and worker nodes, which consist of services deployed on each node of the cluster. Services in the control plane communicate with each other to manage the cluster state, while worker nodes exchange information and commands with the control plane through an API server.

[0147] For each worker node, some basic Kubernetes services and other additional services must be deployed; therefore, the worker node must be a machine-enabled Linux system with sufficient resources for these services. On the other hand, due to its resilient design, the services must be designed to run flexibly on embedded hardware such as IoT gateways and embedded computers. Example worker nodes include PLC-controlled devices (like conveyor belts and other factory automation machines), robots, AGVs and AGHs, as well as infrastructure network components (e.g., IoT gateway service routers, access points, and photonic switches) and cloud instances.

[0148] The main components of the architecture are described in detail below. KPI sampler

[0149] Device Manager is responsible for: a. Enable the data collection endpoint; b. Collect data from the managed system and send the data to the SLA manager (in...) Figure 6 middle); c. Locally compile SLAs and program microservices for evaluating SLA conditions; d. Run local network and system detection functions; e. Manage network connections on the managed systems; f. Performing periodic tasks (e.g., timers) and data-driven events (e.g., actions based on KPI predicates).

[0150] The Device Manager is designed to offload the aforementioned tasks to microservices running in the cloud. To ensure smooth operation even in the event of disconnection and network interruption, the microservices implementing tasks c, d, e, and f can be deployed locally on the Device Manager node. SLA Manager

[0151] Several KPI samplers are provided by the SLA Manager component ( Figure 6 This component manages and provides centralized access to the latest KPIs, network interface counters, system counters, and allocated SLAs. It is responsible for: a. Track the current state of managed KPI sampler instances by storing data related to enabled endpoints, SLA activity, probe functionality, and other device manager active services on a relational database; b. Enable server-side network probing, including one-way ping, and network capacity measurement between the managed device manager and itself; c. Store the requested and permitted SLAs in the system; d. Compile the SLA request and break it down into commands to create network probing capabilities, data collection endpoints, notifications, and actions based on the SLA KPIs. Perform SLA violation checks based on the collected data and the conditions specified in the SLA request. Score Ranking Manager

[0152] Score Ranking Manager Figure 6 The score is calculated periodically for each node and each SLA, and the ranking is continuously updated when a migration is triggered.

[0153] This component is responsible for: a. The services and functions necessary for calculating node scores; b. Calculate the score for each node and each SLA that is allowed access to the system; c. Track the ranking of scores; d. Assess the potential downgrade of KPIs after migration. Migration Processor

[0154] Migration processor ( Figure 6 This component determines when to trigger service migration. It is responsible for: a. Assess the migration time for services affected by SLA violations; b. Decide when to trigger the migration; c. Handle the state transition of services, in the case of stateful services; d. Issue commands for service replication management (input expansion, output merging, vote management, etc.), including virtual network configuration. Violation detection and prediction services

[0155] Violation detection and prediction services ( Figure 6 This component relies on collected KPIs to execute algorithms to predict and detect SLA violations for all SLAs. This component is responsible for: a. Run SLA violation detection and prediction algorithms on the collected KPIs; b. Send an alert to the migration processor in case of a violation; c. Assess the KPIs of the failure and implement on-site mitigation.

[0156] It is worth noting that the violation detection and prediction service must also be distributed across all nodes, analyzing the nodes' own data to reduce latency. This local service then reports to the central service after analysis.

[0157] KPIs that must be analyzed in a very short time are analyzed locally; otherwise, they are analyzed by a service in the control plane that is logically unified but can be physically replicated for scalability reasons. Network orchestration services

[0158] This component is an orchestrator for programmable networks. Its responsibilities include: a. Manage network resources (computing and radio resources); b. Track network channel quality through distributed network monitoring; c. Reconfigure network slices through VNF management when needed; d. Implement node-specific interface switching to increase the probability of SLA compliance; e. Configure a virtual network. Critical Manager

[0159] This component is a plugin for the Kubernetes scheduler component. Its responsibilities include: a. Make the Kube scheduler aware of SLAs and criticality levels; b. Communicates with other introduced components, acting as a bridge; c. Influence or override the decisions of the native kube scheduler and implement the node selection algorithm for migration.

[0160] The criticality manager in the scheduler is responsible for taking the first node in the list, checking if it is available to host a new application, and making that decision. If the first node in the list is not suitable for hosting a new application (e.g., because the list is too full), the criticality manager removes the second node from the list, and so on, until a suitable destination node is found. For example, if the first-ranked node (i.e., the node with the highest score among multiple nodes) has a score that is at least a predetermined threshold higher than the target node, then the first-ranked node is selected as the destination node; otherwise, other nodes are checked in descending order of score until a suitable destination node (e.g., with a score that is at least a predetermined threshold higher than the target node) is found and selected. Feedback to factory floor services

[0161] The service communicates with SLA violation detection and prediction teams to provide feedback to the factory floor when needed to prevent major disruptions. This component is responsible for: a. Track all possible capability reductions / functional eliminations in the equipment shop; b. Select the appropriate countermeasure based on the information received from the SLA violation manager; c. Send orders to the factory workshop to implement countermeasures.

[0162] Below, Table 1 shows an example SLA specification. Table 1: Examples of SLA Specifications Use Case Examples

[0163] Below, examples of use cases for each SLA criticality level are provided, such as... Figures 7 to 10 As shown.

[0164] like Figure 7 As illustrated, as an example of performance SLA, imagine a drone image stream being routed to an edge service for some non-critical inventory applications. Because the drone is moving, the quality of the network channel changes over time, and as the drone moves away from its connected 5G base station (…), the quality will also change. Figure 7 At point B), communication latency increases, and service response times no longer adhere to the SLA, thus pre-predicting SLA violations risks incorrect storage billing. Moving edge services to the IoT gateway itself and switching connectivity to Wi-Fi reduces response time. The orchestration system reserves resources on the gateway to generate services and, when it starts, implements migration and switches network channels. The 5G network is rearranged to free up network resources allocated to drones.

[0165] like Figure 8 As shown, for a high-availability SLA, imagine a fleet manager controlling two AGVs (Automated Guided Vehicles). For example, these robots are used in the freight transport industry, and they move along paths determined point-by-point by the fleet manager in the edge cloud using infrared sensors. The service cannot tolerate even a few seconds of downtime, as this would mean the robots blindly moving several meters, potentially colliding and going out of control. When the robots move away from their connected base stations, the fleet manager must also be migrated to a nearby base station to maintain low response time. In this way, a shadow service is created that follows the AGVs along their paths. Because the AGV paths are known due to the processing information, a cold standby copy can be placed in a nearby base station. When a network switch is performed and the robots connect to gateway B, the cold copy is activated and the previous fleet manager is migrated elsewhere.

[0166] In cases where cold services are in the wrong location and cannot be replaced on time based on migration time estimates, feedback to the workshop can reduce the movement speed of AGVs or stop them, thus saving time for migration. Regularly evaluate KPIs and migrate backup services to so-called nearby base stations to ensure minimal downtime and SLA violations in the event of migration.

[0167] like Figure 9 As shown, for a resilient SLA, a replica on the gateway acts as the master node, sending values ​​to the actuators, while another (on server A) is a hot standby. The actuators receive control values ​​via a Wi-Fi link, but it is also enabled by 5G. The connection to the gateway is degraded due to moving obstacles in the plant.

[0168] To ensure SLA regardless of failures and completely avoid service interruptions, the orchestration system reconfigures the hot replica to take over the role of the primary replica and provide seamless service.

[0169] The primary role is transferred to the hot standby service on server A in the diagram, relying on the 5G channel. After a joint reorganization of network and computing resources, the hot standby replica is regenerated on edge server B in the diagram, ensuring adequate response time for the SLA.

[0170] The two servers have few or no common causes of failure, so the risk of both replicas having common failures is very small.

[0171] like Figure 10 As shown, regarding safety SLAs, a mobile robot is envisioned handling potentially hazardous tools. An operator approaches to check some displayed information, or simply because it crosses the robot's path. Although an unmanned area exists in the workshop, signaled by LEDs, the worker cannot see it and enters that area.

[0172] The presence of a human is detected, and the full production container is immediately replaced by a safe production container whose behavior depends on the action currently taken by the robot that entered the most recent safe state.

[0173] These use cases fall into three general contexts: static environments, where machines are static but redundancy is needed for reliability purposes to handle failures; mobile environments, where redundancy is needed to adapt to a constantly changing environment; and finally, reconfigurable environments, where machinery is semi-static but is periodically moved and reconfigured to change factory production goals and production lines are rented for short periods like a few weeks. An example for the first scenario is a factory with software partitioned into microservices deployed on the edge cloud. An example for the second scenario is a factory heavily utilizing AGVs and drones of varying sizes already deployed in the current factory. An example for the third scenario, as described by the vision of Industry 4.0, is a factory with surface mount technology (SMT) production lines that are periodically rented to customers who own intellectual property (IP) but cannot physically produce chips because it is a very expensive activity. Production lines must be reconfigured, moved, and rearranged according to purpose and customer.

[0174] The following explains how the proposed SLA-driven allocation infrastructure management system for software components can be used to manage components of a 5G core network.

[0175] Four scenarios can be identified, one of which is the previous natural evolution. 1. The core of all clouds 2. Cloud partially in the cloud and partially on edge devices 3. Some core user applications in customer locations are in the cloud. 4. Everything is on the premises.

[0176] The core, entirely in the cloud, is the default architecture for 5G applications, without specific latency or critical requirements. With SLA-driven allocation of components capable of managing a large number of cloud nodes, necessary components are selected for deployment among the most reliable nodes. When new deployment requests with latency-critical requirements arrive, portions of the core network need to be moved to the edge to reduce communication latency. In this case, User Plane Functions (UPF) or other parts of the user plane can be moved to the edge cloud, even onto embedded devices, which can be radio-enabled devices. This minimizes communication latency. In practice, the SLA-driven orchestrator considers only nodes close to the User Equipment (UE) as suitable for adhering to the latency specified in the SLA.

[0177] If a request not only has latency sensitivity requirements but also reliability requirements (such as greater reliability or availability than the internet connection to the cloud), then the core network must be moved to the on-premises customer site. In this case, for example, a deployment request for core network components would specify a criticality that nodes in the cloud cannot meet, thus only nodes at the edge would be selected as suitable for hosting network functions. However, leveraging the scalability of cloud resources, some services that do not have latency or criticality requirements can still be deployed on a remote cloud.

[0178] The different scenario occurs when all networks must benefit from criticality, security, and safety guarantees. In this case, remote connectivity to the cloud cannot be permitted, and the entire infrastructure, including the 5G core network, must be deployed at the customer's premises. In this scenario, core functions are deployed on edge servers within the premises, while user plane functions can be deployed as close as possible to the user equipment, even on embedded devices.

[0179] In cross-scenario scenarios, the orchestration system plays a crucial role because the evolution between one scenario and another is completely seamless and driven by deployment requests and their programmed SLAs. The orchestrator adaptively changes the deployment scenario to comply with the SLA.

[0180] This disclosure provides a unified programmable framework for implementing SLAs using cloud-native technologies such as Kubernetes and Docker. A key objective of the platform in this disclosure is the ability to manage compute resources for network, reliability, and security requirements expressed individually for each connected application. This falls under the category of software orchestration and automation for 6G, specifically in the following aspects: 1. Fog computing and extreme edge computing 2. Decompose RAN (DU / CU) 3. 6G Decentralized PaaS for Vertical Applications 4. Professional 6G services 5. Depth slicing 6. Intent- and SLA-driven networks

[0181] The contributions in this disclosure extend network slicing capabilities (e.g., performance, latency, and reliability implementations) to any connected computing system (including network core services), enabling unified multi-stakeholder orchestration. This allows for SLA-driven end-to-end service orchestration on top of capability orchestration in each domain (including a unified RAN core), employing a layered approach to 6G services and network management. The SLA-driven end-to-end service orchestration vision, central to this disclosure, will also enable the concept of deep slicing to extend the specific slice composition of microservices and allocate dedicated hardware / software stacks to the RAN to increase the specialization level and efficiency of 6G networks by defining and implementing use case / UE-specific SLAs. This will enable flexible feature placement options that consider service requirements and multiple factors such as cloud capabilities, hardware accelerators and network / computer utilization, as well as other network-specific KPIs.

[0182] In summary, this disclosure proposes improved efficiency in mitigation by considering the criticality level of services, predicting failures and migrating workloads elsewhere before violations occur, deciding whether to migrate (a pre-selection between on-site mitigation and migration), when to migrate (long before the failure actually occurs), and where to migrate (the most suitable destination node to meet SLA requirements), as well as the execution of migration, thus making mitigation adaptive. It also proposes implementing KPI-based SLA components and allocating network services across different access networks.

[0183] Violations of prediction and target node selection can be achieved in different ways using the same basic approach as presented in this disclosure.

[0184] Simulations of future scenarios that identify failures can be used instead of KPI monitoring, and simulations that predict migration times can be used to trigger migrations as specified in previous sections of this disclosure.

[0185] Instead of ranking nodes, graphs can also be used to evaluate the conditions for optimal migrations in operation, still using the same methods presented in this disclosure.

[0186] Instead of the criticality level of the application, other relevant parameters can be used to differentiate the methods without changing the nature of the differentiated countermeasures based on the risk of violation, where risk is defined as the probability of failure multiplied by the severity of the consequences, as proposed in this disclosure. List of abbreviations AGV: Automated Guided Vehicle AMH: Automated Material Handling SLA: Service Level Agreement KPI: Key Performance Indicators OT: Operations Technology PLC: Programmable Logic Controller

[0187] As described above, although in the example embodiments described above (refer to the accompanying drawings), messages transmitted / exchanged between network components / elements may appear to have specific / explicit names depending on various implementations (e.g., underscore techniques), these messages may have different names and / or be transmitted / exchanged in different forms / formats, as can be understood and appreciated by those skilled in the art.

[0188] According to some example embodiments, corresponding methods suitable for execution by the apparatus (network element / component) described above, such as UE, CU, DU, etc., are also provided.

[0189] However, it should be noted that the device (equipment) features described above correspond to the corresponding method features, which may not have been explicitly described for the sake of brevity. The disclosure of this document is also intended to extend to such method features. In particular, this disclosure is understood to relate to methods of operating the equipment described above, and / or providing and / or arranging the corresponding elements of such equipment.

[0190] In addition, according to some other example embodiments, corresponding devices (e.g., implementing the UE, CU, DU, etc. as described above) are also provided, which include at least one processing circuit and at least one memory for storing instructions to be executed by the processing circuit, wherein the at least one memory and the instructions are configured such that the corresponding device performs at least the corresponding steps as described above using the at least one processing circuit.

[0191] In some other example embodiments, corresponding apparatus (e.g., implementing the UE, CU, DU, etc. as described above) is provided, which includes: a corresponding component configured to perform at least the corresponding steps as described above.

[0192] It should be noted that the examples of embodiments of this disclosure can be applied to a variety of different network configurations. That is, the examples shown in the accompanying drawings described above, which serve as the basis for the examples discussed above, are merely illustrative and do not limit this disclosure in any way. Specifically, additional existing and proposed new functionalities available in the appropriate operating environment can be used in conjunction with examples of embodiments of this disclosure based on defined principles.

[0193] It should also be noted that the disclosed example embodiments can be implemented in many ways using hardware and / or software configurations. For example, the disclosed embodiments can be implemented using dedicated hardware and / or hardware associated with software executable thereon. The components and / or elements in the drawings are merely examples and do not limit the scope of use or functionality of any hardware, software combined with hardware, firmware, embedded logic components, or combinations of two or more such components that implement particular embodiments of this disclosure.

[0194] It should also be noted that the specification and drawings are merely illustrative of the principles of this disclosure. Those skilled in the art will be able to implement various arrangements, which, although not expressly described or shown herein, embody the principles of this disclosure and are included within its spirit and scope. Furthermore, all examples and embodiments outlined in this disclosure are primarily intended for illustrative purposes only to aid the reader in understanding the principles of the proposed methods. Moreover, all statements and specific examples provided herein regarding the principles, aspects, and embodiments of this disclosure are intended to cover their equivalents.

Claims

1. An apparatus for controlling multiple worker nodes of an application running thereon to provide services, the apparatus comprising a network orchestration unit, a detection and prediction unit, and a migration processing unit, wherein: The network orchestration unit is configured to monitor the plurality of working nodes and configure network resources for the plurality of working nodes; The detection and prediction unit is configured to detect or predict service level agreement (SLA) violations at target nodes included in the plurality of working nodes; as well as The migration processing unit is configured to trigger the migration if the violation time associated with the predicted or detected SLA violation is less than the duration of the migration sensitivity window associated with the migration of the target service provided by the target application running on the target node to the destination node.

2. The apparatus of claim 1, wherein the migration processing unit is configured to trigger the migration if the violation time is greater than the duration of an in-situ mitigation window associated with the in-situ mitigation of the detected or predicted SLA violation.

3. The apparatus of claim 1 or 2, wherein the migration processing unit is configured to trigger on-site mitigation of the detected or predicted SLA violation if the violation time is less than the duration of the on-site mitigation window associated with the on-site mitigation.

4. The apparatus according to any one of claims 1 to 3, wherein the duration of the migration sensitivity window is equal to the estimated migration time multiplied by a criticality factor, and the criticality factor is related to the criticality level of the target service.

5. The apparatus of claim 4, wherein the value of the critical factor increases with respect to criticality levels in the order of a first criticality level, a second criticality level, a third criticality level, and a fourth criticality level, and the apparatus comprises: It is a critical management unit configured to record and output the critical level.

6. The apparatus of claim 5, wherein for the first criticality level: The network orchestration unit is configured to preferably generate a copy of the target service at the destination node when the SLA violation is detected or predicted; and The migration processing unit is configured to migrate the target service to the destination node.

7. The apparatus according to claim 5 or 6, wherein the second critical level is specified as follows: The network orchestration unit is configured to generate a cold standby copy of the target service at the destination node; and The migration processing unit is configured to preferably activate the cold standby copy at the destination node when an SLA violation is detected or predicted, and is configured to migrate the target service from the target node to a node different from the destination node among the plurality of nodes.

8. The apparatus of claim 7, wherein the network orchestration unit is configured to send a delay instruction to the target node to delay the target service if the destination node is unable to mitigate the SLA violation within the estimated migration time.

9. The apparatus according to any one of claims 5 to 8, wherein the network orchestration unit is configured to: for the third critical level: If no SLA violation has been detected or predicted at the target node, the target node is configured as the primary node for the target service, and a hot standby replica for the target service is generated at the destination node; and If an SLA violation is detected or predicted, the destination node is configured as the primary node for the target service, and the target node is configured as the hot standby replica.

10. The apparatus according to any one of claims 5 to 9, wherein the network orchestration unit is configured to: for the fourth key level: If the detected SLA violation is a violation of a security state, the application running on the target node is replaced with a security container used to return to the security state.

11. The apparatus according to any one of claims 1 to 10, wherein the SLA comprises: The mathematical and statistical relationship between the key performance indicators (KPIs) specified for the multiple nodes and the numerical KPI thresholds.

12. The apparatus of any one of claims 1 to 11, wherein the migration processing unit is configured to obtain the violation time for the target node based on a regression model established for sampled KPIs associated with the target node, wherein the violation time is calculated as the time after the regression model intersects with one or more KPI thresholds specified for the target node.

13. The apparatus of any one of claims 1 to 12, wherein the prediction and detection unit is configured to predict and / or detect the SLA violation based on sampled values ​​of the KPIs of the plurality of nodes.

14. The apparatus according to any one of claims 1 to 13, further comprising a scoring and ranking unit configured to calculate a score for each of the plurality of nodes and an SLA associated with each node; The nodes are ranked based on the calculated scores; as well as The critical management unit is configured to examine the plurality of nodes with scores starting from the highest-ranked node in descending order of score, determine whether a node has a score that is at least a predetermined threshold higher than the score of the target node, and select the node as the destination node.

15. The apparatus according to any one of claims 1 to 14, wherein the apparatus is configured to manage factory-related network and computing resources, and / or to manage production lines and to deliver commands to mobile and / or static and / or semi-static industrial components in a timely manner.

16. A method performed by means of an apparatus for controlling a plurality of worker nodes of an application running thereon to provide services, the apparatus comprising a network orchestration unit, a detection and prediction unit, and a migration processing unit, wherein the method comprises: The network orchestration unit monitors the multiple working nodes and configures network resources for the multiple working nodes; The detection and prediction unit detects and predicts service level agreement (SLA) violations at target nodes included in the plurality of working nodes; as well as The migration is triggered by the migration processing unit if the violation time associated with the predicted or detected SLA violation is less than the duration of the migration sensitivity window associated with the migration of the target service provided by the target application running on the target node to the destination node.

17. A system comprising an apparatus according to any one of claims 1 to 15, and a plurality of working nodes controlled by the apparatus, wherein each node includes a KPI sampling unit configured to collect sampled KPI values ​​of KPIs specified in an SLA for each node, and configured to provide the collected values ​​to the apparatus.

18. The system of claim 17, wherein the device and the plurality of working nodes are configured in the cloud.

19. A computer program comprising instructions for causing a device to perform the method according to claim 16.

20. A memory storing computer-readable instructions for causing a device to perform the method according to claim 16.

Citation Information

Patent Citations

  • Consolidation planning services for systems migration

    US10776244B2