First node, second node, first data center and methods performed thereby for handling resources
Patent Information
- Application Number
- PCT/IN2025/050350
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2026-09-17
Smart Images

Figure IN2025050350_17092026_PF_FP_ABST
Abstract
Description
[0001] FIRST NODE, SECOND NODE, FIRST DATA CENTER AND METHODS PERFORMED THEREBY FOR HANDLING RESOURCES
[0002] TECHNICAL FIELD
[0003] The present disclosure relates generally to a first node and methods performed thereby for handling resources. The present disclosure also relates generally to a second node, and methods performed thereby for handling the resources. The present disclosure further relates generally to a first data center, and methods performed thereby for handling the resources. BACKGROUND
[0004] Computer systems in a communications network or communications system may comprise one or more nodes. A node may comprise a processing circuitry which, together with computer program code may perform different functions and actions, a memory, a receiving port, and a sending port. A node may be, for example, a server. Nodes may perform their functions entirely on the cloud.
[0005] The communications system may cover a geographical area which may be divided into cell areas, each cell area being served by a type of node, a network node in the Radio Access Network (RAN), radio network node or Transmission Point (TP), for example, an access node such as a Base Station (BS), e.g., a Radio Base Station (RBS), which sometimes may be referred to as e.g., gNB, evolved Node B (“eNB”), “eNodeB”, “NodeB”, “B node”, or Base Transceiver Station (BTS), depending on the technology and terminology used. The base stations may be of different classes such as e.g., Wide Area Base Stations, Medium Range Base Stations, Local Area Base Stations, and Home Base Stations, based on transmission power and thereby also cell size. A cell may be understood to be the geographical area where radio coverage may be provided by the base station at a base station site. One base station, situated on the base station site, may serve one or several cells. Further, each base station may support one or several communication technologies. The telecommunications network may also comprise network nodes which may serve receiving nodes, such as user equipments, with serving beams.
[0006] The standardization organization Third Generation Partnership Project (3GPP) is currently in the process of specifying a New Radio Interface called Next Generation Radio or New Radio (NR), as well as a Fifth Generation (5G) Packet Core Network, which may be referred to as 5G Core Network (5GC), abbreviated as 5GC.
[0007] Cloud RAN may be understood as a means of deploying RAN functions using a cloud native paradigm, where the compute functionality may be deployed on Commercial off-the-shelf software (COTS) hardware, and these computes may be connected to each other as wellas to external entities via a data center switching infrastructure. Computes may be understood to refer to hardware components providing processing power and / or processing capability, including resources such as Central Processing Unit (CPU), memory, storage, etc.
[0008] Figure 1 is a schematic diagram illustrating a non-limiting example of a network deployment architecture in RAN-Cloud. Particularly, Figure 1 depicts an example of two data centers. A data center may be understood to be a facility that may house servers, storage systems, and networking equipment to store, process, and manage data. A data center may be able to ensure reliable and secure operations for businesses and services. A typical data center network deployment such as that depicted in Figure 1 , may include, as part of the transport 1 section, Provider Edge Gateways (PE-GW) 2, as part of the data fabric 3 Data center Gateways (DC-GWs) 4, spine switches 5, represented as “Spine” in Figure 1 , leaf switches for compute multi-homed connectivity 6, represented as “Leaf” in Figure 1 , client / servers (C / S) 7, represented as “C / S” in Figure 1 , and, as part of the Control Fabric 8 racks 9, control switching (ctrl-SW) 10, and Management Gateways (MgmtGWs) 11. The different components may be interconnected with lines interconnected with links. The capacity of each link may depend on throughput and dimensioning designed for the network.
[0009] The PE-GW 2 may be understood to be a network device that may connect the internal network of the data center to external networks, such as the internet or other data centers.
[0010] The DC-GWs 4 may be understood to be a customer edge network device that may enable to connect the data center to a transport network.
[0011] The spine switches 5 may be understood to be switches that may be able to interconnect leaf switches in a two layer leaf-spine data center.
[0012] The leaf switches 6 may be understood to be access switches that may may be able to aggregate traffic from all the servers and connect to the spine in a leaf-spine data center.
[0013] The C / S 7 may be understood to be compute or storage nodes that may be able to connect to the leaf switch layer in a leaf-spine data center.
[0014] The racks 9 may be understood to be frames in a data center that may be used to mount compute, storage and network equipment.
[0015] The ctrl-SWs 10 may be understood to be switches that may be part of a separate management network, and be may be able to connect the management ports of the computes, network switches, and storage to provide connectivity to a management and operations center.
[0016] The MgmtGWs 11 may be understood to be switches or routers that may be able to connect the management network to external transport networks.
[0017] The Open - Radio Access Network (O-RAN) alliance may be understood to define a standardized architecture for RAN, in which the O-Cloud may be a set of hardware and software to provide cloud computing capabilities to host and run O-RAN network functions.Service Management and Orchestrations (SMOs) may be understood to be service orchestration and management entities responsible for managing RAN workloads across multiple such O-clouds. They may typically include placement algorithms for Centralized Unit Control Plane (CU-CP), Centralized Unit User Plane (CU-UP) and other applications related to cloud RAN, based on application dimensioning and placement constraints with respect to each other. For example, there may be constraints on geographical distances, latencies between functions, and area coverage for the various RAN functions.
[0018] Existing methods to provide cloud services my involve high costs of hardware and maintenance needs.
[0019] SUMMARY
[0020] As part of the development of embodiments herein, one or more challenges with the existing technology will first be identified and discussed.
[0021] Current O-cloud deployments may incorporate hardware overprovisioning principles and may typically deploy double the network capacity than what may be needed so as to provide full High Availability in case of any failures. Network infrastructure may be dimensioned for over-provisioning, and may be deployed in a redundant manner. The extra infrastructure capacity may only be needed when a failure may occur in any of the network infrastructure components, to prevent loss of service. Workloads such as Centralized Unit (CU), Distribution Unit (DU) and other RAN applications may continue to function without any corrective action from the orchestrator in case of infrastructure failures. These may be typically handled reactively with the help of failure alarms and metrics. This may cause wastage of footprint, power and hardware resources under normal operating conditions. These issues may get compounded as cloud RAN deployments may scale, which may be expected to run into hundreds of thousands of small sites.
[0022] A typical data center network deployment may include the following levels of redundancy as shown in Figure 1 : redundant leaf switches for compute multi-homed connectivity, redundant spine switches, redundant DC-GWs, and redundant ports and links between each component to account for link failures.
[0023] Kubernetes scheduling algorithms for the RAN function workloads within a data center may be understood to be primarily based on compute capacity available e.g., Computer Programming Unit (CPU) and memory availability, they may be understood to not consider dynamic network capacity. Instead, the network is statically dimensioned and typically overprovisioned for peak load with redundancy.
[0024] According to the foregoing, it is an object of embodiments herein to improve the handling of resources in a communications system.According to a first aspect of embodiments herein, the object is achieved by a computer-implemented method, performed by a first node. The method is for handling resources. The first node operates in a computer system. The first node obtains one or more first indications indicating a first performance of one or more bridge ports of a first data center of the computer system. The first node also obtains one or more second indications indicating a second performance of one or more Internet Protocol Interfaces of the first data center of the computer system. The first node further obtains one or more third indications indicating a third performance of one or more compute systems of the first data center of the computer system. The first node determines, based on the obtained one or more first indications, one or more second indications, and one or more third indications, and using one or more trained machine learning models, a fourth indication of an availability of resources at the first data center at a future time period. The first node then initiates determining one or more actions based on the determined fourth indication. The one or more actions comprise migrating at least a part of a workload from a first structure comprised in the first data center to another structure comprised in the first data center or in a second data center.
[0025] According to a second aspect of embodiments herein, the object is achieved by a computer-implemented method, performed by the second node. The method is for handling the resources. The second node operates in the computer system. The second node obtains: i) first one or more first indications indicating a fourth performance of one or more bridge ports of a third data center of the computer system, ii) second one or more second indications indicating a fifth performance of one or more Internet Protocol Interfaces of the third data center of the computer system, and iii) third one or more third indications indicating a sixth performance of one or more compute systems of the third data center of the computer system. The second node trains the one or more machine learning models to determine the fourth indication of an availability of resources at the third data center at the future time period, based on the obtained first one or more first indications, second one or more second indications, and third one or more third indications. The second node initiates outputting one or more sixth indications of the one or more trained machine learning models to the first node operating in the computer system.
[0026] According to a third aspect of embodiments herein, the object is achieved by the first node, for handling the resources. The first node is configured to operate in the computer system. The first node is configured to obtain i) the one or more first indications configured to indicate the first performance of the one or more bridge ports of the first data center of the computer system, ii) the one or more second indications configured to indicate the second performance of the one or more Internet Protocol Interfaces of the first data center of the computer system, and iii) the one or more third indications configured to indicate the third performance of the one or more compute systems of the first data center of the computersystem. The first node is also configured to determine, based on the one or more first indications, the one or more second indications, and the one or more third indications configured to be obtained, and using the one or more trained machine learning models, the fourth indication of the availability of resources at the first data center at the future time period. The first node is further configured to initiate determining the one or more actions based on the fourth indication configured to be determined. The one or more actions are configured to comprise migrating at least the part of the workload from the first structure configured to be comprised in the first data center to another structure configured to be comprised in the first data center or in the second data center.
[0027] According to a fourth aspect of embodiments herein, the object is achieved by the second node, for handling the resources. The second node is configured to operate in the computer system. The second node is configured to obtain: i) the first one or more first indications configured to indicate the fourth performance of the one or more bridge ports of the third data center of the computer system, ii) the second one or more second indications configured to indicate the fifth performance of the one or more Internet Protocol Interfaces of the third data center of the computer system, iii) the third one or more third indications configured to indicate the sixth performance of the one or more compute systems of the third data center of the computer system. The second node is also configured to train the one or more machine learning models to determine the fourth indication of the availability of resources at the third data center at the future time period, based on the first one or more first indications, second one or more second indications, and third one or more third indications configured to be obtained. The second node is further configured to initiate outputting the one or more sixth indications of the one or more trained machine learning models to the first node configured to operate in the computer system.
[0028] According to a fourth aspect of embodiments herein, the object is achieved by the first data center. The first data center is configured to lack redundant network structures.
[0029] By obtaining the one or more first indications, the one or more second indications and the one or more third indications, the first node may be enabled to use the obtained indications to monitor the performance and application availability of the lean infrastructure of the first data center with predictive mechanisms and ultimately assess the need to move workload beyond a rack or even beyond the first data center.
[0030] By determining the fourth indication, the first node may rely on AI / ML techniques to monitor the infrastructure and application availability with predictive mechanisms. If the first data center is predicted to have an availability issue at the future time period, the first node may be enabled to mitigate for this situation by for example, migrating workloads to a different compute in the same rack, different rack or a different site based on the predictive feedback provided by the AI / ML components in embodiments herein. The availability of the workload atan application level may thus be improved. The first node may therefore be enabled to ensure the high availability of the first data center, while maintaining its lean structure. Accordingly, the first node may enable to fulfill the goal of embodiments herein to achieve the same level of service availability for an application workload with fewer infrastructure resources, while at the same time having the ability to accommodate failure scenarios and maintenance activities. This scheme may be understood to rely on the highly distributed nature of cloud RAN deployments, due to which the potential availability zone for a workload may be moved beyond a rack or even beyond a site.
[0031] By initiating determining the one or more actions comprising migrating at least a part of the workload from the first structure to another structure, the first node may enable to maintain a high availability (HA) SLA threshold without having a larger hardware (Hw) footprint.
[0032] By obtaining the first one or more first indications, the second one or more second indications and the third one or more third indications, the second node may then be enabled to use the obtained indications to provide input to the one or more ML models and enable that the one or more ML models be used, e.g., by the first node, to ensure day to day operations may be with at par level of performance in terms of availability. The one or more ML models may be provided metrics related to compute performance, interface and bridge port which may enable to identify the port and link health of a data center, such as the third data center, e.g., during the training phase of the one or more ML models and the first data center during the interference phase of the one or more ML models. Based on collected metrics, the one or more ML models may be trained for identification of anomalies and subject matter expert (SME) suggested rule based alarm or faults. These anomalies and faults may be used to train the one or more ML models for classification of e.g., a port as a faulty port or a healthy one.
[0033] By training the one or more ML models, the first node may enable any node, e.g., the second node, to rely on AI / ML techniques to monitor the infrastructure and application availability of lean structures such as the first data center and the third data center, with predictive mechanisms to detect availability problems of such lean structures before they may occur. Predicting the availability problems may then enable to mitigate availability issues by for example, migrating workloads to a different compute in the same rack, different rack or a different site based on the predictive feedback provided by the AI / ML components. The availability of the workload at an application level may thus be improved. The second node may therefore enable to ensure the high availability of a structure such as the first data center, while maintaining its lean structure. Accordingly, the second node may enable to fulfill the goal of embodiments herein to achieve the same level of service availability for an application workload with fewer infrastructure resources, while at the same time having the ability to accommodate failure scenarios and maintenance activities. This scheme may be understoodto rely on the highly distributed nature of cloud RAN deployments, due to which the potential availability zone for a workload may be moved beyond a rack or even beyond a site.
[0034] By outputting the one or more sixth indications, the second node may enable the first node to perform the method described above, and obtain the benefits recited.
[0035] BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Examples of embodiments herein are described in more detail with reference to the accompanying drawings, according to the following description.
[0037] Figure 1 is a schematic diagram illustrating a non-limiting example of a data center, according to existing methods.
[0038] Figure 2 is a schematic diagram illustrating two non-limiting examples, in panel a) and b), respectively, of a computer system, according to embodiments herein.
[0039] Figure 3 is a schematic diagram illustrating a non-limiting example of a data center, according to embodiments herein.
[0040] Figure 4 is a flowchart depicting embodiments of a method in a first node, according to embodiments herein.
[0041] Figure 5 is a flowchart depicting embodiments of a method in a second node, according to embodiments herein.
[0042] Figure 6 is a schematic diagram depicting illustrating a non-limiting example of a method performed by a first node, according to embodiments herein.
[0043] Figure 7 is a schematic diagram depicting illustrating a non-limiting example of a method performed by a second node, according to embodiments herein.
[0044] Figure 8 is a schematic diagram depicting illustrating a non-limiting example of a method performed by a first node, according to embodiments herein.
[0045] Figure 9 is a schematic diagram depicting illustrating a non-limiting example of a method performed by a second node, according to embodiments herein.
[0046] Figure 10 is a schematic diagram depicting illustrating a non-limiting example of a method performed by a first node, according to embodiments herein.
[0047] Figure 11 is a schematic block diagram illustrating a non-limiting example of a first node, according to embodiments herein.
[0048] Figure 12 is a schematic block diagram illustrating a non-limiting example of a second node, according to embodiments herein.
[0049] Figure 13 is a schematic block diagram illustrating a non-limiting example of a first data center, according to embodiments herein.DETAILED DESCRIPTION
[0050] Certain aspects of the present disclosure and their embodiments address one or more of the challenges identified with the existing methods and provide solutions to the challenges discussed.
[0051] Embodiments herein may be understood to address the problems identified with the existing methods and may be understood to relate to optimizing cloud RAN infrastructure from redundant to lean network architectures.
[0052] Embodiments herein may be understood to provide the concept of a thin-network. Such a network may be understood to be a deliberately under-provisioned and non-redundant network with respect to the hardware infrastructure, to save on footprint and power.
[0053] The trade-off of deploying such a network infrastructure may be understood to be that while it may support the full workload under normal operating conditions, any infrastructure fault, such as single link failure or single switch failure may cause loss of service due to the lack of redundancy, and lack of over-provisioning in the fabric. It may be also difficult to perform maintenance activities such as infrastructure upgrades without loss of service.
[0054] Embodiments herein may be understood to enable a mitigation for this situation by relying on AI / ML techniques to monitor the infrastructure and application availability with predictive mechanisms. This scheme may be understood to rely on the highly distributed nature of cloud RAN deployments, due to which, the potential availability zone for a workload may be moved beyond a rack or even beyond a site.
[0055] Examples of embodiments herein may rely on SMO schedulers and RAN applications (rAPPs) to be able to migrate the workloads to a different compute in the same rack, different rack or a different site based on the predictive feedback provided by the AI / ML components in embodiments herein, thus improving the availability of the workload at an application level.
[0056] The overall goal of embodiments herein may be understood to be to achieve the same level of service availability for an application workload with fewer infrastructure resources, while at the same time having the ability to accommodate failure scenarios and maintenance activities.
[0057] The embodiments will now be described more fully hereinafter with reference to the accompanying drawings, in which examples are shown. In this section, embodiments herein are illustrated by exemplary embodiments. It should be noted that these embodiments are not mutually exclusive. Components from one embodiment or example may be tacitly assumed to be present in another embodiment or example and it will be obvious to a person skilled in the art how those components may be used in the other exemplary embodiments. All possible combinations are not described to simplify the description.
[0058] Figure 2 depicts two non-limiting examples, in panels “a” and “b”, respectively, of a computer system 100, in which embodiments herein may be implemented. In some exampleimplementations, such as that depicted in the non-limiting example of Figure 2a, the computer system 100 may be a computer network. In other example implementations, such as that depicted in the non-limiting example of Figure 2b, the computer system 100 may be implemented in a telecommunications system, sometimes also referred to as a telecommunications network, cellular radio system, cellular network, or wireless communications system. In some examples, the telecommunications system may comprise network nodes which may serve receiving nodes, such as wireless devices. The computer system 100 may for example be a network such as a 5G system, or a newer system supporting similar functionality, such as for example, a Sixth Generation (6G) system. In some examples, the computer system 100 may support, additionally or alternatively, a Long-Term Evolution (LTE) network and may support other technologies such as a for example, LTE Frequency Division Duplex (FDD), LTE Time Division Duplex (TDD), LTE Half-Duplex Frequency Division Duplex (HD-FDD), and LTE operating in an unlicensed band. The telecommunications system may also support other technologies, such as Wideband Code Division Multiple Access (WCDMA), Universal Mobile Telecommunications System Terrestrial Radio Access (UTRA) TDD, Global System for Mobile communications (GSM) network, GSM / Enhanced Data Rate for GSM Evolution (EDGE) Radio Access Network (GERAN) network, Ultra-Mobile Broadband (UMB), EDGE network, network comprising any combination of Radio Access Technologies (RATs) such as e.g. Multi-Standard Radio (MSR) base stations, multi-RAT base stations etc., any 3rd Generation Partnership Project (3GPP) cellular network, Wireless Local Area Network / s (WLAN) or WiFi network / s, Worldwide Interoperability for Microwave Access (WiMax), IEEE 802.15.4-based low-power short-range networks such as IPv6 over Low-Power Wireless Personal Area Networks (6LowPAN), Zigbee, Z-Wave, Bluetooth Low Energy (BLE), or any cellular network or system. The telecommunications system may for example support a Low Power Wide Area Network (LPWAN). LPWAN technologies may comprise Long Range physical layer protocol (LoRa), Haystack, SigFox, LTE-M, and Narrow-Band loT (NB-loT).
[0059] The computer system 100 may comprise a plurality of nodes, whereof a first node 111, and a second node 112 are depicted in Figure 2. It may be understood that the computer system 100 may comprise and / or operate in communication with more or less nodes than those represented on Figure 2.
[0060] Any of the first node 111 and the second node 112 may be understood, respectively, as a first computer system and a second computer system. In some examples, any of the first node 111 and the second node 112 may be implemented as a standalone server in e.g., a host computer in the cloud 115, as depicted in the non-limiting example depicted in panel b) of Figure 2. Any of the first node 111 and the second node 112 may in some examples be a distributed node or distributed server, with some of their respective functions beingimplemented locally, e.g., by a client manager, and some of their functions implemented in the cloud 115, by e.g., a server manager. Yet in other examples, any of the first node 111 and the second node 112 may also be implemented as processing resources in a server farm.
[0061] The computer system 100 may also comprise a first data center 121 and a second data center 122. In some embodiments, the computer system 100 may further comprise a third data center 123. Any of the first data center 121, the second data center 122 and the third data center 123 may be understood to lack redundant structures. Particularly, the first data center 121 may lack redundant structures. Further particularly, the first data center 121 lacks redundant network structures. Any of the first data center 121 , the second data center 122 and the third data center 123 may comprise at least one of a single spine structure and a single leaf structure. Any of the first data center 121 , the second data center 122 and the third data center 123 may, respectively, comprise at least one of a first plurality of clients and servers, and a second plurality of racks. Any of the first data center 121 , the second data center 122 and the third data center 123 may be a cloud data center.
[0062] The computer system 100 may also comprise a device 130. The device 130 may be also known as a e.g., UE, wireless device, mobile terminal, wireless terminal and / or mobile station, mobile telephone, cellular telephone, or laptop with wireless capability, an Internet of Things (loT) device, or a Customer Premises Equipment (CPE), just to mention some further examples. The device 130 in the present context may be, for example, portable, pocket-storable, hand-held, computer-comprised, or a vehicle-mounted mobile device, enabled to communicate voice and / or data, via a RAN, with another entity, such as a server, a laptop, a Personal Digital Assistant (PDA), or a tablet, a Machine-to-Machine (M2M) device, an Internet of Things (loT) device, e.g., a sensor or a camera, a device equipped with a wireless interface, such as a printer or a file storage device, modem, Laptop Embedded Equipped (LEE), Laptop Mounted Equipment (LME), USB dongles, CPE or any other radio network unit capable of communicating over a radio link in the computer system 100. The device 130 may be wireless, i.e., it may be enabled to communicate wirelessly in the computer system 100 and, in some particular examples, may be able support transmission using beamforming. The communication may be performed e.g., between two devices, between a device and a radio network node, and / or between a device and a server. The communication may be performed e.g., via a RAN and possibly one or more core networks, comprised, respectively, within the computer system 100.
[0063] The computer system 100 may comprise one or more radio network nodes, whereof a radio network node 140 is depicted in Figure 2b. The radio network node 140 may typically be a base station or Transmission Point (TP), or any other network unit capable to serve a wireless device or a machine type node in the computer system 100. The radio network node 140 may be e.g., a 5G gNB, a 4G eNB, or a radio network node in an alternative 5G radioaccess technology, e.g., fixed or WiFi. The radio network node 140 may be e.g., a Wide Area Base Station, Medium Range Base Station, Local Area Base Station, and Home Base Station, based on transmission power and thereby also coverage size. The radio network node 140 may be a stationary relay node or a mobile relay node. The radio network node 140 may support one or several communication technologies, and its name may depend on the technology and terminology used. The radio network node 140 may be directly connected to one or more networks and / or one or more core networks.
[0064] The computer system 100 covers a geographical area which may be divided into cell areas, wherein each cell area may be served by a radio network node, although, one radio network node may serve one or several cells.
[0065] Any of the first node 111 and the second node 112 may be a radio network node, such as the radio network node 140 just described, a core network node, such as depicted for the first node 111 and the second node 112 in the non-limiting example of Figure 2b, or a device, such as the device 130 described earlier, as depicted for the third node 113 in the non-limiting example of Figure 2b.
[0066] In some examples, the telecommunication network comprised in the computer system 100 may comprise an access network, such as a radio access network (RAN), and a core network, which may include one or more core network nodes. The access network may include one or more access network nodes, such as the radio network node 140, e.g., which may be generally referred to as network nodes, or any other similar 3rd Generation Partnership Project (3GPP) access nodes or non-3GPP access points. Moreover, as will be appreciated by those of skill in the art, a network node is not necessarily limited to an implementation in which a radio portion and a baseband portion are supplied and integrated by a single vendor. Thus, it will be understood that network nodes may include disaggregated implementations or portions thereof. For example, in some embodiments, the telecommunication network may include one or more Open-RAN (ORAN) network nodes. An ORAN network node may be understood as a node in the telecommunication network that may support an ORAN specification, e.g., a specification published by the O-RAN Alliance, or any similar organization, and may operate alone or together with other nodes to implement one or more functionalities of any node in the telecommunication network, including one or more network nodes and / or core network nodes.
[0067] Examples of an ORAN network node include an open radio unit (O-RU), an open distributed unit (O-DU), an open central unit (O-CU), including an O-CU control plane (O-CU-CP) or an O-CU user plane (O-CU-UP), a RAN intelligent controller, near-real time or non-real time, hosting software or software plug-ins, such as a near-real time control application, e.g., xApp, or a non-real time control application, e.g., rApp, or any combination thereof, the adjective “open” designating support of an ORAN specification. The radio network node 140may support a specification by, for example, supporting an interface defined by the ORAN specification, such as an A1 , F1 , W1 , E1 , E2, X2, Xn interface, an open fronthaul user plane interface, or an open fronthaul management plane interface. Moreover, an ORAN access node may be a logical node in a physical node. Furthermore, an ORAN network node may be implemented in a virtualization environment, in which one or more network functions may be virtualized. For example, the virtualization environment may include an O-Cloud computing platform orchestrated by a Service Management and Orchestration Framework via an 0-2 interface defined by the O-RAN Alliance or comparable technologies. The radio network node 140 may facilitate direct or indirect connection of user equipment (UE), such as by connecting the device 130 to the core network over one or more wireless connections.
[0068] The first node 111 may communicate with the second node 112 over a first link 151 , e.g., a radio link or a wired link. The first node 111 may communicate with the first data center 121 over a second link 152, e.g., a wired link. The first node 111 may communicate with the second data center 122 over a third link 153, e.g., a wired link. The second node 112 may communicate with the third data center 123 over a fourth link 154, e.g., a wired link. The second node 112 may communicate with the radio network node 140 over a fifth link 155, e.g., a wired link. The first node 111 may communicate with the radio network node 140 over a sixth link 156, e.g., a wired link. The first data center 121 may communicate with the radio network node 140 over a seventh link 157, e.g., a wired link. Any of the second data center 122 and the third data center 123 may communicate with the radio network node 140 over a respective link, not depicted in Figure 2. The device 130 may communicate with the radio network node 140 over a seventh link 157, e.g., a radio link.
[0069] Any of the links referred to in Figure 2 may be direct links or may each comprise a plurality of links.
[0070] Although terminology from Long Term Evolution (LTE) / 5G has been used in this disclosure to exemplify the embodiments herein, this should not be seen as limiting the scope of the embodiments herein to only the aforementioned system. Other wireless systems supporting similar or equivalent functionality may also benefit from exploiting the ideas covered within this disclosure. In future telecommunication networks, e.g., in the sixth generation (6G), the terms used herein may need to be reinterpreted in view of possible terminology changes in future technologies.
[0071] In general, the usage of “first”, “second”, “third”, etc. herein may be understood to be an arbitrary way to denote different elements or entities and may be understood to not confer a cumulative or chronological character to the nouns they modify.
[0072] Some of the embodiments contemplated herein will now be described more fully with reference to the accompanying drawings. Other embodiments, however, are contained within the scope of the subject matter disclosed herein, the disclosed subject matter should not beconstrued as limited to only the embodiments set forth herein; rather, these embodiments are provided by way of example to convey the scope of the subject matter to those skilled in the art.
[0073] Figure 3 is a schematic diagram depicting a non-limiting example of two data centers, such as the first data center 121 according to embodiments herein, as a so-called thin Network deployment architecture, in RAN-Cloud. Only one of the data centers depicted, the first data center 121 is described in relation to Figure 3. It may be understood that the same description may apply to the other depicted data center. As stated earlier, the first data center 121 may lack redundant structures. Particularly, the first data center 121 lacks redundant network structures. In some embodiments, the first data center 121 may comprise at least one of a single spine structure 301 and a single leaf structure 302. In some embodiments, the first data center 121 may comprise at least one of the first plurality of clients and servers 303, and the second plurality of racks 304. The first data center 121 may be understood to correspond to a thin-network data fabric infrastructure that may be implemented with the following design: single link connectivity from computes 305 to leaf switches 302 instead of dual-homing, that is, e.g., instead of each C / S being connected to two leaf nodes, and fabric and inter-switch link dimensioning may be performed for larger over-subscription ratios. Fabric and inter-switch link dimensioning may be understood to refer to the bandwidth capacity of the links between switches which are represented by black lines in Figure 3. Such a thin-network may be understood to exhibit the following characteristics. One characteristic may be availability of more ports, which may enable double the number of computes to be connected to the leaf switches 302 as compared to a standard network fabric of the same size. In a thin network, each C / S may be understood to be connected to just one leaf node, hence using one port on one leaf node, unlike for example in the data center of Figure 1 , where each C / S is connected to 2 leaf nodes, using one port each on each leaf node. Since the thin network C / S may be understood to use one port, more ports may be understood to be available. Another characteristic may be that fewer Network Interface Cards (NICs) may be needed in the computes. Each NIC may be understood to be needed to connect one port on the server to one port on the switch. If the compute connects to fewer ports on the switches, it may be understood to need fewer NICs. Yet another characteristic may be a better capacity usage on all switches including leaf 302, spine 301 , DC-GW 306 during normal operating conditions. Another characteristic may be that fewer spine switches 301 may need to be deployed for larger number of computes. Yet another characteristic may be that redundant DC-GWs 306 may carry more traffic, giving better capacity utilization. A thin network may be understood to be more cost effective if more capacity may be extracted from the hardware. The number of computes connected to the same number of DC-GWs may be understood to double in anetwork design according to embodiments herein. Another characteristic may be that each link between leaf 302 and spine 301 , and between spine 301 and DC-GW 306 may carry more traffic, as the leaf switches 302 may be understood to host double the number of computes. Also depicted in Figure 3 are the PE-GW 307 in the transport section 308, and in the control fabric 309, the MgmtGW 310, the Ctrl-SW 311 and the Rack 312. The spine 301 , leaf 302 and DC-GW 306 may be understood to be comprised in the data fabric 313.
[0074] Embodiments of a computer-implemented method, performed by the first node 111, will now be described with reference to the flowchart depicted in Figure 4. The method may be understood to be for handling resources. The first node 111 operates in the computer system 100.
[0075] In some examples, the computer system 100 may be a 3GPP network. In some particular examples, the computer system 100 may be a 5G network.
[0076] Several embodiments are comprised herein. In some embodiments, all the actions may be performed. In some embodiments, some actions may be performed. It should be noted that the examples herein are not mutually exclusive. One or more embodiments may be combined, where applicable. All possible combinations are not described to simplify the description. Components from one embodiment may be tacitly assumed to be present in another embodiment and it will be obvious to a person skilled in the art how those components may be used in the other exemplary embodiments. A non-limiting example of the method performed by the first node 111 is depicted in Figure 4. In Figure 4, optional actions are represented with dashed lines.
[0077] Action 401
[0078] Embodiments herein may be understood to aim at decreasing the risk of connectivity loss after deployment of a thin-network based network infrastructure such as that of the first data center 121 in, for example, cloud RAN sites.
[0079] As stated earlier, the first data center 121 may lack redundant structures. Particularly, the first data center 121 may lack redundant network structures.
[0080] In some embodiments, the first data center 121 may comprise at least one of a single spine structure and a single leaf structure.
[0081] In some embodiments, the first data center 121 may comprise at least one of the first plurality of clients and servers, and the second plurality of racks.
[0082] According to embodiments herein, the risk of connectivity loss may be managed by dimensioning and maintaining a thin-network infrastructure within a geographical constraint area. Different geographical areas may be evaluated with respect to workloads, computes or set of computes, switches etc.In order to enable to handle the risk of connectivity loss after deployment of a thin-network based network infrastructure, in this Action 401 , the first node 111 obtains one or more first indications indicating a first performance of one or more bridge ports of the first data center 121 of the computer system 100. In some examples, the one or more first indications may be obtained by e.g., a bridge port data collector. A bridge port may be understood to, in the context of network switches, relate to how a switch may handle data traffic between its interfaces. In Figure 3, for example, various level of hierarchy of leaf / spine switches are depicted. In the architecture, each port that may connect leaf switches to spine switches or to end devices may operate as a bridge port, forwarding traffic based on learned Medium Access Control (MAC) addresses and participating in Layer 2 protocols as required. A bridge port may be understood to perform forwarding, segmentation and Spanning Tree Protocol (STP) functions. For forwarding, bridge ports may forward frames based on MAC address tables. They may learn the MAC addresses of connected devices and use this information to forward traffic only to the appropriate port. In segmentation, by connecting different network segments, bridge ports may help in segmenting collision domains. In many networks, bridge ports may participate in STP to prevent loops. STP may be understood to dynamically disable ports as necessary to ensure there may be a single active path between any two network nodes.
[0083] In this Action 401 , the first node 111 also obtains one or more second indications indicating a second performance of one or more Internet Protocol (IP) Interfaces of the first data center 121 of the computer system 100. In some examples, the one or more second indications may be obtained by e.g., an IP interface data collector. An IP interface may be understood as a logical connection point through which a device may communicate within an IP network. In a leaf-spine network architecture such as that depicted in Figure 3, IP interfaces may enable communication between devices and routing of data packets, e.g., Layer 3 IP-based routing. Each switch, both leaf and spine, may be configured with IP interfaces on its links, as shown in Figure 3 by leaf and spine being connected to each other. This may be understood to allow the network to route packets using IP routing protocols, e.g., Open Shortest Path First (OSPF), Border Gateway Protocol (BGP) or Intermediate System to Intermediate System (IS-IS). Virtual Local Area Networks (VLANs) may be configured across the leaf switches, with IP routing occurring between them. In some designs, certain leaf switches may act as border leaf switches, connecting the data center network to external networks. These switches may have additional IP interfaces configured for external connectivity.
[0084] In this Action 401 , the first node 111 further obtains one or more third indications indicating a third performance of one or more compute systems of the first data center 121 ofthe computer system 100. In some examples, the one or more third indications may be obtained by e.g., a computer system data collector.
[0085] The obtaining in this Action 401 may be performed by streaming of the data.
[0086] Any of the one or more first indications, the one or more second indications and the one or more third indications may comprise, but may not be limited to, any of the following variables: RAN cloud variables, metrics, as explained below, relevant logs in the compute and network switches, and / or link and switch related faults, e.g., alarms. The metrics may comprise workload interface network performance management counters, e.g., bandwidth consumed, bandwidth capacity / bandwidth negotiated, experienced data rate (downlink), experienced data rate (uplink), area traffic capacity (downlink), area traffic capacity (uplink), overall user density, network transmit bytes, network received bytes. The metrics may additionally or alternatively comprise: i) compute interface network performance management counters, e.g., CPU usage, memory size, memory I / O, disk size, disk I / O, ii) network transceiver PHY counters, e.g., transmitter signal strength, jitter, bit error ratio, iii) network switches - CPU, memory, network performance metrics, network resource usage metrics, and / or iv) application latency, e.g., Round-Trip Time (RTT).
[0087] It should be noted that the above list may be understood to be a sample of different parameters that may be used. This list may be understood to not be an exhaustive list of variables as part of this disclosure, but rather demonstrate how such a scheme may be used to de-risk the deployment of a thin-network based network infrastructure in Cloud RAN sites.
[0088] Table 1 shows examples of metrics for different categories related to port-compute performance.
[0089]
[0090] Table 1. Category wise metric sample performance metricsBy obtaining the one or more first indications, the one or more second indications and the one or more third indications in this Action 401 , the first node 111 may then be enabled to use the obtained indications to monitor the performance and application availability of the lean infrastructure of the first data center 121 with predictive mechanisms and ultimately assess the need to move workload beyond a rack or even beyond the first data center 121.
[0091] Action 402
[0092] In some embodiments, in this Action 402, the first node 111 may determine a respective prediction of the obtained one or more first indications, one or more second indications, and one or more third indications at the future time period.
[0093] Determining may be understood as calculating, deriving, estimating or similar.
[0094] As the data flows through various sources of the first data center 121 , such as the bridge port data, IP Interface data and the computer system data, the data may undergo multiple forecasting models. These one or more forecasting models may depict how the data may look like in next timestamps, e.g., a few hours from a certain point in time. In theory, a forecasting model may be a statistical tool designed to predict future trends and outcomes based on historical data. It may involve analyzing past patterns and trends to make informed predictions about future events. Time series models such as ARIMA and variants or RNNs may be used to forecast data points, e.g., bridge port data, system data, etc
[0095] For the determining in this Action 402 an ML module may be used to forecast each group of metrics e.g., Bridge Port, IP Interface and compute related metrics.
[0096] The time horizon of forecasting may be dependent on use cases and type of network deployment area e.g., urban / rural etc. Urban deployment may require very high availability hence, frequent assessment may be required related to leaf-compute combination performance, while rural areas may provide flexibility for less periodic evaluations.
[0097] By determining the respective prediction of the obtained one or more first indications, one or more second indications, and one or more third indications at the future time period in this Action 402, the first node 111 may then be enabled to use the determined respective predictions to monitor the anticipated performance and application availability of the lean infrastructure of the first data center 121 with predictive mechanisms and ultimately assess the need to move workload beyond a rack or even beyond the first data center 121 , ahead of a risk of an event having a negative effect of the overall performance of the first data center 121.
[0098] Action 403
[0099] In some embodiments, in this Action 403, the first node 111 may obtain one or more indications, referred to herein as one or more “sixth indications”, indicating one or more trained machine learning models from the second node 112 operating in the computer system 100.The one or more machine learning models may have been trained to determine a fourth indication of an availability of resources at the first data center 121 at a future time period. In some embodiments, the fourth indication may indicate a probability at the future time period of one of: a capacity of network components of the first data center 121 having a value, a failure happening at the first data center 121 , and a maintenance operation happening at the first data center 121.
[0100] In some embodiments, one or more of the following options may apply. According to one option, at least one of the one or more trained machine learning models may be a decision tree model. In some examples, all of the one or more trained machine learning models may be decision tree models.
[0101] According to another option, the resources may comprise one or more of: compute resources, storage resources and network resources,
[0102] According to another option, the network resources may comprise one or more of: leaf resources, spine resource, switch resources and link resources.
[0103] According to yet another option, network structures of the first data center 121 may comprise one or more of: a spine structure, a leaf structure, switches and links.
[0104] In some examples, once the model training may have been performed by the second node 112, as described in relation to Figure 5, another ML module may be used to forecast each group of metrics e.g., Bridge Port, IP Interface and compute related metrics in Action 402.
[0105] The method described in relation to Figure 4 may be understood to correspond to an inference phase of the one or more machine learning models having been trained by the second node 112, as will be later described in relation to Figure 5. It may be noted that in some examples, the first node 111 and the second node 112 may be the same node. That is, the training phase and inference phase may be performed by the same node.
[0106] By obtaining the one or more sixth indications in this Action 403, the first node 111 may then be enabled to use the one or more trained machine learning models to derive the availability of resources at the first data center 121 at the future time period, and thereby be enabled to initiate remedial action, if needed.
[0107] Action 404
[0108] In some embodiments, in this Action 404, the first node 111 determines, based on the obtained one or more first indications, one or more second indications, and one or more third indications, and using the one or more trained machine learning models, the fourth indication of the availability of resources at the first data center 121 at the future time period.
[0109] As stated earlier, to ensure the high availability, the fourth indication may be determined as a health score that may be derived for each area based upon capacity demand current / forecasted, fault probability and / or operation and maintenance activities along with traffic mix. The current data may be understood to correspond to the obtained one or more first indications, one or more second indications, and one or more third indications, whereas the forecasted data may be understood to correspond to the respective prediction of the obtained one or more first indications, one or more second indications, and one or more third indications at the future time period determined in Action 402.
[0110] Accordingly, in some embodiments, the fourth indication may indicate the probability at the future time period of one of: the capacity of network components of the first data center 121 having the value, the failure happening at the first data center 121, and the maintenance operation happening at the first data center 121. The fourth indication may be understood as a health score of the resources of the first data center 121. In this Action 404, the first node 111 may generate the health score.
[0111] In this Action 404, the first node 111 may use the one or more ML models to ensure that day to day operations in the first data center 121 may be with at par level of performance in terms of availability.
[0112] The first node 111 may use the determined respective predictions from Action 402 as input to the obtained one or more machine learning models. The one or more trained ML models may be provided metrics related to compute performance, interface and bridge port which may be then used to identify the port and link health of the first data center 121. Based on collected metrics, the one or more ML models may have been trained for identification of anomalies and subject matter expert (SME) suggested rule based alarm or faults. These anomalies and faults may have been used to train the one or more ML models for e.g., a classification of a faulty port from healthy one. Since there may be understood to be multiple metrics being monitored / predicted, there may be multiple models trained, respectively, using one or more of these metrics. In some embodiments, the training by the second node 112 of the one or more machine learning models may have comprised one or more of: detecting anomalies in obtained first one or more first indications, second one or more second indications, and third one or more third indications, classifying the detected anomalies as one of known and unknown, and creating one or more rules to determine the availability of the resources at the future time period based on the classified anomalies.
[0113] The determining in this Action 404 of the fourth indication may be based on applying the one or more rules of the obtained one or more machine learning models on the determined respective predictions.
[0114] To determine the fourth indication, these forecasted values may be applied to anomaly rules. Anomaly rules may have been created during a training phase of the one or more machine learning models, as will be described later, in relation to Figure 5. In case any anomaly is detected in the forecasted pattern, the probability of occurrence of that anomalymay be calculated individually, that is, the probability of occurrence of an anomaly for each metric may be calculated independently of other metrics. Further particularly, once forecasting for each category may have been performed, based on the one or more trained ML models, failure probability for each category may be determined. That is, for each of the one or more bridge ports of the first data center 121, the one or more Internet Protocol Interfaces of the first data center 121 and the one or more compute systems of the first data center 121.
[0115] Failure probability may be understood as a composite score considering a bucket of performance metrics. Failure probability may be a composite of one or more anomaly models along with SME defined thresholds. Further, failure probabilities may be generalized in binary score, that is, 0 or 1. For example, if the failure probability is low, e.g., 20%, then it may be assigned a score of 1 , else 0.
[0116] Once all the failure probabilities may have been accumulated, the fourth indication may be derived as a health score. As a non-limiting example, the health score may be determined as follows:
[0117] Health Score= (1 -Failure Probabilities Category 1 (Bridge Port))*(1- Failure Probabilities Category 2 (IP lnterface))*(1- Failure Probabilities Category 3 (Compute System))
[0118] where each failure probabilities category “C” may take a value between (0,1).
[0119] The HealthScore may be a value between 0 and 1 , 1 being healthy, 0 meaning unhealthy.
[0120] For example, if any of the categories has failure probabilities as 1 , that is, high failure probabilities, then the health score may be derived as 0, else 1.
[0121] In case all three turn out to be 1 , that may be understood to mean that the first node 111 may infer there may be absolutely no probability of occurrence of a fault at the first data center 121.
[0122] The failure probabilities may be used to derive the fourth indication as the binary health score, which may be used to make the decision about whether or not the workloads may need to be migrated.
[0123] If the fourth indication, that is, the Health score is 1 , then that leaf to compute redundancy may not be required and the network may continue to operate in lean mode. Else, the redundancy / migration of workload from a certain compute may be required to a different compute, leaf or different site.
[0124] Forecasting metrics and subsequently, the failure probability may be understood to help to proactively prevent network outages. In the event of health score being 0, workloads may be migrated to other healthy computes / sites.This process may continue iteratively for a forecasting time horizon. The effective outcome from iterations may be listed as minimal disruption in service by supporting the same or similar availability characteristics while using much lesser equipment.
[0125] By determining the fourth indication in this Action 404, the first node 111 may rely on AI / ML techniques to monitor the infrastructure and application availability with predictive mechanisms. If the first data center 121 is predicted to have an availability issue at the future time period, the first node 111 may be enabled to mitigate this situation by for example, migrating workloads to a different compute in the same rack, different rack or a different site based on the predictive feedback provided by the AI / ML components in embodiments herein. The availability of the workload at an application level may thus be improved. The first node 111 may therefore be enabled to ensure the high availability of the first data center 121, while maintaining its lean structure. Accordingly, the first node 111 may enable to fulfill the goal of embodiments herein to achieve the same level of service availability for an application workload with fewer infrastructure resources, while at the same time having the ability to accommodate failure scenarios and maintenance activities. This scheme may be understood to rely on the highly distributed nature of cloud RAN deployments, due to which the potential availability zone for a workload may be moved beyond a rack or even beyond a site.
[0126] Action 405
[0127] In this Action 405, the first node 111 may output a fifth indication of the determined fourth indication.
[0128] The fifth indication may be an explicit indication of the fourth indication, e.g., as the health score.
[0129] In some examples, the output from the health score for different areas may be a proactive determination of a redundancy requirement for a given leaf port to compute combination of hierarchy of leaf port to compute, that is, to which leaf port compute it may be associated.
[0130] In some examples, the fifth indication may be a recommendation to move, or not, workload from the first data center 121.
[0131] By outputting the fifth indication in this Action 405, the first node 111 may then be enabled to initiate determining one or more actions based on the determined fourth indication to for example, potentially mitigate availability issues that may have been predicted for the future time period.
[0132] Action 406
[0133] In this Action 406, the first node 111 initiates determining one or more actions based on the determined fourth indication. The one or more actions comprise migrating at least a part ofa workload from a first structure comprised in the first data center 121 to another structure comprised in the first data center 121 or in a second data center 122.
[0134] Initiating may be understood as starting itself, or instructing, enabling or facilitating that another node may determine the one or more actions. The one or more actions may then be performed by the first node 111 itself or by yet another node operating in the computer system 100.
[0135] In some examples, initiating determining may comprise sending the fifth indication, e.g., to a third node operating in the computer system 100, e.g., an SMO.
[0136] In some examples, the output from the health score for different areas may be a proactive determination of a redundancy requirement for a given leaf port to compute combination of hierarchy of leaf port to compute, that is, to which leaf port compute it may be associated, and initiating determining may comprise passing the information to the third node operating in the computer system 100, e.g., an SMO, to take necessary actions.
[0137] In some examples, initiating determining may sending the recommendation to move, or not, workload from the first data center 121 to a scheduler which may decide about the movement of the workload to a different leaf or different site.
[0138] Embodiments herein may be understood to target intra-site, as well as inter-site, high availability, depending on which part of thin network may have been notified as having a low health zone and the first node 111 may flag as problematic. If the first node 111 flags a zone as problematic, a comparable zone may need to be picked by the third node, e.g., the SMO. For example, if the zone is limited to a single compute, health scores for other computes in the same site may be used to migrate the function. If the first node 111 flags a site as problematic, then the first node 111 may recommend an action to the SMO to pick a different site for which the score may be better.
[0139] This process may continue iteratively for a forecasting time horizon. The time horizon of forecasting may be dependent on use cases and type of network deployment area e.g., urban / rural etc. Urban deployment may require very high availability hence, frequent assessment may require related to leaf-compute combination performance, while rural areas may provide flexibility for less periodic evaluations.
[0140] In some embodiments, one or more of the following options may apply. According to one first option, at least one of the one or more trained machine learning models may be a decision tree model. In some examples, all of the one or more trained machine learning models may be decision tree models.
[0141] According to another option, the resources may comprise one or more of: the compute resources, the storage resources and the network resources,
[0142] According to another option, the network resources may comprise one or more of: the leaf resources, the spine resource, the switch resources and the link resources.According to another option, at least one of the first structure and the another structure may be a server.
[0143] According to yet another option, network structures of the first data center 121 may comprise one or more of: the spine structure, the leaf structure, the switches and the links.
[0144] By initiating determining the one or more actions comprising migrating at least a part of the workload from the first structure to another structure in this Action 406, the first node 111 may, for example, enable to maintain a high availability (HA) threshold, e.g., an SLA threshold, to meet a requirement of the network, without having a larger hardware (Hw) footprint.
[0145] Embodiments of a computer-implemented method performed by the second node 112, will now be described with reference to the flowchart depicted in Figure 5. The method may be understood to be for handling the resources. The second node 112 operates in the computer system 100.
[0146] Several embodiments are comprised herein. In some embodiments, all the actions may be performed. In some embodiments, some actions may be performed. It should be noted that the examples herein are not mutually exclusive. One or more embodiments may be combined, where applicable. All possible combinations are not described to simplify the description. Components from one embodiment may be tacitly assumed to be present in another embodiment and it will be obvious to a person skilled in the art how those components may be used in the other exemplary embodiments. The detailed description of some of the following corresponds to the same references provided above, in relation to the actions described for the first node 111 and will thus not be repeated here to simplify the description. For example, failure probabilities may be used to derive the fourth indication as the binary health score.
[0147] Action 501
[0148] The second node 112 may be understood to be a node performing the training of the one or more ML models that may then be used by the first node 111 as described in relation to Figure 4. Figure 5 may be understood to describe the method that the second node 112 may perform to train the one or more ML models.
[0149] As explained earlier, typical data center deployments may include redundancy at multiple levels. It may be expected that the historical data from these deployments may include data for fault occurrences, in addition to data for networks in stable state.
[0150] Initially, historical data from redundant deployments may be used to train the one or more ML models. Subsequently, data from both redundant and non-redundant deployments may be utilized for re-training the one or more ML models.This may be understood to be since embodiments herein may be understood to aim to assist in deployments without redundancy, in lean mode. The objective may be understood to be to make sure a high scale availability, e.g., of 99.99%, may be provided to meet a requirement of the network. In addition to data from redundant deployments, which the second node 112 may be being already trained on, and continuing to get data from such deployments, the second node 112 may start collecting the data for model training from non-redundant deployment(s) as well.
[0151] Accordingly, at least a part of the data that the second node 112 may use to train the one or more ML models, that may later be used by the first node 111 to perform the method described in Figure 4, may originate from the third data center 123, which may have a thin-network based network infrastructure. In agreement with this, the third data center 123 may lack redundant structures. Particularly, the third data center 123 may lack redundant network structures. In some embodiments, the third data center 123 may comprise at least one of a single spine structure and a single leaf structure. In some embodiments, the third data center 123 may comprise at least one of a third plurality of clients and servers, and a fourth plurality of racks.
[0152] In some examples, the third data center 123 may be the same as the first data center 121.
[0153] In particular, in this Action 501 , the second node 112 obtains: i) first one or more first indications indicating a fourth performance of one or more bridge ports of the third data center 123 of the computer system 100, ii) second one or more second indications indicating a fifth performance of one or more Internet Protocol Interfaces of the third data center 123 of the computer system 100, and iii) third one or more third indications indicating a sixth performance of one or more compute systems of the third data center 123 of the computer system 100. These may be understood to correspond to a the same type of information collected in Action 401 , but during the training phase of the one or more ML models.
[0154] The obtaining in this Action 501 may be performed during streaming of the data.
[0155] As stated earlier, any of the first one or more first indications, the first one or more second indications and the first one or more third indications may comprise, but may not be limited to, any of the following variables which may be used as input: RAN cloud variables, metrics, as explained below, relevant logs in the compute and network switches, and / or link and switch related faults, e.g., alarms. The metrics may comprise workload interface network performance management counters, e.g., bandwidth consumed, bandwidth capacity / bandwidth negotiated, experienced data rate (downlink), experienced data rate (uplink), area traffic capacity (downlink), area traffic capacity (uplink), overall user density, network transmit bytes, network received bytes. The metrics may additionally or alternatively comprise: i) compute interface network performance management counters, e.g., CPU usage, memorysize, memory I / O, disk size, disk I / O, ii) network transceiver PHY counters, e.g., transmitter signal strength, jitter, bit error ratio, iii) network switches - CPU, memory, network performance metrics, network resource usage metrics, and / or iv) application latency, e.g., Round-Trip Time (RTT).
[0156] As also noted earlier, it may be noted that the above list may be understood to be a sample of different parameters that may be used as input for training the one or more ML models. This list may be understood to not be an exhaustive list of variables for the one or more models as part of this disclosure, but rather demonstrate how such a scheme may be used to decrease the risk of connectivity loss of the deployment of a thin-network based network infrastructure in Cloud RAN sites.
[0157] By, in this Action 501 , obtaining the first one or more first indications, the second one or more second indications and the third one or more third indications, the second node 112 may then be enabled to use the obtained indications to provide input to the one or more ML models and enable that, one trained, the one or more ML models be used, e.g., by the first node 111, to ensure that day to day operations may be with at par level of performance in terms of availability. The one or more ML models may be provided metrics related to compute performance, interface and bridge port which may enable to identify the port and link health of a data center, such as the third data center 123, e.g., during the training phase of the one or more ML models and the first data center 121 during the interference phase of the one or more ML models. Based on collected metrics, the one or more ML models may be trained for identification of anomalies and subject matter expert (SME) suggested rule based alarm or faults. These anomalies and faults may be used to train the one or more ML models for classification of e.g., a port as a faulty port or a healthy one.
[0158] Action 502
[0159] In this Action 502, the second node 112 may perform a quality check on the obtained first one or more first indications, second one or more second indications, and third one or more third.
[0160] That is, in this Action 502, the data that may later be used to train the one or more ML models for early detection of anomalies, that is, the bridge port data, IP Interface data and computer system data may go through a data quality check, to find out any missing data or to determine whether the data may be meeting an expected standard. In case the data may not be meeting the required standards, the data may be modified or even discarded. This may be understood to be to ensure that the one or more ML models may be built on robust data.
[0161] The variables may then, in this Action 504, be tuned, and one or more models may then be applied on sites with a thin-network infrastructure to provide a health score. Tuning may be understood to mean changing the data set and processing it to refine the one or ML models.Action 503
[0162] In this Action 503, the second node 112 may process the obtained first one or more first indications, second one or more second indications, and third one or more third indications based on the performed quality check, to obtain processed data by one or more of: modifying and discarding data of inadequate quality.
[0163] As explained above, this may be understood to be to ensure that the one or more ML models may be built on robust data.
[0164] Action 504
[0165] Once the data collected may be solid, the data may be aggregated for feature engineering. In this Action 504, the second node 112 may aggregate the processed data from Action 503 based on the obtained first one or more first indications, second one or more second indications, and third one or more third indications and may perform feature engineering on the aggregated data.
[0166] In one scenario, the second node 112 may train the one or more ML models on various types of traditional data centers, including large, small, micro to determine the best set of variables, involved in initial training, as well as predictors, used for forecasting.
[0167] By aggregating the processed data and performing feature engineering in this Action 504, the first node 111 may then process the data in a way that may lead to better model performance, e.g., rules identification, for later stages.
[0168] Action 505
[0169] In this Action 505, the second node 112 trains the one or more machine learning models to determine the fourth indication of an availability of resources at the third data center 123 at the future time period, based on the obtained first one or more first indications, second one or more second indications, and third one or more third indications.
[0170] As stated earlier, the fourth indication may indicate a probability at the first future time period of one of: a capacity of the third data center 123 having a value, a failure happening at the third data center 123, and a maintenance operation happening at the third data center 123.
[0171] As mentioned above, historic data of various variables, and / or Key Performance Indicators (KPIs) involved in Bridge port, interfaces, and compute systems may be used for the training.
[0172] The training in this Action 505 of the one or more ML models may be performed using the aggregated processed data obtained in Action 504 as input.
[0173] The training in this Action 505 may comprise one or more of: detecting anomalies in the obtained first one or more first indications, second one or more second indications, and thirdone or more third indications, classifying the detected anomalies as one of known and unknown, and creating the one or more rules to determine the availability of the resources at the future time period based on the classified anomalies.
[0174] As there may be various known scenarios of anomaly and / or bad health in existing data, the one or more ML models may be a tree based classification model that may be used to identify rules for anomalies.
[0175] In some embodiments, one or more of the following options may apply. According to one first option, at least one of the one or more machine learning models may be a decision tree model. This data tree model may have exhaustive rule generation with root cause and probability.
[0176] According to another option, the resources may comprise one or more of: the compute resources, the storage resources and the network resources,
[0177] According to another option, the network resources may comprise one or more of: the leaf resources, the spine resource, the switch resources and the link resources.
[0178] According to yet another option, network structures of the third data center 123 may comprise one or more of: the spine structure, the leaf structure, the switches and the links.
[0179] Data may be selected on the premise of classifying data instances into 'Known abnormality' and 'Unknown abnormality' . Known abnormalities may be port failures, protocol failures and compute failures. This may typically involve supervised machine learning, where the one or more ML models may be trained on 'Known abnormality' data to discern between normal and anomalous behavior. This technique may offer a fine degree of control and precision. The unknown abnormality may be detected using methods such as the "Local Outlier Factor" (LOF) algorithm, which may be understood to be an unsupervised anomaly detection method which may compute the local density deviation of a given data point with respect to its neighbors. Data points with high LOF score may be termed as abnormal data points, e.g., via annotation, and this annotated data may be used in the training of one or more tree based models to get rules for abnormal condition.
[0180] Inference data may also be added to these data annotations to train the one or more ML models.
[0181] By in this Action 505, training the one or more ML models, the second node 112 may enable any node, e.g., the first node 111 , to rely on AI / ML techniques to monitor the infrastructure and application availability of lean structures such as the first data center 121 and the third data center 123, with predictive mechanisms to detect availability problems of such lean structures before they may occur. Predicting the availability problems may then enable to mitigate availability issues by for example, migrating workloads to a different compute in the same rack, different rack or a different site based on the predictive feedback provided by the AI / ML components. The availability of the workload at an application levelmay thus be improved. The second node 112 may therefore enable to ensure the high availability of a structure such as the first data center 121 , while maintaining its lean structure. Accordingly, the second node 112 may enable to fulfill the goal of embodiments herein to achieve the same level of service availability for an application workload with fewer infrastructure resources, while at the same time having the ability to accommodate failure scenarios and maintenance activities. This scheme may be understood to rely on the highly distributed nature of cloud RAN deployments, due to which the potential availability zone for a workload may be moved beyond a rack or even beyond a site.
[0182] Action 506
[0183] In this Action 506, the second node 112 initiates outputting the one or more sixth indications of the one or more trained machine learning models to the first node 111 operating in the computer system 100.
[0184] The outputting in this Action 506 may be, e.g., sending or providing, and may be performed, e.g., via the first link 151.
[0185] By outputting the one or more sixth indications, the second node 112 may enable the first node 111 to perform the method described in relation to Figure 4, and obtain the benefits recited therein.
[0186] Figure 6 depicts a non-limiting example of a method performed, according to embodiments herein, during the training phase of the one or more ML models by the second node 112 and during the inference phase of the one or more ML models by the first node 111. In this example, the first node 111 is the same node as the second node 112. At 601 , the second node 112 may obtain data from redundant deployments, such as historical and current data. Initially, the historical data from redundant deployments may be used to train the one or more ML models. Subsequently, data from both redundant and non-redundant deployments, obtained according to Action 501 , may be utilized for re-training the models. The training of the one or more ML models may be performed according to Action 505. As there may be various known scenarios of anomaly and / or bad health in existing data, any or all of the one or more ML models may be a tree based classification model that may be used to identify rules for anomalies. Based on the collected metrics, the one or more ML models may be trained in Action 505 for identification of anomalies and SME suggested rule based alarm or faults. These anomalies and faults may be used to train the one or more ML models in Action 505 for classification of, e.g., a port, as a faulty port or a healthy one. These data tree models may have exhaustive rule generation with root cause and probability. As explained earlier, data may be selected on the premise of classifying data instances into 'Known abnormality' and 'Unknown abnormality' . Known abnormalities may for example be port failures, protocolfailures and compute failures. This may typically involve supervised machine learning, where the ML may be trained on 'Known abnormality' data to discern between normal and anomalous behavior. This technique may offer a fine degree of control and precision. The unknown abnormality may be detected using methods such as the LOF algorithm, which may be understood to be an unsupervised anomaly detection method which may compute the local density deviation of a given data point with respect to its neighbors. Data points with high LOF score may be termed as abnormal data points, e.g., via annotation, and this annotated data may be used in the training of a tree based model to obtain rules for abnormal conditions. Inference data may also be added to these data annotations to train the one or more ML models. Once the model training may have been performed by the second node 112, as described in relation to Figure 5, another ML module may be used to forecast each group of metrics e.g., Bridge Port, IP Interface and compute related metrics by the first node 111 , in accordance with Action 402. To determine the fourth indication, these forecasted values may be applied to the anomaly rules that may have been created during the training phase of the one or more machine learning models in Action 505. In case any anomaly is detected in the forecasted pattern, the probability of occurrence of that anomaly may be calculated individually. Further particularly, once forecasting for each category may have been performed, based on the one or more trained ML models, the failure probability for each category may be determined. That is, for each of the one or more bridge ports of the first data center 121 , the one or more Internet Protocol Interfaces of the first data center 121 and the one or more compute systems of the first data center 121. The failure probability may be understood as the composite score considering the bucket of performance metrics. The failure probability may be the composite of the defined thresholds of the one or more anomaly models along with the SME. The failure probabilities determined according to Action 404 may be used to derive the fourth indication, e.g., as the binary health score, which may be used to make the decision about whether or not the workloads may need to be migrated. If the fourth indication, that is, the Health score is 1 , then that leaf to compute redundancy may not be required and the network may continue to operate in lean mode. Else, the redundancy / migration of workload from compute may be required to a different compute, leaf or different site. In some examples, the output from the health score for different areas may be a proactive determination of a redundancy requirement for a given leaf port to compute combination of hierarchy of leaf port to compute, that is, to which leaf port compute it may be associated, and passing the information to a third node operating in the computer system 100, e.g., an SMO, to take necessary actions. The recommendation, in accordance with Action 406, may be sent to a scheduler which may decide about the movement of the workload to a different leaf or different site. The fifth indication may be a recommendation to move or not workload from the first data center 121. Forecasting metrics and subsequently, the failure probability may beunderstood to help to proactively prevent network outages. In the event of health score being 0, workloads may be migrated to other healthy computes / sites.
[0187] Figure 7 depicts a non-limiting example of a method performed by the second node 112 according to embodiments herein, during the training phase of the one or more ML models. In addition to being already trained on data from redundant deployments and continuing to get data from such deployments, the second node 112 may start collecting the data for model training from non-redundant deployment(s) as well. Accordingly, the second node 112, as described in Action 501 , may obtain: i) the first one or more first indications indicating the fourth performance of one or more bridge ports of the third data center 123 of the computer system 100, ii) the second one or more second indications indicating the fifth performance of one or more Internet Protocol Interfaces of the third data center 123 of the computer system 100, and iii) the third one or more third indications indicating the sixth performance of one or more compute systems of the third data center 123 of the computer system 100. The second node 112 may then be enabled to use the obtained indications to provide input to the one or more ML models and use the one or more ML models to ultimately ensure day to day operations may be with at par level of performance in terms of availability. The one or more ML models may be provided the metrics related to compute performance, interface and bridge port which may enable to identify the port and link health of a data center, such as the third data center 123, e.g., during the training phase of the one or more ML models and later of the first data center 121 during the interference phase of the one or more ML models. Based on collected metrics, the one or more ML models may be trained for identification of anomalies and SME suggested rule based alarm or faults. These anomalies and faults may be used to train the one or more ML models for the classification of e.g., a port as a faulty port or a healthy one.
[0188] Figure 8 depicts a non-limiting example of a method performed by the first node 111 according to embodiments herein, during the inference phase of the one or more ML models. Once forecasting for each category may be done, based on one or more trained ML models, failure probability for each category may be determined as shown in Figure 8. The first node 111 may, according to Action 401 , obtain the one or more first indications indicating the first performance of the one or more bridge ports of the first data center 121 of the computer system 100, the one or more second indications indicating the second performance of the one or more Internet Protocol Interfaces of the first data center 121 of the computer system 100 and the one or more third indications indicating the third performance of the one or more compute systems of the first data center 121 of the computer system 100. The first node 111 may then, according to Action 402, determine the respective prediction of the obtained one ormore first indications, one or more second indications, and one or more third indications at the future time period. Next, the first node 111 may, according to Action 404, determine, based on the obtained one or more first indications, one or more second indications, and one or more third indications, and using the one or more trained machine learning models, the fourth indication of the availability of resources at the first data center 121 at the future time period, as, in this case, the probability at the future time period of the failure happening at the first data center 121. The determining in Action 404 of the fourth indication may be based on applying the one or more rules of the obtained one or more machine learning models on the determined respective predictions. That is, to determine the fourth indication, these forecasted values may be applied to anomaly rules. Anomaly rules may have been created during the training phase of the one or more machine learning models, as described in relation to Figure 5 and Figures 6-7. In case any anomaly is detected in the forecasted pattern, the probability of occurrence of that anomaly may be calculated individually. Further particularly, once forecasting for each category may have been performed, based on the one or more trained ML models, failure probability for each category may be determined. That is, for each of the one or more bridge ports of the first data center 121, the one or more Internet Protocol Interfaces of the first data center 121 and the one or more compute systems of the first data center 121. Failure probability may be understood as a composite score considering a bucket of performance metrics, as described earlier. Once all the failure probabilities may have been accumulated, the fourth indication may be derived as the binary health score, which, according to Action 406, may be used to make the decision about whether or not the workloads may need to be migrated. If the fourth indication, that is, the Health score is 1, then that leaf to compute redundancy may not be required and the network may continue to operate in lean mode at 801. Else, the redundancy / migration of workload from compute may be required to a different compute, leaf or different site at 802. Forecasting metrics and subsequently, the failure probability may be understood to help to proactively prevent network outages. In the event of health score being 0, workloads may be migrated to other healthy computes / sites. This process may continue iteratively for a forecasting time horizon. The effective outcome from iterations may be listed as minimal disruption in service by supporting the same or similar availability characteristics while using much lesser equipment.
[0189] Figure 9 depicts a non-limiting example of a method performed by the second node 112 according to embodiments herein, during a multi-model training and policies generation phase. According to Action 501 , the second node 112 may obtain: i) the first one or more first indications indicating the fourth performance of the one or more bridge ports of the third data center 123 of the computer system 100 from a bridge port data collector, II) the second one or more second indications indicating the fifth performance of the one or more Internet ProtocolInterfaces of the third data center 123 of the computer system 100 from an IP interface data collector, and iii) the third one or more third indications indicating the sixth performance of one or more compute systems of the third data center 123 of the computer system 100 from a computer system data collector. According to Action 502, the second node 112 may perform the quality check on the obtained first one or more first indications, second one or more second indications, and third one or more third indications. According to Action 503, the second node 112 may process the obtained first one or more first indications, second one or more second indications, and third one or more third indications based on the performed quality check, to obtain processed data by one or more of: modifying and discarding data of inadequate quality. According to Action 504, the second node 112 may aggregate the processed data from Action 503 based on the obtained first one or more first indications, second one or more second indications, and third one or more third indications and perform feature engineering on the aggregated data. Then, according to Action 505, the second node 112 may train the one or more machine learning models to determine the fourth indication of the availability of resources at the third data center 123 at the future time period, based on the obtained first one or more first indications, second one or more second indications, and third one or more third indications. Data may be selected on the premise of classifying data instances into 'Known abnormality' and 'Unknown abnormality' . The unknown abnormality may be detected using methods such as the LOF algorithm. Data points with high LOF score may be termed as abnormal data points, e.g., via annotation, and this annotated data may be used in the training of one or more tree based models to get rules for abnormal condition. Inference Data may also be added to these data annotations to train the one or more ML models. These anomalies and faults may be used to train the one or more ML models according to Action 505 for classification of a port as a faulty port or a healthy one. These one or more data tree models may have exhaustive rule generation with root cause and probability.
[0190] Figure 10 depicts a non-limiting example of a method performed by the first node 111 according to embodiments herein, during the inference phase of the one or more ML models. Streamed data may be obtained according to Action 401. According to Action 402, the first node 111 may determine the respective prediction of the obtained one or more first indications, one or more second indications, and one or more third indications at the future time period, e.g., n steps ahead. Then, according to Action 404, the first node 111 may determine, based on the obtained one or more first indications, one or more second indications, and one or more third indications, and using the one or more trained machine learning models, the fourth indication of the availability of resources at the first data center 121 at the future time period. The training by the second node 112 of the one or more machine learning models may have comprised training to detect anomalies in the obtained first one ormore first indications, second one or more second indications, and third one or more third indications, with the created one or more rules. In case any anomaly is detected in the forecasted pattern, the probability of occurrence of that anomaly may be calculated individually. Further particularly, once forecasting for each category may have been performed, based on the one or more trained ML models, failure probability for each category may be determined, that is, bridge port related failure probability, IP interface related failure probability and computer system specific failure probability. Once all the failure probabilities may have been accumulated, the fourth indication may be generated as a health score. The failure probabilities may be used to derive the fourth indication as the binary health score, which may be output according to Action 405 and used according to Action 406 to make the decision about whether or not the workloads may need to be migrated. This process may continue iteratively for a forecasting time horizon. The effective outcome from iterations may be listed as minimal disruption in service by supporting the same or similar availability characteristics while using much lesser equipment.
[0191] As mentioned earlier, in some examples, the computer system 100 may support O-RAN. The O-RAN Alliance may be understood to define O-Cloud as a cloud computing platform comprised of a collection of physical infrastructure nodes that may meet O-RAN requirements to host the relevant O-RAN functions, the supporting software components, and the appropriate management and orchestration functions.
[0192] The Non-RT RIC may be understood to enable non-real-time control and optimization of RAN elements and resources, it may include AI / ML workflow including model training and updates, and policy-based guidance of applications / features in the Near-RT RIC. This may be understood to relate to Near RT-RIC with A1 Interface, which may be understood to be between the Non-RT RIC in the SMO and the Near-RT RIC for RAN Optimization. The functionality of the Non-RT RIC may be understood to be directly responsible for driving what may be sent and received across the A1 interface. The Non-RT RIC may be understood to allow applications to run on it. These applications may be called “rApps”, where ‘r’ may be understood to stand for RAN. The Non-RT RIC may be understood to expose SMO Framework functions to “rApps” via a set of “rApps” Services Exposure” functions over the R1 interface. As the R1 interface may be understood to be the only interface between an “rApps” and the functionality of the Non-RT RIC and SMO, and defined to meet all functional needs of rApps, with appropriate interface extensibility capabilities as needed.
[0193] Embodiments herein may be implemented as an rApp. A centralized rApp may collect data from multiple E2 nodes via 02 interface. The centralized rApp may keep on training on this data and may be producing radio unit specific recommendations.Certain embodiments disclosed herein may provide one or more of the following technical advantage(s), which may be summarized as follows.
[0194] Embodiments herein may be understood to help in reducing the required hardware footprint, hence bringing down the Information Technology infrastructure cost with cost efficiency.
[0195] Embodiments herein may be understood to also enable energy saving due to the optimal use of the deployed hardware.
[0196] Embodiments herein may be understood to also enable higher network subscription / traffic capacity as compared to a traditional data center network design, with a similar set of hardware resources.
[0197] In some examples, the one or more AI / ML models may be trained to recommend workload re-distribution for power savings, even for no-fault scenarios, targeting optimal performance of the deployed infrastructure.
[0198] Embodiments herein may be understood to also enable new mechanisms to manage and life-cycle the proposed thin-network deployment while causing minimal disturbance to the workloads running on them.
[0199] Figure 11 depicts an example of the arrangement that the first node 111 may comprise to perform the method described in Figure 4 and / or in any of Figures 6, Figures 8 and / or Figure 10. The first node 111 may be understood to be for handling the resources. The first node 111 is configured to operate in the computer system 100.
[0200] Several embodiments are comprised herein. It should be noted that the examples herein are not mutually exclusive. One or more embodiments may be combined, where applicable. All possible combinations are not described to simplify the description.
[0201] Components from one embodiment may be tacitly assumed to be present in another embodiment and it will be obvious to a person skilled in the art how those components may be used in the other exemplary embodiments. The detailed description of some of the following corresponds to the same references provided above, in relation to the actions described for the first node 111 and will thus not be repeated here to simplify the description. For example, failure probabilities may be configured to be used to derive the fourth indication as the binary health score.
[0202] In Figure 11 , an optional unit is depicted with a dashed box.
[0203] The first node 111 is configured to obtain i) the one or more first indications configured to indicate the first performance of the one or more bridge ports of the first data center 121 of the computer system 100, ii) the one or more second indications configured to indicate the second performance of the one or more Internet Protocol Interfaces of the first data center 121 of the computer system 100, and iii) the one or more third indications configured to indicate the thirdperformance of the one or more compute systems of the first data center 121 of the computer system 100.
[0204] The first node 111 is also configured to determine, based on the one or more first indications, the one or more second indications, and the one or more third indications configured to be obtained, and using the one or more trained machine learning models, the fourth indication of the availability of resources at the first data center 121 at the future time period.
[0205] The first node 111 is further configured to initiate determining the one or more actions based on the fourth indication configured to be determined. The one or more actions are configured to comprise migrating at least the part of the workload from the first structure configured to be comprised in the first data center 121 to another structure configured to be comprised in the first data center 121 or in the second data center 122.
[0206] The first data center 121 may be configured to lack redundant structures. Particularly, the first data center 121 may be configured to lack redundant network structures.
[0207] The first data center 121 may be configured to comprise at least one of a single spine structure and a single leaf structure.
[0208] The first data center 121 may be configured to comprise at least one of the first plurality of clients and servers, and the second plurality of racks.
[0209] In some embodiments, the fourth indication may be configured to indicate the probability at the future time period of one of: the capacity of network components of the first data center 121 having the value, the failure happening at the first data center 121, and the maintenance operation happening at the first data center 121.
[0210] In some embodiments, the first node 111 may be also configured with the following two configurations.
[0211] In some embodiments, the first node 111 may be configured to determine the respective prediction of the one or more first indications, the one or more second indications, and the one or more third indications configured to be obtained at the future time period.
[0212] In some embodiments, the first node 111 may be configured to obtain the one or more sixth indications configured to indicate the one or more trained machine learning models from the second node 112 configured to operate in the computer system 100, and using the respective predictions configured to be determined as input to the one or more machine learning models configured to be obtained. The determining of the fourth indication may be configured to be based on applying the one or more rules of the one or more machine learning models configured to be obtained on the respective predictions configured to be determined.
[0213] In some embodiments, the first node 111 may be further configured to output the fifth indication of the fourth indication configured to be determined.In some embodiments, one or more of the following options may apply. According to one first option, at least one of the one or more machine learning models configured to be trained may be configured to be the decision tree model. According to another option, the resources may be configured to comprise one or more of: the compute resources, the storage resources and the network resources. According to another option, the network resources may be configured to comprise one or more of: the leaf resources, the spine resource, the switch resources and the link resources. According to another option, at least one of the first structure and the another structure may be configured to be a server. According to yet another option, the network structures of the first data center 121 may be configured to comprise one or more of: the spine structure, the leaf structure, the switches and the links.
[0214] The embodiments herein in the first node 111 may be implemented through one or more processors, such as a processing circuitry 1101 in the first node 111 depicted in Figure 11, together with computer program code for performing the functions and actions of the embodiments herein. A processor, as used herein, may be understood to be a hardware component. The program code mentioned above may also be provided as a computer program product, for instance in the form of a data carrier carrying computer program code for performing the embodiments herein when being loaded into the first node 111. One such carrier may be in the form of a CD ROM disc. It is however feasible with other data carriers such as a memory stick. The computer program code may furthermore be provided as pure program code on a server and downloaded to the first node 111.
[0215] The first node 111 may further comprise a memory 1102 comprising one or more memory units. The memory 1102 is arranged to be used to store obtained information, store data, configurations, schedulings, and applications etc. to perform the methods herein when being executed in the first node 111.
[0216] In some embodiments, the first node 111 may receive information from, e.g., the second node 112, the third node, the first data center 121, the second data center 122, the third data center 123, the radio network node 140, the device 130, another node or device, and / or another structure in the computer system 100, through a receiving port 1103. In some embodiments, the receiving port 1103 may be, for example, connected to one or more antennas in first node 111. Since the receiving port 1103 may be in communication with the processing circuitry 1101, the receiving port 1103 may then send the received information to the processing circuitry 1101. The receiving port 1103 may also be configured to receive other information.
[0217] The processing circuitry 1101 in the first node 111 may be further configured to transmit or send information to e.g., the second node 112, the third node, the first data center 121 , the second data center 122, the third data center 123, the radio network node 140, the device 130, another node or device, and / or another structure in the computer system 100, through asending port 1104, which may be in communication with the processing circuitry 1101, and the memory 1102.
[0218] Those skilled in the art will also appreciate that the units comprised within the first node 111 described above as being configured to perform different actions, may refer to a combination of analog and digital circuits, and / or one or more processors configured with software and / or firmware, e.g., stored in memory, that, when executed by the one or more processors such as the processing circuitry 1101, perform as described above. One or more of these processors, as well as the other digital hardware, may be included in a single Application-Specific Integrated Circuit (ASIC), or several processors and various digital hardware may be distributed among several separate components, whether individually packaged or assembled into a System-on-a-Chip (SoC).
[0219] The first node 111 may be configured to perform any of the Actions described in relation to Figure 4 and / or in any of Figures 6, Figures 8 and / or Figure 10, e.g., by means of the processing circuitry 1101 within the first node 111, configured to perform any of such actions.
[0220] Also, in some embodiments, different units comprised within the first node 111 may be configured to perform the different actions described above in relation to Figure 4 and / or in any of Figures 6, Figures 8 and / or Figure 10, implemented as one or more applications running on one or more processors such as the processing circuitry 1101.
[0221] Thus, the methods according to the embodiments described herein for the first node 111 may be respectively implemented by means of a computer program 1105 product, comprising instructions, i.e. , software code portions, which, when executed on at least one processing circuitry 1101, cause the at least one processing circuitry 1101 to carry out the actions described herein, as performed by the first node 111. The computer program 1105 product may be stored on a computer-readable storage medium 1106. The computer-readable storage medium 1106, having stored thereon the computer program 1105, may comprise instructions which, when executed on at least one processing circuitry 1101, cause the at least one processing circuitry 1101 to carry out the actions described herein, as performed by the first node 111. In some embodiments, the computer-readable storage medium 1106 may be a non-transitory computer-readable storage medium, such as a CD ROM disc, or a memory stick. In other embodiments, the computer program 1105 product may be stored on a carrier containing the computer program 1105 just described, wherein the carrier is one of an electronic signal, optical signal, radio signal, or the computer-readable storage medium 1106, as described above.
[0222] The first node 111 may comprise a communication interface configured to facilitate, or an interface unit to facilitate, communications between the first node 111 and other nodes or devices, e.g., the second node 112, the third node, the first data center 121 , the second data center 122, the third data center 123, the radio network node 140, the device 130, anothernode or device, and / or another structure in the computer system 100. The interface may, for example, include a transceiver configured to transmit and receive radio signals over an air interface in accordance with a suitable standard.
[0223] In other embodiments, the first node 111 may comprise a radio circuitry 1107, which may comprise e.g., the receiving port 1103 and the sending port 1104.
[0224] The radio circuitry 1107 may be configured to set up and maintain at least a wireless connection with the second node 112, the third node, the first data center 121 , the second data center 122, the third data center 123, the radio network node 140, the device 130, another node or device, and / or another structure in the computer system 100. Circuitry may be understood herein as a hardware component.
[0225] Hence, embodiments herein also relate to the first node 111, operative to operate in the computer system 100. The first node 111 may comprise the processing circuitry 1101 and the memory 1102, said memory 1102 containing instructions executable by said processing circuitry 1101, whereby the first node 111 is further operative to perform the actions described herein in relation to the first node 111, e.g., in Figure 4 and / or in any of Figures 6, Figures 8 and / or Figure 10.
[0226] It may be understood that in examples wherein the first node 111 may be the same nodes as the second node 112, the first node 111 may be additionally configured as described in Figure 12.
[0227] Figure 12 depicts an example of the arrangement that the second node 112 may comprise to perform the method described in Figure 5, and / or in any of Figures 6-7, and / or Figure 9. The second node 112 may be understood to be for handling the resources. The second node 112 is configured to operate in the computer system 100.
[0228] Several embodiments are comprised herein. It should be noted that the examples herein are not mutually exclusive. One or more embodiments may be combined, where applicable. All possible combinations are not described to simplify the description.
[0229] Components from one embodiment may be tacitly assumed to be present in another embodiment and it will be obvious to a person skilled in the art how those components may be used in the other exemplary embodiments. The detailed description of some of the following corresponds to the same references provided above, in relation to the actions described for the first node 111 and will thus not be repeated here to simplify the description. For example, failure probabilities may be configured to be used to derive the fourth indication as the binary health score.
[0230] In Figure 12, an optional unit is depicted with a dashed box.
[0231] The second node 112 is configured to obtain: i) the first one or more first indications configured to indicate the fourth performance of the one or more bridge ports of the third datacenter 123 of the computer system 100, ii) the second one or more second indications configured to indicate the fifth performance of the one or more Internet Protocol Interfaces of the third data center 123 of the computer system 100, iii) the third one or more third indications configured to indicate the sixth performance of the one or more compute systems of the third data center 123 of the computer system 100.
[0232] The second node 112 is also configured to train the one or more machine learning models to determine the fourth indication of the availability of resources at the third data center 123 at the future time period, based on the first one or more first indications, second one or more second indications, and third one or more third indications configured to be obtained.
[0233] The second node 112 is further configured to initiate outputting the one or more sixth indications of the one or more trained machine learning models to the first node 111 configured to operate in the computer system 100.
[0234] The third data center 123 may be configured to lack redundant structures. Particularly, the third data center 123 may be configured to lack redundant network structures.
[0235] The third data center 123 may be configured to comprise at least one of a single spine structure and a single leaf structure.
[0236] The third data center 123 may be configured to comprise at least one of the third plurality of clients and servers, and the fourth plurality of racks.
[0237] In some embodiments, the fourth indication may be configured to indicate the probability at the first future time period of one of: the capacity of the third data center 123 having the value, the failure happening at the third data center 123, and the maintenance operation happening at the third data center 123.
[0238] In some embodiments, the second node 112 may be also configured with the following three configurations.
[0239] In some embodiments, the second node 112 may be configured to perform a quality check on the first one or more first indications, second one or more second indications, and third one or more third indications configured to be obtained.
[0240] In some embodiments, the second node 112 may be configured to process the first one or more first indications, second one or more second indications, and third one or more third indications configured to be obtained based on the quality check configured to be performed, to obtain processed data by one or more of: modifying and discarding data of inadequate quality.
[0241] In some embodiments, the first node 111 may be further configured to aggregate the data configured to be processed based on the first one or more first indications, second one or more second indications, and third one or more third indications configured to be obtained and perform feature engineering on the aggregated data, and wherein the training is configured to be performed using the aggregated processed data as input.In some embodiments, the training may be configured to comprise one or more of the following: a) detecting anomalies in the first one or more first indications, second one or more second indications, and third one or more third indications configured to be obtained, b) classifying the anomalies configured to be detected as one of known and unknown, and c) creating one or more rules to determine the availability of the resources at the future time period based on the anomalies configured to be classified.
[0242] In some embodiments, one or more of the following options may apply. According to one first option, at least one of machine learning models may be configured to be the decision tree model. According to another option, the resources may be configured to comprise one or more of: the compute resources, the storage resources and the network resources. According to another option, the network resources may be configured to comprise one or more of: the leaf resources, the spine resource, the switch resources and the link resources. According to yet another option, the network structures of the third data center 123 may be configured to comprise one or more of: the spine structure, the leaf structure, the switches and the links.
[0243] The embodiments herein in the second node 112 may be implemented through one or more processors, such as a processing circuitry 1201 in the second node 112 depicted in Figure 12, together with computer program code for performing the functions and actions of the embodiments herein. A processor, as used herein, may be understood to be a hardware component. The program code mentioned above may also be provided as a computer program product, for instance in the form of a data carrier carrying computer program code for performing the embodiments herein when being loaded into the second node 112. One such carrier may be in the form of a CD ROM disc. It is however feasible with other data carriers such as a memory stick. The computer program code may furthermore be provided as pure program code on a server and downloaded to the second node 112.
[0244] The second node 112 may further comprise a memory 1202 comprising one or more memory units. The memory 1202 is arranged to be used to store obtained information, store data, configurations, schedulings, and applications etc. to perform the methods herein when being executed in the second node 112.
[0245] In some embodiments, the second node 112 may receive information from, e.g., the first node 111, the third node, the first data center 121, the second data center 122, the third data center 123, the radio network node 140, the device 130, another node or device, and / or another structure in the computer system 100, through a receiving port 1203. In some embodiments, the receiving port 1203 may be, for example, connected to one or more antennas in second node 112. Since the receiving port 1203 may be in communication with the processing circuitry 1201 , the receiving port 1203 may then send the received information to the processing circuitry 1201. The receiving port 1203 may also be configured to receive other information.The processing circuitry 1201 in the second node 112 may be further configured to transmit or send information to e.g., the first node 111 , the third node, the first data center 121 , the second data center 122, the third data center 123, the radio network node 140, the device 130, another node or device, and / or another structure in the computer system 100, through a sending port 1204, which may be in communication with the processing circuitry 1201, and the memory 1202.
[0246] Those skilled in the art will also appreciate that the units comprised within the second node 112 described above as being configured to perform different actions, may refer to a combination of analog and digital circuits, and / or one or more processors configured with software and / or firmware, e.g., stored in memory, that, when executed by the one or more processors such as the processing circuitry 1201 , perform as described above. One or more of these processors, as well as the other digital hardware, may be included in a single Application-Specific Integrated Circuit (ASIC), or several processors and various digital hardware may be distributed among several separate components, whether individually packaged or assembled into a System-on-a-Chip (SoC).
[0247] The second node 112 may be configured to perform any of the Actions described in relation to Figure 5, and / or in any of Figures 6-7, and / or Figure 9, e.g., by means of the processing circuitry 1201 within the second node 112, configured to perform any of such actions.
[0248] Also, in some embodiments, different units comprised within the second node 112 may be configured to perform different actions described above in relation to Figure 5, and / or in any of Figures 6-7, and / or Figure 9, implemented as one or more applications running on one or more processors such as the processing circuitry 1201.
[0249] Thus, the methods according to the embodiments described herein for the second node 112 may be respectively implemented by means of a computer program 1205 product, comprising instructions, i.e. , software code portions, which, when executed on at least one processing circuitry 1201 , cause the at least one processing circuitry 1201 to carry out the actions described herein, as performed by the second node 112. The computer program 1205 product may be stored on a computer-readable storage medium 1206. The computer-readable storage medium 1206, having stored thereon the computer program 1205, may comprise instructions which, when executed on at least one processing circuitry 1201 , cause the at least one processing circuitry 1201 to carry out the actions described herein, as performed by the second node 112. In some embodiments, the computer-readable storage medium 1206 may be a non-transitory computer-readable storage medium, such as a CD ROM disc, or a memory stick. In other embodiments, the computer program 1205 product may be stored on a carrier containing the computer program 1205 just described, wherein the carrier is one of an electronic signal, optical signal, radio signal, or the computer-readablestorage medium 1206, as described above.
[0250] The second node 112 may comprise a communication interface configured to facilitate, or an interface unit to facilitate, communications between the second node 112 and other nodes or devices, e.g., the first node 111 , the third node, the first data center 121 , the second data center 122, the third data center 123, the radio network node 140, the device 130, another node or device, and / or another structure in the computer system 100. The interface may, for example, include a transceiver configured to transmit and receive radio signals over an air interface in accordance with a suitable standard.
[0251] In other embodiments, the second node 112 may comprise a radio circuitry 1207, which may comprise e.g., the receiving port 1203 and the sending port 1204.
[0252] The radio circuitry 1207 may be configured to set up and maintain at least a wireless connection with the first node 111 , the third node, the first data center 121 , the second data center 122, the third data center 123, the radio network node 140, the device 130, another node or device, and / or another structure in the computer system 100. Circuitry may be understood herein as a hardware component.
[0253] Hence, embodiments herein also relate to the second node 112, operative to operate in the computer system 100. The second node 112 may comprise the processing circuitry 1201 and the memory 1202, said memory 1202 containing instructions executable by said processing circuitry 1201, whereby the second node 112 is further operative to perform the actions described herein in relation to the second node 112, e.g., in Figure 5, and / or in any of Figures 6-7, and / or Figure 9.
[0254] It may be understood that in examples wherein the second node 112 may be the same nodes as the first node 111 , the second node 112 may be additionally configured as described in Figure 11.
[0255] Figure 13 depicts an example of the arrangement that the first data center 121 may comprise according to embodiments herein.
[0256] The first data center 121 may be configured to lack redundant structures. Particularly, the first data center 121 is configured to lack redundant network structures. The first data center 121 may be configured to comprise at least one of the single spine structure 1301 and the single leaf structure 1302. The first data center 121 may be configured to comprise at least one of the first plurality of clients and servers 1303, and the second plurality of racks 1304. The first data center 121 may be configured to operate in the computer system 100.
[0257] The embodiments herein in the first data center 121 may be implemented through one or more processors, such as a processing circuitry 1301 in the first data center 121 depicted in Figure 13, together with computer program code for performing the functions and actions of the embodiments herein. A processor, as used herein, may be understood to be a hardwarecomponent. The program code mentioned above may also be provided as a computer program product, for instance in the form of a data carrier carrying computer program code for performing the embodiments herein when being loaded into the first data center 121. One such carrier may be in the form of a CD ROM disc. It is however feasible with other data carriers such as a memory stick. The computer program code may furthermore be provided as pure program code on a server and downloaded to the first data center 121.
[0258] The first data center 121 may further comprise a memory 1302 comprising one or more memory units. The memory 1302 is arranged to be used to store obtained information, store data, configurations, schedulings, and applications etc. to perform the methods herein when being executed in the first data center 121.
[0259] In some embodiments, the first data center 121 may receive information from, e.g., the first node 111 , the second node 112, the third node, the second data center 122, the third data center 123, the radio network node 140, the device 130, another node or device, and / or another structure in the computer system 100, through a receiving port 1303. In some embodiments, the receiving port 1303 may be, for example, connected to one or more antennas in first data center 121. Since the receiving port 1303 may be in communication with the processing circuitry 1301 , the receiving port 1303 may then send the received information to the processing circuitry 1301. The receiving port 1303 may also be configured to receive other information.
[0260] The processing circuitry 1301 in the first data center 121 may be further configured to transmit or send information to e.g., the first node 111 , the second node 112, the third node, the second data center 122, the third data center 123, the radio network node 140, the device 130, another node or device, and / or another structure in the computer system 100, through a sending port 1304, which may be in communication with the processing circuitry 1301, and the memory 1302.
[0261] Those skilled in the art will also appreciate that the units comprised within the first data center 121 described above as being configured to perform different actions, may refer to a combination of analog and digital circuits, and / or one or more processors configured with software and / or firmware, e.g., stored in memory, that, when executed by the one or more processors such as the processing circuitry 1301 , perform as described above. One or more of these processors, as well as the other digital hardware, may be included in a single Application-Specific Integrated Circuit (ASIC), or several processors and various digital hardware may be distributed among several separate components, whether individually packaged or assembled into a System-on-a-Chip (SoC).
[0262] Thus, the methods according to the embodiments described herein for the first data center 121 may be respectively implemented by means of a computer program 1305 product, comprising instructions, i.e., software code portions, which, when executed on at leastone processing circuitry 1301 , cause the at least one processing circuitry 1301 to carry out the actions described herein, as performed by the first data center 121. The computer program 1305 product may be stored on a computer-readable storage medium 1306. The computer-readable storage medium 1306, having stored thereon the computer program 1305, may comprise instructions which, when executed on at least one processing circuitry 1301 , cause the at least one processing circuitry 1301 to carry out the actions described herein, as performed by the first data center 121. In some embodiments, the computer-readable storage medium 1306 may be a non-transitory computer-readable storage medium, such as a CD ROM disc, or a memory stick. In other embodiments, the computer program 1305 product may be stored on a carrier containing the computer program 1305 just described, wherein the carrier is one of an electronic signal, optical signal, radio signal, or the computer-readable storage medium 1306, as described above.
[0263] The first data center 121 may comprise a communication interface configured to facilitate, or an interface unit to facilitate, communications between the first data center 121 and other nodes or devices, e.g., the first node 111 , the second node 112, the third node, the second data center 122, the third data center 123, the radio network node 140, the device 130, another node or device, and / or another structure in the computer system 100. The interface may, for example, include a transceiver configured to transmit and receive radio signals over an air interface in accordance with a suitable standard.
[0264] In other embodiments, the first data center 121 may comprise a radio circuitry 1307, which may comprise e.g., the receiving port 1303 and the sending port 1304.
[0265] The radio circuitry 1307 may be configured to set up and maintain at least a wireless connection with the first node 111 , the second node 112, the third node, the second data center 122, the third data center 123, the radio network node 140, the device 130, another node or device, and / or another structure in the computer system 100. Circuitry may be understood herein as a hardware component.
[0266] Hence, embodiments herein also relate to the first data center 121 , operative to operate in the computer system 100. The first data center 121 may comprise the processing circuitry 1301 and the memory 1302, said memory 1302 containing instructions executable by said processing circuitry 1301, whereby the first data center 121 is further configured as described herein in relation to the first data center 121.
[0267] When using the word "comprise" or “comprising”, it shall be interpreted as non- limiting, i.e., meaning "consist at least of".
[0268] The embodiments herein are not limited to the above-described preferred embodiments. Various alternatives, modifications and equivalents may be used. Therefore, the above embodiments should not be taken as limiting the scope of the invention.Generally, all terms used herein are to be interpreted according to their ordinary meaning in the relevant technical field, unless a different meaning is clearly given and / or is implied from the context in which it is used. All references to a / an / the element, apparatus, component, means, step, etc. are to be interpreted openly as referring to at least one instance of the element, apparatus, component, means, step, etc., unless explicitly stated otherwise. The steps of any methods disclosed herein do not have to be performed in the exact order disclosed, unless a step is explicitly described as following or preceding another step and / or where it is implicit that a step must follow or precede another step. Any feature of any of the embodiments disclosed herein may be applied to any other embodiment, wherever appropriate. Likewise, any advantage of any of the embodiments may apply to any other embodiments, and vice versa. Other objectives, features and advantages of the enclosed embodiments will be apparent from the following description.
[0269] As used herein, the expression “at least one of:” followed by a list of alternatives separated by commas, and wherein the last alternative is preceded by the “and” term, may be understood to mean that only one of the list of alternatives may apply, more than one of the list of alternatives may apply or all of the list of alternatives may apply. This expression may be understood to be equivalent to the expression “at least one of:” followed by a list of alternatives separated by commas, and wherein the last alternative is preceded by the “or” term.
[0270] Any of the terms processor and circuitry may be understood herein as a hardware component.
[0271] As used herein, the expression “in some embodiments” has been used to indicate that the features of the embodiment described may be combined with any other embodiment or example disclosed herein.
[0272] As used herein, the expression “in some examples” has been used to indicate that the features of the example described may be combined with any other embodiment or example disclosed herein.
[0273] REFERENCES
[0274] 1. Lapukhov P. et al., RFC 7938: Use of BGP for Routing in Large-Scale Data Centers (rfc-editor.org), Internet Engineering Task Force (IETF), August 2016.
[0275] 2. O-RAN ALLIANCE e.V.
[0276] 3. D’Oro, S. et al, OrchestRAN: Network Automation through Orchestrated Intelligence in the Open RAN (2201.05632 (arxiv.org)), IEEE International Conference on Computer Communications (INFOCOM) 2022.
Claims
CLAIMS:
1. A computer-implemented method, performed by a first node (111 ), the method being for handling resources, the first node (111) operating in a computer system (100), the method comprising:- obtaining (401):i. one or more first indications indicating a first performance of one or more bridge ports of a first data center (121) of the computer system (100),ii. one or more second indications indicating a second performance of one or more Internet Protocol Interfaces of the first data center (121) of the computer system (100), andiii. one or more third indications indicating a third performance of one or more compute systems of the first data center (121) of the computer system (100),- determining (404), based on the obtained one or more first indications, one or more second indications, and one or more third indications, and using one or more trained machine learning models, a fourth indication of an availability of resources at the first data center (121) at a future time period, and- initiating (406) determining one or more actions based on the determined fourth indication, the one or more actions comprising migrating at least a part of a workload from a first structure comprised in the first data center (121) to another structure comprised in the first data center (121) or in a second data center (122).
2. The method according to claim 1, wherein the first data center (121) lacks redundant network structures.
3. The method according to claim 2, wherein the first data center (121) comprises at least one of a single spine structure and a single leaf structure.
4. The method according to claim 3, wherein the first data center (121) comprises at least one of a first plurality of clients and servers, and a second plurality of racks.
5. The method according to any of claims 1-4, wherein the fourth indication indicates a probability at the future time period of one of:- a capacity of network components of the first data center (121) having a value,a failure happening at any of the components of the first data center (121), and a maintenance operation happening at the first data center (121).
6. The method according to any of claims 1-5, further comprising:- determining (402) a respective prediction of the obtained one or more first indications, one or more second indications, and one or more third indications at the future time period,- obtaining (403) one or more sixth indications indicating the one or more trained machine learning models from a second node (112) operating in the computer system (100), and using the determined respective predictions as input to the obtained one or more machine learning models, and wherein the determining (404) of the fourth indication is based on applying one or more rules of the obtained one or more machine learning models on the determined respective predictions.
7. The method according to any of claims 1-6, wherein the method further comprises:- outputting (405) a fifth indication of the determined fourth indication.
8. The method according to any of claims 1-7, wherein one or more of:- at least one of the one or more trained machine learning models is a decision tree model,- the resources comprise one or more of: compute resources, storage resources and network resources,- the network resources comprise one or more of: leaf resources, spine resource, switch resources and link resources,- at least one of the first structure and the another structure is a server, and - network structures of the first data center (121 ) comprise one or more of: a spine structure, a leaf structure, switches and links.
9. A computer-implemented method, performed by a second node (112), the method being for handling resources, the second node (112) operating in a computer system (100), the method comprising:- obtaining (501):i. first one or more first indications indicating a fourth performance of one or more bridge ports of a third data center (123) of the computer system (100),ii. second one or more second indications indicating a fifth performance of one or more Internet Protocol Interfaces of the third data center (123) of the computer system (100),iii. third one or more third indications indicating a sixth performance of one or more compute systems of the third data center (123) of the computer system (100),- training (505) one or more machine learning models to determine a fourth indication of an availability of resources at the third data center (123) at a future time period, based on the obtained first one or more first indications, second one or more second indications, and third one or more third indications, and - initiating (506) outputting one or more sixth indications of the one or more trained machine learning models to a first node (111) operating in the computer system (100).
10. The method according to claim 9, wherein the third data center (123) lacks redundant network structures.
11. The method according to claim 10, wherein the third data center (123) comprises at least one of a single spine structure and a single leaf structure.
12. The method according to any of claims 9-11 , wherein the fourth indication indicates a probability at the first future time period of one of:- a capacity of the third data center (123) having a value,- a failure happening at the third data center (123), and- a maintenance operation happening at the third data center (123).
13. The method according to any of claims 9-12, wherein the method further comprises:- performing (502) a quality check on the obtained first one or more first indications, second one or more second indications, and third one or more third, - processing (503) the obtained first one or more first indications, second one or more second indications, and third one or more third indications based on the performed quality check, to obtain processed data by one or more of: modifying and discarding data of inadequate quality, and- aggregating (504) the processed data based on the obtained first one or more first indications, second one or more second indications, and third one or more third indications and performing feature engineering on the aggregated data,and wherein the training (505) is performed using the aggregated processed data as input.
14. A first data center (121) configured to lack redundant network structures.
15. The first data center (121) according to claim 14, wherein the first data center (121) is configured to comprise at least one of a single spine structure and a single leaf structure.
16. A first node (111 ), for handling resources, the first node (111) being configured to operate in a computer system (100), the first node (111) being further configured to:- obtain:i. one or more first indications configured to indicate a first performance of one or more bridge ports of a first data center (121) of the computer system (100),ii. one or more second indications configured to indicate a second performance of one or more Internet Protocol Interfaces of the first data center (121) of the computer system (100), andill. one or more third indications configured to indicate a third performance of one or more compute systems of the first data center (121) of the computer system (100),- determine, based on the one or more first indications, one or more second indications, and one or more third indications configured to be obtained, and using one or more trained machine learning models, a fourth indication of an availability of resources at the first data center (121) at a future time period, and - initiate determining one or more actions based on the fourth indication configured to be determined, the one or more actions being configured to comprise migrating at least a part of a workload from a first structure configured to be comprised in the first data center (121) to another structure configured to be comprised in the first data center (121) or in a second data center (122).
17. A second node (112), for handling resources, the second node (112) being configured to operate in a computer system (100), the second node (112) being further configured to:obtain:i. first one or more first indications configured to indicate a fourth performance of one or more bridge ports of a third data center (123) of the computer system (100),ii. second one or more second indications configured to indicate a fifth performance of one or more Internet Protocol Interfaces of the third data center (123) of the computer system (100),iii. third one or more third indications configured to indicate a sixth performance of one or more compute systems of the third data center (123) of the computer system (100),- train one or more machine learning models to determine a fourth indication of an availability of resources at the third data center (123) at a future time period, based on the first one or more first indications, second one or more second indications, and third one or more third indications configured to be obtained, and- initiate outputting one or more sixth indications of the one or more trained machine learning models to a first node (111) configured to operate in the computer system (100).