Communication Network Fault Detection and Recovery
The system automates the detection and recovery of communication network failures by reallocating workloads to alternative data centers, addressing the inefficiencies of manual intervention and enhancing network reliability.
Patent Information
- Application Number
- JP2025547971
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-05-31
- Publication Date
- 2026-02-20
AI Technical Summary
Communication networks face challenges in efficiently detecting and resolving failures, particularly in data centers, leading to extended downtime due to the lack of automated backup mechanisms and reliance on human intervention.
A system comprising a policy engine, networking controller, and network orchestrator that automatically detects alarm conditions across network nodes and reallocates workloads to alternative data centers, minimizing downtime through automated fault detection and recovery.
The system enables rapid and efficient recovery from network failures by reallocating workloads to compatible nodes, improving network reliability and reducing downtime in complex communication networks.
Smart Images

Figure 2026506162000001_ABST
Abstract
Description
[Technical Field]
[0001] FIELD OF THE DISCLOSURE The present disclosure relates to detecting failures in communication networks and restoring communication networks. [Background technology]
[0002] Network service providers and device manufacturers (e.g., wireless, cellular, etc.) are under constant pressure to deliver value and convenience to consumers by offering compelling network services that are, for example, reliable, configurable, scalable, diverse, and economically operable. Summary of the Invention
[0003] One aspect of the present disclosure is directed to an apparatus comprising a processor and a memory having stored thereon instructions that, when executed by the processor, cause the apparatus to process a first notification received from a first network node to determine a first status of the first network node. The apparatus also processes a second notification received from a second network node based on the instructions to determine a second status of the second network node. The apparatus further causes, based on the instructions, the apparatus to reallocate a workload assigned to a first data center associated with the first network node and the second network node to a second data center different from the first data center in response to determining that the first status and the second status indicate an alarm condition.
[0004] Another aspect herein is directed to a method that includes processing, by a processor, a first notification received from a first network node to determine a first status of the first network node. The method also includes processing a second notification received from a second network node to determine a second status of the second network node. In response to determining that the first status and the second status indicate an alarm condition, the method further includes causing a workload assigned to a first data center associated with the first network node and the second network node to be reallocated to a second data center different from the first data center.
[0005] Another aspect of the present specification is directed to a non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor, cause an apparatus to process a first notification received from a first network node to determine a first status of the first network node. The apparatus also processes a second notification received from a second network node based on the instructions to determine a second status of the second network node. The apparatus further causes, based on the instructions, in response to determining that the first status and the second status indicate an alarm condition, to reallocate a workload assigned to a first data center associated with the first network node and the second network node to a second data center different from the first data center.
[0006] Aspects of the present disclosure are best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be noted that, according to standard industry practice, various features have not been drawn to scale. In fact, dimensions of various features may be arbitrarily increased or decreased for clarity of illustration. [Brief explanation of the drawings]
[0007] [Figure 1] 1 is a diagram of a system that facilitates detection of a communication network failure and restoration of the communication network, according to one or more embodiments.
[0008] [Figure 2] 1 is a flowchart of a process for detecting a communication network failure and restoring a communication network according to one or more embodiments.
[0009] [Figure 3] FIG. 1 is a data flow diagram of a process for detecting a data center failure, according to one or more embodiments.
[0010] [Figure 4] FIG. 10 is a data flow diagram of the subsequent process after a data center failure is detected, according to one or more embodiments.
[0011] [Figure 5] FIG. 10 is a data flow diagram of the subsequent process after a data center failure is detected, according to one or more embodiments.
[0012] [Figure 6] FIG. 10 is a data flow diagram of the subsequent process after a data center failure is detected, according to one or more embodiments.
[0013] [Figure 7] FIG. 10 is a data flow diagram of the subsequent process after a data center failure is detected, according to one or more embodiments.
[0014] [Figure 8] FIG. 10 is a data flow diagram of the subsequent process after a data center failure is detected, according to one or more embodiments.
[0015] [Figure 9] FIG. 10 is a data flow diagram of the subsequent process after a data center failure is detected, according to one or more embodiments.
[0016] [Figure 10]FIG. 10 is a data flow diagram of the subsequent process after a data center failure is detected, according to one or more embodiments.
[0017] [Figure 11] FIG. 1 is a functional block diagram of a computer or processor-based system in which one embodiment may be implemented. DETAILED DESCRIPTION OF THE INVENTION
[0018] The following disclosure provides many different embodiments or examples for implementing different features of the provided subject matter. To simplify the disclosure, specific examples of components, values, operations, materials, arrangements, etc. are described below. Of course, these are merely examples and are not intended to be limiting. Other components, values, operations, materials, arrangements, etc. are included in the present disclosure. For example, when a first component (element, part) is formed over or on a second component (element, part) in the following description, this may include an embodiment in which the first component and the second component are formed in direct contact with each other, or an embodiment in which an additional element (part) may be formed between the first component and the second component such that the first component and the second component are not in direct contact with each other. In addition, the present disclosure may repeat reference numerals and / or symbols in various examples. This repetition is for the purpose of brevity and clarity and does not in itself dictate a relationship between the various embodiments and / or configurations described. Additionally, this disclosure may omit some actions, such as "respond" or "send receipt," that correspond to a previous action, for purposes of brevity and clarity.
[0019] Additionally, spatially relative terms such as "bottom," "lower," "bottom," "upper," and "top" may be used herein to facilitate the description of the relationship of one element or component to another element or component(s), as shown in the figures. Spatially relative terms are intended to encompass different orientations of the device during use or operation in addition to the orientation shown in the figures. The device may be otherwise oriented (rotated 90 degrees or at other orientations), and the spatially relative descriptions and expressions used herein may likewise be interpreted accordingly.
[0020] Some communication networks are provided by employing a network orchestrator that deploys numerous network functions. For example, some communication networks include the deployment of thousands of radio access network (RAN) network functions. Such communication systems are prone to errors due to systematic failures. For example, some cell sites in a communication network are served by radio units that are divided into centralized units (CUs) and distributed units (DUs). When a complete outage occurs at a distributed point of presence (Distri PoP), the CUs running on the corresponding cluster often do not have readily available backups. This often results in a complete outage of the cell site served by the CUs running on the failed Distri PoP, even if the connected DUs are operational. Similar problems arise in the event of a rack failure, power failure, spine switch failure, access gateway switch failure, etc. Traditionally, communication network failures are resolved through human intervention, which is tedious, time-consuming, and can result in extended downtime of the communication network and / or network functions or services before the failure is resolved.
[0021] FIG. 1 is a diagram of a system 100 that facilitates detection of communication network failures and restoration of communication networks, according to one or more embodiments.
[0022] System 100 facilitates automatically detecting and addressing failures in a communications network and / or a data center that at least partially implements the communications network. In some embodiments, system 100 comprises one or more components of a centralized unit (CU).
[0023] The system 100 includes a policy engine 101, a database 103, a networking controller 105, a network orchestrator 107, an observability framework 109, and network nodes 111A, 111B.
[0024] The policy engine 101 is a network assurance policy engine that triggers actions on the communication network managed by the network orchestrator 107 based on matching condition(s) associated with communication network faults defined in the database 103. The networking controller 105 is a centralized, programmable point of automation configured to manage, configure, monitor, and troubleshoot the network infrastructure. The network orchestrator 107 is a network controller that functions in setting up devices, applications, and services in the network to allocate workloads. In some embodiments, the functionality of the above components, including the networking controller 105 and the network orchestrator 107, is software-defined and can be combined into or replaced by a single component. In some embodiments, one or more components included in system 100, such as policy engine 101, networking controller 105, network orchestrator 107, observability framework 109, or network nodes 111A, 111B, include a set of computer-readable instructions that, when executed by a processor, such as processor 1103 (FIG. 11), cause one or more of policy engine 101, networking controller 105, network orchestrator 107, observability framework 109, or network nodes 111A, 111B to perform a process described according to one or more embodiments.
[0025] In some embodiments, database 103 is an inventory of information including node data, node identifiers (IDs), node types, node locations, data center data processing capabilities, node characteristics, node capabilities, node substitution rules, or other suitable information about various network nodes. Database 103 is a memory, such as memory 1105 (FIG. 11), that can be queried or have data stored therein, according to one or more embodiments.
[0026] In some embodiments, network nodes 111A, 111B are connected to data centers 113A, 113B. In some embodiments, system 100 comprises three or more network nodes connected to one or more corresponding data centers 113. In some embodiments, two or more network nodes 111A, 111B (or any additional network nodes 111) are connected to the same data center, such as, for example, data center 113A. Data centers 113A, 113B are, for example, locations (elements) to which workloads associated with operation of a communications network are assigned. In some embodiments, network nodes 111A, 111B comprise one or more of a switch, a border leaf switch, a spine switch, an access gateway switch, a computer, a router, or some other suitable network device or network element. In some embodiments, network nodes 111A, 111B are of the same network node type. For example, in some embodiments, both network nodes 111A, 111B are border leaf switches. In some embodiments, network nodes 111A, 111B are of different network node types. For example, in some embodiments, network node 111A is a spine switch and network node 111B is a border leaf switch.
[0027] The observability framework 109 is communicatively coupled to the network nodes 111A, 111B. The observability framework 109 receives notifications from the network nodes 111A, 111B, including node identification information and hostnames, and forwards the notifications and / or information contained therein to the policy engine 101.
[0028] In some embodiments, the policy engine 101 processes a first notification received from a first network node 111A to determine a first status of the first network node 111A. The first notification includes a first node identifier and a first hostname corresponding to the first network node 111A. The policy engine 101 also processes a second notification received from a second network node 111B to determine a second status of the second network node 111B. The second notification includes a second node identifier and a second hostname corresponding to the second network node 111B.
[0029] Policy engine 101 then determines, based on the first node identifier, the first hostname, the second node identifier, and the second hostname, whether the first notification and the second notification are defined in database 103 as indicating an alarm condition. In response to determining that the first condition and the second condition indicate an alarm condition, policy engine 101 causes the workload assigned to the first data center 113 associated with first network node 111A and second network node 111B to be reallocated to a second data center 113 different from the first data center 113. For example, if the first condition and the second condition indicate an alarm condition, the workload assigned to data center 113A is reallocated to data center 113B or to some other appropriate data center 113 identifiable based on information contained in database 103.
[0030] In some embodiments, the policy engine 101 is configured to identify two or more alternative network nodes 111 that can be used as the first network node 111A or the second network node 111B to facilitate (justify) the workload reallocated to the second data center 113, by searching the database 103 for two or more alternative network nodes 111 of compatible types from among the plurality of alternative network nodes 111 based on the descriptions of the first network node 111A, the second network node 111B, and the plurality of alternative network nodes 111 contained in the database 103. The policy engine 101 then causes the workload to be reallocated to the second data center 113 in response to identifying the two or more alternative network nodes 111.
[0031] In some embodiments, before causing the workload to be reallocated to the second data center 113 in response to identifying two or more alternative network nodes 111, the policy engine 101 is configured to double-check the status of the first network node 111A and the second network node 111B to confirm an alarm condition before reallocating the workload. For example, in some embodiments, the policy engine 101 (1) processes a third notification received from the first network node 111A to determine a third status of the first network node 111A, (2) processes a fourth notification received from the second network node 111B to determine a fourth status of the second network node 111B, and (3) causes the workload to be reallocated to the second data center 113 in response to determining that the third status and the fourth status indicate an alarm condition. The first notification, the second notification, the third notification, and the fourth notification are communicated (sent) to the policy engine 101 via, for example, the observability framework 109.
[0032] In some embodiments, the policy engine 101 is configured to process the third and fourth notifications after a predetermined period of time has elapsed after determining that the first and second conditions indicate an alarm condition. For example, to avoid premature reassignment of the workflow, in some embodiments, the policy engine is configured to wait five minutes or some other suitable period of time to determine whether a condition on one or more of the first network node 111A or the second network node 111B is in an alarm condition based on the third and fourth notifications before reassigning the workflow to the second data center 113.
[0033] In some embodiments, one or more other notifications are received between the first and third notifications, for example, and one or more other notifications are received between the second and fourth notifications, for example. By delaying the processing of the third and fourth notifications for a predetermined period of time, workload reallocation is avoided in situations where the determined alarm condition is in fact a false alarm or where the first network node and / or the second network node self-resolves the failure that caused the alarm condition within the predetermined period of time.
[0034] In some embodiments, policy engine 101 causes network orchestrator 107 to instantiate the workload reallocated to second data center 113. Policy engine 101 then updates database 103 to include information indicating the workload reallocated to second data center 113 and information indicating an association between alternative network node 111 and second data center 113. In some embodiments, to instantiate the workload allocated to second data center 113, network orchestrator 107 pushes at least day 1 and day 2 configurations to alternative network node 111 to facilitate running the workload after reallocation to second data center 113.
[0035] In some embodiments, in the event of a communication network failure, such as a disaster occurring when a distributed point-of-presence (Distri PoP) server outage occurs with no backup centralized unit operating on the failed cluster, system 100 enables the communication network to be automatically repaired. In some embodiments, even if an entire Distri PoP fails or a significant portion of it fails (e.g., several server racks fail), repair occurs by reallocating the centralized unit workload to a different data center and / or by employing different compatible network nodes. Such automated repair helps improve the reliability and performance of the communication network while minimizing downtime. Furthermore, by automating the fault condition detection and reallocation process, system 100 enables efficient repair of failures indicated by alarm conditions, thereby making it possible to manage communication networks with an ever-increasing number and complexity of network devices, data centers, network nodes, etc., with minimal user oversight.
[0036] 2 is a flowchart of a process 200 for detecting a communication network failure and restoring the communication network according to one or more embodiments. In some embodiments, the policy engine 101 (FIG. 1) performs the process 200.
[0037] In step 201, a first notification received from a first network node is processed to determine a first status of the first network node. In some embodiments, the first notification includes a first node identifier and a first hostname corresponding to the first network node.
[0038] In step 203, a second notification received from the second network node is processed to determine a second status of the second network node. In some embodiments, the second notification includes a second node identifier and a second hostname corresponding to the second network node.
[0039] In optional (but not required) step 205, a determination is made as to whether the first notification and the second notification are defined in the database as indicating an alarm condition based on the first node identifier, the first host name, the second node identifier, and the second host name.
[0040] In optional step 207, based on the descriptions of the first network node, the second network node, and the plurality of alternative network nodes contained in the database, two or more alternative network nodes that can be used as the first network node or the second network node are identified by searching the database to find two or more alternative network nodes of a compatible type from among the plurality of alternative network nodes.
[0041] In step 209, in response to determining that the first condition and the second condition indicate an alarm condition, a workload assigned to a first data center associated with the first network node and the second network node is reallocated to a second data center different from the first data center. In some embodiments, the workload is reallocated to the second data center in response to identifying two or more alternative network nodes that can be used as the first network node or the second network node to facilitate the workload reallocated to the second data center.
[0042] In some embodiments, before reallocating the workload to the second data center in response to identifying two or more alternative network nodes, an alarm condition is confirmed before reallocating the workload. For example, in some embodiments, a third notification received from the first network node is processed to determine a third condition of the first network node, a fourth notification received from the second network node is processed to determine a fourth condition of the second network node, and in response to determining that the third condition and the fourth condition indicate an alarm condition, the workload is reallocated to the second data center.
[0043] In some embodiments, the third and fourth notifications are processed after a preset period of time has elapsed since determining that the first and second conditions indicate an alarm condition. For example, to avoid premature reassignment of a workflow, in some embodiments, a preset period of five minutes, or some other suitable period of time, elapses before determining whether a condition in one or more of the first or second network nodes is in an alarm condition based on the third and fourth notifications before reassigning the workflow to the second data center. In some embodiments, the third and fourth notifications immediately follow a series of notifications from the first and second network nodes received via the observability framework. In some embodiments, one or more other notifications are received between the first and third notifications, for example, and one or more other notifications are received between the second and fourth notifications, for example. By delaying the processing of the third and fourth notifications for a predetermined period of time, workload reallocation is avoided in situations where the determined alarm condition is in fact a false alarm or where the first network node and / or the second network node self-resolves the failure that caused the alarm condition within the predetermined period of time.
[0044] In optional step 211, the network orchestrator instantiates the reallocated workload in the second data center.
[0045] In optional step 213, the database is updated to include information indicating the workload reallocated to the second data center and information indicating an association between the alternative network node and the second data center. In some embodiments, to instantiate the workload allocated to the second data center, the network orchestrator pushes at least day 1 and day 2 configurations to the alternative network node to facilitate running the workload after reallocation to the second data center.
[0046] 3-10 are data flow diagrams of processes 300, 400, 500, 600, 700, 800, 900, 1000 for detecting a data center failure and subsequent actions after detecting a failure, according to some embodiments.
[0047] The above processes 300, 400, 500, 600, 700, 800, 900, 1000 are performed by a processor, such as processor 1103, described below with respect to Figure 11. In some embodiments, some or all of the operations of the processes and subsequent operations for detecting a data center failure are performed according to instructions stored in memory 1105, described below with respect to Figure 11.
[0048] Process 300 for detecting a data center failure includes operations 351 through 373. Processes 400, 500, 600, 700, 800, 900, and 1000 include operations 451 through 465, 551 through 567, 651 through 679, 751 through 777, 851 through 869, 951 through 955, and 1051 through 1079. The operations described do not necessarily occur in the order shown. Operations may be added, substituted, reordered, and / or deleted as appropriate, consistent with the spirit and scope of the embodiments. In some embodiments, one or more of the operations in these processes are repeated. In some embodiments, the operations in these processes occur sequentially, unless otherwise specified.
[0049] (Fault)
[0050] FIG. 3 is a data flow diagram of a process 300 for detecting data center failures, according to one or more embodiments.
[0051] At operation 351, a failure is detected at node 311A and / or node 311B. At operation 353, after detecting the network node failure, at least one of network nodes 311A, 311B sends a failure notification to application policy infrastructure controller (APIC) 313. In response to receiving the failure notification, APIC 313 forwards the failure notification to observability framework (OBF) 309 at operation 355. OBF 309 then forwards the failure notification to policy engine 301 at operation 357. In some embodiments, the failure notification includes a node identifier (e.g., "2101, 2102...") and a host name.
[0052] (Bug fix)
[0053] At operation 359, policy engine 301 determines that the fault notification triggers a predefined policy. At operation 361, policy engine 301 sends a request to database 303 for data center details and IPs of all or some of the network nodes identified in database 303. At operation 363, database 303 sends a response to policy engine 301 including the data center details and IPs. In some embodiments, the data center details include one or more of a distributed unit (DU) F1C alarm correlation or a centralized unit (CU) network configuration protocol session status.
[0054] In operation 365, policy engine 301 requests network orchestrator 307 to trigger a ping request on network nodes 311A, 311B. In operation 367, network orchestrator 307 sends an Internet Control Message Protocol (ICMP) / secure shell (SSH) / application programming interface (API) call to network nodes 311A, 311B. In operation 369, a negative response is received from network nodes 311A, 311B by network orchestrator 307. In some embodiments, operation 369 represents no response being received from network nodes 311A, 311B by network orchestrator 307. In some embodiments, operations 365 through 371 are health checks that are repeated two or more times to check for negative responses and / or failure to receive a response. In some embodiments, operations 365 through 371 are repeated five or other suitable predefined number of times.
[0055] In some embodiments, in response to both responses, or the lack of a response, from network node 311A, 311B indicating that the network node is in a failed state, the policy engine triggers a rehoming process to reassign the workload to another data center in operation 373. In some embodiments, the rehoming process includes network orchestrator 307 determining two or more alternative network nodes that can be used as the network node in terms of compatibility, capacity, or other suitable factors according to information available in database 303. Network orchestrator 307 then pushes the day 1 and day 2 configurations to the alternative network nodes, for example, to facilitate carrying the workload.
[0056] (Inventory / DNS Cleanup)
[0057] FIG. 4 is a data flow diagram of a subsequent process 400 following detection of a data center failure, according to one or more embodiments.
[0058] In operation 451, network orchestrator 407 fetches details of existing network services (NSs) and network functions (NFs) from database 403. In operation 453, database 403 sends a response with details of the existing NDs and NFs to the network orchestrator. In operation 455, network orchestrator 407 instructs internet protocol (IP) address manager 415 to release the existing NF IP. Then, in operation 457, network orchestrator 407 receives a success response from IP address manager 415. In operation 459, network orchestrator 407 instructs domain name server (DNS) 417 to deregister the existing fully qualified domain name (FQDN) entry. In operation 461, DNS 417 sends a response to network orchestrator 407 indicating that the existing FQDN entry was successfully deregistered. In operation 463, the network orchestrator 407 sends an instruction to change the state of the NF from instantiated to rehomed to the database 403. Then, in operation 465, the network orchestrator 407 receives a success response message from the database 403.
[0059] (Parameter generation on day 0)
[0060] FIG. 5 is a data flow diagram of a subsequent process 500 following detection of a data center failure, according to one or more embodiments.
[0061] In operation 551, network orchestrator 507 fetches details of a backup cluster data center (CDC) from database 503. In operation 553, network orchestrator 507 receives the backup CDC details from database 503. Then, in operation 555, network orchestrator 507 checks the existing network function topology (NFT) and network service topology (NST) and selects the same NFT / NST for deployment to the CDC. In operation 557, network orchestrator 507 applies business logic and validates the NFT / NST selected for deployment. In operation 559, network orchestrator 507 requests IP address manager 515 to provide a new IP according to the CDC IP schema. In operation 561, IP address manager 515 generates a new IP and responds to network orchestrator 507 with the new IP. At operation 563, network orchestrator 507 sends the new IP to database 503 to update database 503. At operation 565, network orchestrator 507 applies business logic and validation for NF deployment and parameter generation and generates a day 0 payload using the details and new IP received from database 503. In some embodiments, the same node ID and hostname are used.
[0062] In operation 567, the network orchestrator 307 sends the day 0 payload to the vault 519 for loading onto the NF.
[0063] (Focused Unit NF)
[0064] FIG. 6 is a data flow diagram of a process 600 that follows after detecting a data center failure, according to one or more embodiments.
[0065] In operation 651, the network orchestrator 607 sends an instruction to instantiate a centralization unit (CU) on the CDC 621. In operation 653, the CDC 621 sends an instantiation completion message to the network orchestrator 607. In operation 655, the network orchestrator 607 registers the FQDN with the new IP in the Domain Name System (DNS) 617. In operation 657, the DNS 617 sends a success response message to the network orchestrator 607. In operation 659, the CDC 621 performs a transport layer security (TLS) handshake with the configuration manager 625. In operations 661 to 663, the CDC 621 and the configuration manager 625 perform a call home establishment process and a supervision subscription process.
[0066] At operation 665, the configuration manager 625 requests the dynamic day 1 and day 2 parameters from the network orchestrator 607. In response, the network orchestrator 607 sends the dynamic day 1 and day 2 parameters to the configuration manager 625 at operation 667. At operation 669, the configuration manager 625 sends a success response to the network orchestrator 607.
[0067] At operation 671, the configuration manager 625 generates complete Day 1 & Day 2 configurations. At operation 673, the configuration manager 625 sends a configuration generation success message to the network orchestrator 607. At operation 675, the configuration manager 625 pushes the day 1 and day 2 configurations (settings) to the CDC 621. At operation 677, the CDC 621 sends a response to the configuration manager 625 indicating that the configurations have been successfully implemented / committed. At operation 679, the configuration manager 625 sends a day 1 and day 2 configuration push success message to the network orchestrator 607.
[0068] (Distributed Unit NF Option 1)
[0069] FIG. 7 is a data flow diagram of a subsequent process 700 after detecting a centralized unit failure, according to one or more embodiments.
[0070] In operation 751, the network orchestrator 707 instructs the Citi PoP 723 to terminate the distributed unit (DU) NF. In operation 753, the network orchestrator 707 instructs the Citi PoP to re-instantiate the DU NF. In operation 755, the Citi PoP 723 sends a response message to the network orchestrator 707 indicating that the instantiation is complete. In operation 757, the Citi PoP 723 and the configuration manager 725 perform a TLS handshake. In operations 759 to 761, the Citi PoP 723 and the configuration manager 725 perform a call home establishment process and a monitoring subscription process.
[0071] At operation 763, deployment manager (DM) 727 requests network orchestrator 707 to provide the dynamic day 1 and day 2 parameters to configuration manager 725. At operation 765, network orchestrator 707 sends a response with the dynamic day 1 and day 2 parameters to configuration manager 725. At operation 767, configuration manager 725 sends a success message to network orchestrator 707.
[0072] At operation 769, the configuration manager 725 generates the complete day 1 and day 2 configurations. At operation 771, the configuration manager 725 sends a configuration generation success message to the network orchestrator 707. At operation 773, the configuration manager 725 pushes the day 1 and day 2 configurations to the Citi PoP 723. At operation 775, the Citi PoP 723 responds to the configuration manager 725 indicating that the configurations were successfully implemented / committed. At operation 777, the DM 727 sends a day 1 and day 2 configuration push success message to the network orchestrator 707.
[0073] (Distributed Unit NF Option 2)
[0074] FIG. 8 is a data flow diagram of a process 800 that follows after detecting a data center failure, according to one or more embodiments.
[0075] In operation 851, the network orchestrator 807 sends an instruction to the configuration manager 825 to delete existing F1 control plane (F1C) remote endpoints and F1C Internet Protocol Security (IPSEC) remote endpoints in all distributed unit (DU) NFs associated with the centralized unit (CU).
[0076] In operation 853, the configuration manager 825 forwards an instruction to the Citi PoP 823 to delete the existing F1C remote endpoints and F1C IPSEC remote endpoints in all DU NFs associated with the CU.
[0077] In some embodiments, a new F1C remote endpoint and a new F1C IPSEC remote endpoint are added in all DU NFs associated with the CU by the configuration manager 825. At operation 855, the configuration manager 825 sends a message to the network orchestrator 807 requesting the dynamic F1C parameters. At operation 857, the network orchestrator 807 sends the dynamic F1C parameters to the configuration manager 825. At operation 859, the configuration manager 825 sends a success response message to the network orchestrator 807.
[0078] In operation 861, the configuration manager 825 generates the F1C configuration. In operation 863, the configuration manager 825 sends a configuration generation success message to the network orchestrator 807.
[0079] At operation 865, the configuration manager pushes the F1C config to Citi PoP 823. At step 867, Citi PoP 823 returns a response to the configuration manager 825 indicating that the configuration was successfully implemented / committed. At operation 869, the configuration manager 825 sends an F1C config push success message to the network orchestrator 807.
[0080] (Distributed Unit NF Option 3)
[0081] FIG. 9 is a data flow diagram of a subsequent process 900 after detecting a centralized unit failure, according to one or more embodiments.
[0082] In operation 951, the FQDN in the DU configuration for the F1 IP is used.
[0083] The IP address of the FQDN is updated in operation 953. In some embodiments, operation 953 is an API update (patch) operation.
[0084] In operation 955, all DNS and intermediate DNS caching servers, etc. are flushed. In some embodiments, the flushing of DNS and intermediate DNS caching servers occurs with some delay depending on whether there is any negligible negative caching in the DNS setup.
[0085] (Power-on notification for distributed unit NF option 1)
[0086] FIG. 10 is a data flow diagram of a subsequent process 1000 after detecting a centralized unit failure, according to one or more embodiments.
[0087] In operation 1051, an open radio access network (O-RAN) radio unit (ORU) 1029 obtains IP and certification authority (CA) details from a dynamic host configuration protocol (DHCP) server 1031. In operation 1053, the ORU 1029 registers with a CA 1033. In some embodiments, the CA 1033 is an Enterprise JavaBeans Certificate Authority (EJBCA) or some other suitable certification authority. In operation 1055, the network orchestrator 1007 and the ORU 1029 perform a TLS handshake. In operations 1057-1059, the network orchestrator 1007 and the ORU 1029 perform a call-home establishment process and an ORU serial number fetch process.
[0088] In operation 1061, the network orchestrator 1007 sends an ORU parent DU FQDN request to the database 1003. In operation 1065, the database 1003 sends the requested parent DU FQDN to the network orchestrator 1007. In some embodiments, the DU FQDN is the management (MGMT) FQDN of the DU's Rumantsch Grischun (RUMGR)_plain old data (PoD) structure.
[0089] In operation 1067, the network orchestrator 1007 and the ORU 1029 set a parent DU FQDN. In operation 1069, the network orchestrator 1007 and the ORU 1029 fetch session details. In operation 1071, the network orchestrator 1007 and the ORU 1029 set a periodic timer. In operation 1073, a Network Configuration Protocol (Netconf) connection is established between the Citi PoP 1023 and the ORU 1029. In some embodiments, the Citi PoP 1023 is an O-RAN distributed unit (ODU).
[0090] In some embodiments, the ORU 1029 sets up a netconf connection with the Citi PoP 1023 to facilitate the transfer of wireless configuration over the Netconf connection between the ORU 1029 and the Citi PoP 1023. In some embodiments, the Citi PoP 1023 is the parent ODU.
[0091] In operation 1077, the Citi PoP 1023 pushes the wireless configuration to the ORU 1029. In operation 1079, the ORU 1029 sends a configuration push success message to the Citi PoP 1023. After the configuration push, the DU is operational.
[0092] FIG. 11 is a functional block diagram of a computer or processor-based system 1100 in which one embodiment is implemented.
[0093] The processor-based system 1100 is programmed to facilitate detection of communication network failures and restoration of communication networks as described herein and includes components such as a bus 1101, a processor 1103, and a memory 1105.
[0094] In some embodiments, the processor-based system is implemented as a single "system on a chip." The processor-based system 1100, or portions thereof, constitutes a mechanism for detecting a communications network failure and taking one or more steps to restore the communications network.
[0095] In some embodiments, the processor-based system 1100 includes a communication mechanism, such as a bus 1101, for transferring and / or receiving information and / or instructions between components of the processor-based system 1100. The processor 1103 is connected to the bus 1101 to retrieve instructions for execution and process information stored in, for example, a memory 1105. In some embodiments, the processor 1103 also includes one or more specialized components that perform specific processing functions and tasks, such as one or more digital signal processors (DSPs) or one or more application-specific integrated circuits (ASICs). A DSP is typically configured to process real-world signals (e.g., sound) in real time independently of the processor 1103. Similarly, an ASIC can be configured to perform specialized functions not easily performed by a more general-purpose processor. Other specialized components that help perform the functions described herein optionally include one or more field-programmable gate arrays (FPGAs), one or more controllers, or one or more other specialized computer chips.
[0096] In one or more embodiments, the processor(s) 1103 performs a series of operations on information as specified by a set of instructions stored in memory 1105 related to detecting communications network failures and restoring communications networks. Execution of the instructions causes the processor to perform the specified functions.
[0097] The processor 1103 and associated components are connected to memory 1105 via bus 1101. The memory 1105 includes one or more of dynamic memory (e.g., RAM, magnetic disk, writable optical disk, etc.) and static memory (e.g., ROM, CD-ROM, etc.) for storing executable instructions that, when executed, perform steps described herein that facilitate detection of communications network failures and restoration of communications networks. The memory 1105 also stores data associated with or generated by the execution of the steps.
[0098] In one or more embodiments, memory 1105, such as random access memory (RAM) or any other dynamic storage device, stores information, including processor instructions, for detecting a communications network failure and restoring the communications network. Dynamic memory allows the information stored therein to be changed. RAM allows units of information stored at locations, called memory addresses, to be stored and retrieved independently of information at neighboring addresses. Memory 1105 is also used by processor 1103 to store temporary values during execution of processor instructions. In various embodiments, memory 1105 is read-only memory (ROM) or any other static storage device coupled to bus 1101 for storing static information, including instructions, that cannot be changed by processor 1103. Some memories consist of volatile storage, which loses the information stored therein when power is lost. In some embodiments, memory 1105 is a non-volatile (persistent) storage device such as a magnetic disk, optical disk, or flash card for storing information, including instructions, that persists even when system 1100 is turned off or otherwise loses power.
[0099] The term "computer-readable medium" as used herein refers to any medium that participates in providing information, including instructions, to the processor 1103 for execution. Such media take many forms, including, but not limited to, computer-readable storage media (e.g., non-volatile media, volatile media). Non-volatile media include, for example, optical and magnetic disks. Volatile media include, for example, dynamic memory. Common forms of computer-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tape, other magnetic media, CD-ROMs, CDRWs, DVDs, other optical media, punch cards, paper tape, optical mark sheets, other physical media with patterns of holes or other optically recognizable indicia, RAM, PROMs, EPROMs, FLASH-EPROMs, EEPROMs, flash memory, other memory chips or cartridges, or other media from which a computer can read. The term computer-readable storage medium is used herein to refer to a computer-readable medium.
[0100] Various embodiments herein include the following examples.
[0101] Example [1] is an apparatus comprising: a processor; and memory having stored thereon instructions that, when executed by the processor, cause the apparatus to process a first notification received from a first network node to determine a first status of the first network node. The apparatus also processes a second notification received from a second network node to determine a second status of the second network node. In response to determining that the first status and the second status indicate an alarm condition, the apparatus further causes a workload assigned to a first data center associated with the first network node and the second network node to be reallocated to a second data center different from the first data center.
[0102] Example [2] is the apparatus of Example [1], wherein the first notification includes a first node identifier and a first hostname corresponding to the first network node, and the second notification includes a second node identifier and a second hostname corresponding to the second network node, and the apparatus further determines whether the first notification and the second notification are defined in the database as indicating an alarm condition based on the first node identifier, the first hostname, the second node identifier, and the second hostname. The apparatus further identifies two or more alternative network nodes that can be used as the first network node or the second network node for facilitating the workload reallocated to the second data center by searching the database for alternative network nodes of compatible types based on descriptions of the first network node, the second network node, and the alternative network node included in the database. The apparatus also causes the workload to be reallocated to the second data center in response to identifying the two or more alternative network nodes.
[0103] Example [3] is the apparatus of any one of Examples [1] to [2], wherein, before causing the reallocation of the workload to the second data center in response to identifying two or more alternative network nodes, the apparatus processes a third notification received from the first network node to determine a third condition of the first network node. The apparatus also processes a fourth notification received from the second network node to determine a fourth condition of the second network node. The apparatus further causes the reallocation of the workload to the second data center in response to determining that the third condition and the fourth condition indicate an alarm condition.
[0104] Example [4] is the device of any one of Examples [1] to [3], wherein the third notification and the fourth notification are processed a predetermined period of time after determining that the first condition and the second condition indicate an alarm condition.
[0105] Example [5] is the apparatus of any one of Examples [1] to [4], wherein the first notification, the second notification, the third notification, and the fourth notification are received via an observability framework communicatively coupled to the first network node and the second network node.
[0106] Example [6] is the apparatus of any one of Examples [1] to [5], wherein the apparatus further causes the network orchestrator to instantiate the reallocated workload to the second data center. The apparatus also updates the database to include information indicating the reallocated workload to the second data center and information indicating an association between the alternative network node and the second data center.
[0107] Example [7] is the apparatus of any one of Examples [1] to [6], wherein the network orchestrator pushes at least day 1 and day 2 configurations to an alternative network node to instantiate the workload assigned to the second data center, facilitating running the workload after reassignment to the second data center.
[0108] Example [8] is the apparatus of any one of Examples [1] to [7], wherein the first network node and the second network node are switches.
[0109] Example [9] is the apparatus of any one of Examples [1] to [8], wherein the first network node and the second network node are border leaf switches.
[0110] Example
[10] is the apparatus of any one of Examples [1] to [9], wherein the first network node and the second network node are of the same network node type, and the network node type is a switch, a spine switch, an access gateway switch, a computer, or a router.
[0111] Example
[11] is the apparatus of any one of Examples [1] to
[10] , wherein the first network node is a first network node type, the second network node is a second network node type different from the first network node type, and the first network node type and the second network node type include one or more of a switch, a border leaf switch, a spine switch, an access gateway switch, a computer, or a router.
[0112] Example
[12] is a method including, by a processor, processing a first notification received from a first network node to determine a first condition of the first network node. The method also includes processing a second notification received from a second network node to determine a second condition of the second network node. In response to determining that the first condition and the second condition indicate an alarm condition, the method further includes causing a workload assigned to a first data center associated with the first network node and the second network node to be reallocated to a second data center different from the first data center.
[0113] Example
[13] is the method of Example
[12] , wherein the first notification includes a first node identifier and a first hostname corresponding to the first network node, and the second notification includes a second node identifier and a second hostname corresponding to the second network node. The method further includes determining whether the first notification and the second notification are defined in a database as indicating an alarm condition based on the first node identifier, the first hostname, the second node identifier, and the second hostname. The method also includes identifying two or more alternative network nodes that can be used as the first network node or the second network node to facilitate the workload reassigned to the second data center by searching the database for alternative network nodes of compatible types based on descriptions of the first network node, the second network node, and the alternative network node included in the database. The method further includes causing the workload to be reassigned to the second data center in response to identifying the two or more alternative network nodes.
[0114] Example
[14] is the method of any one of Examples
[12] to
[13] , wherein the method further includes processing a third notification received from the first network node to determine a third condition of the first network node before causing the workload to be reallocated to the second data center in response to identifying two or more alternative network nodes. The method also includes processing a fourth notification received from the second network node to determine a fourth condition of the second network node. The method further includes causing the workload to be reallocated to the second data center in response to determining that the third condition and the fourth condition indicate an alarm condition.
[0115] Example
[15] is the method of any one of Examples
[12] to
[14] , wherein the third notification and the fourth notification are processed after a predetermined period of time has elapsed since determining that the first condition and the second condition indicate an alarm condition.
[0116] Example
[16] is the method of any one of Examples
[12] to
[15] , further including causing a network orchestrator to instantiate the reallocated workload at the second data center. The method further includes updating the database to include information indicating the reallocated workload at the second data center and information indicating an association between the replacement network node and the second data center.
[0117] Example
[17] is the method of any one of Examples
[12] to
[16] , wherein the first network node and the second network node are border leaf switches.
[0118] Example
[18] is the method of any one of Examples
[12] to
[17] , wherein the first network node and the second network node are of the same network node type, and the network node type is a switch, a spine switch, an access gateway switch, a computer, or a router.
[0119] Example
[19] is the method of any one of Examples
[12] to
[18] , wherein the first network node is a first network node type, the second network node is a second network node type different from the first network node type, and the first network node type and the second network node type include one or more of a switch, a border leaf switch, a spine switch, an access gateway switch, a computer, or a router.
[0120]
[0020] is a non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor, cause an apparatus to process a first notification received from a first network node to determine a first status of the first network node. The apparatus also processes a second notification received from a second network node to determine a second status of the second network node. In response to determining that the first status and the second status indicate an alarm condition, the apparatus further causes a workload assigned to a first data center associated with the first network node and the second network node to be reallocated to a second data center different from the first data center.
[0121] The foregoing outlines features of several embodiments so that those skilled in the art may better understand aspects of the present disclosure. Those skilled in the art will appreciate that they may readily use this disclosure as a basis for designing or modifying other processes and structures which carry out the same purposes and / or achieve the same advantages as the embodiments presented herein. Those skilled in the art will also recognize that such equivalent constructions do not depart from the spirit and scope of the present disclosure, and that various changes, substitutions, and alterations can be made in the present disclosure without departing from the spirit and scope of the present disclosure. While features of the present disclosure have been expressed in particular combinations, it is contemplated that these features can be arranged in any combination and order without departing from the spirit and scope of the present disclosure.
Claims
1. 1. An apparatus comprising: a processor; a memory having stored thereon instructions that, when executed by the processor, cause the device to perform operations, the operations including: processing a first notification received from a first network node to determine a first status of the first network node; processing a second notification received from a second network node to determine a second status of the second network node; in response to determining that the first condition and the second condition indicate an alarm condition, reallocating workload assigned to a first data center associated with the first network node and the second network node to a second data center different from the first data center; 1. An apparatus comprising:
2. the first notification includes a first node identifier and a first hostname corresponding to the first network node, the second notification includes a second node identifier and a second hostname corresponding to the second network node, and the actions the instructions cause the device to perform include: determining whether the first notification and the second notification are defined in a database as indicating the alarm condition based on the first node identifier, the first host name, the second node identifier, and the second host name; identifying two or more alternative network nodes that can be used as the first network node or the second network node to facilitate the workload reallocated to the second data center by searching the database for two or more alternative network nodes of compatible types from among the plurality of alternative network nodes based on descriptions of the first network node, the second network node, and a plurality of alternative network nodes contained in the database; reassigning the workload to the second data center in response to identifying the two or more alternative network nodes; The apparatus of claim 1 further comprising:
3. The instructions cause the device to perform, before reallocating the workload to the second data center in response to identifying the two or more alternative network nodes, processing a third notification received from the first network node to determine a third status of the first network node; processing a fourth notification received from the second network node to determine a fourth status of the second network node; in response to determining that the third condition and the fourth condition indicate the alarm condition, reallocating the workload to the second data center; The apparatus of claim 2 further comprising:
4. 4. The device of claim 3, wherein the third notification and the fourth notification are processed a predetermined time period after determining that the first condition and the second condition indicate the alarm condition.
5. 4. The apparatus of claim 3, wherein the first notification, the second notification, the third notification, and the fourth notification are received via an observability framework communicatively coupled to the first network node and the second network node.
6. The actions that the instructions cause the device to perform further include: instantiating the reassigned workload to the second data center by a network orchestrator; updating the database to include information indicating the workload reallocated to the second data center and information indicating an association between the alternative network node and the second data center; 3. The apparatus of claim 2, comprising:
7. 7. The apparatus of claim 6, wherein the network orchestrator pushes at least day 1 and day 2 configurations to the alternative network node to instantiate the workload assigned to the second data center to facilitate running the workload after reassignment to the second data center.
8. The apparatus of claim 1 , wherein the first network node and the second network node are switches.
9. The apparatus of claim 8 , wherein the first network node and the second network node are border leaf switches.
10. 10. The apparatus of claim 1, wherein the first network node and the second network node are of the same network node type, and the network node type is a switch, a spine switch, an access gateway switch, a computer, or a router.
11. the first network node is of a first network node type and the second network node is of a second network node type different from the first network node type; the first network node type and the second network node type include one or more of a switch, a border leaf switch, a spine switch, an access gateway switch, a computer, or a router; 10. The apparatus of claim 1.
12. processing, by a processor, a first notification received from a first network node to determine a first status of the first network node; processing a second notification received from a second network node to determine a second status of the second network node; in response to determining that the first condition and the second condition indicate an alarm condition, causing a workload assigned to a first data center associated with the first network node and the second network node to be reallocated to a second data center different from the first data center; A method comprising:
13. the first notification includes a first node identifier and a first hostname corresponding to the first network node, and the second notification includes a second node identifier and a second hostname corresponding to the second network node, and the method further comprises: determining whether the first notification and the second notification are defined in a database as indicating the alarm condition based on the first node identifier, the first host name, the second node identifier, and the second host name; identifying two or more alternative network nodes that can be used as the first network node or the second network node to facilitate the workload reassigned to the second data center by searching the database for the alternative network nodes of a compatible type based on descriptions of the first network node, the second network node, and the alternative network node contained in the database; reassigning the workload to the second data center in response to identifying the two or more alternative network nodes; The method of claim 12 further comprising:
14. Prior to reallocating the workload to the second data center in response to identifying the two or more alternative network nodes, the method further comprises: processing a third notification received from the first network node to determine a third status of the first network node; processing a fourth notification received from the second network node to determine a fourth status of the second network node; in response to determining that the third condition and the fourth condition indicate the alarm condition, causing the workload to be reallocated to the second data center; and 14. The method of claim 13, further comprising:
15. 15. The method of claim 14, wherein the third notification and the fourth notification are processed a predetermined time period after determining that the first condition and the second condition indicate the alarm condition.
16. causing a network orchestrator to instantiate the reassigned workload at the second data center; updating the database to include information indicating the workload reallocated to the second data center and information indicating an association between the alternative network node and the second data center; 14. The method of claim 13, further comprising:
17. 13. The method of claim 12, wherein the first network node and the second network node are border leaf switches.
18. 13. The method of claim 12, wherein the first network node and the second network node are of the same network node type, and the network node type is a switch, a spine switch, an access gateway switch, a computer, or a router.
19. the first network node is of a first network node type and the second network node is of a second network node type different from the first network node type; the first network node type and the second network node type include one or more of a switch, a border leaf switch, a spine switch, an access gateway switch, a computer, or a router; The method of claim 12.
20. A non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor, cause an apparatus to perform an operation, the operation including: processing a first notification received from a first network node to determine a first status of the first network node; processing a second notification received from a second network node to determine a second status of the second network node; in response to determining that the first condition and the second condition indicate an alarm condition, reallocating workload assigned to a first data center associated with the first network node and the second network node to a second data center different from the first data center; 1. A non-transitory computer-readable medium comprising:
Citation Information
Patent Citations
Information reporting method and information processing method, and device
EP4047886A1
Distributed Workload Reallocation after Communication Failure
JP2017530437A
Management method, system, and device for master and standby databases
JP2020502686A
Computer implementation method, computer program and system (distributed multi-environment stream computing)
JP2022094947A
Systems and methods for resource utilization analysis in information management environments
US20020152305A1