System and method for split-brain handling in a communication network

The system addresses split-brain scenarios in clustered networks by storing and replicating delta data changes to maintain data consistency and enhance network resilience, reducing service disruptions.

WO2026062710A1PCT designated stage Publication Date: 2026-03-26JIO PLATFORMS LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

In clustered communication networks, split-brain scenarios occur when clusters disconnect, leading to inconsistent data synchronization and potential service disruptions due to autonomous operation of disconnected clusters, causing network management and performance issues.

Method used

A system and method for managing split-brain conditions by detecting cluster disconnections, storing delta representations of data changes during the split-brain scenario, and replicating these changes upon reconnection to ensure data consistency across NRF clusters.

Benefits of technology

Ensures seamless network operation by reducing data mismatches and discrepancies, enhancing fault tolerance, and minimizing service interruptions through efficient data synchronization and consistency management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IN2025051540_26032026_PF_FP_ABST
    Figure IN2025051540_26032026_PF_FP_ABST
Patent Text Reader

Abstract

A system (108) and a method (400) for managing split-brain conditions in a network (106) is disclosed. Upon detecting network repository function (NRF) clusters (112) that have lost connectivity and are operating independently due to split-brain conditions, a split-brain handling process is triggered to cause the clusters (112) to store delta representations of data changes during the split-brain conditions in a data storage. Once connectivity is restored, the delta representations are replicated across all clusters to ensure data synchronization and seamless data recovery in the event of network disruptions. By replicating the delta representations between the clusters, the risk of data inconsistencies is minimized, and consistency is maintained across all clusters.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR SPLIT-BRAIN HANDLING IN A COMMUNICATION NETWORKRESERVATION OF RIGHTS

[0001] A portion of the disclosure of this patent document contains material, which is subject to intellectual property rights such as, but are not limited to, copyright, design, trademark, Integrated Circuit (IC) layout design, and / or trade dress protection, belonging to Jio Platforms Limited (JPL) or its affiliates (hereinafter referred as owner). The owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure, as it appears in the Patent and Trademark Office patent files or records, but otherwise reserves all rights whatsoever. All rights to such intellectual property are fully reserved by the owner.TECHNICAL FIELD

[0002] The present disclosure relates generally to the field of communication systems. More particularly, the present disclosure relates to systems and methods for split-brain handling between a plurality of network repository function (NRF) clusters in a communication network.DEFINITION

[0003] As used in the present disclosure, the following terms are generally intended to have the meaning as set forth below, except to the extent that the context in which they are used to indicate otherwise.

[0004] The term ‘Network Repository Function (NRF)’ used hereinafter in the specification refers to a centralized repository that maintains comprehensive records about various available network functions (NFs), including their capabilities, configurations, and current status.

[0005] The term ‘Network Function (NF)’ used hereinafter in the specification refers to a specific software or hardware component within a communication network(interchangeably referred to as network) and is designed to perform a particular function, such as routing, switching, firewalling, load balancing, traffic optimization, and the like, to enable network operations and enhance performance.

[0006] The term “NRF cluster” refers to a group of interconnected NRFs that work together as a single system to achieve high availability and load balancing. By clustering, these NRFs can collaborate to handle tasks more efficiently and provide continuous service even if one or more NRFs fail.

[0007] The term “Split-brain condition / scenario” refers to a scenario that occurs when the connection between clusters is disrupted, causing the disconnected cluster to operate autonomously without synchronizing data with the other cluster.

[0008] The term “Disconnected cluster” refers to a situation where a cluster becomes disconnected or isolated from the rest of the system (or the rest of the clusters) due to a network partition, communication failure, or other disruptions.

[0009] The term “Network Function Set (NFSet)” refers to a collection of NFs that are grouped based on certain criteria, such as their role, functionality, or service requirements. The NFs work together to provide a specific set of services or manage particular aspects of the network.

[0010] The term “Split-brain handling process” refers to a mechanism used to detect and resolve situations, i.e., split-brain conditions, where network nodes (e.g., NRF clusters) lose communication with each other but continue to operate independently, leading to potential inconsistent or conflicting network behavior.

[0011] The term “Delta representations” refers to a method of capturing and storing deltas, i.e., data changes during the split-brain conditions, rather than the entire dataset.

[0012] The term “Temporary data storage” refers to a data storage used to store data temporarily (i.e., a short time period) during the split-brain conditions.

[0013] The term “Data synchronization” refers to a process of detecting and resolving data conflicts when the NRF clusters operated independently due to a network partition or failure and are later reconnected.

[0014] The term “Communication timeout” refers to a condition that occurs when a response is not received within a predefined time limit during a communication attempt between the NRFs.

[0015] The term “Timestamp” refers to time-related information (i.e., time and date) indicating when the delta representations of data changes are captured and stored in the data storage.

[0016] The term “Connection status” refers to the current state of network connectivity and communication links between NRF clusters, indicating whether they are able to communicate and coordinate with each other. The connection status comprises a connected state (i.e., operational state) and a disconnected state (i.e., non- operational state).

[0017] The term “Call-back indication” refers to a notification to inform about the connection state of the NRF clusters in the network.

[0018] These definitions are in addition to those expressed in the art.BACKGROUND

[0019] The following description of related art is intended to provide background information pertaining to the field of the disclosure. This section may include certain aspects of the art that may be related to various features of the present disclosure. However, it should be appreciated that this section be used only to enhancethe understanding of the reader with respect to the present disclosure, and not as admissions of prior art.

[0020] In a clustered network, multiple network repository functions (NRFs) are deployed and continuously replicate data between clusters. A split-brain scenario occurs when the connection between clusters is disrupted, causing the disconnected cluster to operate autonomously without synchronizing data with the other clusters. As a result, the disconnected cluster continues to function independently, while the remaining clusters continue to share data among themselves. This lack of synchronization leads to inconsistencies and divergence in network states and configurations, resulting in conflicting information being maintained by each cluster. Such discrepancies cause network management and operation issues, potentially leading to service disruptions and degraded network performance.

[0021] Therefore, there is a need for systems and methods to recover from the split-brain scenarios and reduce the risk of data mismatches and discrepancies.OBJECTIVES

[0022] Some of the objectives of the present disclosure, which at least one embodiment herein satisfies, are as follows:

[0023] An objective of the present disclosure is to provide a system and a method for managing split-brain conditions in a network.

[0024] Another objective of the present disclosure is to handle the split- brain conditions between clusters of network repository functions (NRFs) in the network.

[0025] Another objective of the present disclosure is to store changes in data (i.e., delta representations of data changes) during the split-brain conditions.

[0026] Yet another objective of the present disclosure is to compare and update the states of the NRF clusters when connectivity is restored.

[0027] Yet another objective of the present disclosure is to perform replication of delta representations of data changes between the plurality of NRF clusters after the split-brain situation ends.

[0028] Yet another objective of the present disclosure is to ensure all clusters converge to same state and maintain consistent data across a network function set (NFset).

[0029] Yet another objective of the present disclosure is to enhance fault tolerance by allowing the network to continue operating seamlessly, even if individual NRFs or entire clusters experience disruptions.

[0030] Yet another objective of the present disclosure is to reduce the amount of data transferred and enhance the overall performance of the NFs.

[0031] Yet another objective of the present disclosure is to reduce the complexity involved in ensuring data consistency by simplifying the management of clustered NRFs for handling split-brain conditions in the network.

[0032] Yet another objective of the present disclosure is to mitigate the risk of data mismatch and discrepancies and provide more reliable network operations by tracking and synchronizing the data changes.

[0033] Yet another objective of the present disclosure is to proactively resolve potential conflicts by replicating the latest changes with timestamps.

[0034] Yet another objective of the present disclosure is to ensure the most upto-date information across all clusters.

[0035] Yet another objective of the present disclosure is to share updated information among all NRFs, restoring data consistency and maintaining data integrity across the NFSet.

[0036] Yet another objective of the present disclosure is to reduce downtime and minimize service interruptions.

[0037] Other objectives and advantages of the present disclosure will be more apparent from the following description, which is not intended to limit the scope of the present disclosure.SUMMARY

[0038] In an exemplary embodiment, a method for managing split-brain conditions in a network is disclosed. The method comprises detecting a disconnection of at least one cluster within the network. The method comprises triggering a splitbrain handling process in response to the detected disconnection. Triggering of the split-brain handling process causes each cluster to store delta representations of data changes in a data storage during the split-brain conditions. The method comprises monitoring a connection status of the at least one disconnected cluster to detect reestablishment of the connection. The method comprises terminating the split-brain handling process upon detection of the re-establishment of the connection. The method comprises initiating replication of the delta representations among a plurality of clusters to synchronize the data changes and restore consistency across the plurality of clusters.

[0039] In some embodiments, the plurality of clusters comprises a plurality of network repository function (NRF) clusters.

[0040] In some embodiments, the split-brain conditions comprise at least one of the disconnection of the at least one cluster in the network and each cluster operating independently of data synchronization.

[0041] In some embodiments, the cluster disconnection is detected based on a first indication received by other clusters from at least one NRF of the at least one disconnected cluster in the network. The connection reestablishment is detected based on a second indication by the other clusters from the at least one NRF.

[0042] In some embodiments, the first indication of cluster disconnection is generated based on a loss of connection or communication timeout between at least two NRFs. The second indication of connection reestablishment is generated based on resumption of connection between the at least two NRFs. The resumption of connection is detected through at least one of network-level communication or control signaling.

[0043] In some embodiments, the data storage used to store data changes during the split-brain conditions is configured as a temporary or intermediate delta storage module.

[0044] In some embodiments, the data changes across the plurality of clusters are synchronized based on a timestamp.

[0045] In some embodiments, the replication of the delta representations comprises replicating, by each cluster, the delta representations from other clusters.

[0046] In another exemplary embodiment, a system for managing split-brain conditions in a network is disclosed. The system comprises a detection unit configured to detect a disconnection of at least one cluster within the network. The detection unit is configured to trigger a split-brain handling process in response to the detected disconnection. Triggering of the split-brain handling process causes each cluster to store delta representation of data changes in a data storage during the split-brainconditions. A monitoring unit is configured to monitor a connection status of the at least one disconnected cluster to detect re-establishment of the connection and terminate the split-brain handling process upon detection of the re-establishment of the connection. A replication unit is configured to initiate replication of delta representations among the plurality of clusters to synchronize the data changes and restore consistency across the plurality of clusters.

[0047] In yet another exemplary embodiment, a computer program product comprising a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to execute a method for managing split-brain conditions in a network is disclosed. The method comprises detecting a disconnection of at least one cluster within the network. The method comprises triggering a split-brain handling process in response to the detected disconnection. Triggering of the split-brain handling process causes each cluster to store delta representations of data changes in a data storage during the split-brain conditions. The method comprises monitoring a connection status of the at least one disconnected cluster to detect re-establishment of the connection. The method comprises terminating the split-brain handling process upon detection of the reestablishment of the connection. The method comprises initiating replication of the delta representations among a plurality of clusters to synchronize the data changes and restore consistency across the plurality of clusters.BRIEF DESCRIPTION OF THE ACCOMPANYING DRAWING

[0048] The accompanying drawings, which are incorporated herein, and constitute a part of this disclosure, illustrate exemplary embodiments of the disclosed methods and systems in which like reference numerals refer to the same parts throughout the different drawings. Components in the drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the present disclosure. Some drawings may indicate the components using block diagramsand may not represent the internal circuitry of each component. It will be appreciated by those skilled in the art that disclosure of such drawings includes disclosure of electrical components, electronic components or circuitry commonly used to implement such components.

[0049] FIG. 1 illustrates an exemplary network architecture for implementing a system for managing split-brain conditions in a network, in accordance with an embodiment of the present disclosure.

[0050] FIG. 2A illustrates an exemplary system architecture of the system for managing split-brain conditions in the network, in accordance with an embodiment of the present disclosure.

[0051] FIG. 2B illustrates an exemplary block diagram of the system for managing split-brain conditions in the network, in accordance with an embodiment of the present disclosure.

[0052] FIG. 3 illustrates an exemplary flow diagram of a method for managing split-brain conditions in the network, in accordance with an embodiment of the present disclosure.

[0053] FIG. 4 illustrates another flow diagram of a method for managing splitbrain conditions in the network, in accordance with an embodiment of the present disclosure.

[0054] FIG. 5 illustrates an exemplary block diagram of a computer system in which or with which embodiments of the present disclosure may be implemented.

[0055] The foregoing shall be more apparent from the following more detailed description of the disclosure.LIST OF REFERENCE NUMERALS100 Network Architecture102 User104 User Equipment106 Network 108 System110 Network Function (NF)112 Network Repository Function (NRF) clusters / NRFs200A System Architecture200B Block Diagram 202 Processor204 Memory206 Interface(s)208 Detection Unit210 Monitoring Unit 212 Replication Unit214 Receiving Unit216 Database300 Flow Diagram400 Flow Diagram500 Computer System510 External Storage Device520 Bus530 Main Memory540 Read-Only Memory550 Mass Storage Device560 Communication Ports570 ProcessorDETAILED DESCRIPTION

[0056] In the following description, for the purposes of explanation, various specific details are set forth in order to provide a thorough understanding of embodiments of the present disclosure. It will be apparent, however, that embodiments of the present disclosure may be practiced without these specific details. Several features described hereafter can each be used independently of one another or with any combination of other features. An individual feature may not address any of the problems discussed above or might address only some of the problems discussed above. Some of the problems discussed above might not be fully addressed by any of the features described herein. Example embodiments of the present disclosure are described below, as illustrated in various drawings in which like reference numerals refer to the same parts throughout the different drawings.

[0057] The ensuing description provides exemplary embodiments only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the exemplary embodiments will provide those skilled in the art with an enabling description for implementing an exemplary embodiment. It shouldbe understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the disclosure as set forth.

[0058] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it will be understood by one of ordinary skill in the art that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.

[0059] Also, it is noted that individual embodiments may be described as a process that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.

[0060] The word “exemplary” and / or “demonstrative” is used herein to mean serving as an example, instance, or illustration. For the avoidance of doubt, the subject matter disclosed herein is not limited by such examples. In addition, any aspect or design described herein as “exemplary” and / or “demonstrative” is not necessarily to be construed as preferred or advantageous over other aspects or designs, nor is it meant to preclude equivalent exemplary structures and techniques known to those of ordinary skill in the art. Furthermore, to the extent that the terms “includes,” “has,” “contains,” and other similar words are used in either the detailed description or the claims, suchterms are intended to be inclusive like the term “comprising” as an open transition word without precluding any additional or other elements.

[0061] Reference throughout this specification to “one embodiment” or “an embodiment” or “an instance” or “one instance” means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. Thus, the appearances of the phrases “in one embodiment” or “in an embodiment” in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0062] The terminology used herein is to describe particular embodiments only and is not intended to be limiting the disclosure. As used herein, the singular forms “a”, “an”, and “the” are intended to include the plural forms as well, unless the context indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any combinations of one or more of the associated listed items. It should be noted that the terms “mobile device”, “user equipment”, “user device”, “communication device”, “device” and similar terms are used interchangeably for the purpose of describing the invention. These terms are not intended to limit the scope of the invention or imply any specific functionality or limitations on the described embodiments. The use of these terms is solely for convenience and clarity of description. The invention is not limited to any particular type of device or equipment, and it should be understood that other equivalent terms or variations thereof may be used interchangeably without departing from the scope of the invention as defined herein.

[0063] While considerable emphasis has been placed herein on the components and component parts of the preferred embodiments, it will be appreciated that many embodiments can be made and that many changes can be made in the preferred embodiments without departing from the principles of the disclosure. These and other changes in the preferred embodiment as well as other embodiments of the disclosure will be apparent to those skilled in the art from the disclosure herein, whereby it is to be distinctly understood that the foregoing descriptive matter is to be interpreted merely as illustrative of the disclosure and not as a limitation.

[0064] In a clustered network, multiple network repository functions (NRFs) are deployed and continuously replicate data between clusters. A split-brain scenario can occur when the connection between clusters is disrupted, causing the disconnected cluster to operate autonomously without synchronizing data with the other clusters. As a result, the disconnected cluster continues to function independently, while the remaining clusters continue to share data among themselves. This lack of synchronization can lead to inconsistencies and divergence in network states and configurations, resulting in conflicting information being maintained by each cluster. Such discrepancies can cause network management and operation issues, potentially leading to service disruptions and degraded network performance.

[0065] Therefore, there is a need for systems and methods to recover from the split-brain scenarios and reduce the risk of data mismatches and discrepancies.

[0066] The present disclosure aims to overcome the above-mentioned and other existing problems in this field of technology by providing a system and a method for managing split-brain conditions in a network. The split-brain conditions between a plurality of clusters of network repository functions (NRFs) in a network are handled. The method comprises receiving a cluster disconnection indication from at least one cluster. Upon receiving the indication of cluster disconnection from at least one cluster, a split-brain handling process (e.g., split-brain logic) is triggered. The method furthercomprises, upon triggering the split-brain handling process, storing delta representations of data changes by the plurality of clusters of NRFs. The method comprises receiving a connection re-establishment indication from at least one NRF of the at least one cluster. Upon receiving the indication, pause the split-brain handling process and perform replication of the delta representations of data changes between the plurality of clusters.

[0067] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the accompanying drawings FIGs.1-4.

[0068] FIG. 1 illustrates an exemplary network architecture (100) for implementing a system (108) for managing split-brain conditions in a network (106), in accordance with an embodiment of the present disclosure.

[0069] As illustrated in FIG. 1 , one or more user equipments (UEs) (104- 1 , 104- 2... 104-N) may be connected to the network (106). A person of ordinary skill in the art will understand that the one or more UEs (104-1, 104-2... 104-N) may be collectively referred to as UEs (104) and individually referred to as a UE (104). One or more users (102-1, 102-2... 102-N) may provide one or more requests to the end server through the network (106). A person of ordinary skill in the art will understand that the one or more users (102-1, 102-2... 102-N) may be collectively referred to as users (102) and individually referred to as a user (102).

[0070] In an embodiment, the UE (104) may include, but not be limited to, a mobile, a laptop, etc. Further, the UE (104) may include one or more in-built or externally coupled accessories including, but not limited to, a visual aid device such as a camera, audio aid, microphone, or keyboard. Furthermore, the UE (104) may include a mobile phone, smartphone, virtual reality (VR) devices, augmented reality (AR) devices, a laptop, a general-purpose computer, a desktop, a personal digital assistant, a tablet computer, and a mainframe computer. Additionally, input devices for receivinginput from the user (102) such as a touchpad, touch-enabled screen, electronic pen, and the like may be used. A person of ordinary skill in the art will appreciate that the UE (104) may not be restricted to the mentioned devices and various other devices may be used.

[0071] Referring to FIG. 1, the UE (104) is configured to communicate with the network (106). In an embodiment, the network (106) may include at least one of a Third Generation (3G) Network, Fourth Generation (4G) network, Fifth Generation (5G) network, 6G network, WiFi Network, Fiber to the Home (FTTx) network, or the like which provides internet service to individual or to the subscriber home. The network (106) may enable the UE (104) to communicate with other devices in the network architecture (100). The network (106) may include a wireless card or some other transceiver connection to facilitate this communication. In another embodiment, the network (106) may be implemented as, or include any of a variety of different communication technologies such as a wide area network (WAN), a local area network (LAN), a wireless network, a mobile network, a Virtual Private Network (VPN), the Internet, the Public Switched Telephone Network (PSTN), or the like.

[0072] In an embodiment, the network (106) may include, by way of example but not limitation, at least a portion of one or more networks having one or more nodes that transmit, receive, forward, generate, buffer, store, route, switch, process, or a combination thereof, etc. one or more messages, packets, signals, waves, voltage or current levels, some combination thereof, or so forth. The network (106) may also include, by way of example but not limitation, one or more of a wireless network, a wired network, an internet, an intranet, a public network, a private network, a packet- switched network, a circuit-switched network, an ad hoc network, an infrastructure network, a Public Switched Telephone Network (PSTN), a cable network, a cellular network, a satellite network, a fiber optic network, or some combination thereof.

[0073] In an embodiment, the network architecture (100) further includes a system (108). The system (108) comprises a plurality of network functions (NFs) (110- 1, 110-2... .110-N) and a plurality of network repository function (NRF) clusters (112- 1, 112-2....112-N). A person of ordinary skill in the art will understand that the one or more NRF clusters (112-1, 112-2... 112-N) may be collectively referred to as NRF clusters (112) and individually referred to as NRF cluster (112). In an aspect, the NRF cluster is also referred to as a cluster. The plurality of NRF clusters (112) is also referred to as a plurality of clusters. A person of ordinary skill in the art will understand that the one or more NFs (110-1, 110-2... 110-N) may be collectively referred to as NFs (110) and individually referred to as NF (110).

[0074] In an embodiment, the plurality of NRF clusters (112) may be deployed such that data is continuously replicated between the plurality of NRF clusters (112). The data may include, but is not limited to, registration information (e.g., NF instance identifiers (ID), NF type, supported interfaces, service information, etc.), deregistration information (e.g., NF instance ID, reason for deregistration, etc.), service discovery requests and responses, health monitoring information (e.g., operational status, load information, etc.), configuration updates (e.g., configuration changes, updated service information, etc.), failure notification (e.g., failure reason, recovery actions, etc.).

[0075] According to an implementation, if one of the NRF clusters (e.g., NRF cluster 2 (112-2)) goes down (or disconnects) due to a network issue, the disconnected NRF cluster (e.g., NRF cluster 2 (112-2)) may send a call back indication to all other NRF clusters (e.g., NRF cluster 1 (112-1), NRF cluster 3 (112-3), to NRF cluster N (112-N)), to inform the clusters about the cluster failure. Upon receiving the call-back indication, a split-brain logic or handling process is triggered by the system (108). During the split-brain handling process, all the NRF clusters (112) store data changes in a delta till the split-brain handling process (or split-brain logic) is being performed or until the split-brain logic stops. Upon connection re-establishment, one of the NRFsof the disconnected NRF cluster (e.g., NRF cluster 2 (112-2)) sends a call back indication to all other NRF clusters (e.g., NRF cluster 1 (112-1), NRF cluster 3 (112- 3), to NRF cluster N (112-N)), to inform the clusters about the connection reestablishment. The split-brain handling process is halted / paused on receiving the indication of connection re-establishment. The replication of delta data is performed between the plurality of NRF clusters (112). The delta data replication is performed to ensure that the latest information is shared among all NRF clusters (112), restoring data consistency and integrity across the network function set (NFSet). The use of delta ensures that only the changes in data are replicated, rather than the entire data set. This optimization reduces the amount of data transferred and enhances the overall performance of the network functions.

[0076] Although FIG. 1 shows exemplary components of the network architecture (100), in other embodiments, the network architecture (100) may include fewer components, different components, differently arranged components, or additional functional components than depicted in FIG. 1. Additionally, or alternatively, one or more components of the network architecture (100) may perform functions described as being performed by one or more other components of the system (108).

[0077] FIG. 2A illustrates an exemplary system architecture (200A) of the system (108) for managing split- brain conditions in the network (106), in accordance with an embodiment of the present disclosure.

[0078] The system architecture (200A) includes the plurality of NFs (110-1, 110-2... 110-N) and the plurality of NRF clusters (112-1, 112-2....112-N).

[0079] In an aspect, the network repository function (NRF) is a network function used in 5G and network function virtualization (NFV) environments. The NRF is used in managing and coordinating network functions by providingmechanisms for service discovery, registration, configuration management, and communication. Further, the NRF functions as a centralized repository that maintains comprehensive records about various available network functions, including their capabilities, configurations, and current status.

[0080] In an aspect, the NRF cluster refers to a group of interconnected NRFs that work together as a single system to achieve high availability and load balancing. By clustering, these NRFs collaborate to handle tasks more efficiently and provide continuous service even if one or more NRFs fail. In an operative aspect, the NRF cluster (112) consists of a plurality of NRFs working together. Each NRF performs functions related to service discovery, registration, and management. Incoming requests are distributed across the plurality of NRFs in the NRF cluster (112) to ensure an even distribution of load and avoid overloading any NRF. Each NRF comprises data storage. The data storage stores the data corresponding to the functions. Data and state information are synchronized across the plurality of NRFs to ensure consistency. The health and performance of each NRF are monitored to ensure that any failures or issues are detected and addressed promptly.

[0081] In an aspect, the network functions (NFs) (110) are used to handle specific tasks within the network (106). The plurality of NFs (110) includes, but is not limited to, an access and mobility management function (AMF), a session management function (SMF), a user plane function (UPF), a network slice selection function (NSSF), a policy control function (PCF), an application function (AF) and a network exposure function (NEF).

[0082] The AMF manages user registration, mobility, and connection setup in the network. The SMF handles the establishment, modification, and release of user sessions and manages session-related data. The UPF manages user plane data traffic, including routing and forwarding data packets. The PCF defines, distributes, and enforces policies related to network behavior and service management. This includespolicies for Quality of Service (QoS), traffic management, and access control. The AF manages application-level services and interacts with other network functions. The AF provides specific application services, such as content delivery or application analytics. The NEF exposes network capabilities and services to external applications and third- party services.

[0083] In an embodiment, the plurality of NRF clusters (112) are deployed in the network (106). The data is replicated continuously between the plurality of NRF clusters (112). If one or more NRFs in the NRF cluster (112) or the entire NRF cluster (112) go down due to network issues or disconnection, the disconnected NRF cluster (e.g., NRF (112-2)) sends a call back indication to all other NRF clusters (e.g., NRF cluster 1 (112-1), NRF cluster 3 (112-3), to NRF cluster N (112-N)), to inform about the cluster failure due to disconnection. On receiving the disconnection indication from the disconnected NRF cluster (e.g., NRF cluster (112-2)), a split-brain logic or handling process is triggered by the system (108). During the split-brain handling process, all the NRF clusters (112) store data changes in the data storage of NRFs till the splitbrain handling process is being performed or until the split-brain handling process stops. After connection re-establishment, one of the NRFs of the previously disconnected NRF cluster (e.g., NRF cluster (112-2)) sends an indication of connection re-establishment to all other NRF clusters (e.g., NRF cluster 1 (112-1), NRF cluster 3 (112-3), to NRF cluster N (112-N)). On receiving the connection re-establishment indication, the split-brain handling process is paused. After the split-brain handling process ends, delta data replication occurs between all NRF clusters (112). The delta data replication is done to ensure that the latest information is shared among all NRF clusters (112). In this way, data replication after the split-brain handling process ensures that all NRF clusters (112) eventually converge to the same state, maintaining data consistency across the NFSet. Furthermore, data replication improves the fault tolerance of the NFSet. This redundancy allows the network (106) to continueoperating seamlessly, even if individual NRF or entire NRF clusters (112) experience disruptions.

[0084] In an aspect, when one of NRF clusters (112) loses connectivity and operates independently, then all the NRF clusters (112) store changes of data in delta. In an aspect, when the NRF clusters operate independently, the NRF clusters maintain a local copy of the service registry (e.g., a list of registered NRFs and NFs with their profiles). The NFs connected to the NRF cluster to register, deregister, and query the NRF as usual, but only within that cluster. Changes to NF registrations or subscriptions are updated locally and stored in the data storage (e.g., temporary delta storage module)

[0085] Once the connection is restored / re-established, the stored delta data is replicated between all the NRF clusters (112) to ensure that the latest data is synchronized among all the NRF clusters (112). This redundancy facilitates seamless data recovery in case of network issues. The risk of data mismatch is reduced by replicating data between the NRF clusters (112). The data consistency is ensured across all the NRFs (112) within the same NFSet. This enhances the reliability and resilience of the network (106), minimizing the impact of split-brain scenarios on network operations and service delivery. Storing data changes in the delta allows efficient data recovery. When connectivity is restored, the NRF clusters (112) quickly compare and update their states. This reduces downtime and minimizes service interruptions.

[0086] Although FIG. 2A shows exemplary components of the system architecture (200A), in other embodiments, the system architecture (200A) may include fewer components, different components, differently arranged components, or additional functional components than depicted in FIG. 2A. Additionally, or alternatively, one or more components of the system architecture (200A) may perform functions described as being performed by one or more other components of the system (108).

[0087] FIG. 2B illustrates an exemplary block diagram (200B) of the system (108) for managing split-brain conditions in the network (106), in accordance with an embodiment of the present disclosure.

[0088] Referring to FIG. 2B, in an embodiment, the system (108) may include one or more processor(s) (202). The one or more processor(s) (202) may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, logic circuitries, and / or any devices that process data based on operational instructions. Among other capabilities, the one or more processor(s) (202) may be configured to fetch and execute computer-readable instructions stored in a memory (204) of the system (108). The memory (204) may be configured to store one or more computer-readable instructions or routines in a non- transitory computer readable storage medium, which may be fetched and executed to create or share data packets over a network service. The memory (204) may comprise any non-transitory storage device including, for example, volatile memory such as random-access memory (RAM), or non-volatile memory such as erasable programmable read only memory (EPROM), flash memory, and the like.

[0089] In an embodiment, the system (108) may include an interface(s) (206). The interface(s) (206) may comprise a variety of interfaces, for example, interfaces for data input and output devices (VO), storage devices, and the like. The interface(s) (206) may facilitate communication through the system (108). The interface(s) (206) may also provide a communication pathway for one or more components of the system (108).

[0090] The system (108) may include a detection unit (208), a monitoring unit (210), a replication unit (212), a receiving unit (214) and a database (216).

[0091] In an aspect, the database (216) is configured to store program instructions. The database (216) is configured to store the data received from thedetection unit (208), the monitoring unit (210), the replication unit (212) and the receiving unit (214). The program instructions include a program that implements the system (108) for split-brain handling between the plurality of NRF clusters (112) in the network (106) in accordance with embodiments of the present disclosure and may implement other embodiments described in this specification. The database (216) may include any computer-readable medium known in the art including, for example, volatile memory, such as Static Random Access Memory (SRAM) and Dynamic Random Access Memory (DRAM), and / or nonvolatile memory, such as Read Only Memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes.

[0092] In an aspect, a plurality of clusters is deployed in the network (106). The plurality of clusters comprises, but is not limited to, the plurality of network repository function (NRF) clusters (112). The data is replicated continuously between the plurality of NRF clusters (112). In an aspect, the plurality of clusters may comprise, but is not limited to, an access and mobility management function (AMF) cluster, a session management function (SMF) cluster, a user plane function (UPF), a policy control function (PCF) cluster, a network slice selection function (NSSF) cluster, and an application cluster.

[0093] The detection unit (208) may detect at least one of NRFs or any NRF cluster (e.g., NRF cluster (112-2)) goes down. The at least one of NRFs or any NRF cluster (e.g., NRF cluster (112-2)) may go down due to at least one of the network issues, communication timeout or disconnection between at least two NRFs of the NRF cluster (e.g., NRF cluster (112-2)). When the at least one of NRFs or any cluster goes down, the split-brain conditions occur. In the split-brain conditions, each NRF cluster operates autonomously or independently of data synchronization (i.e., the disconnected NRF cluster (e.g., NRF cluster (112-2)) continues to function on its own data and does not replicate data between the plurality of NRF clusters (112)). In an operative aspect,the detection unit (208) may detect the cluster disconnection based on a first indication received from at least one NRF from the disconnected NRF cluster (e.g., NRF cluster (112-2)) in the network (106). The at least one NRF may generate the first indication of cluster disconnection based on a loss of connection or communication timeout between at least two NRFs. In an aspect, the first indication may include, but is not limited to, a call-back indication. The call-back indication may refer to a notification by one NRF to inform other NRF clusters (e.g., NRF cluster (112-1), NRF cluster (112- 3) and NRF cluster (112-N)) in the network (106) about the disconnection or loss of connection.

[0094] In another aspect, the at least one NRF of the disconnected NRF cluster (e.g., NRF cluster (112-2)) may generate the first indication based on one or more mechanisms. The one or more mechanisms comprise, but are not limited to, heartbeat signal monitoring, request timeout monitoring, error responses or negative acknowledgment monitoring, health check application programming interfaces (APIs), or integration with a network monitoring system.

[0095] In heartbeat monitoring, each NRF continuously sends heartbeat signals to indicate its operational state after a predefined time period (e.g., 30 sec). If the heartbeat signals are not received from the NRF after the predefined time period for a predefined number of times (e.g., 3 times), then the NRF is down. In request timeout monitoring, if any NF (110) tries to register with or query the NRF and gets no response or the request times out, repeated failures exceeding a certain threshold may indicate that the NRF or NRF cluster (e.g., NRF cluster (112-2)) is down. In an example, the NF (110) tries to discover a service from the NRF cluster. If three attempts fail due to a timeout, the detection unit (208) flags the cluster as non-operational. The error responses or negative acknowledgment monitoring refers to the process of tracking failed or rejected requests (e.g., service registrations, updates, queries, etc.) between the NFs (110) and the NRFs. When the NRF responds with error codes or explicitlysends negative acknowledgments (NACKs), this indicates problems (e.g., unavailability, conflicts, or overload. By monitoring the frequency and patterns of the error responses, the detection unit (208) may quickly detect NRF issues (i.e., NRF is down or non-operational). In an aspect, the health check APIs are endpoints provided by the NRF that allow other components or monitoring systems to regularly verify their operational status. The APIs respond with the current health information of the services (i.e., whether the NRF is running correctly, overloaded, or facing internal errors). By periodically calling the health check APIs, the detection unit (208) may quickly detect failures or unavailability of the NRF. In an aspect, integration with the network monitoring system for detecting NRF failures / disconnection involves continuously collecting and analyzing health and performance data from the NRF using methods (e.g., periodic polling of health check APIs, monitoring heartbeat signals, and gathering logs or telemetry data). The monitoring system processes information to identify anomalies (e.g., missed heartbeats, error responses, or degraded performance indicators). When thresholds are exceeded, the detection unit (208) detects the NRF disconnection / failure. This integration ensures the timely detection of NRF outages or malfunctions, enabling proactive maintenance and minimizing disruption in 5G core network operations.

[0096] The at least one NRF may send the first indication of cluster disconnection. The receiving unit (214) may receive the first indication from the at least one NRF of the disconnected NRF cluster (e.g., NRF cluster (112-2)).

[0097] Upon receiving the first indication of cluster disconnection from at least one NRF of the disconnected NRF cluster (e.g., NRF cluster (112-2)), the detection unit (208) may trigger a split-brain handling process (also referred to as a split-brain logic). In an aspect, the split-brain handling process refers to a mechanism to manage and resolve the split-brain conditions upon detecting the at least one NRF cluster goes down or disconnects.

[0098] In an aspect, the triggering of the split-brain handling process comprises initiating the split-brain handling process, changing the state of each cluster from a clustered state to an isolated state, and then initiating storing of data changes in the data storage in each cluster.

[0099] In an operative aspect, triggering of the split-brain handling process causes each cluster (112) to store delta representations of data changes in a data storage during the split-brain conditions. Upon triggering of the split-brain handling process, delta representations of data changes during the split-brain conditions are stored in the data storage of each NRF cluster of the plurality of NRF clusters (112). In an aspect, the delta representations of data changes refer to differences (i.e., deltas) in data during the split- brain conditions (i.e., when the NRF is unavailable or disconnected). In an aspect, the delta representations of data changes comprise incremental changes of data rather than the full data. The incremental changes comprise, but are not limited to, field addition, field updation, field deletion, modified structures, and incremental counters. Each delta representation is versioned by associating it with a version identifier to maintain order and detect conflicts. The versioning of the delta representations is performed based on at least one of timestamp, version numbers, vector clocks, or universally unique identifier (UUID)-based delta representations. In the timestampbased versioning, the timestamp of each delta representation is maintained. In the version numbers-based versioning, the version numbers for each delta representation of data changes are incremented upon detecting the data changes. In the vector clocks, maintaining per-cluster version vectors to track split-brain conditions and concurrent data changes. Further, the delta representations of data changes are maintained in logs.[000100] In an aspect, the delta representations may be structured as, but are not limited to, field-level differences. In the field-level differences, each delta representation includes only field changes / updates. Further, the delta representation may be structured as full object-level and operation logs. In the full object-level, each1 cluster is configured to store the delta representation of the entire data field. When any update happens, the entire data field is rewritten and versioned as a new data change. In the operation logs, each cluster logs operations (e.g., increment counter, set field X, append item). The operations logs include semantic-level changes rather than the entire dataset.[000101] In the split-brain handling process, the delta representations (i.e., changes in the data) are stored in the data storage of each NRF cluster during the splitbrain conditions. In an operative aspect, the data storage used to store data changes during the split-brain conditions is configured as a temporary or intermediate delta storage module. In an aspect, the delta storage module refers to a storage module that stores data changes, corresponding to sessions, configurations, and state information corresponding to the network functions, over time, rather than full data. For example, the delta representations of data changes correspond to the network function (e.g., session management function (SMF)), such as the user starts a data session. The delta stored: {"event": "session_start", "session_id": "abcl23", "timestamp": "12:00"}. After 10 minutes, the session is handed over to another network slice, delta stored: {"event": "session_update", "session_id": "abcl23", "new_slice": "slice-2", "timestamp": "12: 10"}. After some time, the user disconnects, the delta stored: {"event": "session_end", "session_id": "abcl23", "timestamp": "12:30"}.[000102] In an aspect, the temporary or intermediate delta storage module refers to a storage module that locally stores the delta representations of data changes during the split-brain conditions. Configuring the data storage for the temporary or intermediate delta storage comprises configuring separate data storage in the database (e.g., database (216)) to store delta representations of data changes during the splitbrain conditions. A start indication corresponding to the split-brain conditions (i.e., disconnection of the cluster in the network) is set to inform the data storage (i.e., temporary or intermediate delta storage module) upon detecting the split-brainconditions in the network. Upon receiving the indication corresponding to the splitbrain conditions in the network, the data storage (i.e., temporary or intermediate delta storage module) initiates storing of delta representations of data changes in the data storage (i.e., temporary or intermediate delta storage module). Further, upon resumption of the connection of the cluster in the network, an end indication is received by the data storage from the cluster to stop the storage of the delta representations of data changes in the data storage.[000103] In an aspect, the temporary or intermediate delta storage module buffers the delta representations of the data changes that are not immediately sent to the other NRF clusters (e.g., NRF cluster 1 (112-1), NRF cluster 3 (112-3), to NRF cluster N (112-N)) due to disconnection or failure. The delta representations are stored temporarily until the NRF re-establishes the connection or becomes reachable again, at which point the temporary or intermediate module forwards the changes for synchronization. The storing of delta representations of data changes allows data recovery.[000104] The monitoring unit (210) is configured to monitor a connection status of the at least one disconnected cluster (e.g., NRF cluster (112-2)) to detect reestablishment of the connection. In an aspect, the connection status may comprise a connected state or a disconnected state. The connected state refers to an operational state of the NRF cluster (112), i.e., NRFs actively perform communication or data synchronization by replicating data between the NRF clusters (112). The disconnected state refers to a non-operational state of the NRF cluster (112), i.e., NRFs operate autonomously or independently (i.e., function on their own) without data synchronization. In an aspect, the monitoring unit (210) is configured to continuously monitor the connection status of the disconnected NRF cluster (e.g., NRF cluster (112- 2)) to detect whether the connection of the disconnected NRF cluster is re-established or resumption.[000105] In an aspect, the re-establishment or resumption of connection is detected through at least one of network-level communication or control signaling. The network-level communication or control signaling refers to a process of exchanging messages between the clusters in the network to inform about the resumption of the connection. The network-level communication or control signaling comprises, but is not limited to, health check API monitoring, a network- level connectivity check, heartbeat signal monitoring, event or log-based monitoring, and a retry mechanism.[000106] In the health check API monitoring, the monitoring unit (210) sends periodic Hypertext Transfer Protocol Secure (HTTPS) GET requests to a health check endpoint exposed by the NRF cluster (e.g., health, status). A successful response (e.g., HTTP 200 OK) from the NRF indicates that the NRF is healthy and reachable. In the network-level connectivity checks (also referred to as Ping or Port Probing), the monitoring unit (210) uses low-level network tools (e.g., Internet Control Message Protocol (ICMP) ping or Transmission Control Protocol (TCP) port checks) to verify if the NRF is network-reachable. This detects whether the network path to the NRF cluster (112) has been restored, even if the NRF is not fully operational. In the heartbeat signal monitoring, the NRF may exchange heartbeat messages at regular intervals. The monitoring unit (210) tracks whether heartbeat messages are received by checking whether the heartbeat messages arrive on time and at expected time intervals. Further, the validity of the heartbeat messages is checked by verifying formats, timestamps, cryptographic signatures or checksums. If heartbeats resume after the disconnection, it signals that the NRF cluster is back online. In the event / log-based monitoring, the monitoring unit (210) subscribes to logs or events generated by the NRF cluster (112). Reconnection events, service recovery logs, or successful registration logs from NFs (110) can be used as signals to detect that the NRF is functioning again. In the retry mechanism in NF requests, the monitoring unit (210) may retry failed NF registration or discovery requests at intervals. A successful retry after repeated failures is interpreted as evidence that the NRF cluster (112) has reconnected.[000107] In this way, the monitoring unit (210) may use at least one of health checks, network probes, heartbeats, telemetry, event logs, or retry logic to continuously assess whether the disconnected NRF cluster (e.g., NRF cluster (112-2)) has reestablished the connection. Once reconnection is detected, follow-up actions (e.g., delta representations of data changes synchronization) may be triggered.[000108] After connection re-establishment, the receiving unit (214) may receive a second indication of connection re-establishment from at least one NRF of the at least one NRF cluster (e.g., NRF cluster (112-2). The second indication of connection reestablishment is generated based on the resumption of connection between the at least two NRFs. In an aspect, the second indication may be, but is not limited to, a heartbeat signal indicating the operational state of the NRF cluster (112).[000109] Upon receiving the second indication, the monitoring unit (210) may terminate or pause the split- brain handling process (i.e., split-brain logic). Termination of the split-brain handling process comprises the detection of resumption of the connection of at least one disconnected cluster in the network, and pausing storage of the delta representations of data changes in the data storage.[000110] Upon pausing the split-brain handling process, the replication unit (212) initiates the triggering of recovery action. The triggering of recovery action comprises, but is not limited to, replication of the delta representations of data changes among the plurality of NRF clusters (112) to synchronize the data changes and restore consistency across the plurality of NRF clusters (112). In an aspect, the replication of the delta representations comprises steps such as identifying by each cluster the locally stored delta representations during the split-brain conditions and initiating the process of exchanging the delta representations in the network.[000111] Replication of delta representations refers to a process of replicating only the delta representations of data changes across the plurality of NRF clusters (112)rather than the entire data set. Upon receiving the second indication, each cluster is configured to replicate the delta representations from other clusters. The use of delta representations ensures that only the data changes are replicated, rather than the entire data set. This optimization reduces the amount of data transferred and enhances the overall performance of the NRF clusters (112).[000112] The replication of the delta representations of data changes synchronizes the data changes and restores the data consistency across the plurality of clusters (112). In an aspect, the synchronization of data changes and restoration of data consistency is performed based on a conflict resolution policy. The conflict resolution policy is applied based on at least one of timestamp, last-writer wins, quorum-based merge or use of version vector. In an aspect, in the timestamp-based conflict resolution policy, changes are resolved by comparing timestamps assigned to each delta representation of data changes. The data changes / updates with the latest timestamp are considered the correct version. In an aspect, the last-writer wins-based conflict resolution policy, data changes from the writer with the most recent update are kept. Clock synchronization is used and overwrites earlier changes. In the quorum-based merge conflict resolution policy, the majority of clusters need to agree on the data changes (i.e., data changes of the majority of clusters need to be the same). In the version vector-based conflict resolution policy, each data change is tracked with a vector representing history across the clusters. If any data conflicts are detected, the conflicts are resolved based on order and origin of data changes.[000113] In an operative aspect, the data changes across the plurality of NRF clusters (112) are synchronized based on a timestamp, i.e., when the NRFs store the delta representations of data changes during the split-brain conditions, the delta representations of data changes are tagged with the timestamp that records the exact time of the delta representations. In an aspect, timestamp” refers to time-related information (i.e., time and date) indicating when the delta representations of datachanges during the split-brain conditions are captured and stored in the data storage. In an example, during split-brain condition, on NRF 2 (112-2) at 12:05 PM, a user updates a customer record = Change: {"customer name": "Alice Johnson"} and Timestamp: 2023-08-24T 12:05:00Z. On NRF 1 (112-1), at 12:06 PM, the user updates the same customer record differently = Change: {"customer name": "Alice J."} and Timestamp: 2023-08-24T 12:06:00Z. Each NRF stores the data in its temporary data storage with respective timestamps.[000114] When each NRF cluster (112) replicates the delta representations from the other clusters, each NRF cluster (112) may compare the timestamps of the delta representations of any conflicting or missing data changes. Based on the comparison, data changes of the most recent change (i.e., the latest timestamp) take priority or precedence. By replicating the latest changes with timestamps, any potential conflicts in data (e.g., data mismatch or discrepancies) may be resolved proactively, ensuring that the most up-to-date information is used across all clusters (112). This leads to more reliable network operations. In an example, after resumption of the connection, NRF 1 (112-1) has data changes at the timestamp 12:06, and NRF 2 (112-2) has data changes at the timestamp 12:05. So, the data changes from the NRF 1 (112-1) are kept because the NRF 1 (112-1) has the later timestamp.[000115] The consistency is restored by initiating replication of delta representations among the plurality of clusters, involves detecting data changes that occurred independently in each cluster during the split-brain conditions, and then exchanging only those changes (delta representations) between the clusters once connectivity is restored. Each cluster sends and receives delta representations (e.g., updates such as modified sessions, configurations, or state data) rather than full datasets, enabling efficient synchronization. These delta representations of data changes are then replicated by using the mechanism (e.g., conflict resolution strategies such as timestamp comparison or vector version) to ensure that all clusters reconciledifferences and reach the consistent, unified state. This process ensures eventual consistency across all clusters with minimal data transfer and faster convergence.[000116] In an aspect, the system (108) may perform authentication or verification of the integrity of replication (or exchange) of the delta representations of data changes during the synchronization. The authentication or verification is performed to ensure that the delta representations of data changes are authenticated and not tampered with during the replication (or exchange). The system (108) performs authentication or verification using one or more authentication mechanisms. The one or more authentication mechanisms comprise, but are not limited to, a transport layer security (TLS) encryption, message authentication codes (MAC), and hash checksums. In TLS encryption, a secure communication channel is established between the clusters using public / private key cryptography, and the clusters are authenticated using digital certificates. In the MAC mechanism, a cryptographic MAC is computed over the delta representations of data changes by the sending cluster using a shared secret key. The computed MAC is then attached to the delta data. Upon receiving the data, the receiving cluster recalculates the MAC and verifies it against the received MAC to ensure data integrity and authenticity. In the hash checksums, each delta representation includes a cryptographic hash during transmission of the delta representations. The receiving cluster recomputes the hash to verify that the content has not been altered.[000117] The replication of delta representations ensures that all NRF clusters (112) eventually converge to the same state (i.e., all the clusters have synchronized data), maintaining consistent data across all NRF clusters (112). This simplifies the management of clustered NRFs and reduces the complexity involved in data consistency and manages the network operation during connectivity issues (i.e., NRF failure or disconnection in the split-brain conditions) without a significant impact on service quality.[000118] Although FIG. 2B shows exemplary components of the system (108), in other embodiments, the system (108) may include fewer components, different components, differently arranged components, or additional functional components than depicted in FIG. 2B. Additionally, or alternatively, one or more components of the system (108) may perform functions described as being performed by one or more other components of the system (108).[000119] FIG. 3 illustrates an exemplary flow diagram of a method (300) for managing split-brain conditions in the network (106), in accordance with an embodiment of the present disclosure.[000120] At step (302), the method (300) includes deploying the plurality of NRF clusters (112) in the network (106). In an aspect, deploying the plurality of NRF clusters (112) includes determining the network's requirements (e.g., availability, redundancy, connectivity, integration requirements, security, etc.), configuring settings for the NRF clusters (112) (e.g., network parameters, database configuration, assigning identifiers, registrations, discovery, etc.), and integration (e.g., ensuring that the NRF clusters (112) may interact with other network functions). Further, deployment of the NRF clusters (112) involves planning, installation, configuration, integration, testing, and ongoing management.[000121] At step (304), the method (300) includes checking whether any of the NRF clusters (112) is down. On detecting that an NRF cluster (e.g., NRF cluster-2 (112-2)) goes down, a call-back is received by all other NRF clusters (e.g., NRF cluster (112-1) and NRF cluster (112-N)) from the NRF cluster (e.g., NRF cluster-2 (112-2)). In an aspect, the NRF cluster (e.g., NRF cluster-2 (112-2)) may go down due to various reasons (e.g., hardware failure, software issues, network problems, resource exhaustion, storage problems, power loss, security issues, cluster management issues, etc.).[000122] At step (306), the method (300) includes, if the callback is not received, no action is performed. In an aspect, if the call-back indicating failure or disconnection of the NRF cluster (e.g., NRF cluster-2 (112-2)) is not received, it indicates that the NRF clusters (112) are functioning normally.[000123] At step (308), the method (300) includes, upon receiving the call-back (i.e., indication of the cluster failure / disconnection) from the disconnected NRF cluster (e.g., NRF cluster-2 (112-2)), a split-brain logic is initiated or triggered by the system (108). In an aspect, the split-brain scenario can occur when the connection between clusters (112) is disrupted, causing the disconnected cluster (e.g., NRF cluster-2 (112- 2)) to operate autonomously without synchronizing data with the other clusters (112). So, the split-brain logic is triggered to inform about the disconnection of one of the NRF clusters (e.g., NRF cluster-2 (112-2)).[000124] At step (310), the method (300) includes, upon triggering of the splitbrain logic, all the NRF clusters (112) store data changes in data storage as delta representations. The data changes are stored as the delta representations till the splitbrain logic stops. Also, it is checked whether the re-establishment of the disconnected cluster (e.g., NRF cluster-2 (112-2)) has occurred. In an aspect, the stored data may correspond to the NFs (110). The data includes, but is not limited to, NF IDs, NF type, supported interfaces, service information, service query, service response, location, health monitoring data, operational status, configuration changes, failure reason, recovery action, etc.[000125] At step (312), the method (300) includes, upon re-establishment of connection from one NRF within the disconnected NRF cluster (e.g., NRF cluster-2 (112-2)), a call back of re-establishment from the previously disconnected NRF cluster (e.g., NRF cluster-2 (112-2)) is received by all other NRF clusters (e.g., NRF cluster 1 (112-1), NRF cluster 3 (112-3), to NRF cluster N (112-N)). In an aspect, when the disconnected NRF cluster (e.g., NRF cluster-2 (112-2)) re-establishes the connection,the NRF cluster (e.g., NRF cluster-2 (112-2)) sends an indication to inform about the establishment of the connection to all other NRF clusters (e.g., NRF cluster 1 (112-1), NRF cluster 3 (112-3), to NRF cluster N (112-N)).[000126] At step (314), the method (300) includes, upon receiving a call back of re-establishment from the NRF cluster (e.g., NRF cluster-2 (112-2)), the split-brain logic is halted or paused. In an aspect, after the re-establishment of connection, the split-brain logic is paused, terminated, or halted.[000127] At step (316), the method (300) includes, after the split-brain phase ends, delta data replication occurs between all the NRF clusters (112). The delta data replication causes the replication of the latest data with all the NRF clusters (112). The delta data replication ensures that the latest updates are synchronized across all the NRF clusters (112). Storing changes in data in the delta allows efficient data recovery. When connectivity is restored, the NRF clusters (112) can quickly compare and update their states, reducing downtime and minimizing service interruptions. Also, by ensuring that data changes are tracked and synchronized, the risk of data mismatch and discrepancies is reduced, which leads to more reliable network operations.[000128] FIG. 4 illustrates another flow diagram of a method (400) for managing split-brain conditions in the network (106), in accordance with an embodiment of the present disclosure.[000129] The plurality of clusters (e.g., NRF clusters) (112) is deployed in the network (106). The data is replicated continuously between the plurality of NRF clusters (112).[000130] At step (402), the method (400) includes detecting a disconnection of at least one cluster (e.g., NRF cluster (112-2)) within the network (106). The disconnection of the at least one cluster may be due to at least one of the network issues, the communication timeout or the disconnection between the two NRFs. Thedisconnection of the at least one NRF cluster (e.g., NRF cluster (112-2)) causes the split-brain conditions in the network. In the split-brain conditions, each cluster operates independently of data synchronization.[000131] The cluster disconnection is detected based on a first indication received from at least one NRF of the disconnected NRF cluster (e.g., NRF cluster (112-2)) in the network (106). The first indication of cluster disconnection is generated based on a loss of connection or communication timeout between at least two NRFs. In an aspect, the first indication may be a call-back indication to inform other NRF clusters (e.g., NRF cluster 1 (112-1), NRF cluster 3 (112-3), to NRF cluster N (112-N)) about the disconnection of the at least one NRF cluster (e.g., NRF cluster (112-2)) in the network (106).[000132] At step (404), the method (400) includes triggering a split-brain handling process in response to the detected disconnection. Triggering of the split-brain handling process causes each cluster (112) to store delta representations of data changes in the data storage during the split-brain conditions. The data storage is configured as a temporary or intermediate delta storage module that temporarily stores the delta representations of data changes during the split-brain conditions until the disconnected NRF cluster (e.g., NRF cluster (112-2)) re-establishes the connection.[000133] At step (406), the method (400) includes monitoring the connection status of the at least one disconnected cluster (e.g., NRF cluster (112-2)) to detect reestablishment of the connection. The resumption of connection is detected through at least one of network-level communication or control signaling that comprises, but is not limited to, health check API monitoring, a network- level connectivity check, heartbeat signal monitoring, event or log-based monitoring, and a retry mechanism. In an aspect, the connection reestablishment is detected based on a second indication from the at least one NRF. The second indication of connection reestablishment is generated based on resumption of connection between the at least two NRFs. In an aspect, thesecond indication may be, but is not limited to, the heartbeat signal or call-back indication to indicate re-establishment or resumption of the connection.[000134] At step (408), the method (400) includes terminating the split-brain handling process upon detection of the reestablishment of the connection. In an aspect, upon receiving the second indication (e.g., heartbeat signal or call-back indication) from the at least one NRF of the NRF cluster (e.g., NRF cluster (112-2)), which reestablishes the connection, the split-brain handling process is terminated or paused.[000135] At step (410), the method (400) includes initiating replication of the delta representations among the plurality of clusters (112). Upon terminating the splitbrain handling process, the replication of the delta representations is initiated to synchronize the data changes and restore consistency across the plurality of clusters (112). The data changes across the plurality of NRF clusters (112) are synchronized based on a timestamp (i.e., based on the time and date of data changes stored in the temporary data storage). Data changes with the latest timestamp are replicated. This helps in reducing data mismatch or discrepancies to ensure the most up-to-date information is replicated across all clusters (112). Further, only the delta representations of data changes are replicated, rather than the entire data set. This reduces the amount of data transferred and enhances the overall performance of the NRF clusters (112).[000136] FIG. 5 illustrates an exemplary computer system (500) in which or with which embodiments of the present disclosure may be implemented.[000137] As shown in FIG. 5, the computer system (500) may include an external storage device (510), a bus (520), a main memory (530), a read-only memory (540), a mass storage device (550), communication port(s) (560), and a processor (570). A person skilled in the art will appreciate that the computer system may include more than one processor and communication ports. The processor (570) may include variousmodules associated with embodiments of the present disclosure. The communication port(s) (560) may be any of an RS-232 port for use with a modem-based dialup connection, a 10 / 100 Ethernet port, a Gigabit or 10 Gigabit port using copper or fiber, a serial port, a parallel port, or other existing or future ports. The communication port(s) (560) may be chosen depending on a network, such a Local Area Network (LAN), Wide Area Network (WAN), or any network to which the computer system connects.[000138] The main memory (530) may be random access memory (RAM), or any other dynamic storage device commonly known in the art. The read-only memory (540) may be any static storage device(s) e.g., but not limited to, a Programmable Read Only Memory (PROM) chips for storing static information e.g., start-up or Basic Input / Output System (BIOS) instructions for the processor (570). The mass storage device (550) may be any current or future mass storage solution, which can be used to store information and / or instructions. Exemplary mass storage device (550) includes, but is not limited to, Parallel Advanced Technology Attachment (PATA) or Serial Advanced Technology Attachment (SATA) hard disk drives or solid-state drives (internal or external, e.g., having Universal Serial Bus (USB) and / or Lirewire interfaces), one or more optical discs, Redundant Array of Independent Disks (RAID) storage, e.g., an array of disks.[000139] The bus (520) communicatively couples the processor (570) with the other memory, storage, and communication blocks. The bus (520) may be, e.g., a Peripheral Component Interconnect (PCI) / PCI Extended (PCI-X) bus, Small Computer System Interface (SCSI), Universal Serial Bus (USB), or the like, for connecting expansion cards, drives, and other subsystems as well as other buses, such a front side bus (FSB), which connects the processor (570) to the computer system.[000140] Optionally, operator and administrative interfaces, e.g., a display, keyboard, joystick, and a cursor control device, may also be coupled to the bus (520) to support direct operator interaction with the computer system. Other operator andadministrative interfaces can be provided through network connections connected through the communication port(s) (560). Components described above are meant only to exemplify various possibilities. In no way should the aforementioned exemplary computer system limit the scope of the present disclosure.[000141] The exemplary computer system (500) is configured to execute a computer program product comprising a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform a method for managing split-brain conditions in a network is disclosed. The method comprises detecting a disconnection of at least one cluster within the network. The method comprises triggering a split-brain handling process in response to the detected disconnection. Triggering of the split-brain handling process causes each cluster to store delta representations of data changes in a data storage during the split-brain conditions. The method comprises monitoring a connection status of the at least one disconnected cluster to detect re-establishment of the connection. The method comprises terminating the split-brain handling process upon detection of the reestablishment of the connection. The method comprises initiating replication of the delta representations among a plurality of clusters to synchronize the data changes and restore consistency across the plurality of clusters.[000142] The present disclosure provides technical advancements related to splitbrain handling in the network. The advancement addresses the limitations of existing solutions by replicating delta representations of data changes during the split-brain conditions in the network. Upon detecting the disconnection in the cluster, storing changes in the data (i.e., delta representations of data changes) in temporary data storage. Replicating the delta representation among the clusters upon resumption of the connection in the clusters. Network performance is optimized as only delta representations of data to changes are replicated, rather than the entire data set. This optimization reduces the amount of data transferred and enhances the overallperformance of the network functions. Data recovery is improved by storing changes in data (i.e., delta representations of data changes). When connectivity is restored, the NRF clusters quickly compare and update their states. This reduces downtime and minimizes service interruptions. Further, the risk of data mismatch is reduced by ensuring that the data changes are tracked and synchronized. This mitigates the risk of data mismatches and discrepancies.[000143] While the foregoing describes various embodiments of the invention, other and further embodiments of the invention may be devised without departing from the basic scope thereof. The scope of the invention is determined by the claims that follow. The invention is not limited to the described embodiments, versions or examples, which are included to enable a person having ordinary skill in the art to make and use the invention when combined with information and knowledge available to the person having ordinary skill in the art.ADVANTAGES OF THE PRESENT INVENTION[000144] The present disclosure described herein above has several technical advantages including,• Enhancing data consistency by implementing data replication after a split-brain scenario. This ensures that all NRF clusters eventually converge to the same state, maintaining consistent data across the NFset.• Improving fault tolerance by implementing data replication. This enhances the fault tolerance of the NFset. This allows the network to continue operating seamlessly, even if individual NRFs or entire NRF cluster experiences disruptions.Improving data recovery by storing changes in data (i.e., delta data). When connectivity is restored, the NRF clusters can quickly compare and update their states. This reduces downtime and minimizes service interruptions.• Optimizing performance by using the delta data to ensure that only the changes are replicated, rather than the entire data set. This optimization reduces the amount of data transferred and enhances the overall performance of the network functions.• Simplifying the management of the clustered NRFs by handling split-brain scenarios. This reduces the complexity involved in ensuring data consistency and managing network operations during connectivity issues.• Reducing the risk of data mismatch by ensuring that the data changes are tracked and synchronized. This mitigates the risk of data mismatches and discrepancies, leading to more reliable network operations.• Increasing network resilience by recovering from split-brain scenarios and maintaining data integrity enhances overall network robustness and ensures that disruptions are managed with minimal impact on service quality.• Providing proactive conflict resolution by replicating the latest changes with timestamps. This resolves any potential conflicts in data proactively and ensures that the most up-to-date information is used across all NRF clusters.

Claims

We claim:

1. A method (400) for managing split-brain conditions in a network (106), the method (400) comprising: detecting (402) a disconnection of at least one cluster within the network (106); triggering (404) a split-brain handling process in response to the detected disconnection, wherein triggering of the split-brain handling process causes each cluster (112) to store delta representations of data changes in a data storage during the split-brain conditions; monitoring (406) a connection status of the at least one disconnected cluster to detect re-establishment of the connection; terminating (408) the split-brain handling process upon detection of the re-establishment of the connection; and initiating (410) replication of the delta representations among a plurality of clusters (112) to synchronize the data changes and restore consistency across the plurality of clusters (112).

2. The method (400) as claimed in claim 1 , wherein the plurality of clusters (112) comprises a plurality of network repository function (NRF) clusters.

3. The method (400) as claimed in claim 1, wherein the split-brain conditions comprise at least one of the disconnection of the at least one cluster (112) in the network (106) and each cluster (112) operating independently of data synchronization.

4. The method (400) as claimed in claim 1, wherein the cluster disconnection is detected based on a first indication received by other clusters from at least one NRF of the at least one disconnected cluster in the network (106), and whereinthe connection re-establishment is detected based on a second indication received by the other clusters from the at least one NRF.

5. The method (400) as claimed in claim 4, wherein the first indication of cluster disconnection is generated based on a loss of connection or communication timeout between at least two NRFs, wherein the second indication of connection reestablishment is generated based on resumption of connection between the at least two NRFs, and wherein the resumption of connection is detected through at least one of network-level communication or control signaling.

6. The method (400) as claimed in claim 1 , wherein the data storage used to store data changes during the split-brain conditions is configured as a temporary or intermediate delta storage module.

7. The method (400) as claimed in claim 1, wherein the data changes across the plurality of clusters (112) are synchronized based on a timestamp.

8. The method (400) as claimed in claim 1, comprising: wherein the replication of the delta representations comprising: replicating, by each cluster (112), the delta representations from the other clusters (112).

9. A system (108) for managing split-brain conditions in a network (106), the system (108) comprising: a detection unit (208) configured to: detect a disconnection of at least one cluster within the network (106); andtrigger a split-brain handling process in response to the detected disconnection, wherein triggering of the split-brain handling process causes each cluster (112) to store delta representation of data changes in a data storage during the split-brain conditions; a monitoring unit (210) configured to: monitor a connection status of the at least one disconnected cluster to detect re-establishment of the connection; and terminate the split-brain handling process upon detection of the reestablishment of the connection; and a replication unit (212) configured to: initiate replication of delta representations among plurality of clusters (112) to synchronize the data changes and restore consistency across the plurality of clusters (112).

10. The system (108) as claimed in claim 9, wherein the plurality of clusters (112) comprises a plurality of network repository function (NRF) clusters.

11. The system (108) as claimed in claim 9, wherein the split-brain conditions comprise at least one of the disconnection of the at least one cluster (112) in the network (106) and each cluster (112) operating independently of data synchronization.

12. The system (108) as claimed in claim 9, wherein a receiving unit is configured to receive a first indication and a second indication from the at least one NRF, wherein the cluster disconnection is detected based on the first indication received by other clusters from the at least one NRF of the at least one disconnected cluster in the network (106), and wherein the connection re-establishment is detected based on the second indication received by the other clusters from the at least one NRF.

13. The system (108) as claimed in claim 12, wherein the first indication of cluster disconnection is generated based on a loss of connection or communication timeout between at least two NRFs, wherein the second indication of connection reestablishment is generated based on resumption of connection between the at least two NRFs, and wherein the resumption of connection is detected through at least one of network-level communication or control signaling.

14. The system (108) as claimed in claim 9, wherein the data storage used to store data changes during the split-brain conditions is configured as a temporary or intermediate delta storage module.

15. The system (108) as claimed in claim 9, wherein the data changes across the plurality of clusters (112) are synchronized based on a timestamp.

16. The system (108) as claimed in claim 9, wherein the replication of the delta representations comprising: each cluster (112) is configured to replicate the delta representations from other clusters (112).

17. A computer program product comprising a non-transitory computer- readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to execute a method (400) for managing split-brain conditions in a network (106), the method (400) comprising:detecting (402) a disconnection of at least one cluster within the network (106); triggering (404) a split-brain handling process in response to the detected disconnection, wherein triggering of the split-brain handling process causes each cluster to store delta representations of data changes in a data storage during the split-brain conditions; monitoring (406) a connection status of the at least one disconnected cluster to detect re-establishment of the connection; terminating (408) the split-brain handling process upon detection of the reestablishment of the connection; and initiating (410) replication of the delta representations among a plurality of clusters (112) to synchronize the data changes and restore consistency across the plurality of clusters (112).

Citation Information

Patent Citations

  • HA split brain over network

    US9537739B1