System and method for managing malfunctioning of network nodes in a communication network

WO2026202938A1PCT designated stage Publication Date: 2026-10-01JIO PLATFORMS LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/IN2026/050509
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-22
Filing Date
2026-03-20
Publication Date
2026-10-01

Smart Images

  • Figure IN2026050509_01102026_PF_FP_ABST
    Figure IN2026050509_01102026_PF_FP_ABST
Patent Text Reader

Abstract

SYSTEM AND METHOD FOR MANAGING MALFUNCTIONING OF NETWORK NODES IN A COMMUNICATION NETWORK Disclosed is a method and a system for managing malfunctioning of network nodes in a communication network. For managing the malfunction of the network nodes, data 5 of one or more nodes may be fetched from an Element Management System (EMS). Further, presence of alarm information may be determined in the data of the one or more nodes. When the alarm information is present, one or more parameters related to configuration of restart may be obtained. The one or more parameters are configured by an operator of the system. Furthermore, the restart of the node may be executed 10 based on the one or more parameters.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR MANAGING MALFUNCTIONING OF NETWORK NODES IN A COMMUNICATION NETWORK TECHNICAL FIELD

[0001] The embodiments of the present disclosure generally relate to the field of wireless communication network. More particularly, the present disclosure relates to a system and method for managing malfunctioning of network nodes in a communication network.BACKGROUND OF THE INVENTION

[0002] The subject matter disclosed in the background section should not be assumed or construed to be prior art merely because of its mention in the background section. Similarly, any problem statement mentioned in the background section or its association with the subject matter of the background section should not be assumed or construed to have been previously recognized in the prior art.

[0003] In a 3 GPP Long Term Evolution (LTE) network, service outages and degradation can be difficult to detect and will require considerable manual effort for troubleshooting. These service outages or degradations are difficult to detect because of manual intervention. Hence, the failed or degraded cell site can remain in a failed or degraded state without being noticed for a period of time.

[0004] In conventional approach, network automation aimed to minimize the repetitive manual tasks of network operator engineers. Generally, a domain technician may find faulty Remote Radio Heads (RRHs) by visiting the site where the RRHs are installed. This type of service outage or degradation may go undetected for hours or even days. Furthermore, troubleshooting can require manual analysis and unplanned site visits that will increase network maintenance costs for the service provider. In a vast mobile network having multitude of nodes, it is humanly difficult and time consuming to attend to all the service impacting failures leading to large restoration time.

[0005] To restore the services quickly and improve the customer experience, there is a need to automate the clearance of certain failures by continuously monitoring the functioning of the nodes. In light of the aforementioned challenges, there is a need for a solution that can address the issue of managing malfunctioning of network nodes in a communication network.SUMMARY

[0006] The following embodiments present a simplified summary in order to provide a basic understanding of some aspects of the disclosed invention. This summary is not an extensive overview, and it is not intended to identify key / critical elements or to delineate the scope thereof. Its sole purpose is to present some concepts in a simplified form as a prelude to the more detailed description that is presented later.

[0007] In an embodiment, a method for managing malfunctioning of network nodes in a communication network is disclosed. The method includes fetching, by a data collection module, alarm data in real-time from one or more network nodes. The method further includes determining, by a determination module based on the alarm data, one or more alarms that are defined to identify one or more faulty nodes among the one or more network nodes. Further, the method includes determining, by the determination module based on the alarm data, whether at least one alarm among the one or more alarms associated with at least one faulty node among the one or more faulty nodes remains active for a predefined duration. Furthermore, the method includes obtaining, by an input module, one or more parameters related to configuration of restart of the at least one faulty node. Thereafter, the method includes executing, by an execution module, restarting at least one faulty node based on the one or more parameters.

[0008] In some aspects of the present disclosure, the one or more parameters comprise a restart interval to restart alarm not cleared by auto-restart, a number of retries to restartalarm not cleared by auto-restart, or a schedule time of execution. The one or more parameters are configurable by a network operator.

[0009] In some aspects of the present disclosure, the method further includes updating, by the execution module, status of the at least one faulty node based on a result of the execution of the restart of the at least one faulty node.

[0010] In some aspects of the present disclosure, the method further includes notifying, by a notification module, information of the at least one faulty node and a result of the execution of the restart of the at least one faulty node to the network operator via a user dashboard or a message.

[0011] In some aspects of the present disclosure, the method further includes determining, by the determination module, a number of alarm events based on the determined at least one alarm. Further, the method includes tagging, by the determination module, the number of alarm events as in-progress alarms.

[0012] In some aspects of the present disclosure, the method further includes executing, by the execution module, the restart of the in-progress alarms based on a schedule time of execution defined in the one or more parameters. Further, the method includes updating, by the execution module, the tagging of the in-progress alarms to cleared-by-auto-restart alarms upon successful execution of the restart of the inprogress alarms. Furthermore, the method includes updating, by the execution module, the tagging of the in-progress alarms to not-cleared-by-auto-restart alarms upon unsuccessful execution of the restart of the in-progress alarms.

[0013] In some aspects of the present disclosure, the method further includes retrying executing, by the execution module, the restart of the not-cleared-by-auto-restart alarms based on a restart interval and a number of retries to restart alarm defined in the one or more parameters. Further, the method includes updating, by the execution module, the tagging of the not-cleared-by-auto-restart alarms to cleared-by-auto-restart alarms upon successful execution of the restart of the not-cleared-by-auto-restartalarms. Furthermore, the method includes updating, by the execution module, the tagging of the not-cleared-by-auto-restart alarms to hardware-to-be-replaced upon unsuccessful execution of the restart of the not-cleared-by-auto-restart alarms.

[0014] In some aspects of the present disclosure, the method further includes determining, by the determination module, whether a serial number of the at least one faulty node is changed when the at least one alarm is cleared automatically. Further, the method includes updating, by the execution module, the tagging of the hardware-to-be-replaced to cleared-by- hardware-replacement upon determination that the serial number of the at least one faulty node is changed. Furthermore, the method includes updating, by the execution module, the tagging of the hardware-to-be-replaced to autocleared alarm upon determination that the serial number of the at least one faulty node is not changed.

[0015] In another embodiment, a system for managing malfunctioning of network nodes in a communication network is disclosed. The system includes a data collection module configured to fetch alarm data in real-time from one or more network nodes. The system further includes a determination module configured to determine, based on the alarm data, one or more alarms that are defined to identify one or more faulty nodes among the one or more network nodes. The determination module is further configured to determine, based on the alarm data, whether at least one alarm among the one or more alarms associated with at least one faulty node among the one or more faulty nodes remains active for a predefined duration. Further, the system includes an input module configured to obtain one or more parameters related to configuration of restart of the at least one faulty node. Furthermore, the system includes an execution module configured to execute restart of the at least one faulty node based on the one or more parameters.BRIEF DESCRIPTION OF DRAWINGS

[0016] Various embodiments disclosed herein will become better understood from the following detailed description when read with the accompanying drawings. The accompanying drawings constitute a part of the present disclosure and illustrate certain non-limiting embodiments of inventive concepts disclosed herein. Further, components and elements shown in the drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the present disclosure. For the purpose of consistency and ease of understanding, similar components and elements are annotated by reference numerals in the exemplary drawings.

[0017] FIG. 1 illustrates a diagram depicting an exemplary communication network, in accordance with an embodiment of the present disclosure.

[0018] FIG.2 is a block diagram of a system for managing malfunctioning of network nodes in the communication network, in accordance with an embodiment of the present disclosure.

[0019] FIG. 3 illustrates a call flow between one or more components of the system, in accordance with an embodiment of the present disclosure.

[0020] FIG. 4 illustrates a call flow for displaying real-time status of the nodes, in accordance with an embodiment of the present disclosure.

[0021] FIG. 5 illustrates a dashboard for selecting one or more attributes to identify faulty nodes for auto-restart mechanism, in accordance with an embodiment of the present disclosure.

[0022] FIGs. 6A-6D illustrate a dashboard for rendering live status of the network nodes, in accordance with an embodiment of the present disclosure.

[0023] FIGs.7A-7G illustrate a call flow for managing malfunctioning of the network nodes in the communication network, in accordance with an embodiment of the present disclosure.

[0024] FIG. 8 illustrates a flow diagram of a method for managing malfunctioning of the network nodes in the communication network, in accordance with an embodiment of the present disclosure.DETAILED DESCRIPTION OF THE INVENTION

[0025] Inventive concepts of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, in which examples of one or more embodiments of inventive concepts are shown. Inventive concepts may, however, be embodied in different forms and should not be construed as limited to the embodiments set forth herein. Further, the one or more embodiments disclosed herein are provided to describe the inventive concept thoroughly and completely, and to fully convey the scope of each of the present inventive concepts to those skilled in the art. Furthermore, it should be noted that the embodiments disclosed herein are not mutually exclusive concepts. Accordingly, one or more components from one embodiment may be tacitly assumed to be present or used in any other embodiment.

[0026] The following description presents various embodiments of the present disclosure. The embodiments disclosed herein are presented as teaching examples and are not to be construed as limiting the scope of the present disclosure. The present disclosure should in no way be limited to the illustrative implementations, drawings, and techniques illustrated below, including the exemplary design and implementation illustrated and described herein, but may be modified, omitted, or expanded upon without departing from the scope of the present disclosure.

[0027] The following description contains specific information pertaining to embodiments in the present disclosure. The detailed description uses the phrases “in some embodiments” which may each refer to one or more or all of the same or different embodiments. The term “some” as used herein is defined as “one, or more than one, or all.” Accordingly, the terms “one,” “more than one,” “more than one, but not all” or “all” would all fall under the definition of “some.” In view of the same, the terms, forexample, “in an embodiment” refers to one embodiment and the term, for example, “in one or more embodiments” refers to “at least one embodiment, or more than one embodiment, or all embodiments.”

[0028] The term “comprising,” when utilized, means “including, but not necessarily limited to;” it specifically indicates open-ended inclusion in the so-described one or more listed features, elements in a combination, unless otherwise stated with limiting language. Furthermore, to the extent that the terms “includes,” “has,” “have,” “contains,” and other similar words are used in either the detailed description, such terms are intended to be inclusive in a manner similar to the term “comprising.”

[0029] In the following description, for the purposes of explanation, various specific details are set forth to provide a thorough understanding of embodiments of the present disclosure. It will be apparent, however, that embodiments of the present disclosure may be practiced without these specific details. Several features described hereafter can each be used independently of one another or with any combination of other features.

[0030] The description provided herein discloses exemplary embodiments only and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the foregoing description of the exemplary embodiments will provide those skilled in the art with an enabling description for implementing any of the exemplary embodiments. Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it may be understood by one of the ordinary skilled in the art that the embodiments disclosed herein may be practiced without these specific details.

[0031] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein the description, the singular forms "a", "an", and "the" include plural forms unless the context of the invention indicates otherwise.

[0032] The terminology and structure employed herein are for describing, teaching, and illuminating some embodiments and their specific features and elements and do not limit, restrict, or reduce the scope of the present disclosure. Accordingly, unless otherwise defined, all terms, and especially any technical and / or scientific terms, used herein may be taken to have the same meaning as commonly understood by one having ordinary skill in the art.

[0033] Embodiments of the present disclosure will be described below in detail with reference to the accompanying drawings. FIG. 1 through FIG. 8, discussed below, and the one or more embodiments used to describe the principles of the present disclosure are by way of illustration only and should not be construed in any way to limit the scope of the present disclosure. Those skilled in the art will understand that the principles of the present disclosure may be implemented in any suitably arranged system or device.

[0034] FIG. 1 illustrates a diagram depicting an exemplary communication network 100, in accordance with an embodiment of the present disclosure. The communication network 100 may provide a wireless communication service by using wireless communication technologies, such as Long-Term Evolution (LTE), LTE-advanced (LTE-A), and Fifth Generation (5G). Wireless communication technologies to be used for the communication service are not limited to the exemplified technologies.

[0035] The communication network 100 may include a data processing device 102, an Element Management System (EMS) 104, a network 106, and multiple Remote Radio Heads (RRHs) 108-1 through 108-n (hereinafter may also be referred to as ‘RRHs 108’ or “nodes 108” or “one or more network nodes 108” or “network nodes 108”). The data processing device 102, the EMS 104, and the RRHs 108 are communicably connected with each other by way of the network 106.

[0036] The EMS 104 is connected to multiple RRHs 108 and manages the multiple RRHs 108 connected to multiple base stations. The EMS 104 may store an alarm log.An alarm may be generated by the RRHs 108 when operation parameter values of the RRHs 108 are beyond an allowable range. The alarm may indicate that there are likely to be faults in the RRHs 108.

[0037] Each base station may be referred to as a node B or an eNB. One or more RRHs 108 may be connected to one base station. All the base stations are not required to be necessarily connected to the RRHs 108. Each base station may provide a communication service to a User Equipment (UE) directly or through the RRHs 108 connected thereto. The base stations may transfer, to the EMS 104, information regarding operations of the RRHs 108 connected to the base stations, which information may include alarms for the RRHs 108 and operation parameters of the RRHs 108. Each base station may store, as a hardware log, information regarding operation parameters of one or more RRHs 108 connected thereto.

[0038] The RRHs 108 may be connected to the EMS 104 in a wireless or wired manner. The RRHs 108 may issue an alarm when the operation parameters of the RRHs 108 are beyond an allowable range. The RRHs 108 may transfer the operation parameters of the RRHs 108 to the EMS 104 connected to the RRHs 108. The EMS 104 may receive an alarm when values of the operation parameters of the RRHs 108 are beyond an allowable range.

[0039] The data processing device 102 may be configured as an individual entity and may detect or predict a fault in each RRH 108 on the basis of information received from the EMS 104. In one embodiment, the data processing device 102 may be installed separately from the EMS 104. In another embodiment, the data processing device 102 may be included in the EMS 104. The data processing device 102 may include processor(s) (comprising data processing engines) configured with suitable logic, instructions, circuitry, interfaces, and / or codes for executing operations of various operations performed by the data processing device 102 for computations and data processing related to detection and management of malfunctioning network nodes.

[0040] The network 106 may include suitable logic, circuitry, and interfaces that may be configured to provide several network ports and several communication channels for transmission and reception of data related to operations of various entities of the communication network 100. Each network port may correspond to a virtual address (or a physical machine address) for transmission and reception of the communication data. For example, the virtual address may be an Internet Protocol Version 4 (IPV4) (or an IPV6 address) and the physical address may be a Media Access Control (MAC) address. The network 106 may be associated with an application layer for implementation of communication protocols based on communication requests from the various entities of the communication network 100. The communication data may be transmitted or received via the communication protocols. Examples of the communication protocols may include, but are not limited to, Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Simple Mail Transfer Protocol (SMTP), Domain Network System (DNS) protocol, Common Management Interface Protocol (CMIP), Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Long Term Evolution (LTE) communication protocols, or any combination thereof. In some aspects of the present disclosure, the communication data may be transmitted or received via at least one communication channel of several communication channels in the network 106. The communication channels may include, but are not limited to, a wireless channel, a wired channel, a combination of wireless and wired channel thereof. The wireless or wired channel may be associated with a data standard which may be defined by one of a Local Area Network (LAN), a Personal Area Network (PAN), a Wireless Local Area Network (WLAN), a Wireless Sensor Network (WSN), Wireless Area Network (WAN), Wireless Wide Area Network (WWAN), a metropolitan area network (MAN), a satellite network, the Internet, an optical fiber network, a coaxial cable network, an infrared (IR) network, a radio frequency (RF) network, and a combination thereof. Aspects of the present disclosureare intended to include or otherwise cover any type of communication channel, including known, related art, and / or later developed technologies.

[0041] In one or more embodiments, the data processing device 102 may identify the failures of the mobile network nodes, such as the RRHs 108, taking autonomous action to resolve the issue. If it is not resolved, alert an operation team about the root cause of the failure & initiate end to end closed loop action. Alarms may be triggered by the data processing device 102 whenever the malfunctioning of the network nodes 108 happen. In such situations, many alarms can be rectified remotely without manual interventions. In this disclosure, the service impacting alarms can be cleared through automation by auto restarting the impacted Radio entity. Automation of restarting any radio nodes or part of the nodes improves the network uptime and the customer experience. Further, automation of the restarting of the impacted radio nodes may be achieved by continuous monitoring of alarm messages and automatically taking appropriate remedial action to resolve issues in the impacted radio nodes. The data processing device 102 will also provide the transaction logs and a detailed report on the clearance of faults in addition to a dashboard on a User Interface (UI).

[0042] FIG.2 is a block diagram of a system 200 for managing malfunctioning of the network nodes 108 in the communication network 100, in accordance with an embodiment of the present disclosure. The embodiment of the system 200 as shown in FIG. 2 is for illustration only. However, the system 200 may come in a wide variety of configurations, and FIG. 2 does not limit the scope of the present disclosure to any particular implementation of the system 200.

[0043] As shown in FIG. 2, the system 200 includes the data processing device 102 which includes an Input-Output (I / O) interface 202, one or more processors 204 (hereinafter may also be referred to as “processor 204”), a memory 206, a network communication manager 208, the console host 210, a database 212, and one or moreprocessing modules 214. Components of the data processing device 102 are coupled to each other via a communication bus 226.

[0044] The I / O interface 202 may include suitable logic, circuitry, interfaces, and / or codes that may be configured to receive input(s) and present (or display) output(s) on the data processing device 102. For example, the I / O interface 202 may have an input interface and an output interface. The input interface may be configured to enable a network operator user to provide input(s) to trigger (or configure) the data processing device 102 to perform various operations. Examples of the input interface may include, but are not limited to, a User Interface (UI), a touch interface, a touch screen, a mouse, a keyboard, a motion recognition unit, a gesture recognition unit, a voice recognition unit, or the like. The output interface is configured to control the data processing device 102 to display a dashboard to the network operator user. Examples of the output interface of the I / O interface 202 may include, but are not limited to, the UI, a digital display, an analog display, a touch screen display, an appearance of a desktop, and / or illuminated characters.

[0045] The processor 204 may include various processing circuitry and communicates with the memory 206, the network communication manager 208, the console host 210, and the database 212 via the communication bus 226. The processor 204 is configured to execute instructions 206A (hereinafter also referred to as “a set of instructions 206A”) stored in the memory 206 and to perform various processes. The processor 204 may include one or a plurality of processors, including a general-purpose processor, such as, for example, and without limitation, a central processing unit (CPU), an application processor (AP), a dedicated processor, a graphics-only processing unit such as a graphics processing unit (GPU) or the like, a programmable logic device, or any combination thereof.

[0046] The memory 206 stores the set of instructions 206A required by the processor 204 of the data processing device 102 for controlling its overall operations. Thememory 206 may include non-volatile storage elements. Examples of such non-volatile storage elements may include magnetic hard discs, optical discs, floppy discs, flash memories, or forms of electrically programmable memories (EPROM) or electrically erasable and programmable (EEPROM) memories. In addition, the memory 206 may, in some examples, be considered a non-transitory storage medium. The "non-transitory" storage medium is not embodied in a carrier wave or a propagated signal. However, the term "non-transitory" should not be interpreted as the memory 206 is non-movable. In some examples, the memory 206 may be configured to store larger amounts of information. In certain examples, a non-transitory storage medium may store data that can, over time, change (e.g., in Random Access Memory (RAM) or cache). The memory 206 may be an internal storage unit or an external storage unit of the data processing device 102, cloud storage, or any other type of external storage. In certain examples, the memory 206 configured as the non-transitory storage medium may include hard drives, solid-state drives, flash drives, Compact Disk (CD), Digital Video Disk (DVD), and the like. Further, the memory 206 may include any type of non-transitory storage medium, without deviating from the scope of the present disclosure.

[0047] More specifically, the memory 206 may store computer-readable instructions 206A including instructions that, when executed by a processor (e.g., the processor 204) cause the data processing device 102 to perform various functions described herein. In some cases, the memory 206 may contain, among other things, a BIOS which may control basic hardware or software operation such as the interaction with peripheral components or devices.

[0048] The network communication manager 208 may manage communications with a core network 106 (e.g., via one or more wired backhaul links). For example, the network communications manager 208 may manage the transfer of data communications for the data processing device 102 and the EMS 104. The networkcommunication manager 208 may include an electronic circuit specific to a standard that enables wired or wireless communication. The network communication manager 208 is configured for communicating with external devices via one or more networks.

[0049] The console host 210 may include suitable logic, instructions, and / or codes for executing various operations of one or more computer executable applications to host the console on an external user device, by way of which the network operator user can trigger the data processing device 102 to receive input from the EMS 104. In some other aspects of the present disclosure, the console host 210 may provide a Graphical User Interface (GUI) for the data processing device 102 for user interaction.

[0050] The data processing device 102 may be communicatively coupled with the database 212. The database 212 may store alarm data received from the EMS 104 or the base stations. The alarm data includes information of alarm or faults associated with the network nodes 108 and operation parameters of the network nodes 108. Aspects of the present disclosure are intended to include and / or otherwise cover any type of data associated with the data processing device 102, without deviating from the scope of the present disclosure.

[0051] The processing module(s) 214 may be implemented as a combination of hardware and programming (for example, programmable instructions) to implement one or more functionalities of the data processing device 102. In non-limiting examples, described herein, such combinations of hardware and programming may be implemented in several different ways. For example, the programming for the processing modules(s) 214 may be processor-executable instructions stored on a non-transitory machine-readable storage medium and the hardware for the processor 204 may comprise a processing resource (for example, one or more processors), to execute such instructions. In the present examples, the machine-readable storage medium may store instructions that, when executed by the processing resource, implement the processing module(s) 214. In such examples, the data processing device 102 may alsocomprise the machine-readable storage medium storing the instructions and the processing resource to execute the instructions, or the machine-readable storage medium may be separate but accessible to the data processing device 102 and the processing resource. In other examples, the processing module(s) 214 may be implemented using an electronic circuitry.

[0052] In one or more embodiments, the processing module(s) 214 may include a data collection module 216, a determination module 218, an input module 220, an execution module 222, and a notification module 224.

[0053] The EMS 104 continuously monitors the operation parameter values of the network nodes 108 and when the operation parameter values are beyond an allowable range, an alarm may be generated by the EMS 104. The alarm indicates that there are likely to be faults in the network nodes 108.

[0054] The data collection module 216 fetches the alarm data in real-time from the network nodes 108 via the EMS 104. The alarm data includes information of alarms set / issued by EMS 104 or the network nodes 108 when the operation parameters of the network nodes 108 are beyond the allowable range. Further, the determination module 218 determines one or more alarms included in the alarm data that are defined to identify faulty nodes among the network nodes 108. In a non-limiting example, the data collection module 216 may control display of the alarm data on a dashboard. Further, a user may filter the alarm data based on one or more attributes selectable by the user.

[0055] Further, the determination module 218 determines whether at least one alarm among the one or more alarms associated with at least one faulty node among the faulty nodes remains active for a predefined duration. For instance, the determination module 218 may determine or identify a number of alarm events (a number of the at least one alarm) that are active for the predefined duration and tags that number of alarm eventsas “in-progress” alarms. In a non-limiting example, the predefined duration may be a threshold duration or may be configurable by a network operator.

[0056] Further, the input module 220 obtains one or more parameters (may also be referred to as “one or more attributes”) related to configuration of restart of the each of the at least one faulty node for which the alarm remains active for the predefined duration. In a non-limiting example, the one or more parameters include a schedule time of execution, a restart interval to restart alarm not cleared by auto-restart, and a number of retries to restart alarm not cleared by auto-restart. For instance, the one or more parameters are also configurable by the network operator.

[0057] Further, the execution module 222 executes restart of the at least one faulty node based on the one or more parameters. For example, the execution module 222 executes the restart of the “in-progress” alarms based on the schedule time of execution defined in the one or more parameters. Further, the execution module 222 updates the tagging of the alarms based on a result of the execution of the restart. For instance, the execution module 222 may update the tagging of the “in-progress” alarms to “cleared-by-auto-restart” alarms upon successful execution of the restart of the “in-progress” alarms. Also, the execution module 222 may update the tagging of the “in-progress” alarms to “not-cleared-by-auto-restart” alarms upon unsuccessful execution of the restart of the in-progress alarms.

[0058] Thereafter, upon updating the tagging, the execution module 222 retries the restart of the “not-cleared-by-auto-restart” alarms based on the restart interval and the number of retries to restart alarm not cleared by auto-restart defined in the one or more parameters. Further, the execution module 222 again updates the tagging of the alarms based on the result of the re-execution of the restart. For instance, the execution module 222 may update the tagging of the “not-cleared-by-auto-restart” alarms to “cleared-by-auto-restart” alarms upon successful execution of the restart of the “not-cleared-by-auto-restart” alarms. Further, the execution module 222 may update the tagging of theY1“not-cleared-by-auto-restart” alarms to “hardware-to-be-replaced” upon unsuccessful execution of the restart of the not-cleared-by-auto-restart alarms.

[0059] Further, the determination module 218 determines whether a serial number of the at least one faulty node is changed when the at least one alarm is cleared automatically. The execution module 222 may update the tagging of the “hardware-to-be-replaced” to “cleared-by- hardware-replacement” upon determination that the serial number of the at least one faulty node is changed. Further, the execution module 222 may update the tagging of the “hardware-to-be-replaced” to “auto-cleared” alarm upon determination that the serial number of the at least one faulty node is not changed.

[0060] The execution module 222 performs retrying the restart of the alarms in an order as defined in the current disclosure based on the configured one or more parameters. The one or more parameters further include next day execution for alarm not cleared by auto-restart, a restart interval to restart execution failed, a number of retries to restart execution failed, next day execution for execution failed. The priority for the execution of the alarm is set as explained in FIGS. 7A-7G.

[0061] Further, the determination module 218 determines whether the restart of the alarms is completed by the execution module 222 based on each of one or more parameters. Further, the execution module 222 may update the tagging of the “not-cleared-by-auto-restart” alarms to “execution failed” upon unsuccessful execution of the restart.

[0062] Further, notification module 224 notifies information of the at least one faulty node and a result of the execution of the restart of the at least one faulty node to the network operator via a message or the dashboard.

[0063] In one or more embodiments, the execution module 222 may priorities the restarting of the impacted node based on a priority logic given by the network operator / admin. The execution module 222 may priorities the restarting of the impacted node based on severity level defined by the network operator / admin in advance.

[0064] FIG. 3 illustrates a call flow 300 between one or more components of the system 200, in accordance with an embodiment of the present disclosure. The system 200 may be implemented using a microservice architecture. In such implementation, the microservice architecture includes a microservice 302 which may be executed by the processor 204 and may utilize the memory 206, the network communication manager 208, and the database 212. The microservice 302 communicates with other layers using one or more Application Programing Interfaces (APIs). The microservice architecture further includes a read engine layer 304 (may also be referred to as “read engine 304”), a change engine layer 306 (may also be referred to as “change engine 306”), and the EMS 104. The read engine 304 reads the status of nodes from the EMS 104 and change engine 306 provides a response to the EMS 104 to restart the network nodes 108. The read engine 304 may correspond to the data collection module 216, the determination module 218, and the input module 220. The change engine 306 may correspond to the execution module 222.

[0065] At block 308, the microservice 302 may read files and may call PRE read request Application Programming Interface (API). The files may include the one or more parameters or attributes corresponding to restart.

[0066] At block 310, the read engine 204 may accept and process the PRE read request. Further, the read engine 204 may provide a response in call back API.

[0067] At block 312, the EMS 104 may read the status of the network nodes 108 in response to the call back API. Further, the EMS 104 may provide the response to the read engine 204. At block 314, the microservice 302 may accept and process the call back. Furthermore, the microservice 302 may call the change engine 206.

[0068] At block 316, the change engine 206 may provide the request to the EMS 104. Further, the change engine 206 may provide response in call back API. At block 318, the EMS 104 may restart the node and provide the response to the change engine 206.

[0069] At block 320, the microservice 302 may process and store the response. Further, the microservice 302 may call post read request API. At block 322, the read engine 204 may accept and process the request. Further, the read engine 204 may provide response in call back API.

[0070] At block 324, the EMS 104 may read the status of the node and may provide the response to the read engine 204. At block 326, the microservice 302 may accept and process the call back.

[0071] FIG. 4 illustrates a call flow 400 for displaying real-time status of the nodes 108, in accordance with an embodiment of the present disclosure. The real-time status may be displayed through a User Interface (UI) 402 associated with the I / O interface 202 of the data processing device 102. The UI 402 may be communicatively coupled with the database 212.

[0072] At block 404, the UI 402 may provide a query to the microservice 302 to render the data associated with at least one node among the nodes 108. The request may include identifiers of the at least one node for which the status is required.

[0073] At block 406, the microservice 302 may receive the query from the UI 402 and may forward the query to the database 212. At block 408, the database 212 may provide the response on the basis of the query.

[0074] At block 410, the microservice 302 may fetch the response from the database 212. Further, the microservice 302 may transfer the response to the UI 402. At block 412, the UI 402 may receive the response from the microservice 302. Further, the UI 402 may include the dashboard for rendering the response.

[0075] FIG. 5 illustrates a dashboard 500 for selecting the one or more attributes to identify faulty nodes for auto-restart mechanism, in accordance with an embodiment of the present disclosure.

[0076] The dashboard 500 may comprise selection panes for selecting the one or more attributes for identifying faulty nodes. The one or more attributes may comprise technology, vendor, alarm ID, alarm name, geography, hardware type, restart type, steady active alarm duration, schedule time of execution, retry interval (or restart interval) to restart alarm not cleared by auto-restart, a number of retries to restart alarm not cleared by auto-restart, next day execution for alarm not cleared by auto-restart, retry interval (or restart interval) to restart execution failed, number of retries to restart execution failed, next day execution for execution failed, and band / carrier.

[0077] In an implementation, the one or more attributes may have at least one option for selection. For example, technology (4G or 5G), vendor (name of company), alarm ID (identifier of the alarm), alarm name (unique name of the alarm), hardware type (RRH), restart type (RESTART COLD & HWRESETREQ), steady active alarm duration (minimum time for which the alarm should be active. (UI should show the minimum time as per the active alarm fetch time)), restart interval to restart alarm not cleared by auto-restart (restart interval time to auto-restart the impacted RRH), No. of retries to restart alarm not cleared by auto-restart (restart retries to auto-restart the impacted RRH), next day execution for alarm not cleared by auto-restart (Yes / No (Yes-Means node will restart next day, No- Means node will not restart next day)), restart interval to restart execution failed (restart interval time to auto-restart the impacted RRH), No. of retries to restart execution failed (restart retries to auto-restart the impacted RRH), next day execution for execution failed (Yes / No (Yes- Means node will restart next day, No- Means node will not restart next day)), schedule time (instant or hour selection for example (01:00 hrs to 03:00 hrs)), and band (carrier: All / 700_l / 3500_l / 3500_2).

[0078] FIGs.6A-6D illustrate a dashboard 600 for rendering live status of the network nodes 108, in accordance with an embodiment of the present disclosure. The dashboard 600 may be an auto-restart dashboard that is a near real-time live dashboard whichgives a single pane of glass view for the restart action being performed on the faulty nodes basis the alarms defined in admin panel.

[0079] The dashboard 600 displays following tags for the faulty nodes. The tag “Total” shows all the status count, i.e., “in-progress”, “cleared-by-auto-restart”, “not-cleared-by-auto-restart”, “hardware-to-be-replaced”, “cleared-by- hardware-replacement”, “auto-cleared”, “execution failed”, and “request failure” status.

[0080] The tag “in-progress” means all the alarm events which are in process of restart. The tag “auto-cleared” means the alarm is auto cleared before auto-restart execution. In another option, the alarm is auto cleared after the event has been into any of the below stages and automatically the alarm got cleared without restart. The tag “cleared-by-auto-restart” means the alarm is cleared for by restart mechanism. The tag “not-cleared-by-auto-restart” means the alarm is not cleared by auto-restart mechanism.

[0081] The tag “hardware-to-be-replaced” means the alarm is not cleared for the faulty RRH after multiple restarts defined as per admin settings and thereafter the event will be moved to hardware-to-be-replaced. The tag “cleared-by- hardware-replacement” means the alarm is cleared by Hardware Replacement by checking the serial number of the impacted RRH (serial number of the RRH before the alarm is auto cleared and serial number of the RRH after the alarm is auto cleared).

[0082] The tag “execution failed” means the execution module 222 is not able to restart the impacted RRH. The tag “response failure” means system 200 has not been able to read any attribute related to restart.

[0083] Further, FIGs 6A-6D illustrate site details including multiple attributes, such as technology, vendor, geography, maintenance, business, site ID, cell ID, execution status, execution attempts, execution timings, alarm ID, alarm name, alarming object, severity, alarm event time (or event time), EMS hostname, hardware type, hardware status, and hardware serial number.

[0084] FIGS. 7A-7G illustrate a call flow 700 for managing malfunctioning of the network nodes 108 in the communication network 100, in accordance with an embodiment of the present disclosure.

[0085] At block 702, the data processing device 102 may determine whether autorestart function is enabled or not. The auto-restart may be enabled or disabled by an admin (network administrator). If the auto-restart function is disabled, the flow of the method 700 may proceed to block 704. At block 704, the process may terminate.

[0086] If auto-restart function is enabled, the flow of the method 700 may proceed to block 706. At block 706, the data processing device 102 may fetch the details from the dashboard 600. The details comprise settings for geography for which the nodes 108 to be monitored.

[0087] At block 708, the data processing device 102 may determine whether a Service Access Point (SAP) ID (or device ID of the corresponding node) is present in an exclusion list uploaded by the admin. If the device ID is present in the exclusion list, the flow of the method 700 may proceed to block 704 and may terminate the process. If the device ID is not present in the exclusion list, the flow of the method 700 may proceed to block 710.

[0088] At block 710, the data processing device 102 may determine whether the alarming object for the node corresponding to the device ID is pending for execution. For example, if the status is “in-progress”, “execution failed”, and “alarm-not-cleared-by-auto-restart”. If the alarming object is pending, the method 700 may proceed to block 704 and may terminate the process. If the alarming object is not pending, the method 700 may proceed to block 712.

[0089] At block 712, the data processing device 102 may determine whether the node is restarted for a first pre-defined number of times in a day. If the node is restarted for the first pre-defined number of times, the method 700 may proceed to block 704 andmay terminate the process. If the node is not restarted for the pre-defined number of times, the method 700 may proceed to block 714.

[0090] At block 714, the data processing device 102 may determine whether Heartbeat failure alarm is active for the node. If the Heartbeat failure alarm is active for the node, the method 700 may proceed to block 704 and may terminate the process. If the Heartbeat failure alarm is not active for the node, the method 700 may proceed to block 716.

[0091] At block 716, the data processing device 102 may determine whether any alarm is active for alarming object - event time combination. If any alarm is not active for the alarming object - event time combination, the method 700 may proceed to block 704 and may terminate the process. If any alarm is active for the alarming object - event time combination, the method 700 may proceed to block 718.

[0092] At block 718, the data processing device 102 may wait for a second pre-defined time period. At block 720, the data processing device 102 may determine whether alarm is active for a third pre-defined time period. If the alarm is not active for the third predefined time period, the method 700 may proceed to block 704 and may terminate the process. If the alarm is active for the third pre-defined time period, the method 700 may proceed to block 722.

[0093] At block 722, the data processing device 102 may determine whether the admin has scheduled the execution for scheduled time or instant time. If the admin has scheduled the execution for scheduled time, the method 700 may proceed to block 724 and wait as per scheduled time. For example, the data processing device 102 may wait for 1 - 4 hours as defined in admin control to restart the faulty hardware. If the admin has scheduled the execution for instant time, the method 700 may proceed to block 726.

[0094] At block 726, the data processing device 102 may capture details of each impacted entity (faulty node). The details may include at least one of EMS hostname, status of the node, product name, and serial number.

[0095] At block 728, the data processing device 102 may wait for a fourth pre-defined time period. At block 730, the data processing device 102 may determine whether a wait time of the fourth pre-defined time period is completed or not. If the wait time of the fourth pre-defined time period is completed, the method 700 may proceed to block 732.

[0096] At block 732, the data processing device 102 may check a history table for clear events. At block 734, the data processing device 102 may determine whether the alarm events are active for any of the alarming object - event time combination before restart the execution. If the alarm events are active, the method 700 may proceed to block 736 and may combine multiple alarm events for the alarming object - event time combination in single file for each vendor EMS hostname separately.

[0097] If the wait time of the fourth pre-defined time period is not completed, the method 700 may proceed to block 738. At block 738, the data processing device 102 may determine whether there is any other alarm on same or another node and another alarming object - event time combination as per admin control. If there is any other alarm on same or another node and another alarming object - event time combination as per admin control, the method 700 may proceed to block 730. If there is no any other alarm on same or another node and another alarming object - event time combination as per admin control, the method 700 may proceed to block 740.

[0098] At block 740, the data processing device 102 may combine N alarming object - event time combinations from multiple alarm events in single file for each vendor EMS hostname separately and push to the execution module 222. At block 742, the data processing device 102 may call the execution module and may restart the impacted node.

[0099] At block 744, the data processing device 102 may determine whether the alarming object - event time combination is restarted as per the file shared to the execution module. If the alarming object - event time combination is restarted as per the file shared to the execution module, the method 700 may proceed to block 746. If the alarming object - event time combination is not restarted as per the file shared to the execution module, the method 700 may proceed to block 748.

[0100] At block 748, the data processing device 102 may change the status from “inprogress” to “execution failed” against each Alarm ID for the alarming object - event time combination. At block 750, the data processing device 102 may wait for fifth threshold time period. At block 752, the data processing device 102 may retry the restart for defined times.

[0101] At block 754, the data processing device 102 may determine whether the restart for defined times is completed. If the restart for the defined times is completed, the method 700 may proceed to block 704 and may terminate the process. If the restart for the defined times is not completed, the method 700 may proceed to block 756.

[0102] At block 756, the data processing device 102 may determine whether the alarm is cleared for the alarming object - event time combination. If the alarm is not cleared, the method 700 may proceed to block 758. At block 758, the data processing device 102 may determine whether the restart is executed in next day based on the admin control. If the restart is executed in next day, the method 700 may proceed to block 702. If the restart is not executed in next day, the method 700 may proceed to block 756. If the alarm is cleared, the method 700 may proceed to block 760.

[0103] At block 746, the data processing device 102 may determine whether the alarm is cleared for the alarming object - event time combination. If the alarm is cleared, the method 700 may proceed to block 762. If the alarm is not cleared, the method 700 may proceed to block 767.

[0104] At block 762, the data processing device 102 may determine whether a pre serial number is same as a new serial number. For instance, pre serial number is a serial number of impacted node before the alarm is cleared. If the pre serial number is same as the new serial number, the method 700 may proceed to block 764 and may change the status of the alarm event for the alarming object - event time combination from “inprogress” to “alarm-cleared-by-auto-restart”. If the pre serial number is not same as the new serial number, the method 700 may proceed to block 766 and may change the status of the alarm event for the alarming object - event time combination from “inprogress” to “alarm-cleared-by-hardware-replacemenf ’.

[0105] At block 767, the data processing device 102 may change the status of the alarm event for the alarming object - event time combination from “in-progress” to “alarm-not-cleared-by-auto-restart”. At block 768, the data processing device 102 may enable storing of the status of alarm information and not to try to restart for alarm ID-alarming object - event time combination till XI hours and should retry Y1 times, if the alarm is not cleared. At block 770, the data processing device 102 may wait for a pre-defined time period. At block 772, the data processing device 102 may retry the restart of the node.

[0106] At block 774, the data processing device 102 may determine whether the restart is retried for Y1 times for alarming object for the day. If the node is restarted, the method 700 may proceed to block 776.

[0107] At block 776, the data processing device 102 pauses restart for the alarm event for same set of alarming object - event time combination defined in the admin control and keeps the status as “alarm-not-cleared-by-auto-restart” till the impacted RU {RHH} / entity is replaced by a field engineer or the alarm gets auto cleared.

[0108] At block 778, the data processing device 102 may change the status of the alarm event for alarming object - event time combination to “hardware-to-be-replaced” category. At block 780, the data processing device 102 may determine whether thealarm is cleared for the alarming object event time. If the alarm is cleared, the method 700 may proceed to block 760. At block 760, the data processing device 102 may determine whether the pre serial number is same as current serial number of the node. If the pre serial number is same as the current serial number of the node, the method 700 may proceed with block 782.

[0109] At block 782, the data processing device 102 may change the status of the alarm event for the alarming object - event time combination from “in-progress” to “autocleared” against alarm ID for alarming object - event time.

[0110] At block 784, the data processing device 102 may change the status of the alarm event for the alarming object - event time combination from “in-progress” to “alarm-cleared-by-hardware-replacement”.

[0111] If the alarm is not cleared, the method 700 may proceed to block 786. At block 786, the data processing device 102 may determine whether the day is complete and next day is executed. If not, the method 700 may proceed to block 776. If yes, the method 700 may proceed to block 702.

[0112] FIG.8 illustrates a flow diagram of a method 800 for managing malfunctioning of the network nodes 108 in the communication network 100, in accordance with an embodiment of the present disclosure. The method 800 comprises a series of operation steps indicated by blocks 802 through 810.

[0113] At block 802, the data collection module 216 fetches the alarm data in real-time from the network nodes 108. The alarm data includes information of the alarms set by EMS 104 or the network nodes 108 when the operation parameters of the network nodes 108 are beyond the allowable range.

[0114] At block 804, the determination module 218 determines, based on the alarm data, the one or more alarms that are defined to identify the faulty nodes among the network nodes 108.

[0115] At block 806, the determination module 218 determines whether at least one alarm among the one or more alarms associated with at least one faulty node among the faulty nodes remains active for the predefined duration.

[0116] At block 808, the input module 220 obtains the one or more parameters related to configuration of restart of the at least one faulty node which remains active for the predefined duration. The one or more parameters comprise the restart interval to restart alarm not cleared by auto-restart, the number of retries to restart alarm not cleared by auto-restart, or the schedule time of execution.

[0117] At block 810, the execution module 222 executes the restart of the at least one faulty node based on the one or more parameters. Further, the execution module 222 updates the status of the at least one faulty node based on the result of the execution of the restart of the at least one faulty node.

[0118] Now, referring to the technical abilities and advantageous effect of the present disclosure, the system uses real-time data collection of the alarms from radio network entities to monitor the malfunctioning of radio hardware. The system communicates with orchestration tools to manage dependencies and prevent cascading failures during restarts. After the restart, the system evaluates the success of the action and collects data for continuous improvement.

[0119] Further, the system detects faults in the network nodes and automatically initiates restarting the network nodes, thus minimizes the service disruption, the system ensures that critical services remain operational with minimal disruption, improving customer satisfaction. The system eliminates the need for human technicians to manually restart components, saving time and reducing operational overhead. Automation of the system ensures consistent and precise handling of restart procedures across the network. The system can isolate affected components, restart only the necessary parts, and prevent widespread disruptions. The system can initiate restarts as a temporary solution while triggering alerts for deeper fault analysis. The system needsfewer on-site visits for manual restarts, leading to significant cost savings. The system ensures faster restoration of services which minimizes revenue loss caused by outages.

[0120] Further, implementing the disclosed system using the microservice architecture offers several advantages. Microservice architecture improves scalability, resilience, and agility by breaking a system into independently deployable, distributable, and fault-isolated services.

[0121] Those skilled in the art will appreciate that the methodology described herein in the present disclosure may be carried out in other specific ways than those set forth herein in the above disclosed embodiments without departing from essential characteristics and features of the present invention. The above-described embodiments are therefore to be construed in all aspects as illustrative and not restrictive. The automation is designed to work across diverse telecom entities, from legacy equipment to modern 5G infrastructure. The system supports large, multilayered telecom networks.

[0122] The drawings and the forgoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein. Any combination of the above features and functionalities may be used in accordance with one or more embodiments.

[0123] In the present disclosure, each of the embodiments has been described with reference to numerous specific details which may vary from embodiment to embodiment. The foregoing description of the specific embodiments disclosed herein may reveal the general nature of the embodiments herein that others may, by applying current knowledge, readily modify and / or adapt for various applications such specificembodiments without departing from the generic concept, and, therefore, such adaptations and modifications are intended to be comprehended within the meaning of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and is not limited in scope.LIST OF REFERENCE NUMERALS

[0124] The following list is provided for convenience and in support of the drawing figures and as part of the text of the specification, which describe innovations by reference to multiple items. Items not listed here may nonetheless be part of a given embodiment. For better legibility of the text, a given reference number is recited near some, but not all, recitations of the referenced item in the text. The same reference number may be used with reference to different examples or different instances of a given item. The list of reference numerals is:100 - Communication network102 - Data processing device104 - Element Management System (EMS)106 - Network108 - Network nodes / Remote Radio Heads (RRHs)200 - System for managing malfunctioning of the network nodes202 - Input-Output (I / O) interface204 - Processor206 - Memory206 A - Set of instructions208 - network communication manager210 - Console Host212 - Database214 - Processing modules216 - Data collection module218 - Determination module220 - Input module222 - Execution module224 - Notification module226 - Communication bus300 - Call flow between one or more components of the system 200 302 - Microservice304 - Read engine306 - Change engine308-326 - Operation steps for call flow 300400 - Call flow for displaying real-time status of the nodes402 -UI404-412 - Operation steps for call flow 400500 - Dashboard for selecting the one or more attributes600 - Dashboard for rendering live status of the network nodes 700 - Call flow for managing malfunctioning of the network nodes 702-786 - Operation steps for call flow 700800 - Method for managing malfunctioning of the network nodes 802-810 - Operation steps for the method 800

Claims

I / We claim:

1. A method (800) for managing malfunctioning of network nodes in a communication network (100), the method (800) comprising:fetching, by a data collection module (216), alarm data in real-time from one or more network nodes (108);determining, by a determination module (218) based on the alarm data, one or more alarms that are defined to identify one or more faulty nodes among the one or more network nodes (108);determining, by the determination module (218) based on the alarm data, whether at least one alarm among the one or more alarms associated with at least one faulty node among the one or more faulty nodes remains active for a predefined duration;obtaining, by an input module (220), one or more parameters related to configuration of restart of the at least one faulty node; andexecuting, by an execution module (222), restart of the at least one faulty node based on the one or more parameters.

2. The method (800) as claimed in claim 1, whereinthe one or more parameters comprise a restart interval to restart alarm not cleared by auto-restart, a number of retries to restart alarm not cleared by auto-restart, or a schedule time of execution, andthe one or more parameters are configurable by a network operator.

3. The method (800) as claimed in claim 1, further comprising updating, by the execution module (222), status of the at least one faulty node based on a result of the execution of the restart of the at least one faulty node.

4. The method (800) as claimed in claim 1, further comprising notifying, by a notification module (224), information of the at least one faulty node and a result of the execution of the restart of the at least one faulty node to the network operator via a user dashboard or a message.

5. The method (800) as claimed in claim 1, further comprising:determining, by the determination module (218), a number of alarm events based on the determined at least one alarm; andtagging, by the determination module (218), the number of alarm events as inprogress alarms.

6. The method (800) as claimed in claim 5, further comprising :executing, by the execution module (222), the restart of the in-progress alarms based on a schedule time of execution defined in the one or more parameters;updating, by the execution module (222), the tagging of the in-progress alarms to cleared-by-auto-restart alarms upon successful execution of the restart of the inprogress alarms; andupdating, by the execution module (222), the tagging of the in-progress alarms to not-cleared-by-auto-restart alarms upon unsuccessful execution of the restart of the in-progress alarms.

7. The method (800) as claimed in claim 6, further comprising:retrying executing, by the execution module (222), the restart of the not-cleared-by-auto-restart alarms based on a restart interval and a number of retries to restart alarm defined in the one or more parameters;updating, by the execution module (222), the tagging of the not-cleared-by-auto-restart alarms to cleared-by-auto-restart alarms upon successful execution of the restart of the not-cleared-by-auto-restart alarms; andupdating, by the execution module (222), the tagging of the not-cleared-by-auto-restart alarms to hardware-to-be-replaced upon unsuccessful execution of the restart of the not-cleared-by-auto-restart alarms.

8. The method (800) as claimed in claim 7, further comprising:determining, by the determination module (218), whether a serial number of the at least one faulty node is changed when the at least one alarm is cleared automatically;updating, by the execution module (222), the tagging of the hardware-to-be-replaced to cleared-by- hardware-replacement upon determination that the serial number of the at least one faulty node is changed; andupdating, by the execution module (222), the tagging of the hardware-to-be-replaced to auto-cleared alarm upon determination that the serial number of the at least one faulty node is not changed.

9. A system (200) for managing malfunctioning of network nodes in a communication network (100), the system (200) comprising:a data collection module (216) configured to fetch alarm data in real-time from one or more network nodes (108);a determination module (218) configured to:determine, based on the alarm data, one or more alarms that are defined to identify one or more faulty nodes among the one or more network nodes (108); anddetermine, based on the alarm data, whether at least one alarm among the one or more alarms associated with at least one faulty node among the one or more faulty nodes remains active for a predefined duration;an input module (220) configured to obtain one or more parameters related to configuration of restart of the at least one faulty node; andan execution module (222) configured to execute restart of the at least one faulty node based on the one or more parameters.

10. The system (200) as claimed in claim 9, whereinthe one or more parameters comprise a restart interval to restart alarm not cleared by auto-restart, a number of retries to restart alarm not cleared by auto-restart, or a schedule time of execution, andthe one or more parameters are configurable by a network operator.

11. The system (200) as claimed in claim 9, wherein the execution module (222) is further configured to update status of the at least one faulty node based on a result of the execution of the restart of the at least one faulty node.

12. The system (200) as claimed in claim 9, further comprising a notification module (224) configured to notify information of the at least one faulty node and a result of the execution of the restart of the at least one faulty node to the network operator via a user dashboard or a message.

13. The system (200) as claimed in claim 9, wherein the determination module (218) is further configured to:determine a number of alarm events based on the determined at least one alarm; andtag the number of alarm events as in-progress alarms.

14. The system (200) as claimed in claim 13, wherein the execution module (222) is further configured to:execute the restart of the in-progress alarms based on a schedule time of execution defined in the one or more parameters;update the tagging of the in-progress alarms to cleared-by-auto-restart alarms upon successful execution of the restart of the in-progress alarms; andupdate the tagging of the in-progress alarms to not-cleared-by-auto-restart alarms upon unsuccessful execution of the restart of the in-progress alarms.

15. The system (200) as claimed in claim 14, wherein the execution module (222) is further configured to:retry execution of the restart of the not-cleared-by-auto-restart alarms based on a restart interval and a number of retries to restart alarm defined in the one or more parameters;update the tagging of the not-cleared-by-auto-restart alarms to cleared-by-auto-restart alarms upon successful execution of the restart of the not-cleared-by-auto-restart alarms; andupdate the tagging of the not-cleared-by-auto-restart alarms to hardware-to-be-replaced upon unsuccessful execution of the restart of the not-cleared-by-auto-restart alarms.

16. The system (200) as claimed in claim 15, whereinthe determination module (218) is further configured to determine whether a serial number of the at least one faulty node is changed when the at least one alarm is cleared automatically, andthe execution module (222) is further configured to:update the tagging of the hardware-to-be-replaced to cleared-by- hardware-replacement upon determination that the serial number of the at least one faulty node is changed; andupdate the tagging of the hardware-to-be-replaced to auto-cleared alarm upon determination that the serial number of the at least one faulty node is not changed.

17. A computer program product comprising computer-executable instructions that are stored on a non-transitory computer-readable medium and that, when executed by at least one processor performs operations, comprising:fetching alarm data in real-time from one or more network nodes (108); determining, based on the alarm data, one or more alarms that are defined to identify one or more faulty nodes among the one or more network nodes (108);determining, based on the alarm data, whether at least one alarm among the one or more alarms associated with at least one faulty node among the one or more faulty nodes remains active for a predefined duration;obtaining one or more parameters related to configuration of restart of the at least one faulty node; andexecuting restart of the at least one faulty node based on the one or more parameters.