Mass outage detection in a communication network
Patent Information
- Application Number
- EP2023828220
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-17
- Filing Date
- 2023-12-15
- Publication Date
- 2025-12-24
AI Technical Summary
Current wireless communication networks face challenges in managing complex alarms generated from multiple cell sites due to faults, leading to increased maintenance costs and reduced network quality, as network devices struggle to correlate alarms from different sources effectively, resulting in multiple service tickets for similar root causes.
A computer-implemented method for detecting mass outages by analyzing a stream of alarm messages from cell sites, grouping alarms based on spatial proximity and temporal correlation, and outputting a single indication of a mass outage, which reduces the complexity of managing alarms and generates fewer service tickets.
This approach enables faster root cause analysis, reduces the number of individual alarms, and creates fewer trouble tickets by identifying correlated outages across multiple cell sites, improving network management efficiency and user experience.
Smart Images

Figure FI2023050697_22082024_PF_FP
Abstract
Description
MASS OUTAGE DETECTION IN A COMMUNICATION NETWORKTECHNICAE FIEED
[0001] Various example embodiments generally relate to the field of wireless communications. Some example embodiments relate to detection of a mass outage associated with multiple cell sites of a wireless communication network.BACKGROUND
[0002] Wireless communication may be implemented with a cellular radio network comprising multiple cell sites, each cell site offering communication services via one or more cells corresponding to certain geographical coverage area(s). Outages may occur due to many different reasons in such a network, including transmission, power, and configuration failures. Power failures and power outages may occur for example as a consequence of storms or other weather conditions, or simply due to human errors. Transmission failures may occur for example due to device malfunctions.SUMMARY
[0003] This summary is provided to introduce a selection of concepts in a simplified form that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0004] Example embodiments of the present disclosure enable to detection of mass outage and therefore reduction in complexity of managing alarms in a communication network. This and other benefits may be achieved by the features of the independent claims. Further example embodiments are provided in the dependent claims, the description, and the drawings.
[0005] According to a first aspect, a computer-implemented method is disclosed. The method may comprise: receiving a stream of alarm messages associated with cell sites of a communication network; detecting a mass outage associated with a group of the cell sites, in response to detecting, from the stream of alarm messages,a group of alarm messages originated at the group of the cell sites, wherein each of the group of the cell sites is located within a threshold distance from at least one of the group of the cell sites; and outputting an indication of the mass outage associated with the group of cell sites.
[0006] According to an example embodiment of the first aspect, the mass outage is detected based on alarm messages originated at the cell sites within a predetermined or configurable time window and / or alarm messages originated at cell sites located at a predetermined geographical area.
[0007] According to an example embodiment of the first aspect, the method may further comprise: determining the alarm messages originated at the cell sites within the predetermined or configurable time window based on time stamps included in the alarm messages.
[0008] According to an example embodiment of the first aspect, each of the group of alarm messages is associated with a particular type of alarm indicated in each of the group of alarm messages.
[0009] According to an example embodiment of the first aspect, the threshold distance is dependent on cell site density at neighbourhood of a respective cell site.
[0010] According to an example embodiment of the first aspect, the method may further comprise: determining, for each of the cell sites, the threshold distance based on distances of nearest neighbouring cell sites from the respective cell site.
[0011] According to an example embodiment of the first aspect, the threshold distance is determined based on an average distance of the nearest neighbouring cell sites from the respective cell site.
[0012] According to an example embodiment of the first aspect, the threshold distance is determined based on multiplying the average distance of the nearest neighbouring cell sites by a predetermined or configurable factor.
[0013] According to an example embodiment of the first aspect, the predetermined or configurable factor is selected based on the particular type of alarm.
[0014] According to an example embodiment of the first aspect, the predetermined or configurable factor is common for the cell sites.
[0015] According to an example embodiment of the first aspect, the threshold distance is configured to be lower or equal to a maximum allowed threshold distance.
[0016] According to an example embodiment of the first aspect, a number of the nearest neighbouring cell sites is predetermined or configurable.
[0017] According to an example embodiment of the first aspect, the threshold distance is different for at least two of the cell sites, or the threshold distance is unique for each of the cell sites.
[0018] According to an example embodiment of the first aspect, the method may further comprise: retrieving locations of the cell sites from a database based on identifiers of the cell sites, wherein each alarm message comprises an identifier of a cell site associated with the alarm message; and determining distances between the cell sites based on the locations of the cell sites.
[0019] According to an example embodiment of the first aspect, the indication of the mass outage comprises identifiers of the cell sites of the group of cell sites, the indication of the mass outage comprises an indication of the particular type of alarm, and / or the indication of the mass outage is provided as an automated service ticket.
[0020] According to an example embodiment of the first aspect, the method may further comprise: receiving a further alarm message of the stream of alarm messages; assigning a cell site associated with the further alarm message to the group of the cell sites, in response to determining that the cell site associated with the further alarm message is located within the threshold distance from at least one of the group of the cell sites.
[0021] According to an example embodiment of the first aspect, the method may further comprise: detecting a mass outage of another group of the cell sites; and merging the other group of the cell sites with the group of the cell sites, in response to determining that the further cell site is located within the threshold distance from at least one of the other group of the cell sites.
[0022] According to a second aspect, an apparatus may comprise means for performing any example embodiment of the method of the first aspect.
[0023] According to a third aspect, computer program or a computer program product may comprise program code configured to, when executed by a processor, cause an apparatus at least to perform any example embodiment of the method of the first aspect.
[0024] According to a fourth aspect, an apparatus may comprise at least one processor; and at least one memory including computer program code; the at least one memory and the computer code configured to, with the at least one processor, cause the apparatus at least to perform any example embodiment of the method of the first aspect.
[0025] Any example embodiment may be combined with one or more other example embodiments. Many of the attendant features will be more readily appreciated as they become better understood by reference to the following detailed description considered in connection with the accompanying drawings.DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings, which are included to provide a further understanding of the example embodiments and constitute a part of this specification, illustrate example embodiments and together with the description help to understand the example embodiments. In the drawings:
[0027] FIG. 1 illustrates an example of a wireless communication network;
[0028] FIG. 2 illustrates an example of an apparatus configured to practise one or more example embodiments;
[0029] FIG. 3 illustrates an example of a flow chart for mass outage detection;
[0030] FIG. 4 illustrates an example of successfully detected mass outage;
[0031] FIG. 5 illustrates an example of missing sites with correlated outages in mass outage detection;
[0032] FIG. 6 illustrates an example of merging groups of alarms;
[0033] FIG. 7 illustrates an example of an algorithm for merging alarms in groups;
[0034] FIG. 8 illustrates an example an algorithm formass outage detection; and
[0035] FIG. 9 illustrates an example of a computer-implemented method for mass outage detection.
[0036] Like references are used to designate like parts in the accompanying drawings.DETAILED DESCRIPTION
[0037] Reference will now be made in detail to example embodiments, examples of which are illustrated in the accompanying drawings. The detailed description provided below in connection with the appended drawings is intended as a description of the present examples and is not intended to represent the only forms in which the present example may be constructed or utilized. The description sets forth the functions of the example and the sequence of steps for constructing and operating the example. However, the same or equivalent functions and sequences may be accomplished by different examples.
[0038] Mobile radio access networks (RAN) are facing ever growing challenges in their fault management due to increased number of alarms and the growing complexity of heterogeneous networks comprising access nodes with cells of different sizes. Faults may be configured to be constantly detected in the network and alarms may be generated at various parts of the network, e.g., at different cell sites, to gather information about such failures.
[0039] Performance of the RAN affects user experience. For example, the network may have degraded quality of service (QoS) or it may not be available at all. When a failure is detected, a large number of alarms may be generated from multiple different network elements within a short period of time. Alarms may be sent as messages including an indication of a type of the associated failure, but network devices may have limited knowledge of the network and the cause of the error. Network devices may be configured to record the alarms and report them to a network controller, which may be configured to centrally process the generated alarms. For example, an element management system (EMS) of the network may receive the alarms and may forward them higher to a network management system (NMS). The NMS may be for example configured to display inventory and alarm data and to manage multiple networks. The NMS may receive alert messages fromdetected events. The alarms may be processed by network operation center(s) (NOC), where trouble tickets may be created for issues that require further actions. NOC can receive multiple alarms from different access nodes and other network elements and create a trouble ticket for each of them. However, in the case of wider scale outages (e.g., power outage), some of these alarms and access nodes may be part of same incident and have the same root cause. In general, fault management in telecommunication may include detection and analysis of the issues together with fixing network problems.
[0040] FIG. 1 illustrates an example of a wireless communication network. Communication network 100 may comprise one or more devices, which may be also referred to as client nodes, user nodes, or user equipment (UE). An example of a device is UE 110, which may communicate with one or more access nodes of a radio access network (RAN) 120. Signals transmitted by an access node to UE 110 may be referred to as downlink signals. Signals transmitted by UE 110 to an access node may be referred to as uplink signals. An access node may be also referred to as an access point or a base station. Communication network 100 may be configured for example in accordance with the 4thor 5thgeneration (4G, 5G) digital cellular communication networks, as defined by the 3rdGeneration Partnership Project (3GPP). In one example, communication network 100 may operate according to 3GPP (4G) LTE (Long-Term Evolution) or 3GPP 5G NR (New Radio). Communication network 100 may hence comprise a cellular radio network. Access nodes 122, 124, 126 of RAN 120 may comprise 5thgeneration access nodes (gNB) or 4thgeneration access nodes (eNodeB). It is however appreciated that example embodiments presented herein are not limited to these example networks and may be applied in any present or future wireless communication networks, or combinations thereof, for example other type of cellular networks, short-range wireless networks (e.g., Wi-Fi), multicast networks, broadcast networks, wired communication networks, or the like.
[0041] An access node may provide communication services within one or more cells, illustrated with dotted circles in FIG. 1. Cells may correspond to geographical area(s) covered by signals transmitted by the access node. A cell site may refer to a geographical location and / or equipment that serves one or more ofsuch cells. For example, each of the cell sites associated with access nodes 122, 124, 126 serve three cells from a single location. Coverage, capacity, and throughput provided by an access node may vary. For example, larger cells in terms of coverage area may be configured at rural areas, where customer density may be smaller, and smaller cells may be configured at urban areas to serve a larger number of customers within a smaller coverage area. Cell density and cell site density may be therefore higher at urban areas and lower at rural areas.
[0042] Different types of access nodes (e.g., macro, micro, pico or femto) may have two main component blocks. A first block may include a cooling system and / or a microwave link. A second block may include a power amplifier (PA), a transceiver (TX / RX), and / or a digital signal processor (DSP). The access node(s) may be powered by electricity and if an access node is without a power source, it will not be able to communicate data traffic. An alarm may be triggered, in response to detecting a power failure or device malfunction at a cell site, for example malfunction of any of the abovementioned blocks or sub-blocks of an access node or auxiliary equipment, such as for example cabling.
[0043] Communication network 100 may further comprise a core network 130, which may comprise various network functions (NF) for establishing, configuring, and controlling data communication sessions of users, for example UE 110. The data communication sessions may carry data traffic, for example application data associated with one or more applications running on UE 110. Communication network 100 may further comprise a network controller 140, which may be responsible of configuring various operations of RAN 120 and / or core network 130. Network controller 140 may for example comprise the element management system (EMS), the network management system (NMS), and / or the network operation center(s) (NOC). Even though illustrated as a separate entity, network controller 140 may be also embodied as part of core network 130. Even though some operations have been described as being performed by network controller 140, it is understood that similar functions may be performed alternatively by other network device(s) or network function(s) of communication network 100. One task of network controller 140 may be to detect mass outages within RAN 120.
[0044] A fault or a failure may comprise one kind of a problem. A fault may be permanent or intermittent. Network controller 140 may be informed about occurrence of faults as alarms, for example in the form of alarm messages. A single fault may cause multiple alarms. Examples of a fault include power outage or connectivity loss. An error may be a discrepancy between assumed to be correct condition and an observed condition. A fault may cause multiple errors but may not need to be directly corrected and can be invisible. An example of an error is when internet protocol (IP) packets are discarded in a router because of an erroneous header.
[0045] An event may be an occurrence of an exceptional condition at certain time in the managed communication network. There can be fault events indicative of the start of a problem or clear events notifying of the end of the problem. Alarms may comprise information of the faults generated by network devices. Alarms may be transmitted by the relevant network devices to network controller 140 (e.g., NOC). An alarm may contain multiple different states and events. Alarm correlation may refer to a process of finding relationships and correlations between alarms and grouping alarms referring to the same problem.
[0046] The number of generated alarms may increase with complexity of the network and one goal for operators may be to reduce maintenance costs. Targets may also include improving network quality, predicting future incidents, and improving fault resolution, while reducing operational costs. It may be therefore desired to reduce the number of service tickets generated based on the alarms. Therefore, example embodiments of the present disclosure enable grouping individual alarms to detect a mass outage, for which a single service ticket may be output. For example, only one trouble ticket for field maintenance from one high level failure may be generated by network controller 140. For example when malfunctioning equipment causes correlated alarms at different cell sites, only one service ticket may be generated instead of multiple service tickets.
[0047] FIG. 2 illustrates an example embodiment of an apparatus 200 configured to perform one or more example embodiments. Apparatus 200 may be for example used to implement network controller 140. Apparatus 200 may comprise at least one processor 202. The at least one processor 202 may comprise,for example, one or more of various processing devices or processor circuitry, such as for example a co-processor, a microprocessor, a controller, a digital signal processor (DSP), a processing circuitry with or without an accompanying DSP, or various other processing devices including integrated circuits such as, for example, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a microcontroller unit (MCU), a hardware accelerator, a special-purpose computer chip, or the like.
[0048] Apparatus 200 may further comprise at least one memory 204. The at least one memory 204 may be configured to store, for example, computer program code or the like, for example operating system software and application software. The at least one memory 204 may comprise one or more volatile memory devices, one or more non-volatile memory devices, and / or a combination thereof. For example, the at least one memory 204 may be embodied as magnetic storage devices (such as hard disk drives, floppy disks, magnetic tapes, etc.), optical magnetic storage devices, or semiconductor memories (such as mask ROM, PROM (programmable ROM), EPROM (erasable PROM), flash ROM, RAM (random access memory), etc.).
[0049] Apparatus 200 may further comprise a communication interface 208 configured to enable apparatus 200 to transmit and / or receive information to / from other devices, functions, or entities. In one example, apparatus 200 may use communication interface 208 to transmit or receive information over a service based interface (SBI) message bus of core network 130, for example to core network 130 and / or RAN 120 about detected alarms, to output indication(s) of detected mass alarms, for example to a human user or an automated service ticket system. Apparatus 200 may further comprise a user interface, for example for configuring apparatus 200 or for providing user output by the apparatus, such as for example visual and / or audible signal(s), for example by speaker(s), display(s), light(s), or the like.
[0050] When apparatus 200 is configured to implement some functionality, some component and / or components of apparatus 200, such as for example the at least one processor 202 and / or the at least one memory 204, may be configured to implement this functionality. Furthermore, when the at least one processor 202 isconfigured to implement some functionality, this functionality may be implemented using program code 206 comprised, for example, in the at least one memory 204.
[0051] The functionality described herein may be performed, at least in part, by one or more computer program product components such as for example software components. According to an embodiment, the apparatus comprises a processor or processor circuitry, such as for example a microcontroller, configured by the program code when executed to execute the embodiments of the operations and functionality described. A computer program or a computer program product may therefore comprise instructions for causing, when executed, apparatus 200 to perform the method(s) described herein. Alternatively, or in addition, the functionality described herein can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), application-specific Integrated Circuits (ASICs), applicationspecific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), Graphics Processing Units (GPUs).
[0052] Apparatus 200 comprises means for performing at least one method described herein. In one example, the means comprises the at least one processor 202, the at least one memory 204 including program code 206 configured to, when executed by the at least one processor, cause the apparatus 200 to perform the method.
[0053] Apparatus 200 may comprise a computing device such as for example an access point, a base station, a server, a network device, a network function device, a network controller, or the like. Although apparatus 200 is illustrated as a single device it is appreciated that, wherever applicable, functions of apparatus 200 may be distributed to a plurality of devices, for example to implement example embodiments as a cloud computing service.
[0054] FIG. 3 illustrates an example of a method for mass outage detection. The method may be computer-implemented and be performed for example by network controller 140. The method may be however performed by any suitable network device. A purpose of the method may be to find correlated cell sites that containrelevant active alarms and thereby to detect a mass outage at communication network 100.
[0055] Methods described with reference to FIG. 3 enable detecting failures, e.g., power or transmission failures, which occur in more than one affected sites and are correlated with a high probability. For example, at some unknown time, a large power or transmission incident may occur, affecting a number of access nodes in the network, which may indicate the detected failures in the form of alarm events (messages). The incident may be dynamic and it may propagate along the network affecting more sites within a certain time-window, for example in the case of power mass outages. Examples of FIG. 3 provide a density-based stream clustering algorithm, which may be used for detecting mass outages for such failures. The disclosed algorithm may be sequential over a data stream as it may monitor and gain knowledge of the current situation in real-time. An objective may be to detect clusters where more than M sites are affected and which can grow with time. Additionally, there is provided a mass outage detection service that can be used to detect mass outages in real time, which enables fast root cause analysis. Detecting mass outage clusters with correlated sites enables reducing the number of the created trouble tickets. A mass outage may be determined to occur when alarms of the same type arrive from a certain number (M) of network elements, for example within a certain time window. The alarms may be determined to have the same root cause for failure with high probability. Correlated power or transmission outages may not typically occur on cell sites that are located spatially far from each other, for example such that there are many other cell sites in the intervening area that do not have active alarms of the same type. Also, if alarms are temporally far from each other, they may not be typically correlated. These observations may be used to determine correlated outages. Examples of such procedures are described below with reference to operations 301 to 308.
[0056] At operation 301, network controller 140 may receive an alarm, e.g., an alarm message, from cell site Sn. The alarm message may include an identifier of this cell site. The alarm may be part of a stream of alarm messages that are associated with, e.g., originated at, different cell sites, which may be indexed by n =1 . ..N, where N is the number of cell sites from which alarms are received. Thealarm stream may include alarm messages from multiple network operators. Network controller 140 may be configured to parse the relevant information, e.g., cell site identifier, alarm type, etc., from different types of alarm messages, which may be received from devices of different manufacturers. Hence, detecting mass outages associated with multiple network operators and / or different device manufacturers is enabled.
[0057] An alarm message may include various information about the alarm, for example the associated cell site a type of the alarm, or the like. An alarm message may for example include one or more of the following fields: alarmCSN'. this field may indicate an identifier of the alarm, e.g., a serial number of the alarm that uniquely identifies it, alarmCategory-. this field may indicate a category of the alarm message, for example whether the alarm message is associated with a raised alarm or a cleared alarm, alarmOccurTime'. this field may include a time stamp indicative of when the alarm was triggered, alarmMOname'. this field may include the name of the device (MO, managed object) where the alarm is generated, alarmNEDevID'. this field may include an identifier of the device where the alarm is generated, alarmID'. this field may include a unique identifier of the alarm, alarmType'. this field may indicate a type of the alarm, e.g., power, environment, signalling, trunk, hardware, software, running system, communication system, quality-of-service, processing error, or the like. alarmLevek this field may indicate a level of the alarm, e.g., critical, major, minor, warning, indeterminate, or cleared. alarmAckTime'. this field may indicate acknowledgement time of the alarm, alarmRestoreTime'. this field may indicate clearance time of the alarm, alarmProbaleCause'. this field may indicate an identifier of the cause of the alarm,alarmAdditionallnfo'. this field may indicate additional information, for example the radio access technology (RAT) associated with the alarm, e.g., LTE, or alarmExtendlnfa'. this field may indicate location information (e.g. coordinates, for example as latitude and longitude) of the alarm, for example a location of the device where / when the alarm was generated.
[0058] An alarm message may be classified based on any suitable field of the alarm message. A mass outage may be detected among alarm messages belonging to the same class, e.g., having a particular type and / or occurring during some time window. In general, a data stream, such as for example an alarm message stream, may comprise data that is being produced incrementally over time such that it is not available fully as a static data image before processing it. Network fault management alarm data from various network elements is one example of a continuous data stream. This data stream may be collected and analysed by network controller 140 to detect mass outage(s). Certain type(s) of data may be periodically saved for offline analysis but in order to react in real-time to alarms, network controller 140 may be configured to continuously monitor and analyse the alarm stream.
[0059] At operation 302, network controller 140 may determine whether there are other alarms that are associated with cells sites located within distance dn(threshold distance) from cell site Sn. Note that distance dn, also referred to as radius or search radius, may be different for different cell sites Sn. Distance dnmay be therefore cell site specific. For example, distance dnmay be unique for each cell site, or it may be different for at least two of the cell sites.
[0060] Network controller 140 may be configured to determine distance dnwith a / .'-nearest neighbours (kNN) algorithm, where distance dnis determined based on distances between site Snand each of its k nearest neighbouring cell sites. Distance dnmay be therefore dependent on cell site density at neighbourhood of cell site Sn. In one example, distance dnmay be determined based on, or be equal to, the average distance to the k nearest neighbouring cell sites. Parameter k may be configurable, for example selected from the range of 5 to 7. Configurability of parameter k enables mass outage detection to be configured (e.g., optimised) for differentnetwork topologies and individual networks. Such a density-based spatiotemporal clustering algorithm exploits a uniquely defined distance value for each alarm or cell site based on cell site density calculation at neighbourhood of cell site Sn.
[0061] Network controller 140 may be configured to retrieve, for example from a database, locations of the cell sites associated with the alarm messages, in order to determine distances between relevant cell sites. The database may include location information of the cell sites associated with identifiers of the cell sites. The locations of the cell sites may be therefore retrieved from the database based on identifiers of the cell sites included in the alarm messages. As described above, each alarm message may comprise an identifier of a cell site associated with the alarm message. Network controller 140 may then determine distances between the cell sites based on the locations of the cell sites.
[0062] A scaling factor may be applied when determining distance dn. For example, the distance obtained based on the distances to the k nearest neighbouring cells may be scaled to obtain distance dn. This enables to adjust the threshold distance used for grouping the alarms. This may improve clustering accuracy. The scaling factor may be predetermined, for example preconfigured at network controller 140 at installation of communication network 100, or configurable during operation of communication network 100. Furthermore, the scaling factor may be dependent on the type of alarm. For example, different scaling factors may be (pre)configured for alarms having different alarm types. This enables same algorithm to be used, e.g., in parallel, for detecting different types of mass outages, for example to separately detect power or transmission outages. This may be beneficial, because different types of outages may have different geographical correlations. For example, transmission outages may spread according to topology of communication network 100, while power outages may spread according to topology of the underlying electricity network, which may be different from the topology of communication network 100. The scaling factor may be however common for different cell sites. This enables to prevent the scaling factor from affecting the relationships between the cell site density based threshold distances of different cell sites.
[0063] Distance dnmay be upper bounded by a maximum allowed threshold distance dmax, which may be common for different cell sites. Hence, distance dnmay be limited such that it is lower than or equal to dmax, regardless of the value calculated based on the distances to the k nearest neighbouring cell sites. This enables to prevent spatially distant alarms to be treated as correlated alarms and being reported as belonging to same mass outage. The value of dmax may be based on (e.g., equal to) a maximum coverage distance of an access node within the different cell sites, for example 25 km.
[0064] The threshold distance may be alternatively determined separately for determining whether a certain pair of cell sites is to be assigned to the same group. For example, distance dn,i may be initially determined for a first cell site based on distances to its k nearest neighbouring cell sites, as described above. Similarly, distance dn,2 may be initially determined for a second cell site that is one of the k nearest neighbouring cell sites of the first cell site. When determining whether to assign an alarm associated with the second cell site to the same group as an alarm associated with the first cell site, network controller 140 may use the average of the two threshold distances (dn,i, dn,2) of these two cell sites. The threshold distance may be therefore different for one or more of the k nearest neighbouring cell sites. The threshold distance may be adjusted for each of the k nearest neighbouring cell sites of the first cell site based on the (non-adjusted) threshold distance of the respective site (e.g., dn,2), for example by taking their average. This enables the cell site density also at the neighbourhoods of the k neighbouring cell sites to be considered when grouping the alarms.
[0065] If there are no other alarms from cell sites located within distance dnfrom site Sn, network controller 140 may move to execution of operation 303. If there is at least one other alarm from cell sites located within distance dnfrom site Sn, network controller 140 may move to execution of operation 304. Note that when determining whether there are other alarms, network controller 140 may consider alarms generated or received during a certain time window, which may be predetermined (e.g., past two hours) or configurable during operation of communication network 100. This enables to prevent temporally distant alarms to be treated as correlated alarms and being reported as belonging to same massoutage. Alternatively, or additionally, network controller 140 may consider alarms originated at cell sites located at a predetermined or configurable geographical area. This enables to restrict detection of mass outage to a specific geographical area, for example corresponding to topology of electricity network.
[0066] At operation 303, network controller 140 may assign the alarm received at operation 301 to a new group. Since detecting a mass outage requires more than one alarm, network controller 140 may move back to execution of operation 301 to receive another alarm, which may be subsequently merged to the newly generated group. It is noted that instead of individual alarms or alarm messages, a group may include identifiers of cell sites associated with the alarms or alarm messages.
[0067] At operation 304, network controller 140 may merge the received alarm to a same group with alarm(s) detected within distance dn. Merging an alarm within a group may comprise associating the alarm (e.g., alarm identifier, or respective cell site identifier or device identifier) at memory of network controller 140 to earlier members of the group. Network controller 140 may move from execution of operation 304 to execution of operation 305. It is however noted that network controller 140 may in some example embodiments be configured to perform mass outage detection at operation 307 without operations 305 or 306. Therefore, network controller 140 may alternatively move from execution of operation 304 to execution of operation 307 directly.
[0068] At operation 305, network controller 140 may determine whether there are other group(s), which include alarms within distance dnfrom site Sn. If yes, network controller 140 may merge the groups at operation 306. This enables to avoid provision of separate indications for mass outages for groups, for which their correlation is detected after forming the separate groups. Thus, unnecessary alarms may be avoided regardless of the order of incoming alarm messages. Due to the threshold distance applied for each cell site separately, merging the alarms to a group (cf. operation 304) and / or merging different groups (cf. operation 306) results in a group of alarm messages originated at a group of cell sites, where each of the group of cell sites is located within (its own) threshold distance from at least one other cell site of the group of cell sites. Determining the correlated sites this way enables to detect large outages occurring for example in a chain of cell sites. Themost distant cell sites within the group may be located far away from each other, if there are intermediate cell sites that cause condition(s) for the grouping to be met. Hence, outages associated with even irregular geographical topologies may be effectively detected. One example of merging groups is provided in FIG. 6.
[0069] At operation 307, network controller 140 may determine whether a mass outage is detected. Network controller 307 may determine a mass outage to be detected for a group of cell sites, if the number of cell sites associated with a group of alarms is higher than or equal to a threshold (M). Threshold M may be for example equal to two, which results in the mass outage to be detected if a group of alarms is associated with at least two cell sites. To detect the mass outage, it may be required that each of the group of cells sites is associated with at least one of the group of alarm messages. Alternatively, the threshold may be set higher, depending on the degree of reduction in the number of service tickets that desired to be achieved. The threshold may be for example three, four, five, ten, or twenty cell sites, but higher thresholds are possible as well.
[0070] If no mass outage is detected (with any groups of alarms messages), network controller 140 may move back to execution of operation 301 to receive a further alarm message. If a mass outage is detected, network controller 140 may move to execution of operation 308. Examples of detected mass outages are described below with reference to FIG. 4 and FIG. 5.
[0071] At operation 308, network controller 140 may output an indication of the detected mass outage. The indication of the mass outage may comprise identifiers of the group of cell sites associated with the group of alarms. This enables to identify the sites affected by the mass outage. The indication of the mass outage may comprise an indication of a particular type of alarm that is common to the group of alarm messages. This enables correct service action(s) to be determined in order to recover from the mass outage. The indication of the mass outage may be provided for example as an automated service ticket, which may be indicative of the cell sites affected by the mass outage and / or the type of alarm associated with the mass outage. The indication of the mass outage may comprise an identifier of the mass outage, for example to uniquely identify the detected mass outage fromother detected mass outages. A group may comprise a plurality of members, for example a plurality of alarms, alarm messages, cell sites, or cell site identifiers.
[0072] FIG. 4 illustrates an example a successfully detected mass outage. Black dots represent cell sites from which alarm messages have been received at network controller 140. White dots represent other cell sites of communication network 100. It is noted that density of cell sites varies at different geographical regions of the network.
[0073] Two examples of cell sites, namely Si and Sz are considered in more detail. Cell site Si is associated with distance di (search radius) and cell site & is associated with distance dz. For Si, four other alarming sites can be identified within di, including cell site 2. For S2, three other alarming sites can be identified within dz. Since cell site density is higher at the neighbourhood of Sz, distance dz is shorter than distance d . By analysing the alarming cell sites with respective distances dnfor each of the alarming cell sites, network controller 140 may determine a group of alarms associated with a group cell sites denoted by the black dots within the dashed area. FIG. 4 also provides an example of an outlier that is located far away from the detected group, that is, cell sites included in the detected mass outage are outside the threshold distance of the outlier cell site. Network controller 140 may accordingly determine to output an indication of the detected mass outage.
[0074] FIG. 5 illustrates an example of missing sites with correlated outages in mass outage detection. In this example, distance dz has been set to be two low, which results in three alarming cell sites to be missed when detecting the mass outage. Distance dz may be too low for example because of not applying a sufficiently high scaling factor. Note that Sz still belongs to the same group as Si, even though distance dz has been scaled similarly. This is because of the alarming cell site located between Si and Sz. It is therefore noted that enabling configuration of threshold distance dnimproves performance of mass outage detection, because the algorithm may be tuned based on how well it finds the correlated sites in a particular network.
[0075] FIG. 6 illustrates an example of merging groups of alarms. In this example, a first group (Group 1) of alarms associated with certain cell sites has been detected similar to FIG. 4. Another group (Group 2) has been detected as well, butthe distances between cell sites associated with these two groups of alarms have been longer than threshold distance dnfor any of the cell sites associated with the groups. When receiving a further alarm message associated with cell site S3, network controller 140 may determine (cf. operation 304) that there are cell sites associated with Group 1 within distance ch from S3 and assign the received alarm to Group 1. However, network controller 140 may also determine (cf. operation305) that there are cell sites associated with Group 2 within distance ch from S3. Consequently, network controller 140 may merge Groups 1 and 2 (cf. operation306), detect mass outage for the merged group (cf. operation 307), and output an indication of the mass outage. Merging the groups may trigger the detection and indication of the mass outage, if the threshold for the number of cell sites is higher than the number of cell sites associated with Group 1 but lower than the number of cell sites associated with the merged group.
[0076] FIG. 7 illustrates an example (pseudocode) of an algorithm for merging alarms in groups. Merging each alarm into a group may be implemented by defining a density value for each alarm and cell site, which may be used as the search radius, optionally with a scaling (multiplication) factor. The value of the search radius may be calculated based on the configured number (A) of neighbouring cell sites. When an alarm is attempted to be merged into a specific group, as seen in Algorithm 1, the search radius may be defined based on found distance value of the respective cell site, optionally with a configurable scaling factor. In the configurations, also a maximum search radius may be set to avoid too large search areas in sparse locations. Other active alarms may be searched based on the defined radius, optionally within a specified time-window and / or for certain alarm type(s), which may be configurable parameters. If alarms match the search, the received alarm (a) may be merged to an existing group. If no alarms matching the search criteria are found, a new group may be created for the arriving alarm. If several groups are found for the arriving alarm, groups may be combined. More specifically:
[0077] On line 2, the search radius (cf. distance <7«) is defined by scaling the average distance of neighbouring cell sites for the cell site associated with an incoming alarm (incomingAlarm).
[0078] On lines 3-5, the search radius is limited such that it is at most equal to the maximum allowed threshold distance (definedMaxRadius).
[0079] On line 6, alarms within the radius from the cell site, having a particular type, and occurred within a certain time window are found and assigned to a list of alarms (alarmList).
[0080] On lines 7-9, if the alarmList is not an empty list, the incomingAlarm is assigned to same group as the first alarm of this list.
[0081] On lines 10-12, the incomingAlarm is assigned to a new group, if the alarmList was an empty list.
[0082] On lines 14-18, the algorithm goes through the alarmList and possible separate groups are merged if there exists another group in the found area.
[0083] FIG. 8 illustrates an example an algorithm (pseudocode) for mass outage detection. Mass outages may be determined from the groups of alarms based on a configurable threshold value. Different network operators may have diverse views on the minimum number of different cell sites having active alarms in order to determine presence of a mass outage. If the threshold is crossed, network controller 140 may attempt to find already determined mass outages and add the arriving alarm to the existing mass outage group. Otherwise a new mass outage group, containing all alarms already belonging to the defined group, e.g., with the same alarm identifier, may be created, as can be seen in Algorithm 2. More specifically:
[0084] On line 2, alarms of a certain group (identifier by groupID) are listed in alarmsln GroupList.
[0085] On line 3, cell sites associated with alarmsInGroupList are listed in sitesList.
[0086] On lines 4-5, if the number of cell sites in sitesList is above or equal to a threshold (siteTreshold), the algorithm tries to find already detected mass outage(s). The condition of line 4 will not be met if there are multiple alarms (e.g., ten) from the same cell site. The condition will be met only if there are alarms from multiple cell sites, which results in checking (cf. line 6) whether there already exists detected mass outage(s).
[0087] On lines 6-7, if no existing mass outages were found, a new mass outage may be saved.
[0088] On lines 10-11, if at least one existing mass outage was found, the alarm may be assigned to the existing mass outage.
[0089] On lines 16-20, alarms may be merged similar to Algorithm 1 to form a mass outage. The mass outage may be assigned a mass outage identifier.
[0090] FIG. 9 illustrates an example of a computer-implemented method mass outage detection.
[0091] At 901, the method may comprise receiving a stream of alarm messages associated with cell sites of a communication network.
[0092] At 902, the method may comprise detecting a mass outage associated with a group of the cell sites, in response to detecting, from the stream of alarm messages, a group of alarm messages originated at the group of the cell sites, wherein each of the group of the cell sites is located within a threshold distance from at least one of the group of the cell sites.
[0093] At 903, the method may comprise outputting an indication of the mass outage associated with the group of cell sites.
[0094] Further features of the method directly result for example from the functionalities of network controller 140 or in general apparatus 200, as described throughout the specification and in the appended claims, and are therefore not repeated here. Different variations of the method may be also applied, as described in connection with the various example embodiments.
[0095] One or more of the following benefits may be provided by various example embodiments: faster automatic failure analysis and detection of alarms in sites which have the same root cause, less individual alarms to react, more effective ticket creation (e.g., creating only one trouble ticket for the mass outage instead of one ticket for each cell site or alarm), the disclosed method for grouping alarms is suitable for detecting different kind of failures by selecting suitable type of alarms, better cell site density detection between dense and sparse areas with cell-specific or unique search radius, automatic selection of density threshold, search radius may be defined for each cell site and calculated for each alarm at run-time, more accurate clustering of alarms / cell sites, or capability to maintain information of ongoing / active mass outages and their states.
[0096] An apparatus, such as for example a network controller device configured to implement one or more network functions or entities, may be configured to perform or cause performance of any aspect of the method(s) described herein. Further, a computer program or a computer program product may comprise instructions for causing, when executed, an apparatus to perform any aspect of the method(s) described herein. Further, an apparatus may comprise means for performing any aspect of the method(s) described herein. According to an example embodiment, the means comprises at least one processor, and memory including program code, the at least one processor, and program code configured to, when executed by the at least one processor, cause performance of any aspect of the method(s). In general, computer program instructions may be executed on means providing generic processing functions. Such means may be embedded for example in a computer, a server, or the like. The method(s) may be thus computer- implemented, for example based algorithm(s) executable by the generic processing functions, an example of which is the at least one processor 202.
[0097] Any range or device value given herein may be extended or altered without losing the effect sought. Also, any embodiment may be combined with another embodiment unless explicitly disallowed.
[0098] Although the subject matter has been described in language specific to structural features and / or acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as examples of implementing the claims and other equivalent features and acts are intended to be within the scope of the claims.
[0099] It will be understood that the benefits and advantages described above may relate to one embodiment or may relate to several embodiments. The embodiments are not limited to those that solve any or all of the stated problems or those that have any or all of the stated benefits and advantages. It will further be understood that reference to 'an' item may refer to one or more of those items.
[0100] The steps or operations of the methods described herein may be carried out in any suitable order, or simultaneously where appropriate. Additionally, individual blocks may be deleted from any of the methods without departing fromthe scope of the subject matter described herein. Aspects of any of the example embodiments described above may be combined with aspects of any of the other example embodiments described to form further example embodiments without losing the effect sought.
[0101] The term 'comprising' is used herein to mean including the method, blocks, or elements identified, but that such blocks or elements do not comprise an exclusive list and a method or apparatus may contain additional blocks or elements.
[0102] Although subjects may be referred to as ‘first’ or ‘second’ subjects, this does not necessarily indicate any order or importance of the subjects. Instead, such attributes may be used solely for the purpose of making a difference between subjects.
[0103] It will be understood that the above description is given by way of example only and that various modifications may be made by those skilled in the art. The above specification, examples and data provide a complete description of the structure and use of exemplary embodiments. Although various embodiments have been described above with a certain degree of particularity, or with reference to one or more individual embodiments, those skilled in the art could make numerous alterations to the disclosed embodiments without departing from scope of this specification.
Claims
CLAIMS1. A computer-implemented method, comprising: receiving a stream of alarm messages associated with cell sites of a communication network; detecting a mass outage associated with a group of the cell sites, in response to detecting, from the stream of alarm messages, a group of alarm messages originated at the group of the cell sites, wherein each of the group of the cell sites is located within a threshold distance from at least one of the group of the cell sites; and outputting an indication of the mass outage associated with the group of cell sites.
2. The method according to claim 1, wherein the mass outage is detected based on alarm messages originated at the cell sites within a predetermined or configurable time window and / or alarm messages originated at cell sites located at a predetermined geographical area.
3. The method according to claim 2, further comprising: determining the alarm messages originated at the cell sites within the predetermined or configurable time window based on time stamps included in the alarm messages.
4. The method according to any preceding claim, wherein each of the group of alarm messages is associated with a particular type of alarm indicated in the each of the group of alarm messages.
5. The method according to any preceding claim, wherein the threshold distance is dependent on cell site density at neighbourhood of a respective cell site.
6. The method according to claim 5, further comprising:determining, for each of the cell sites, the threshold distance based on distances of nearest neighbouring cell sites from the respective cell site.
7. The method according to claim 6, wherein the threshold distance is determined based on an average distance of the nearest neighbouring cell sites from the respective cell site.
8. The method according to claim 7, wherein the threshold distance is determined based on multiplying the average distance of the nearest neighbouring cell sites by a predetermined or configurable factor.
9. The method according to claims 4 and 8, wherein the predetermined or configurable factor is selected based on the particular type of alarm.
10. The method according to claim 9 or 10, wherein the predetermined or configurable factor is common for the cell sites.
11. The method according to any of claims 7 to 10, wherein the threshold distance is configured to be lower or equal to a maximum allowed threshold distance.
12. The method according to any of claims 6 to 11, wherein a number of the nearest neighbouring cell sites is predetermined or configurable.
13. The method according to any preceding claim, wherein the threshold distance is different for at least two of the cell sites, or wherein the threshold distance is unique for each of the cell sites.
14. The method according to any preceding claim, further comprising: retrieving locations of the cell sites from a database based on identifiers of the cell sites, wherein each alarm message comprises an identifier of a cell site associated with the alarm message; anddetermining distances between the cell sites based on the locations of the cell sites.
15. The method according to any of claims 4 to 14, wherein the indication of the mass outage comprises identifiers of the cell sites of the group of cell sites, wherein the indication of the mass outage comprises an indication of the particular type of alarm, and / or wherein the indication of the mass outage is provided as an automated service ticket.
16. The method according to any preceding claim, further comprising: receiving a further alarm message of the stream of alarm messages; assigning a cell site associated with the further alarm message to the group of the cell sites, in response to determining that the cell site associated with the further alarm message is located within the threshold distance from at least one of the group of the cell sites.
17. The method according to claim 16, further comprising: detecting a mass outage of another group of the cell sites; and merging the other group of the cell sites with the group of the cell sites, in response to determining that the further cell site is located within the threshold distance from at least one of the other group of the cell sites.
18. An apparatus comprising means for performing the method according to any of claims 1 to 17.
19. A computer program comprising program code configured to, when executed by a processor, cause an apparatus at least to perform the method according to any of claims 1 to 17.