Transmission network fault location method, system, device and storage medium
By building a multi-grained network topology model and reverse hierarchical sorting algorithm, the problem of inefficient fault positioning in traditional networks is solved, and fast and accurate fault positioning is achieved.
Patent Information
- Application Number
- CN202411647133.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2044-11-18
AI Technical Summary
Traditional network failure location relies on manual analysis, is inefficient and susceptible to human factors, making it difficult to cope with complex and changeable network environments.
A multi-grained network topology model is constructed, and the topological relationship between the alarms is analyzed by collecting physical connections, logical connections and business bearer relationship data, and the topological relationship between alarms is analyzed, and the alarm is compressed using graph clustering algorithm, and the root alarm is located through the reverse hierarchical sorting algorithm.
Improves the efficiency and accuracy of fault location, and can quickly lock the most suspicious root cause alarms in massive alarm data, reduces the data scale and provides diagnostic confidence.
Smart Images

Figure CN119520248B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communications, and in particular to a method, system, device, and storage medium for locating transmission network faults. Background Art
[0002] With the continuous expansion of networks and the increasing variety of services, transmission network fault diagnosis and location face unprecedented challenges. Traditional network fault location relies primarily on the experience of operations and maintenance personnel, who manually analyze massive amounts of alarm data to gradually identify the root cause of the problem. This approach is not only inefficient but also susceptible to human error, making it difficult to cope with today's complex and changing network environment. Summary of the Invention
[0003] The present application provides a transmission network fault location method and system for improving the efficiency and accuracy of fault troubleshooting.
[0004] In a first aspect, the present application provides a method for locating a transmission network fault, the method comprising:
[0005] Collect physical connection data, logical connection data, and service bearer relationship data in the transmission network, and build a multi-granularity network topology model based on the physical connection data, logical connection data, and service bearer relationship data;
[0006] Obtain raw alarm data from the transmission network and analyze the topological relationships between different alarms based on the raw alarm data and a multi-granularity network topology model;
[0007] Mining the topological associations between different alarms to obtain horizontal peer associations and vertical upstream and downstream associations between the mined alarm objects. Based on the identification of horizontal peer associations and vertical upstream and downstream associations, the multi-alarm data of each alarm object within a specific time window is determined.
[0008] Constructing a correlated alarm directed graph based on multiple alarm data, and compressing the correlated alarm directed graph based on the correlated alarm directed graph to obtain a compressed alarm set;
[0009] An alarm propagation graph is constructed based on the compressed alarm set and the multi-granularity network topology model. The alarm propagation graph is then queried using a reverse hierarchical sorting algorithm to determine the root alarm.
[0010] By employing the above technical solution, a multi-granularity network topology model was constructed by collecting physical connection data, logical connection data, and service bearer relationship data from the transmission network. This model not only reflects the physical connection relationships between network devices, but also depicts the protocol communication relationships between logical network elements and the service data bearer mapping relationships in the network. Compared with a single physical topology, the multi-granularity network topology model comprehensively expresses the network connection semantics at different semantic granularities, such as device, network element, and service, providing richer contextual information for subsequent refined alarm correlation analysis.
[0011] Based on the multi-granularity network topology model, the present invention performs a topological correlation analysis on the acquired original alarm data. By mapping the alarms with the corresponding entities in the topology model, extracting the topological correlation features of the alarms under multi-granularity semantics, and calculating the topological correlations between different alarms, the semantic dependencies between the alarms are revealed. Compared with traditional methods, the correlation analysis of the present invention is not limited to the alarm content itself, but makes full use of the topological context of the alarm, and mines the horizontal same-level correlations and vertical upstream and downstream correlations between alarms. This multi-dimensional correlation mechanism can deeply characterize the propagation impact of alarms in complex network topologies, and provide more comprehensive semantic constraints for fault tracing.
[0012] By introducing a time window, the present invention can identify multiple alarm sequences for each alarm object. These multiple alarm data are constructed into a directed graph of associated alarms, which intuitively expresses the causal semantics between alarms. On this basis, an innovative associated alarm compression method is proposed. This method uses a graph clustering algorithm to discover the associated cluster structure between alarms, and through homogeneity analysis and redundant filtering, it realizes double-layer compression of intra-cluster and inter-cluster alarms, while retaining the key alarm semantics to the maximum extent, greatly reducing the data scale for subsequent processing. Compared with existing alarm compression methods, the present invention fully considers the global correlation of the graph structure, and the accuracy and efficiency of alarm compression are higher.
[0013] Finally, the compressed alarm set is mapped onto a multi-granularity network topology to construct an alarm propagation graph. This graph model intuitively depicts the propagation paths and diffusion range of alarms within the network. Based on this, a reverse hierarchical sorting algorithm was designed. This algorithm leverages the hierarchical semantics of the multi-granularity topology. By tracing back the key propagation paths, starting from the leaf alarm nodes, it recursively selects the root cause alarms with the greatest propagation strength. Compared to forward search, reverse sorting can quickly identify the most suspicious root cause alarms within massive amounts of alarm data and assign diagnostic confidence levels through a scoring mechanism, significantly improving the efficiency and accuracy of fault location.
[0014] In a second aspect of the present application, a transmission network fault location system is provided, the system comprising:
[0015] A multi-granularity network topology model building module is used to collect physical connection data, logical connection data, and service bearer relationship data in the transmission network, and build a multi-granularity network topology model based on the physical connection data, logical connection data, and service bearer relationship data;
[0016] The topology correlation analysis module is used to obtain the original alarm data in the transmission network and analyze the topology correlation between different alarms based on the original alarm data and the multi-granularity network topology model;
[0017] The multi-alarm data determination module is used to mine the topological associations between different alarms, obtain the horizontal peer associations and vertical upstream and downstream associations between the mined alarm objects, and determine the multi-alarm data of each alarm object within a specific time window based on the identification of horizontal peer associations and vertical upstream and downstream associations;
[0018] The associated alarm compression module is used to construct an associated alarm directed graph based on multiple alarm data, and compress the associated alarm directed graph based on the associated alarm directed graph to obtain a compressed alarm set;
[0019] The root alarm determination module is used to construct an alarm propagation graph based on the compressed alarm set and the multi-granularity network topology model, and query the alarm propagation graph through a reverse hierarchical sorting algorithm to determine the root alarm.
[0020] In a third aspect of the present application, a computer storage medium is provided. The computer storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the above method steps.
[0021] In the fourth aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface and a network interface, the memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs the above method.
[0022] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0023] 1. By collecting physical connection data, logical connection data, and service bearer relationship data from the transmission network, a multi-granularity network topology model is constructed. This model not only reflects the physical connection relationships between network devices, but also depicts the protocol communication relationships between logical network elements and the service data bearer mapping relationships in the network. Compared with a single physical topology, the multi-granularity network topology model comprehensively expresses the network connection semantics at different semantic granularities, such as device, network element, and service, providing richer context for subsequent refined alarm correlation analysis.
[0024] 2. This application first performs baseline calibration on the real-time response signal data to obtain preliminary response signal data. Baseline calibration is a critical step in signal processing. Its purpose is to eliminate interference such as DC components and drift in the signal, ensuring that the signal consistently fluctuates around zero. Through baseline calibration, interference such as low-frequency noise and baseline drift in the real-time response signal data is effectively suppressed, expanding the signal's dynamic range and significantly improving the signal-to-noise ratio. The preliminary response signal data is "purer" than the original data, laying a good foundation for subsequent processing. Secondly, the preliminary response signal data is filtered to obtain the target real-time response signal. Filtering is a key method for further improving signal quality based on baseline calibration. Through filtering, impurities such as high-frequency noise and random interference remaining in the preliminary response signal data are "filtered out," enhancing the smoothness and continuity of the signal. The target real-time response signal is the "result" of filtering. It faithfully reflects the true response characteristics of the measured object and provides high-quality data support for subsequent feature extraction, pattern recognition, and other analyses.
[0025] 3. This application clusters the directed graph of associated alarms through a graph clustering algorithm to obtain multiple associated clusters. Graph clustering is an important method for graph data analysis. Its purpose is to divide closely associated nodes into the same cluster based on the association relationship between nodes, forming several internally compact and externally isolated associated clusters. Through graph clustering, the directed graph of associated alarms is divided into multiple "small groups". The correlation between members within each group is strong, and the correlation between different groups is weak. This process is like "birds of a feather flock together" for alarms, making subsequent intra-cluster compression and inter-cluster compression more targeted. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 A schematic diagram of a flow chart of a transmission network fault location method provided in an embodiment of the present application;
[0027] Figure 2 An architectural diagram of a transmission network fault location system provided in an embodiment of the present application;
[0028] Figure 3 This is a schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments.
[0030] In the description of the embodiments of this application, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a concrete manner.
[0031] In the description of the embodiments of the present application, the term "multiple" means two or more. For example, multiple devices refer to two or more devices, and multiple screen terminals refer to two or more screen terminals. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized.
[0032] In order to facilitate understanding of the method and system provided by the embodiments of the present application, before introducing the embodiments of the present application, the background of the embodiments of the present application is first introduced.
[0033] In today's rapidly developing information age, networks have become critical infrastructure supporting every industry. With the continuous expansion of network scale and the increasing diversity of services, the complexity of transmission networks is growing exponentially. In such a complex network environment, failures seem to be commonplace. However, network failures not only cause service interruptions, but also result in significant economic losses and even negatively impact social stability. Therefore, quickly and accurately diagnosing and locating network failures has become a daunting challenge for operations and maintenance personnel.
[0034] Traditional network fault location has long relied primarily on the professional experience and intuition of operations personnel. When a network failure occurs, operators must manually collect and analyze massive amounts of alarm data, relying on their personal experience to determine the possible cause and gradually troubleshoot the root cause. While this manual diagnostic approach can solve problems to a certain extent, its low efficiency and high labor costs make it difficult to meet the current needs of network operations and maintenance.
[0035] After the background introduction of the above content, those skilled in the art can understand the problems existing in the prior art. The technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.
[0036] On this basis, the present invention provides a method for locating a transmission network fault. Figure 1 , Figure 1 This is a flow chart of a transmission network fault location method provided in an embodiment of the present application. The method can be implemented by a computer program or run as an independent tool application. The method includes:
[0037] S101, collecting physical connection data, logical connection data, and service bearer relationship data in the transmission network, and building a multi-granularity network topology model based on the physical connection data, logical connection data, and service bearer relationship data;
[0038] Specifically, to comprehensively characterize the connection semantics of a transport network, the present invention first collects physical connection data, logical connection data, and service bearer relationship data from the transport network. Physical connection data reflects the physical link connections between network devices, such as optical fibers and cables. Logical connection data represents the logical communication relationships between network devices based on specific protocols, such as routing relationships at the IP layer and tunnel relationships at the MPLS layer. Service bearer relationship data describes the mapping relationship between network service data and network entities, such as the network nodes and links through which service data flows.
[0039] After completing data collection, the present invention constructs a multi-granularity network topology model based on physical connection data, logical connection data, and service bearer relationship data. Specifically, the underlying physical topology map is first generated based on the physical connection data, with nodes representing network devices and edges representing physical links between devices. Then, based on the physical topology map, logical connection relationships are added based on the logical connection data to form a logical topology map. Nodes in the logical topology map represent logical network elements on network devices, such as router interfaces at the IP layer and LSP endpoints at the MPLS layer. Edges in the logical topology map represent communication relationships between logical network elements, such as routes at the IP layer and tunnels at the MPLS layer. Finally, based on the service bearer relationship data, service data flows are mapped within the logical topology map to form a service bearer topology map. The service bearer topology map intuitively displays the transmission paths and bearer relationships of service data across logical network elements and physical devices. By integrating information from multiple dimensions—physical connections, logical connections, and service bearer—the multi-granularity network topology model forms a comprehensive, three-dimensional view of network connectivity.
[0040] Based on the above embodiment, as an optional embodiment, a multi-granularity network topology model is constructed based on physical connection data, logical connection data, and service bearer relationship data, including:
[0041] S201, constructing a physical connection topology based on physical connection data;
[0042] Specifically, the physical connection topology reflects the actual physical connection relationships between network devices and represents the lowest-level view of the network's infrastructure. The specific process for constructing the physical connection topology is as follows: First, key elements from the physical connection data are extracted, including device name, device type, interface information, and link information. Next, an undirected graph is constructed, with devices as nodes and the physical links between them as edges. Attributes for each node include device name, device type, and computer room location, while attributes for each edge include link bandwidth, link type, and fiber length. Next, connectivity analysis is performed on the constructed undirected graph to identify physically connected subnets within the network. Subnets are then classified into hierarchies based on characteristics such as subnet size and node type distribution. Finally, the physical connection topology is optimized, taking into account factors such as network scale, topology, and geographic distribution, to intuitively and systematically present the hierarchical structure of the network's physical components.
[0043] S202, constructing a logical connection topology based on the logical connection data;
[0044] Specifically, the device nodes in the physical connection topology are first divided into several logical device groups based on logical rules such as location proximity, functional similarity, and role consistency. The physical devices within each logical device group are logically treated as a single entity. Then, key attributes (such as bandwidth, number of ports, protocol type, and IP address range) of the physical device nodes within each group are extracted and aggregated to generate a logical network element with comprehensive attributes. Each logical network element corresponds to a logical device group and inherits the underlying characteristics of the physical device nodes within that group. This abstraction reduces the number of physical device nodes into a single logical unit, significantly simplifying the granularity of network representation. Next, the connectivity between different logical network elements is analyzed. If there are hierarchical physical links between the physical devices contained in two logical network elements, a logical connection is established between the two logical network elements. The newly generated logical connection carries the aggregated traffic of these physical links, and its attributes (such as bandwidth) are determined by the superposition of the underlying physical link attributes. This step achieves a high-level description of the interconnection between network areas. Furthermore, considering the hierarchical distribution of logical network elements in the network, logical network elements are divided into different layers, such as the core layer, aggregation layer, and access layer, forming a hierarchical topology of the logical network. This hierarchical semantics helps to understand the functional positioning and boundaries of each part of the network.
[0045] S203, constructing a service bearer relationship topology based on the service bearer relationship data;
[0046] Specifically, we first aggregate and collect metadata from each service node in the network, including key attributes such as service name, type, user group, service content, and performance requirements. This metadata outlines the characteristics of each service. We then conduct in-depth analysis of the interactions between different services to extract data on service dependencies. Specifically, we analyze the resource demands that a service places on other services to implement its own services, as well as the services that a service provides to other services. This data reflects the service supply and consumption relationships between services and reveals the directionality of service interactions.
[0047] Based on service attribute metadata and service dependency data, services are abstracted into topological nodes, and the bearer relationships between services are abstracted into directed edges. This constructs a service bearer relationship topology. In this topology, the starting point of each directed edge represents the service provider, and the end point represents the service demander. The edge weight reflects the strength of the service dependency. Key parameters of the service bearer relationship topology are further analyzed and calculated. For example, the in-degree and out-degree of each service node are calculated. The in-degree reflects the service's dependence on external services, while the out-degree reflects the popularity of the services it provides. The betweenness centrality of each service node is calculated to reflect the criticality of the service in the overall service network. Quality attributes such as latency, packet loss rate, and jitter are calculated for each service link to characterize the health of the service interaction pathway. Next, the service bearer relationship topology is used to analyze the propagation of service anomalies. When a service node or link fails, the propagation path of the failure's impact in the topology is deduced, and the affected related services are identified. Targeted anomaly demarcation and isolation measures are then implemented to curb the spread of the failure and ensure the continuity of core services.
[0048] S204 , vertically integrating the physical connection topology, the logical connection topology, and the service bearer topology to obtain a multi-granularity network topology model.
[0049] Specifically, the present invention builds upon the physical connection topology, logical connection topology, and service bearer topology by further vertically integrating these three topologies, ultimately resulting in a multi-granular network topology model. This multi-granular network topology model aims to bridge the physical, logical, and service perspectives of network management, forming a panoramic information map of the network and enabling full-dimensional visualization, global awareness, and a closed-loop process for network operations and maintenance.
[0050] S102, obtaining original alarm data in the transmission network, and analyzing the topological association between different alarms based on the original alarm data and a multi-granularity network topology model;
[0051] Specifically, after constructing a multi-granular network topology model, the present invention further acquires raw alarm data from the transmission network. Raw alarm data is log information automatically generated by network equipment or monitoring systems when an anomaly is detected. It contains various attributes such as the time of alarm occurrence, alarm source, alarm type, alarm level, and alarm content. This alarm data contains important clues to the occurrence of network failures and is the basis for fault diagnosis.
[0052] After obtaining the original alarm data, the present invention analyzes the topological association relationship between different alarms based on the original alarm data and the multi-granularity network topology model. Specifically, first, according to the alarm source information in the alarm data, each alarm is mapped to the corresponding network entity in the multi-granularity network topology model to determine the location coordinates of the alarm in the physical layer, logical layer and business layer. Then, based on the multi-granularity network topology model, the topological features of the alarm at different semantic granularities are extracted, such as the connection relationship of the alarm source in the physical topology, the dependency relationship in the logical topology, the transmission relationship in the business topology, etc. Next, the topological correlation between different alarms is calculated, considering the semantic association of the alarm source in multiple dimensions such as physical topology, logical topology, and business topology, and characterizing the horizontal peer relationship and vertical upstream and downstream relationship between alarms.
[0053] By analyzing the topological relationships between alarms, this paper leverages the semantic properties of alarm data and the structural semantics of network topology to reveal the mutual influence and propagation patterns of alarms in complex network environments. Compared to methods that correlate alarms based solely on alarm content, topological correlation analysis considers the network environment and contextual semantics of the alarms. This allows for a deeper understanding of the inherent connections between alarms and uncovers correlation patterns hidden within massive amounts of alarm data.
[0054] Based on the above embodiment, as an optional embodiment, the topological association relationship between different alarms is analyzed based on the original alarm data and the multi-granularity network topology model, including:
[0055] S301: Extract the original alarm data attributes to obtain key alarm attributes. Based on the key alarm attributes, associate each alarm with the topological entity of the corresponding granularity level in the multi-granularity network topology model to form an "alarm-topology entity" mapping table.
[0056] Specifically, based on the multi-granularity network topology model constructed in this invention, the topological relationships between different alarms are further analyzed based on the original alarm data and the multi-granularity network topology model. Analyzing the topological relationships of alarms aims to reveal the network semantic basis for the generation of alarms and explore the role of alarms in the overall network. This allows for a more comprehensive understanding of the essential meaning of alarms and more accurate traceability analysis and impact assessment of alarms. This is a key step in achieving the transition from alarm management to network management.
[0057] The first step in analyzing the topological relationships between different alarms is to extract attributes from the raw alarm data to obtain key alarm properties. Specifically, key alarm metadata, such as the time the alarm was generated, the network element where the alarm originated, the alarm name, the alarm severity, and the alarm description, are identified and filtered from the massive and complex alarm information stream. These key attributes outline the basic structure of each alarm and serve as the foundation for further correlation analysis.
[0058] Once key alarm attributes are extracted, the multi-granular network topology model can be used to associate each alarm with topological entities at the corresponding granularity level. Specifically, key alarm attributes are analyzed to determine whether the alarm is caused by a physical, logical, or service layer fault. The alarm is then mapped to the corresponding network element node or link in the physical connection topology, logical connection topology, or service bearer topology. This semantic association of "alarm-topology entity" imbues each alarm with network topology semantics, establishing a direct mapping between the physical, logical, and service worlds of the network.
[0059] When implementing the mapping between alarms and topological entities, it's important to fully utilize the unified mapping identification mechanism within the multi-granular network topology model. Each alarm's key attributes often carry identification information that indicates its network affiliation, such as device name, port number, IP address, VLAN ID, subnet ID, and service ID. This identification information corresponds to the mapping identifiers for each node and link in the multi-granular network topology model. By comparing and analyzing the alarm attribute identifiers with the topological mapping identifiers, alarms can be quickly and accurately associated with the corresponding topological entities.
[0060] After the mapping between alarms and topological entities is complete, an "alarm-topological entity" mapping table is generated. This table records the topological node or link to which each alarm belongs, as well as the topological granularity level at which this affiliation occurs. By querying the "alarm-topological entity" mapping table, you can quickly locate the "identity coordinates" of each alarm in the physical, logical, and business worlds, tracing its network semantic origin.
[0061] S302, traversing each alarm in the "alarm-topology entity" mapping table to extract the associated features of each alarm in the multi-granularity topology environment to obtain an alarm topology feature vector;
[0062] Specifically, based on the "alarm-topology entity" mapping table, the present invention further traverses each alarm in the table, extracting the associated features of each alarm in a multi-granular topological environment, and ultimately obtaining an alarm topology feature vector. The purpose of extracting the alarm topology feature vector is to depict the network semantics of the alarm from a multi-dimensional perspective, quantify the status of the alarm in the global network, and lay the mathematical foundation for in-depth analysis of the internal connections between different alarms.
[0063] The specific steps for extracting alarm topology feature vectors are as follows: First, traverse the "alarm-topology entity" mapping table and analyze the topological affiliation of each alarm one by one. By querying the mapping table, the topological granularity level (physical layer, logical layer, or service layer) to which the current alarm is mapped is determined, as well as the specific topological node or link to which it belongs.
[0064] After determining the topological location of an alarm, we can extract its associated features from the corresponding granularity of the network topology. These features fall into two categories: First, the attributes of the alarm itself, such as its level, type, occurrence time, and duration, which characterize its severity and urgency. Second, the attributes of the topological entity to which the alarm belongs, such as its centrality, clustering coefficient, and load capacity, which reflect the importance of the network location where the alarm occurred.
[0065] When extracting alarm attributes, focus on the OID (Object Identifier) information generated by the alarm. The OID uniquely identifies the alarm and carries many of its key attributes. By parsing the alarm OID, you can accurately extract characteristic parameters such as the alarm's level, type, and occurrence time.
[0066] When extracting the attribute characteristics of the entity to which the alarm belongs, we fully utilize the multi-granularity network topology model to perform multi-level, multi-perspective topological calculations. For example, in the physical connection topology, we calculate the degree centrality of the device node where the alarm is located to assess the connectivity importance of this device in the network. In the logical connection topology, we calculate the betweenness centrality of the logical link where the alarm is located to assess the critical influence of this logical link on the entire network operation. In the service bearer topology, we analyze the impact of the health status of the service link where the alarm is located on the service level of the related core business. By comprehensively evaluating the correlation characteristics of alarms at different topological granularities, we can fully reflect the network semantic status of the alarm.
[0067] After extracting the associated features of an alarm, an alarm topology feature vector can be constructed. From a mathematical perspective, an alarm topology feature vector is a multidimensional quantity, with each dimension corresponding to a topologically associated feature parameter of the alarm. For example, the feature vector of an alarm can be represented as (Level = Important, Type = Link Down, Occurrence Time = 2022-01-01 12:00:00, Physical Connectivity = 5, Logical Betweenness = 0.8, Service Health = 0.6, ...). Through this vectorized representation, alarms, which were originally semantically ambiguous, are imbued with rigorous mathematical meaning. Mathematical operations such as similarity calculation and cluster analysis can be performed on the feature vectors of different alarms, revealing the inherent connections between alarms.
[0068] S303 : Based on the alarm topology feature vectors, calculate the topological semantic correlation between different alarms, and determine the topological association relationship between the different alarms based on the correlation.
[0069] Specifically, after extracting the topological feature vectors for each alarm, the present invention further calculates the topological semantic correlations between different alarms based on these topological feature vectors and, based on these correlations, determines the topological associations between them. Analyzing the topological associations between different alarms aims to reveal the underlying connections between seemingly independent alarms and identify the culprits behind the alarm surge. This allows for a comprehensive and coordinated approach to addressing these alarms from a global perspective, addressing their root causes and avoiding patchy, localized treatments.
[0070] The specific steps for calculating the topological semantic relevance between different alarms are as follows: First, based on the alarm topological feature vectors constructed above, calculate the similarity between different alarm feature vectors. Common similarity calculation methods include Euclidean distance, cosine similarity, and Pearson correlation coefficient. Among them, Euclidean distance describes the proximity of two vectors in spatial position. The smaller the distance, the more similar the two alarms are in topological features; cosine similarity describes the consistency of the two vectors in direction. The greater the similarity, the more consistent the feature distribution trends of the two alarms; and the Pearson correlation coefficient describes the degree of linear correlation between the two vectors. The larger the coefficient, the positive correlation between the features of the two alarms.
[0071] In actual calculations, a combination of multiple similarity metrics is used to comprehensively assess the topological semantic relevance of alarms. For example, if the spatial distance between alarms A and B is 0.1, the directional similarity is 0.9, and the correlation coefficient is 0.8, then alarms A and B can be determined to have a high level of topological relevance. Their roles in the network topology are very similar, and their occurrence times, causes, and impacts are likely similar. Conversely, if the topological characteristics of two alarms differ significantly, it means that their locations in the network are far apart, and the likelihood of their association is low.
[0072] When calculating alarm similarity, we must fully leverage the hierarchical nature of multi-granularity network topologies and examine the correlation of features at different granularities. The lower the granularity level, the higher the weight of similarity should be. For example, if two alarms have very similar features in the physical layer topology, relatively similar features in the logical layer topology, but significantly different features in the service layer topology, the similarity of the physical layer topology features should be prioritized. This is because the network's physical connections are the foundation of logical connections and service bearer, and correlations at the physical layer can often be transmitted to upper-layer networks. Therefore, physical topology similarity should be the primary basis for determining alarm relevance.
[0073] By calculating the similarity of the alarm topology feature vectors, we can characterize the semantic relevance of different alarms in terms of network semantics. Generally speaking, alarm pairs whose feature similarity exceeds a preset threshold are considered to have significant topological correlations. These correlated alarms are located close together in the network topology, and their fault causes and impacts are likely to be the same. These correlated alarms require coordinated analysis and joint processing to identify the "culprit" behind the incidents and avoid multiple, erratic, and inconsistent responses.
[0074] S103, mining the topological associations between different alarms to obtain horizontal peer associations and vertical upstream and downstream associations between the mined alarm objects, and determining multiple alarm data for each alarm object within a specific time window based on the identification of the horizontal peer associations and vertical upstream and downstream associations;
[0075] Specifically, based on a multi-granular network topology model, an adjacency matrix is constructed between network entities, characterizing the connectivity between them at the physical, logical, and business layers. Alarm data is then mapped into the adjacency matrix to form an association matrix between alarm objects. Next, association rule mining algorithms, such as Apriori and FP-Growth, are applied to discover frequently co-occurring alarm association patterns in the association matrix and calculate their support and confidence. Support represents the frequency of the association pattern across all alarms, while confidence represents the conditional probability that, given the presence of one alarm, another alarm will also occur. By setting thresholds for support and confidence, strongly associated alarm patterns are screened, forming a candidate set of horizontal sibling associations and vertical upstream and downstream associations. Finally, based on the network topology and business application semantics, the candidate associations are further evaluated and filtered to remove redundant or invalid associations, resulting in the final set of horizontal sibling associations and vertical upstream and downstream associations.
[0076] By exploring horizontal peer-level and vertical upstream-downstream correlations between alarm objects, we can gain a deeper understanding of alarm propagation patterns and impact spheres within the network topology. Horizontal peer-level correlations reveal the mutual influence and causal relationships between alarms at the same level, helping to identify large-scale, batch-prone faults of the same type. Vertical upstream-downstream correlations depict the dependencies and transmission relationships between alarms at different levels, helping to trace the propagation paths of faults across various network layers. Operations and maintenance personnel can use this correlation information to quickly determine the severity and impact of faults, identify the common root causes of batch alarms, and infer the fault propagation path within the network, enabling them to implement targeted fault isolation and remediation measures.
[0077] After identifying the associations between alarm objects, the present invention further identifies multiple alarm data for each alarm object within a specific time window based on horizontal peer associations and vertical upstream and downstream associations. By setting a sliding time window, all alarm data falling within that window is collected, and the alarms are clustered using association information, multiple alarms with strong associations are aggregated into a multi-alarm sequence for a single alarm object. This step effectively compresses the size of alarm data, transforming scattered alarm entries into a compact alarm object representation while preserving the alarm's temporal characteristics and association semantics.
[0078] It should be noted that "multi-alarm data" refers to a set of alarm data associated with a specific alarm object within a specific time window. This alarm data is obtained by mining horizontal peer relationships and vertical upstream and downstream relationships between alarm objects.
[0079] Specifically, horizontal peer correlation refers to the association between different alarm objects within the same layer of the network topology. For example, within a transmission network layer (such as the IP layer or optical layer), if two network elements are connected and both generate alarms within a similar timeframe, a horizontal peer correlation may exist between these two alarm objects. Identifying horizontal peer correlations helps discover related alarms within the same layer and reveals the scope of a fault within a local area.
[0080] Vertical upstream and downstream correlation refers to the association between alarm objects at different layers of the network topology. For example, if an outage alarm occurs on a port on the underlying optical network, an outage alarm may also occur on a port on the IP network router that carries it, creating a vertical upstream and downstream correlation between the two. Identifying vertical upstream and downstream correlations helps discover cross-layer alarm propagation paths and reveals the vertical impact of a fault.
[0081] By mining horizontal peer associations and vertical upstream and downstream associations, we can identify other alarm objects closely related to a particular alarm object, and then determine multiple alarm data for that alarm object within a specific time window (such as the period before and after a fault occurs). This multi-alarm data reflects the temporal correlation between the alarm object and its related alarms, helping to fully understand the fault context in which the alarm object exists.
[0082] For example, suppose within a specific time window, association mining identifies alarm object A as having horizontal, peer-level associations with alarm objects B, C, and D, and vertical, upstream-downstream associations with alarm objects E and F. Then, A's multi-alarm data includes A's own alarms as well as the related alarms of B, C, D, E, and F. Together, these alarm data constitute A's "family tree," reflecting A's "fault family tree" and providing rich clues for in-depth analysis of the causes and impacts of A's faults.
[0083] S104, constructing a correlated alarm directed graph based on the multiple alarm data, and compressing the correlated alarm directed graph based on the correlated alarm directed graph to obtain a compressed alarm set;
[0084] Specifically, first, the multiple alarm sequences of each alarm object are represented as a node in the associated alarm directed graph. Then, the association relationships between the alarm sequences of different alarm objects are analyzed, including horizontal peer associations and vertical upstream and downstream associations. For horizontal peer associations, that is, the concurrent or sequential relationship between alarms between different network elements, undirected edges are added between corresponding nodes to represent the peer impact between these alarm objects. For vertical upstream and downstream associations, that is, the causal or transitive relationship between alarms at different network levels, directed edges are added between corresponding nodes, with the direction of the edge pointing from the cause alarm to the result alarm, indicating the propagation path of the alarm in the network. The weight value of the edge is calculated based on the association strength or the degree of dependence. The greater the association strength, the higher the weight value. In this way, the scattered multi-alarm data is integrated into a compact associated alarm directed graph, which fully preserves the semantic characteristics and temporal characteristics of the alarm data, while revealing the complex topological associations and logical dependencies between alarms.
[0085] Based on the above embodiment, as an optional embodiment, constructing a directed graph of associated alarms based on multiple alarm data includes:
[0086] S401, extracting nodes from multiple alarm data to obtain alarm nodes;
[0087] Specifically, we traverse multiple alarm records and analyze the basic information of each one. Alarm records typically contain key fields such as the alarm's unique identifier, alarm name, alarm severity, alarm description, alarm occurrence time, and alarm clearing time. This information acts as an "identity card" for the alarm, providing crucial clues for identifying and distinguishing different alarms.
[0088] When analyzing alarm records, we use the alarm ID as a unique identifier to extract distinct alarm instances. Each alarm instance is considered a basic node in the directed graph of associated alarms. Node attributes, including the alarm ID, alarm name, and alarm severity, serve as node "labels" for subsequent association analysis and graph algorithm design.
[0089] To ensure the accuracy and completeness of node extraction, the raw alarm data must be cleaned and normalized. Alarm data cleaning primarily addresses "dirty data" such as noise, missing values, and outliers in alarm records, improving data quality. Common cleaning methods include field completion, outlier removal, and duplicate merging. Alarm data normalization extracts the format and content of alarms, eliminating heterogeneous differences between different alarm sources and creating a unified representation of alarm instances.
[0090] Through the above process, each alarm instance in the multi-alarm data is extracted as an independent node. These nodes are the basic "building blocks" of the directed graph of related alarms and contain rich semantic information.
[0091] S402, performing node association on the alarm nodes based on the multiple alarm data to obtain association edges between the alarm nodes;
[0092] Based on the alarm nodes obtained above, the association strength between different nodes is calculated. Association strength characterizes the semantic relevance between two nodes and reveals the influence of one node on another. The calculation of association strength comprehensively considers the multidimensional attributes of nodes, such as alarm level, alarm type, alarm frequency, and alarm clearing time, reflecting the static characteristics of node associations. Furthermore, the temporal evolution of nodes is analyzed, such as the order in which alarms occur, the duration of alarms, and the occurrence cycle of alarms, thereby depicting the dynamic characteristics of node associations.
[0093] For example, if the alarm level for node A is higher than that for node B, and the alarm for node A occurs earlier than that for node B, it can be inferred that node A may be the cause or precursor of node B, and the correlation between the two nodes is strong. For another example, if the alarm types for nodes C and D are the same, with similar frequency and duration, it can be inferred that the correlation between nodes C and D is strong, possibly due to a common root cause. Of course, determining node associations also relies on expert experience and domain knowledge. Simple data calculations can only provide a reference; human intelligence is required to train, calibrate, and interpret the association model.
[0094] When calculating the node association strength, the present invention innovatively introduces network topology features. The above steps have obtained the topological feature vector and topological correlation of each alarm, and these factors are also included in the calculation of the association strength. Intuitively speaking, if two nodes are close in position in the network topology, the connected devices and links are similar, and the services and customers involved are highly overlapping, then their association strength will be greatly improved. This is because the network topology is a carrier where elements such as devices, links, protocols, and services are interwoven, and the associations in the topology can often transmit and amplify the semantic connections between nodes. The topological perspective allows the analysis of node associations to stand on the global commanding heights and see the "barrel effect" behind local alarms.
[0095] By comprehensively considering the static properties, dynamic characteristics and topological factors of the alarm nodes, the association strength between nodes can be comprehensively evaluated.
[0096] S403: Construct an associated alarm directed graph based on the associated edges between the alarm nodes and the alarm nodes.
[0097] Specifically, first, using alarm nodes as the basic elements and associated edges as the connecting links, we apply classic definitions from graph theory to standardize the mathematical form of the associated alarm directed graph. Let V be the set of alarm nodes, E be the set of associated edges, and G = (V, E). Here, V = {v1, v2, ..., vn}, where vi represents the i-th alarm node and n is the total number of nodes; E = {e1, e2, ..., em}, where ej represents the j-th directed edge and m is the total number of edges. Each edge ej has a starting point and an end point, corresponding to the two nodes it connects, and the direction of the edge is from the starting point to the end point. Thus, through elements such as nodes, edges, directions, and weights, the mathematical model of the associated alarm directed graph is rigorously defined. Secondly, starting from the alarm nodes and associated edges, the data structure of the associated alarm directed graph is constructed. Common graph data structures include adjacency matrices and adjacency lists. An adjacency matrix is a two-dimensional array whose rows and columns correspond to the nodes of the graph. Matrix elements indicate whether there are edges connecting the nodes and the weights of the edges. An adjacency list is a linked list array, where each element represents a node, and the edges connected to that node form a linked list. The advantages and disadvantages of these two data structures depend on the scale and sparsity of the graph. When the number of nodes is large and the number of edges is relatively small, the adjacency list is more memory-efficient; when the graph is dense, the adjacency matrix provides a more natural representation. Finally, based on the defined data structure, the extracted alarm nodes and associated edges are loaded into it, thereby constructing an in-memory directed graph of associated alarms.
[0098] Based on the above embodiment, as an optional embodiment, the associated alarm directed graph is compressed based on the associated alarm directed graph to obtain a compressed alarm set, including:
[0099] S601, clustering the directed graph of associated alarms using a graph clustering algorithm to obtain multiple associated clusters;
[0100] Specifically, starting from the graph's topological structure, we examine the closeness of nodes within the graph. Closeness can be measured using metrics such as node degree, inter-node distance, and node betweenness. Generally speaking, nodes within the same cluster tend to have a high degree of closeness; they are closer together in the graph and have denser connections. Based on this prior knowledge, clustering methods based on connected subgraphs, such as the well-known Girvan-Newman algorithm, can be employed.
[0101] S602, analyzing the internal alarms of each associated cluster, determining the semantic redundancy of the internal alarm nodes, and determining compressible homogeneous alarms in the internal alarms based on the semantic redundancy;
[0102] Specifically, a semantic analysis is performed on the internal alarm nodes to determine whether there is redundancy in the semantic information in the internal alarm nodes. Semantic redundancy refers to the situation where the semantic meanings of network faults reflected by multiple alarm nodes are repeated or similar. The purpose of performing semantic redundancy analysis is to find compressible homogeneous alarms. Then, based on the results of the semantic redundancy analysis, compressible homogeneous alarms in the internal alarms are determined. Homogeneous alarms refer to alarms with the same or similar semantic information. These alarms reflect faults of the same type or at the same location in the network and can be compressed without losing the key information required for fault location. The method for determining compressible homogeneous alarms is to extract semantic features for each alarm node, calculate the semantic similarity between different alarm nodes, and if the semantic similarity is higher than a threshold, determine these alarm nodes as compressible homogeneous alarms.
[0103] S603, compressing internal alarms of the associated cluster based on the compressible homogeneous alarms to obtain a compressed simplified alarm sequence within the cluster;
[0104] Specifically, a compressible set of homogeneous alarms is obtained for each associated cluster. These homogeneous alarms are highly semantically similar and carry a large amount of overlapping information, making them safe for compression. The goal of compression is to minimize information loss while maximizing the number of alarms, thereby reducing the complexity of alarm processing.
[0105] Next, cluster compressible homogeneous alarms. Because homogeneous alarms share similar semantic features, clustering can be highly efficient. Common clustering algorithms such as K-Means and DBSCAN can be applied. The goal of clustering is to group homogeneous alarms with high semantic redundancy and strong substitutability into the same cluster. Each cluster can be represented by its central alarm. The number of clusters can be adaptively determined based on the scale and distribution of homogeneous alarms, or an upper threshold can be set.
[0106] Then, for each homogeneous alarm cluster, the central alarm is selected as the representative of the cluster, forming a compressed and streamlined alarm sequence within the cluster. The central alarm is the alarm with the highest average similarity to other alarms in the cluster and best represents the semantic characteristics of the cluster. The process of selecting the central alarm is actually to further refine the essence of the homogeneous alarm cluster and remove relatively marginal alarms. Through this step, each homogeneous alarm cluster is condensed into a single, most representative alarm that carries the core semantic information of the cluster.
[0107] Finally, the condensed alarm sequences of all associated clusters are merged in the order of timestamps to form the final compressed alarm sequence of the associated cluster.
[0108] S604, performing homogeneity condition analysis on each correlation cluster, determining homogeneous correlation clusters among multiple correlation clusters, and compressing the homogeneous correlation clusters to obtain an aggregated alarm cluster;
[0109] First, we need to extract the feature vector for each cluster. This feature vector integrates information from multiple dimensions, such as the cluster's size, duration, alarm type distribution, and alarm attribute distribution. These features reflect the cluster's aggregation patterns across time, space, type, and attributes. Using aggregation functions, statistical indicators, and distribution fitting, we can encode the cluster's multidimensional features into a compact vector representation, paving the way for subsequent homogeneity analysis.
[0110] Based on the feature vectors of the associated clusters, the similarity between them is calculated. Common similarity metrics include Euclidean distance, cosine similarity, and the Jaccard coefficient. Similarity quantifies the proximity of associated clusters in a multidimensional feature space. Higher similarity indicates that the failure modes and business scenarios represented by the associated clusters are likely to be highly correlated. The similarity matrix can be used to characterize the homogeneity of the entire set of associated clusters.
[0111] Homogeneous clusters are identified based on a preset homogeneity threshold. The homogeneity threshold comprehensively considers the feature similarity and business relevance of clusters. Clusters that meet both the feature similarity threshold (e.g., 0.8) and the business relevance requirement (e.g., "same business system") are identified as homogeneous. These clusters exhibit high consistency in failure modes and business semantics and can be further compressed and merged. Setting the homogeneity threshold requires a balance between compression rate and information fidelity, aiming to maximize the merging of homogeneous clusters while avoiding the loss of critical information due to over-compression.
[0112] Identified homogeneous clusters are compressed to form aggregated alarm clusters. The compression strategy can select the central cluster of a homogeneous cluster as the representative cluster, or generate a new cluster to integrate the characteristics of the homogeneous cluster. The alarm sequence of the aggregated alarm cluster is obtained by merging the simplified alarm sequences of the homogeneous clusters. The timestamp can be the earliest or latest timestamp of the homogeneous cluster. At this point, the original massive alarm data is reduced to a small number of highly summarized and information-dense aggregated alarm clusters, completing the alarm reduction process. Compression of homogeneous clusters can be performed recursively, building a cluster structure in a fragmented manner through multiple rounds and at multiple granularities.
[0113] S605 , integrating the compressed simplified alarm sequence within the cluster with the aggregated alarm cluster to obtain a compressed alarm set.
[0114] Specifically, a compressed, condensed alarm sequence is obtained within each association cluster. These condensed alarm sequences are obtained by compressing homogeneous alarms within the association clusters, with each association cluster corresponding to a condensed alarm sequence. The condensed alarm sequence retains the core alarm information within the association cluster, eliminates redundant homogeneous alarms, and significantly reduces the number of alarms. Next, the aggregated alarm cluster formed in step S604 is obtained. An aggregated alarm cluster is obtained by compressing multiple homogeneous association clusters, and each aggregated alarm cluster can be considered a representative of multiple homogeneous association clusters. Aggregated alarm clusters achieve alarm simplicity across association clusters, aggregating homogeneous alarm patterns scattered across different association clusters into a high-level alarm cluster. Aggregated alarm clusters further refine the core patterns of the alarm data and embody a higher-level fault summary. The condensed alarm sequences and aggregated alarm clusters are then integrated in timestamp order to form a compressed alarm set. Specifically, the condensed alarm sequence of each association cluster is traversed and inserted into the compressed alarm set in timestamp order. For aggregated alarm clusters, select their timestamps (either the earliest or the latest) and insert them into the compressed alarm set in timestamp order. This timestamp alignment allows the compressed alarm set to retain the chronological order of the condensed alarm sequence while also reflecting the cross-cluster summarization of the aggregated alarm clusters.
[0115] S105 , constructing an alarm propagation graph based on the compressed alarm set and the multi-granularity network topology model, and querying the alarm propagation graph through a reverse hierarchical sorting algorithm to determine the root alarm.
[0116] Specifically, after completing the compression of associated alarms, the present invention further constructs an alarm propagation graph based on the compressed alarm set and a multi-granularity network topology model. The alarm propagation graph is a multi-level, multi-perspective fault propagation model that integrates the associated semantics and temporal features of the compressed alarm set, building on the multi-granularity network topology model. It uses network entities as nodes and the physical connections, logical dependencies, and business associations between entities as edges. It maps compressed alarms to corresponding entity nodes and timestamps the nodes and edges according to the order in which the alarms occurred, ultimately forming a heterogeneous information network that is unified in time and space and rich in semantics.
[0117] The key to constructing an alarm propagation graph lies in fully integrating the semantic characteristics of the compressed alarm set with the structural characteristics of the network topology to uniformly characterize the dynamic process of fault propagation across the three dimensions of time, space, and semantics. Specifically, based on a multi-granular network topology model, network entities at the physical, logical, and service layers are abstracted as unified nodes, and connections at different granularity levels are abstracted as unified edges. Alarm nodes in the compressed alarm set are then mapped to corresponding physical nodes, and the nodes are labeled with alarm types and levels based on the semantic attributes of the alarms. Next, logical edges are established between physical nodes based on the associated edges and directions between alarm nodes, characterizing the alarm propagation path and impact range within the network. Furthermore, alarm occurrence timestamps are attached to alarm nodes and propagation edges to characterize the temporal characteristics of fault propagation. Finally, based on the end-to-end path of the service data flow, cross-layer propagation edges are correlated and merged to construct a unified alarm propagation graph, visually presenting the full picture of fault propagation under multi-granular semantics.
[0118] Based on the alarm propagation graph, the present invention uses a reverse hierarchical sorting algorithm to query the graph and ultimately identify the root alarm. The root alarm is the source of all subsequent alarms and represents the ultimate goal of fault diagnosis. The basic concept of the reverse hierarchical sorting algorithm is to start from the leaf nodes of the alarm propagation graph and hierarchically backtrack and sort the alarm nodes in the reverse direction of alarm propagation until the source alarm located farthest upstream and without a parent node is found. Specifically, the leaf nodes are first sorted based on the timestamp of the alarm node, with the latest leaf node being considered the starting point for backtracking. Then, based on the direction of the alarm propagation edge, all parent nodes of the current node are identified and sorted based on the semantic attributes and weights of the edge, prioritizing parent nodes with high alarm levels and strong association strengths. The sorting process is then recursively repeated for the selected parent node and its subgraph until a source alarm without a parent node is found or the entire propagation graph is traversed. During the backtracking process, a pruning strategy is used to dynamically maintain an optimal set of root alarms, avoiding repeated and ineffective searches and improving query efficiency. Finally, the root alarm set is further screened and evaluated based on business semantics, and the final root alarm is determined by comprehensively considering factors such as the fault impact scope, fault duration, and alarm level.
[0119] Based on the above embodiment, as an optional embodiment, querying the alarm propagation graph using a reverse hierarchical sorting algorithm to determine the root alarm includes:
[0120] S701, calculating the weight of each directed edge in the alarm propagation graph to obtain the propagation edge weight of each directed edge;
[0121] Specifically, the propagation edge weight takes into account multiple factors, such as the correlation of alarm types, the time difference between alarm occurrences, and the similarity of alarm attributes. These factors reflect the degree of influence of the antecedent alarm on the consequent alarm from different perspectives. For example, the higher the correlation of alarm types, the smaller the time difference, and the greater the attribute similarity, the stronger the influence of the antecedent alarm on the consequent alarm, and the larger the propagation edge weight should be.
[0122] In a specific embodiment, each directed edge in the alarm propagation graph is traversed to extract corresponding alarm attribute information, such as alarm type, timestamp, and alarm attributes. Using this attribute information, the propagation edge weight of each directed edge is calculated according to a predefined formula or algorithm. Common calculation methods include weighted summation, geometric mean, and function mapping. For example, a weighted summation of indicators such as alarm type relevance, inverse time difference, and attribute similarity can be used to obtain the propagation edge weight. The distribution of weights needs to be adjusted based on business scenarios and expert experience to accurately reflect the impact mechanism of alarm propagation.
[0123] During the calculation process, it's important to consider special cases. For example, directed edges with significant time differences may have a weak propagation impact and should be given a lower weight or ignored. Similarly, certain critical alarm types may have an exceptionally significant propagation impact and require a higher weight. These special cases can be implemented by setting thresholds or introducing business rules.
[0124] By calculating edge by edge, we ultimately obtain the propagation edge weight for each directed edge in the alarm propagation graph. These weights numerically represent the impact strength between alarms, acting as a "weight" for alarm propagation. A larger propagation edge weight indicates a stronger influence of the antecedent alarm on the consequent alarm, and a more critical role in alarm propagation. Conversely, a smaller weight indicates a weaker influence of the antecedent alarm on the consequent alarm, and a less important role in alarm propagation.
[0125] The introduction of propagation edge weights transforms the alarm propagation graph from an unweighted graph to a weighted graph, increasing the precision and depth of alarm propagation analysis. Based on propagation edge weights, importance metrics for alarm nodes, such as in-degree weight and out-degree weight, can be calculated to characterize the different roles of alarms in propagation, such as "source," "transit," and "aggregation." This provides a quantitative basis for subsequent root alarm identification, making root alarm determination more accurate and objective.
[0126] S702, starting from the terminal leaf node of the alarm propagation graph, traverse the directed edges in reverse order to determine the alarm propagation node in the alarm propagation graph, and determine the terminal alarm node and all upstream alarm propagation nodes upstream of the terminal alarm node based on the alarm propagation node;
[0127] Specifically, we need to clarify the topology of the alarm propagation graph and identify all terminal leaf nodes. Terminal leaf nodes are nodes in the alarm propagation graph that have only incoming edges and no outgoing edges. They represent the endpoints of alarm propagation. These terminal leaf nodes are the starting points for reverse traversal, acting as the "downstream exits" of alarm propagation.
[0128] Starting from each terminal leaf node, traverse the alarm propagation graph in reverse along the incoming edges. During the traversal process, determine the direct upstream node of the current node based on the direction of the directed edge. The direct upstream node refers to the node that is directly connected to the current node by an incoming edge, representing the direct "predecessor" alarm of the current node. Mark the direct upstream node as a propagation alarm node, indicating that it plays a "propagation" role in the alarm propagation. Continue to traverse in reverse along the incoming edge of the propagation alarm node, recursively determine the propagation alarm node further upstream, until reaching a node with no incoming edge (i.e., the source node) or a node that has been traversed (i.e., forming a loop). During the traversal process, record each propagation alarm node and the upstream and downstream relationships between them to generate a directed subgraph of the propagation alarm node. This subgraph depicts the complete propagation path from the terminal leaf node to the source node.
[0129] Repeating the reverse traversal process for each terminal leaf node yields multiple directed subgraphs of alarm propagation nodes. These subgraphs are then merged to form a single directed graph of alarm propagation nodes covering all terminal leaf nodes. This directed graph reveals the overall trajectory of alarm propagation, presenting a complete causal chain from the terminal to the source.
[0130] During reverse traversal, special handling is important. For example, when encountering loops, repeated traversals must be avoided by using marking or pruning. Alternatively, when multiple upstream nodes converge on the same node, the impact of each upstream node must be comprehensively considered to select the propagation path with the strongest influence. These special handling procedures can be implemented by setting traversal rules and introducing priorities.
[0131] By traversing backward, we ultimately identified the alarm propagation nodes in the alarm propagation graph, as well as the terminal alarm nodes and all their upstream alarm propagation nodes. These nodes and their upstream and downstream relationships outline the backbone network of alarm propagation, revealing the key links in the causal transmission of alarms. The terminal alarm node represents the surface of the problem, while the upstream alarm propagation nodes reveal the underlying causes. Through this causal chain, operations and maintenance personnel can analyze the root causes of alarms in a more comprehensive and systematic manner, from the surface to the core, tracing cause to effect.
[0132] S703 , sorting the upstream alarm propagation nodes according to the propagation edge weights, taking the upstream alarm propagation node with the largest propagation edge weight as the root alarm node, and determining the root alarm according to the alarm node.
[0133] Specifically, we first need to extract the edge weights for each upstream alarm propagation node. These edge weights, calculated in step S701, reflect the influence between alarm nodes. For each upstream alarm propagation node, we find its corresponding incoming edge and extract the corresponding edge weight. These weights act like "importance indicators" for upstream nodes, quantifying their influence on alarm propagation.
[0134] All upstream alarm propagation nodes are sorted based on the extracted propagation edge weights. Sorting is done in descending order, meaning that nodes with larger propagation edge weights are ranked higher. Nodes with larger propagation edge weights are more influential, play a key role in alarm propagation, and are more likely to be the root cause of the problem. This sorting process arranges upstream alarm propagation nodes in descending order of importance, forming an orderly "importance queue."
[0135] The first node in the sorted upstream alarm propagation node queue, that is, the node with the highest propagation edge weight, is identified as the root alarm node. The root alarm node is the node with the greatest influence and the highest likelihood of causing an alarm among all upstream alarm propagation nodes, acting like a "key hub" in the alarm propagation network. It carries the most alarm traffic and is the source of the problem. Identifying the root alarm node pinpoints the root cause of the problem and lays the foundation for subsequent root cause analysis and fault diagnosis.
[0136] Based on the identified root alarm node, the root alarm is further identified. The root alarm is the initial alarm corresponding to the root alarm node, causing the entire alarm to propagate and representing the root cause of the problem. By correlating the root alarm node's attributes, such as the alarm type, alarm target, and alarm details, a complete description of the root alarm is generated. This root alarm acts like a "diagnostic report," identifying the essence of the problem and providing O&M personnel with a basis for prescribing the right solution.
[0137] During the sorting and root alarm determination process, special considerations must be taken. For example, when multiple upstream alarm propagation nodes share the same edge weight and are tied for first place, additional evaluation metrics, such as node degree and time sequence, should be introduced to further refine the sorting until a single root alarm node is determined. Furthermore, when a root alarm node corresponds to multiple alarms, the impact of each alarm must be comprehensively analyzed to select the most critical and representative alarm as the root alarm.
[0138] See also Figure 2 , Figure 2 This is an architecture diagram of a transmission network fault location system provided in an embodiment of the present application. The transmission network fault location system may include:
[0139] A multi-granularity network topology model building module 1 is used to collect physical connection data, logical connection data, and service bearer relationship data in the transmission network, and build a multi-granularity network topology model based on the physical connection data, logical connection data, and service bearer relationship data;
[0140] Topology correlation analysis module 2, used to obtain original alarm data in the transmission network and analyze the topology correlation between different alarms based on the original alarm data and the multi-granularity network topology model;
[0141] The multi-alarm data determination module 3 is used to mine the topological association relationship between different alarms, obtain the horizontal peer association and vertical upstream and downstream association between the mined alarm objects, and determine the multi-alarm data of each alarm object within a specific time window based on the identification of the horizontal peer association and vertical upstream and downstream association;
[0142] The associated alarm compression module 4 is used to construct an associated alarm directed graph based on multiple alarm data, and compress the associated alarm directed graph based on the associated alarm directed graph to obtain a compressed alarm set;
[0143] The root alarm determination module 5 is used to construct an alarm propagation graph based on the compressed alarm set and the multi-granularity network topology model, and query the alarm propagation graph through a reverse hierarchical sorting algorithm to determine the root alarm.
[0144] Please refer to Figure 3 The present application also discloses an electronic device. Figure 3 The electronic device 300 may include: at least one processor 301 , at least one network interface 304 , a user interface 303 , a memory 305 , and at least one communication bus 302 .
[0145] The communication bus 302 is used to implement the connection and communication between these components.
[0146] The user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.
[0147] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0148] The processor 301 may include one or more processing cores. Using various interfaces and circuits, the processor 301 connects to various components within the server. It executes instructions, programs, code sets, or instruction sets stored in the memory 305, as well as accesses data stored in the memory 305, to perform various server functions and process data. Optionally, the processor 301 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 301 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 301 but implemented as a separate chip.
[0149] Among them, the memory 305 may include a random access memory (Random Access Memory, RAM) and may also include a read-only memory (Read~Only Memory). The memory 305 includes a non-transitory computer-readable medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 305 may optionally be at least one storage system located away from the aforementioned processor 301. Refer to Figure 3 The memory 305 as a computer storage medium may include an operating system, a network communication module, a user interface module and an application program for a transmission network fault location method.
[0150] exist Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an input interface for the user and obtain the data input by the user; and the processor 301 can be used to call the application program storing the road assessment method in the memory 305. When executed by one or more processors 301, the electronic device 300 executes one or more methods in the above-mentioned embodiments. It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that this application is not limited to the described order of actions, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for this application. In the above-mentioned embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0151] In the several embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some service interface, and the indirect coupling or communication connection of the system or unit can be electrical or other forms.
[0152] The present application also provides a computer storage medium that can store multiple instructions, which are suitable for being loaded and executed by a processor as described above. Figure 1 The road assessment method of the embodiment shown, the specific execution process can be found in Figure 1 The detailed description of the illustrated embodiment will not be repeated here.
[0153] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0154] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes N instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of this application. The aforementioned memory includes various media that can store program code, such as USB flash drives, mobile hard drives, magnetic disks, or optical disks.
[0155] The above are merely exemplary embodiments of the present disclosure and are not intended to limit the scope of the present disclosure. In other words, any equivalent variations and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the disclosure and the practical implications thereof.
[0156] This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not described herein. The description and examples are to be considered as exemplary only, and the scope and spirit of the present disclosure are to be defined by the claims.
Claims
1. A method for locating a transmission network fault, characterized in that: The method includes: collecting physical connection data, logical connection data and service bearing relationship data in the transmission network, and constructing a multi-granularity network topology model based on the physical connection data, the logical connection data and the service bearing relationship data; obtaining original alarm data in the transmission network, and analyzing the topological association relationship between different alarms based on the original alarm data and the multi-granularity network topology model; mining the topological association relationship between the different alarms to obtain horizontal same-level associations and vertical upstream and downstream associations between the mined alarm objects, and determining multiple alarm data of each alarm object within a specific time window based on the horizontal same-level associations and the vertical upstream and downstream associations; constructing an associated alarm directed graph based on the multiple alarm data, and performing associated alarm compression on the associated alarm directed graph based on the associated alarm directed graph to obtain a compressed alarm set; constructing an alarm propagation graph based on the compressed alarm set and the multi-granularity network topology model, and querying the alarm propagation graph through a reverse hierarchical sorting algorithm to determine a root alarm; The method of compressing the associated alarm directed graph based on the associated alarm directed graph to obtain a compressed alarm set includes: clustering the associated alarm directed graph by a graph clustering algorithm to obtain a plurality of associated clusters; analyzing the internal alarms of each associated cluster to determine the semantic redundancy of the internal alarm nodes, and determining the compressible homogeneous alarms in the internal alarms based on the semantic redundancy; compressing the internal alarms of the associated cluster based on the compressible homogeneous alarms to obtain a compressed simplified alarm sequence within the cluster; performing homogeneity condition analysis on each associated cluster to determine a homogeneous associated cluster among a plurality of associated clusters, and compressing the homogeneous associated clusters to obtain an aggregated alarm cluster; integrating the compressed simplified alarm sequence within the cluster with the aggregated alarm cluster to obtain the compressed alarm set; The alarm propagation graph is queried through a reverse hierarchical sorting algorithm to determine the root alarm, including: performing weight calculation on each directed edge of the alarm propagation graph to obtain the propagation edge weight of each directed edge; starting from the terminal leaf node of the alarm propagation graph, traversing the directed edges in reverse to determine the propagation alarm nodes in the alarm propagation graph, and determining the terminal alarm node and all upstream propagation alarm nodes upstream of the terminal alarm node based on the propagation alarm nodes; sorting the upstream propagation alarm nodes according to the propagation edge weights, taking the upstream propagation alarm node with the largest propagation edge weight as the root alarm node, and determining the root alarm based on the root alarm node.
2. The method according to claim 1, characterized in that The multi-granularity network topology model is constructed based on the physical connection data, the logical connection data and the service bearer relationship data, including: constructing a physical connection topology based on the physical connection data; constructing a logical connection topology based on the logical connection data; constructing a service bearer relationship topology based on the service bearer relationship data; and vertically integrating the physical connection topology, the logical connection topology and the service bearer topology to obtain the multi-granularity network topology model.
3. The method according to claim 1, characterized in that The topological association relationship between different alarms is analyzed based on the original alarm data and the multi-granularity network topology model, including: extracting the attributes of the original alarm data to obtain key alarm attributes, and associating each alarm with the topological entity of the corresponding granularity level in the multi-granularity network topology model according to the key alarm attributes to form an "alarm-topology entity" mapping table; traversing each alarm in the "alarm-topology entity" mapping table to extract the association features of each alarm in the multi-granularity topology environment to obtain an alarm topology feature vector; based on the alarm topology feature vector, calculating the topological semantic correlation between different alarms, and determining the topological association relationship between different alarms based on the correlation.
4. The method according to claim 1, wherein The constructing of the associated alarm directed graph based on the multiple alarm data includes: extracting nodes from the multiple alarm data to obtain alarm nodes; performing node association on the alarm nodes based on the multiple alarm data to obtain associated edges between the alarm nodes; and constructing the associated alarm directed graph based on the associated edges between the alarm nodes and the alarm nodes.
5. The method according to claim 1, wherein The alarm propagation graph is constructed based on the compressed alarm set and the multi-granularity network topology model, including: mapping the compressed alarm set to the multi-granularity network topology model to obtain a multi-granularity mapping relationship table between the compressed alarm set and the multi-granularity network topology model; and constructing the alarm propagation graph based on the multi-granularity mapping relationship table and the compressed alarm set.
6. A transmission network fault location system, configured to execute a transmission network fault location method according to any one of claims 1 to 5, characterized in that: Applied to the user end, the system includes: a multi-granularity network topology model establishment module, which is used to collect physical connection data, logical connection data and service bearing relationship data in the transmission network, and build a multi-granularity network topology model based on the physical connection data, the logical connection data and the service bearing relationship data; a topology association analysis module, which is used to obtain original alarm data in the transmission network, and analyze the topology association relationship between different alarms based on the original alarm data and the multi-granularity network topology model; a multi-alarm data determination module, which is used to mine the topology association relationship between the different alarms and obtain the mined alarm data. Horizontal same-level associations and vertical upstream and downstream associations between alarm objects, and based on the identification of the horizontal same-level associations and the vertical upstream and downstream associations, determine the multiple alarm data of each alarm object within a specific time window; an associated alarm compression module, used to construct an associated alarm directed graph based on the multiple alarm data, and perform associated alarm compression on the associated alarm directed graph based on the associated alarm directed graph to obtain a compressed alarm set; a root alarm determination module, used to construct an alarm propagation graph based on the compressed alarm set and the multi-granularity network topology model, and query the alarm propagation graph through a reverse hierarchical sorting algorithm to determine the root alarm.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor and executed by a method according to any one of claims 1 to 5.
8. An electronic device, characterized in that: The electronic device comprises a processor, a memory and a transceiver, wherein the memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Electric power communication network alarm association mining method based on improved GSP
CN110445665A
Data center operation and maintenance alarm information merging method and system
CN113849513A