Group obstacle identification method and device and electronic equipment

By obtaining and analyzing alarm information, combining the equipment map and alarm map, using the fault identification model to determine the correlation rules of the fault equipment, the problem of inefficient group fault identification in the existing technology is solved, and rapid and accurate fault positioning and network operation and maintenance efficiency are achieved.

CN120075857AActive Publication Date: 2025-05-30CHINA TELECOM CORP LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510240774.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

In the prior art, group fault identification methods are inefficient and difficult to quickly and accurately locate the cause of failure, which affects the timely detection and processing of network faults.

Method used

By obtaining the alarm information within the preset time period, determining the device address and uplink network equipment, determining the correlation rules of the faulty equipment based on the alarm map and the fault identification model, and then identifying group obstacles and their impact range.

Benefits of technology

It realizes the rapid and accurate identification of network group obstacles and their impact range, improves network operation and maintenance efficiency, shortens failure recovery time, and improves service quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075857A_ABST
    Figure CN120075857A_ABST
Patent Text Reader

Abstract

The invention discloses a group obstacle identification method and device and electronic equipment. The method comprises the following steps: acquiring alarm information in a preset time period; an equipment address corresponding to the alarm information is determined, uplink network equipment corresponding to the equipment address is determined according to an equipment map, and the equipment map is used for representing the incidence relation between the network equipment; the alarm state of the uplink network equipment is determined according to an alarm atlas, fault equipment is determined from the uplink network equipment according to the alarm state, and the alarm atlas is used for representing the incidence relation between the alarm events; and determining an association rule of the fault equipment in an alarm process through the fault identification model, and determining a group fault identification result according to the association rule. According to the method and the device, the technical problem that timely discovery and processing of network faults are affected due to the fact that a group fault identification method in the related technology is low in efficiency and particularly when a large number of equipment alarms and user declaration are faced, fault causes are difficult to locate quickly and accurately is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a method, apparatus, and electronic device for group fault identification. Background Art

[0002] With the expansion of network coverage and the surge in the number of users, the complexity of the network and the challenges of operation and maintenance are gradually increasing. In particular, the occurrence of group faults often involves multiple devices reporting alarms simultaneously or a large number of users reporting service interruptions at the same time, which not only tests the robustness of the network but also poses higher requirements for the operation and maintenance capabilities of operators.

[0003] However, in the current field of network operation and maintenance, the group fault identification methods in related technologies have significant deficiencies in terms of efficiency, accuracy, response speed, and resource utilization. For example, when dealing with a large number of alarms, the group fault identification method relying on manual analysis not only takes a long time but also easily misses key information. In the face of complex network structures and unknown fault patterns, it may lead to misjudgments or missed judgments. At the same time, from receiving the alarm to officially confirming the group fault and then locating the root cause of the fault, this process may take several hours or even longer, resulting in an extended service recovery time and impaired user experience. In addition, ineffective fault location and repair attempts will waste valuable operation and maintenance resources and affect the efficiency of subsequent fault handling.

[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of this application provide a method, apparatus, and electronic device for group fault identification to at least solve the technical problem that due to the low efficiency of group fault identification methods in related technologies, especially when facing a large number of device alarms and user reports, it is difficult to quickly and accurately locate the cause of the fault, affecting the timely discovery and handling of network faults.

[0006] According to one aspect of the embodiments of this application, a method for group fault identification is provided, including: obtaining alarm information within a preset time period, where the alarm information is used to indicate that a network device is in an abnormal state; determining the device address corresponding to the alarm information, and determining the upstream network device corresponding to the device address according to a device map, where the device map is used to represent the association relationship between each network device; determining the alarm status of the upstream network device according to an alarm map, and determining the faulty device from the upstream network devices according to the alarm status, where the alarm map is used to represent the association relationship between each alarm event; determining the association rule of the faulty device during the alarm process through a fault identification model, and determining the group fault identification result according to the association rule, where the association rule is used to reflect the fault propagation mode and historical processing efficiency of the faulty device.

[0007] Optionally, obtain the alarm information within a preset time period, including: obtaining the alarm information within a preset time period from the operation and maintenance support center through a preset call mechanism, where the operation and maintenance support center stores a device map, an alarm map, and an operation and maintenance map, and the operation and maintenance map is used to provide repair guidance for faulty devices.

[0008] Optionally, the method further includes: converting the group fault identification result into a preset format, and providing a link to a corresponding plan document, where the link to the plan document contains the fault handling measures determined based on the operation and maintenance map.

[0009] Optionally, the device map is determined in the following manner: obtaining the link data of all network devices, and determining an initial device map based on the link data, where the link data includes the device information of the network devices and the link information between the network devices; obtaining the station and computer room data and power environment data of the network devices, where the station and computer room data is used to represent the station information and computer room information with clear geographical coordinates and type attributes, and the power environment data is used to represent the environmental monitoring data in the station and computer room; optimizing the initial device map based on the station and computer room data and the power environment data to obtain the device map.

[0010] Optionally, the device map is constructed with the resource objects corresponding to the network devices as nodes and the logical relationships between the resource objects as edges, where the logical relationships at least include the spatial position relationship, physical connection relationship, and bearing relationship between the resource objects.

[0011] Optionally, the alarm map is constructed with the alarm events corresponding to the network devices as nodes and the causal relationships between the alarm events as edges.

[0012] Optionally, the fault identification model is trained in the following manner: obtaining historical alarm data and historical work order data; performing feature extraction on the historical alarm data and historical work order data to obtain a first feature sequence corresponding to the historical alarm data and a second feature sequence corresponding to the historical work order data, where the first feature sequence includes the time series feature and spatial distribution feature corresponding to the historical alarm data, and the second feature sequence is used to reflect the relationship between the historical alarm events and the historical fault handling measures; training an initial model based on the first feature sequence and the second feature sequence in a target application scenario to obtain the fault identification model, where the target application scenario is the cluster offline fault scenario of the optical line terminal.

[0013] Optionally, the association rules of the faulty devices during the alarm process include: the frequent occurrence periods, durations of various fault types of the faulty devices, and the collinear patterns with other network events, where the fault types at least include: the cluster offline fault of the optical line terminal.

[0014] According to another aspect of the embodiments of the present application, a group fault identification device is further provided, including: an acquisition module, configured to acquire alarm information within a preset time period, where the alarm information is used to indicate that a network device is in an abnormal state; a first determination module, configured to determine a device address corresponding to the alarm information, and determine an upstream network device corresponding to the device address according to a device map, where the device map is used to represent the association relationship between each network device; a second determination module, configured to determine the alarm status of the upstream network device according to an alarm map, and determine a faulty device from the upstream network devices according to the alarm status, where the alarm map is used to represent the association relationship between each alarm event; a third determination module, configured to determine the association rule of the faulty device during the alarm process through a fault identification model, and determine a group fault identification result according to the association rule, where the association rule is used to reflect the fault propagation mode and historical processing efficiency of the faulty device.

[0015] According to yet another aspect of the embodiments of the present application, an electronic device is further provided, including: a memory and a processor, where the memory is configured to store program instructions; the processor is connected to the memory and is configured to execute to implement the above-mentioned group fault identification method.

[0016] According to still another aspect of the embodiments of the present application, a non-volatile storage medium is further provided. The non-volatile storage medium includes a stored computer program, where the device where the non-volatile storage medium is located executes the above-mentioned group fault identification method by running the computer program.

[0017] According to still another aspect of the embodiments of the present application, a computer program product is further provided, including computer instructions, and the computer program product implements the above-mentioned group fault identification method when the computer instructions are executed by a processor.

[0018] In the embodiments of the present application, by obtaining the alarm information within a preset time period, where the alarm information is used to indicate that the network device is in an abnormal state; determining the device address corresponding to the alarm information, and determining the upstream network device corresponding to the device address according to the device map, where the device map is used to represent the association relationship between each network device; determining the alarm status of the upstream network device according to the alarm map, and determining the faulty device from the upstream network devices according to the alarm status, where the alarm map is used to represent the association relationship between each alarm event; determining the association rule of the faulty device during the alarm process through the fault identification model, and determining the group fault identification result according to the association rule, where the association rule is used to reflect the fault propagation mode and historical processing efficiency of the faulty device, the purpose of quickly and accurately identifying the network group fault and its influence range is achieved, thereby realizing the technical effects of improving the network operation and maintenance efficiency, shortening the fault recovery time, and improving the service quality, and further solving the technical problem that the group fault identification method in the related technology is inefficient, especially when facing a large number of device alarms and user complaints, it is difficult to quickly and accurately locate the cause of the fault, affecting the timely discovery and handling of network faults. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments and descriptions thereof of the present application are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0020] Figure 1 is a hardware structure diagram of a computer terminal for implementing a group fault identification method according to an embodiment of the present application;

[0021] Figure 2 is a flowchart of a group fault identification method according to an embodiment of the present application;

[0022] Figure 3 is a schematic architecture diagram of a group fault identification system according to an embodiment of the present application;

[0023] Figure 4 is a structure diagram of a group fault identification device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0025] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] First, some nouns or terms that appear in the process of explaining the embodiments of this application are applicable to the following explanations:

[0027] OLT (Optical Line Terminal): The core device in the fiber access network, responsible for converting the signals from the user side into a format suitable for fiber transmission, and then sending them to the network center or upper-layer devices. At the same time, the OLT can also convert the data signals from the center into a format that can be received by the user side, realizing two-way data transmission. In the FTTH (Fiber To The Home) or FTTC (Fiber To The Curb) network, the OLT plays the role of a gateway, connecting the fiber distribution network and the local area network.

[0028] OTN (Optical Transport Network): A high-bandwidth, long-distance transmission network based on WDM (Wavelength Division Multiplexing) technology. The OTN is mainly responsible for efficiently transmitting large-granularity services, such as high-speed Ethernet, SDH signals, etc. on the fiber medium, and providing strong network protection and restoration capabilities to ensure the stability of data transmission and service quality.

[0029] BAS (Broadband Access Server): Used to manage and control the process of broadband users accessing the network. Its main functions include user authentication, authorization and accounting (AAA service). BAS devices are usually located at the network edge, connected to user terminals (such as modems), responsible for processing user data packets and sending them to the core network or the Internet.

[0030] Function Calling: In software development, function calling refers to the ability of one function to call another to perform a specific task. In the field of AI, especially in large language models, Function Calling allows the model to call external functions or services when generating responses to query or update databases, perform specific calculations, or obtain real-time data, thereby enhancing the interactivity and practicality of the model.

[0031] Knowledge Graph: A data structure that graphically represents entities and the relationships between them, aiming to simulate the knowledge network in human cognition. In the operation and maintenance of communication networks, the Knowledge Graph can include device graphs, alarm graphs, operation and maintenance graphs, etc., which are used to record the attributes, locations, connection relationships of network devices, as well as historical fault data and handling experiences, thereby assisting in fault diagnosis and problem solving.

[0032] Alarm Graph: A type of Knowledge Graph specific to network operation and maintenance that records alarm information in the network, including the type, level, time, location, description of the alarm, and the causal relationships between them. The construction and update of the Alarm Graph help quickly locate the root cause of faults and reduce the group fault response time.

[0033] Device Graph: A Knowledge Graph that shows the internal logical relationships between various resource objects in the network, including spatial location, physical connection, logical connection, and bearing relationship, etc. The construction of the Device Graph is crucial for understanding the network structure and resource distribution, and helps in fault analysis and resource scheduling.

[0034] Operation and Maintenance Graph: Established through manufacturer manuals and fixed rules, it records the operation and maintenance strategies, operation manuals, and emergency plans of various devices, and is used to guide the decision-making and actions of operation and maintenance personnel when dealing with faults. The Operation and Maintenance Graph can recommend repair solutions according to the type of fault, improving the operation and maintenance efficiency.

[0035] To solve the problem of poor efficiency in identifying group faults in related technologies, the embodiments of this application provide a method for identifying group faults, which can run on Figure 1 the computer terminal shown below, and the computer terminal will be described as follows.

[0036] The embodiments of the method for identifying group faults provided by the embodiments of this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 The following shows a hardware structure block diagram of a computer terminal for implementing the method for identifying group faults. As Figure 1As shown, the computer terminal 10 may include one or more processors (illustrated as 102a, 102b, ……, 102n in the figure) (the processor may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions connected via wired and / or wireless networks. In addition, it may further include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown therein, or have a different configuration from Figure 1 that shown.

[0037] It should be noted that the above one or more processors and / or other data processing circuits may generally be referred to as "data processing circuits" herein. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0038] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the group fault identification method in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned group fault identification method. The memory 104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0039] The transmission module 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission module 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0040] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10.

[0041] It should be noted here that in some alternative embodiments, the above Figure 1 illustrated computer terminal may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 1 is only an example of a specific specific instance and is intended to show the types of components that may exist in the above computer terminal.

[0042] Under the above operating environment, an embodiment of a group fault identification method is provided in an embodiment of the present application. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0043] Figure 2 is a flowchart of a group fault identification method according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:

[0044] Step S202, obtain alarm information within a preset time period, where the alarm information is used to indicate that the network device is in an abnormal state.

[0045] In the above step S202, the alarm information includes real-time feedback that the network device is in an abnormal state, such as OLT off-network alarm, PON port no optical reception, dynamic environment system alarm (such as battery hidden danger, power outage, high temperature, water immersion, smoke sensor, etc.), wavelength division system failure, OLT uplink single-sided alarm, OLT uplink device alarm (SW, BAS), and OLT-related alarms such as abnormal device link status, which is the starting point of group fault identification.

[0046] Among them, each alarm message includes real-time alarm event information (such as alarm ID, alarm type, alarm level, alarm time, alarm location, and alarm description) and real-time status information (device operating status, device performance metrics, and link status), which are used for real-time analysis and alarm generation.

[0047] Step S204: Determine the device address corresponding to the alarm message, and determine the upstream network device corresponding to the device address according to the device atlas, where the device atlas is used to represent the association relationship between each network device.

[0048] In the above step S204, by analyzing the obtained alarm message, the device address can be identified therefrom, such as the IP address of the OLT device. Subsequently, according to the device atlas, the upstream link is traced to identify all upstream network devices related to the device address. Among them, the device atlas is a comprehensive knowledge graph that details the physical and logical connection relationships between devices in the network, including but not limited to the connections between devices such as cities, counties, bureaus, computer rooms, OLTs, SWs, BASs, and OTNs, as well as their configuration information and geographical layout. Through the device atlas, the upstream link structure of the alarm device can be quickly located, thereby locking in potential fault points.

[0049] Step S206: Determine the alarm status of the upstream network device according to the alarm atlas, and determine the faulty device from the upstream network devices according to the alarm status, where the alarm atlas is used to represent the association relationship between each alarm event.

[0050] In the above step S206, by further analyzing the real-time alarm status of the upstream network device and with the help of the alarm atlas, the faulty device experiencing a fault can be identified therefrom. Among them, the alarm atlas is another knowledge graph that not only records each alarm event but also shows the mutual causal relationship and time series characteristics between each alarm event. By analyzing the alarm atlas, the model can understand how the fault propagates between different devices, which helps to accurately determine which upstream devices are the direct sources of the fault and lays a foundation for subsequent fault root cause location.

[0051] Step S208: Determine the association rule of the faulty device during the alarm process through the fault identification model, and determine the group fault identification result according to the association rule, where the association rule is used to reflect the fault propagation mode and historical processing efficiency of the faulty device.

[0052] In the above step S208, a pre-trained fault identification model can be utilized to deeply analyze the correlation rules of faulty devices during the warning process. These correlation rules incorporate fault propagation patterns and historical processing efficiency, enabling the model to understand how specific faults spread across devices and the network, as well as how similar past faults were resolved. Based on these correlation rules, the model can comprehensively judge whether a group fault has occurred, as well as the scope of influence and potential root causes of the group fault, forming a group fault identification result.

[0053] Through the above steps S202 to S208, the purpose of quickly and accurately identifying network group faults and their scope of influence is achieved, thereby realizing the technical effects of improving network operation and maintenance efficiency, shortening fault recovery time, and enhancing service quality. Furthermore, it solves the technical problem that the group fault identification method in the related technology is inefficient, especially when facing a large number of device warnings and user complaints, it is difficult to quickly and accurately locate the cause of the fault, affecting the timely discovery and handling of network faults. The following is a detailed description.

[0054] In the above step S202, warning information within a preset time period is obtained, including: obtaining warning information within a preset time period through a preset call mechanism from an operation and maintenance support center, where the operation and maintenance support center stores a device map, a warning map, and an operation and maintenance map, and the operation and maintenance map is used to provide repair guidance for faulty devices.

[0055] In the embodiment of the present application, in order to monitor the network health status in real time, through the Function Calling mechanism (i.e., the above-mentioned preset call mechanism), the large model can be given the ability to dynamically query the current status of network devices. For example, when a warning signal is received, such as multiple OLT uplink anomalies, the model immediately activates Function Calling and sends a query request to the network management system to collect real-time warning information of key devices including but not limited to OLTs, switches, and wavelength division multiplexing systems. This process is not limited to top-level devices but also penetrates down to lower-level network components to ensure the comprehensiveness of the evaluation.

[0056] Figure 3 It is a schematic diagram of the architecture of a group fault identification system according to an embodiment of the present application. As Figure 3As shown in the figure, the group fault identification system includes three core components: the network operation and maintenance platform 30 (collection and control platform), the network management system 32 (resource system), and the intelligent operation and maintenance support center 34 (referred to as the operation and maintenance support center), as well as the resource relational database 36 that supports the entire system. Among them, the operation and maintenance support center 34 is a newly added processing center in the group fault identification system. The operation and maintenance support center includes: an alarm graph construction and update module 302, a device graph construction and update module 304, an operation and maintenance graph construction and update module 306, a fault identification model 308, a root cause location module 310, a repair suggestion (intelligent) recommendation model 312, and a repair suggestion push module 314.

[0057] Specifically, first, device alarms are generated on the network operation and maintenance platform 30, and the network management system 32 provides real-time alarm information to the operation and maintenance support center 34. Among them, the resource relational database 36 provides detailed device resource information, which is the cornerstone of device graph construction. As the intelligent brain of the system, the alarm graph construction and update module 302 and the device graph construction and update module 304 under the operation and maintenance support center are responsible for constructing and maintaining dynamically updated graphs based on real-time alarm information and device resource information respectively; the fault identification model 308 deeply analyzes the alarm information based on the constructed alarm graph and device graph, combines the training ability of the large model, identifies group fault events and evaluates potential risks; the root cause location module 310 further analyzes the root cause of the fault to provide accurate guidance for fault handling; the operation and maintenance graph construction module 308 integrates operation and maintenance manuals and rules to establish a knowledge framework for fault repair; the repair suggestion push module 314 and the repair suggestion recommendation model 312 generate and push customized repair suggestions based on the fault type, root cause location, and operation and maintenance graph to ensure the efficiency and intelligence of fault handling. The entire system constructs an all-round and intelligent operation and maintenance support system covering fault prediction, identification, location, and repair guidance through close data interaction and logical analysis, significantly improving the efficiency and accuracy of network operation and maintenance.

[0058] In the above step S204, the device graph is determined in the following manner: Obtain the link data of all network devices, and determine the initial device graph based on the link data, where the link data includes the device information of the network devices and the link information between the network devices; obtain the station and computer room data and power environment data of the network devices, where the station and computer room data is used to represent the station information and computer room information with clear geographical coordinates and type attributes, and the power environment data is used to represent the environmental monitoring data in the station and computer room; optimize the initial device graph based on the station and computer room data and the power environment data to obtain the device graph.

[0059] Optionally, the device graph is constructed with resource objects corresponding to network devices as nodes and the logical relationships between resource objects as edges. Among them, the logical relationships include at least the spatial location relationship, physical connection relationship, and bearing relationship between resource objects.

[0060] In the embodiments of the present application, the device graph not only reflects the direct physical and logical connections between network devices, but also comprehensively considers the environmental factors where the devices are located and their positional relationships in the global network. It is a multi-dimensional and multi-level knowledge graph. The construction method of this device graph is as follows:

[0061] First, collect the link data of all network devices. These data not only include the detailed information of the devices (such as device ID, device type, device model, device location, etc.), but also cover the physical and logical connection information between devices (such as link ID, start device ID, end device ID, port number, transmission rate, and transmission medium, etc.). By integrating these link data, the system can construct the basic connection structure of the devices in the network, form an initial device graph, and lay a foundation for subsequent fault location and root cause analysis.

[0062] Second, in addition to the link information of the devices themselves, it is also necessary to collect the detailed data of the station and the machine room, as well as the power environment data in these places. Among them, the station and machine room data provide the geographical location and type attributes of the devices, such as station ID, machine room ID, geographical locations (latitude and longitude) of the station and the machine room, and environmental monitoring data (such as temperature, humidity, power status, etc.); the power environment data covers key information such as the temperature, humidity, and power supply status of the environment, such as power supply status (such as the status of the main power supply and the standby power supply), air conditioning system status (such as the operating status and parameters of the air conditioning equipment), and environmental alarm data (such as the alarm information of the dynamic environment system, such as power failure, abnormal temperature and humidity, etc.). These data can help the system understand the environmental conditions where the devices are located and how these conditions affect the normal operation of the devices.

[0063] Finally, integrate the station and machine room data and the power environment data on the basis of the initial device graph to form the final device graph. This operation not only enhances the geographical information of the device graph, making the geographical redundancy of the network layout and potential regional risks visually displayed, but also supplements the analysis dimension of the impact of environmental factors on the devices. The optimized device graph can more comprehensively reflect the spatial location relationship, physical connection relationship, and bearing relationship between network devices, and provide richer information for fault analysis.

[0064] It should be noted that in the construction of the device map, resource objects corresponding to network devices are used as nodes in the map. Among them, the resource objects cover various entities in the network, such as OLT, SW, BAS, OTN devices, and their environments, such as cities, counties, bureau stations, computer rooms, etc.; the edges between nodes represent the logical relationships between these resource objects, including but not limited to:

[0065] Spatial location relationship: such as the geographical location association of city - county - bureau station - computer room - device.

[0066] Physical connection relationship: such as the direct connection between OLT ports and SW ports.

[0067] Carrying relationship: such as the logical association that the OLT device is carried on the OTN circuit through the link.

[0068] This device map combined with geographical information can not only help analysts quickly identify and locate network faults, but also predict potential regional risks, providing a more comprehensive and in - depth perspective for network operation and maintenance.

[0069] In addition, the device map also deeply records key configuration information such as the dual - uplink routing configuration of each link, including but not limited to primary - backup routing path information (such as path ID, start device ID, end device ID, path node list, and path edge list) and routing attribute information (such as routing transmission rate, path delay, and path error rate). These configuration details, such as the existence of single - side or non - redundant link paths, are crucial for identifying potential network vulnerabilities. By analyzing the configuration data in the device map, the system can intelligently judge the dependency relationships between devices. Especially when facing complex faults such as OLT group faults, it can quickly locate possible single - point faults or configuration defects, such as the absence or improper setting of dual - uplink routing, providing a forward - looking means of hidden - danger identification and risk assessment for network operation and maintenance, and greatly enhancing the stability and fault - prevention ability of the network.

[0070] In the above step S206, the alarm map is constructed with alarm events corresponding to network devices as nodes and the causal relationships between alarm events as edges.

[0071] In the embodiments of the present application, in the construction of the alarm graph, various alarm events occurring in the network are used as nodes in the alarm graph. Each node includes, but is not limited to, information such as alarm ID, alarm status, alarm time, location, type, and description. These information together constitute a comprehensive description of the alarm event. The edges between the nodes represent the potential causal relationships between the alarm events. By analyzing this causal relationship, the system can identify subsequent alarms that may be triggered by an alarm event, or the potential root cause pointed to by multiple alarm events. The construction of this causal relationship is not only based on the physical or logical connections between devices, but also takes into account the propagation law of faults and the actual experience of network operation and maintenance, providing a powerful analysis tool for the identification of group fault phenomena, root cause analysis, and fault prediction.

[0072] In the above step S208, the fault identification model is trained in the following manner: Obtain historical alarm data and historical work order data; perform feature extraction on the historical alarm data and historical work order data to obtain a first feature sequence corresponding to the historical alarm data and a second feature sequence corresponding to the historical work order data. Among them, the first feature sequence includes time series features and spatial distribution features corresponding to the historical alarm data, and the second feature sequence is used to reflect the relationship between historical alarm events and historical fault handling measures; Based on the first feature sequence and the second feature sequence, train the initial model in the target application scenario to obtain the fault identification model, where the target application scenario is the cluster offline fault scenario of the optical line terminal.

[0073] In the embodiments of the present application, the training and application of the fault identification model are the key links of the group fault identification method, especially for the target application scenario of the cluster offline fault of the optical line terminal (OLT). The specific process can be as follows:

[0074] First, obtain historical alarm data and historical work order data from the resource relational database. Among them, the historical work order data includes historical work orders and merged work order records. Specifically, the historical alarm data contains various alarm events that occurred in the past network operation, including but not limited to information such as historical alarm ID, type, level, time, location, and description. The historical work order data records the detailed description of the fault event, including but not limited to historical work order ID, fault description, handling measures, handling time, handling results, and associated alarm ID, device ID, and path ID, etc.

[0075] Secondly, perform feature extraction on the historical alarm data to obtain the corresponding first feature sequence. This first feature sequence not only includes time series features (such as frequent fault occurrence periods, alarm duration, etc.), but also covers spatial distribution features (such as the distribution pattern of faults in the geographical space).

[0076] Meanwhile, analyze the historical work order data to extract the corresponding second feature sequence, which mainly reflects the correlation law between the fault handling measures and the alarm events, including the priority of fault handling, the efficiency and effectiveness of the handling measures, etc.

[0077] Finally, according to the extracted first feature sequence and second feature sequence, train the initial model in the target application scenario (in this application, the offline fault of the optical line terminal (OLT) cluster is used as a specific training scenario). The model training process can be achieved by fine-tuning the sequence-to-sequence (SFT) model, aiming to enable the model to learn and master the time series pattern, spatial distribution law of the fault events, and the relationship between the fault handling measures and the fault types. During the model training process, the parameters will be continuously adjusted, the feature weights will be optimized, and new learning strategies will be introduced to improve the recognition accuracy and processing efficiency of the model for the cluster offline fault, and finally obtain the final fault recognition model.

[0078] Optionally, the correlation law of the faulty device during the alarm process includes: the frequent occurrence time period, duration of various fault types that occur in the faulty device, and the collinear pattern with other network events, where the fault types at least include: the cluster offline fault of the optical line terminal.

[0079] In the embodiment of the present application, after the model training is completed, the correlation law of the faulty device during the alarm process can be determined through it. These correlation laws include but are not limited to the frequent occurrence time period in time of various fault types that occur in the faulty device, the average time of fault duration, and the collinear pattern between the fault and other network events. For example, when the model identifies that multiple OLT devices simultaneously experience offline faults in a similar time period, and there is a common alarm pattern in the uplink or the area where these OLT devices are located, the model will determine it as an OLT cluster offline fault and form a group fault recognition result.

[0080] Based on the determined correlation law, the model can further analyze the root cause of the fault to obtain the root cause location result, including but not limited to identifying the initial fault point that triggers the group fault and how the fault propagates in the network. For example, if multiple OLT devices are offline simultaneously, the model will analyze the uplink status of these OLT devices, the status of the dynamic environment system in the computer room where they are located, and the status of the associated SW and BAS devices to locate the root cause of the fault, such as whether there is a dynamic environment system fault, an upper-layer device fault, or an optical cable interruption. After determining the group fault recognition result and the root cause of the fault, the analysis results can be integrated to provide decision-making support for the operation and maintenance personnel.

[0081] After the above step S208, the method may further include: converting the group fault recognition result into a preset format and providing a link to the corresponding plan document, where the link to the plan document contains the fault handling measures determined according to the operation and maintenance map.

[0082] In the embodiments of the present application, after completing the identification of group faults and root cause location, the analysis results can be converted into a preset JSON format, which contains all important identification information, such as keywords fields like root alarm ID, root alarm type, judgment basis, group fault judgment, hidden danger identification, etc., and can ensure that all necessary fault details are clearly recorded and presented. Subsequently, according to the fault identification results, the operation and maintenance map is automatically associated to find treatment measures that match the fault type and root cause. Among them, the operation and maintenance map contains detailed equipment operation manuals, fixed rules, and historical fault cases, and this information provides a basis for generating repair suggestions.

[0083] Specifically, the system will generate a (urgent) plan document link, which directly points to the document containing specific fault treatment measures, helping operation and maintenance personnel quickly access and execute. For example, based on the description of the current fault and the device status, the model can recommend the most appropriate treatment steps and preventive measures, and this ability significantly improves the pertinence and efficiency of fault handling.

[0084] With the handling of each group fault event, the system will collect feedback information, including the actual effect of fault handling and the operation experience of operation and maintenance personnel, and these data will be used for the continuous training and optimization of the model. Through continuous learning, the model can more accurately predict fault patterns, provide more intelligent fault handling suggestions, and form a closed-loop process from fault identification to plan generation and then to feedback optimization.

[0085] While outputting the standardized group fault identification results, the system will also provide multi-dimensional fault analysis, including the frequent occurrence period, duration, impact range of the fault, and the correlation with other network events. This information can not only help operation and maintenance personnel understand the full picture of the fault, but also guide them to make more reasonable repair decisions to ensure the stability of the network.

[0086] In the embodiments of the present application, an interface call function is also provided to achieve the timely identification and analysis of OLT group faults. Specifically, by specifying a time point (for example, every 5 minutes, analyzing the alarm data within the previous and next 5 minutes), a series of intelligent operations can be automatically triggered, including database query, upstream device analysis, alarm root cause location, etc., and finally return easy-to-understand fault diagnosis information and a link to the coping strategy.

[0087] Specifically, for example, select the alarm information within 10 minutes before and after 12:50:00 on February 24, 2025 after entering the OLT group fault ability scenario for group fault determination and identification. Based on the provided call time period, after the system executes the automated analysis process, the expected JSON format data to be returned includes but is not limited to the following fields, and the output parameters (API response example) are shown in Table 1:

[0088]

[0089]

[0090] In the embodiments of the present application, the construction of the knowledge graph is ingeniously combined with the real-time data processing technology of Function Calling, providing a brand-new group fault identification method for the field of communication network operation and maintenance. Specifically, by deeply mining the time series features, spatial distribution rules, and fault handling measures in historical alarm data and work order records, a detailed device graph and alarm graph are constructed. During the model training and fault identification process, the Function Calling mechanism allows the large model to call the status of network devices in real time, realizing the dynamic association of alarm information and resource information, effectively improving the accuracy and response speed of group fault identification. In addition, the system also has the ability of intelligent decision-making and pre-plan recommendation, and can quickly generate repair suggestions according to the fault identification results, further optimizing the fault handling process. That is, through its unique data processing and analysis strategy, the present application significantly enhances the intelligence and efficiency of the operation and maintenance support system, providing a revolutionary tool for the prevention and solution of network faults.

[0091] According to the embodiments of the present application, a group fault identification device is provided. It should be noted that the group fault identification device in the embodiments of the present application can be used to execute the group fault identification method provided in the embodiments of the present application. The following introduces the group fault identification device provided in the embodiments of the present application.

[0092] Figure 4 is a structural diagram of a group fault identification device provided according to the embodiments of the present application. As Figure 4 shown, the device includes:

[0093] An acquisition module 40, configured to acquire alarm information within a preset time period, where the alarm information is used to indicate that a network device is in an abnormal state;

[0094] A first determination module 42, configured to determine the device address corresponding to the alarm information, and determine the upstream network device corresponding to the device address according to the device graph, where the device graph is used to represent the association relationship between each network device;

[0095] A second determination module 44, configured to determine the alarm status of the upstream network device according to the alarm graph, and determine the faulty device from the upstream network devices according to the alarm status, where the alarm graph is used to represent the association relationship between each alarm event;

[0096] A third determination module 46, configured to determine the association rule of a faulty device during an alarm process through a fault identification model, and determine a group fault identification result according to the association rule, where the association rule is used to reflect the fault propagation mode and historical processing efficiency of the faulty device.

[0097] Through the acquisition module, the first determination module, the second determination module, and the third determination module in the above group fault identification device, the purpose of quickly and accurately identifying network group faults and their influence ranges is achieved, thereby realizing the technical effects of improving network operation and maintenance efficiency, shortening fault recovery time, and improving service quality. Furthermore, the technical problem that the group fault identification method in the related technology is inefficient, especially when facing a large number of device alarms and user complaints, it is difficult to quickly and accurately locate the cause of the fault, affecting the timely discovery and handling of network faults, is solved.

[0098] In the group fault identification device provided in the embodiment of the present application, the acquisition module is further configured to obtain alarm information within a preset time period from an operation and maintenance support center through a preset call mechanism, where the operation and maintenance support center stores a device map, an alarm map, and an operation and maintenance map, and the operation and maintenance map is used to provide repair guidance for faulty devices.

[0099] In the group fault identification device provided in the embodiment of the present application, a processing module 48 is further included, and the processing module is configured to convert the group fault identification result into a preset format, and provide a link to a preplan document corresponding to the group fault identification result, where the link to the preplan document includes fault handling measures determined according to the operation and maintenance map.

[0100] In the group fault identification device provided in the embodiment of the present application, the processing module is further configured to obtain link data of all network devices, and determine an initial device map according to the link data, where the link data includes device information of network devices and link information between network devices; obtain station and machine room data and power environment data of network devices, where the station and machine room data is used to represent station information and machine room information with clear geographical coordinates and type attributes, and the power environment data is used to represent environmental monitoring data in stations and machine rooms; optimize the initial device map according to the station and machine room data and the power environment data to obtain a device map.

[0101] In the group fault identification device provided in the embodiments of the present application, the processing module is further configured to obtain historical alarm data and historical work order data; extract features from the historical alarm data and the historical work order data to obtain a first feature sequence corresponding to the historical alarm data and a second feature sequence corresponding to the historical work order data, where the first feature sequence includes time series features and spatial distribution features corresponding to the historical alarm data, and the second feature sequence is used to reflect the relationship between historical alarm events and historical fault handling measures; train an initial model in a target application scenario according to the first feature sequence and the second feature sequence to obtain a fault identification model, where the target application scenario is a cluster offline fault scenario of an optical line terminal.

[0102] The embodiments of the present application further provide an electronic device, including: a memory and a processor, where the memory is used to store program instructions; the processor is connected to the memory and is used to execute the above-mentioned group fault identification method.

[0103] It should be noted that the above-mentioned electronic device is used to execute Figure 2 the group fault identification method shown, so the relevant explanations in the above-mentioned group fault identification method also apply to this electronic device and will not be elaborated here.

[0104] The embodiments of the present application further provide a non-volatile storage medium, which includes a stored computer program, where the device where the non-volatile storage medium is located executes the above-mentioned group fault identification method by running the computer program.

[0105] It should be noted that the above-mentioned non-volatile storage medium is used to execute Figure 2 the group fault identification method shown, so the relevant explanations in the above-mentioned group fault identification method also apply to this non-volatile storage medium and will not be elaborated here.

[0106] The embodiments of the present application further provide a computer program product, including computer instructions, which implement the above-mentioned group fault identification method when executed by a processor.

[0107] It should be noted that the above-mentioned computer program product is used to execute Figure 2 the group fault identification method shown, so the relevant explanations in the above-mentioned group fault identification method also apply to this computer program product and will not be elaborated here.

[0108] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0109] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0110] In several embodiments provided by this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0111] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0112] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0113] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of this application. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0114] The above is only the preferred implementation manner of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A method for identifying group disorders, characterized in that: include: Acquire alarm information within a preset time period, wherein the alarm information is used to prompt that the network device is in an abnormal state; Determine a device address corresponding to the alarm information, and determine an uplink network device corresponding to the device address according to a device map, wherein the device map is used to represent an association relationship between various network devices; Determine the alarm status of the upstream network device according to the alarm map, and determine the faulty device from the upstream network device according to the alarm status, wherein the alarm map is used to represent the correlation between various alarm events; The association law of the faulty device in the alarm process is determined through the fault identification model, and the group fault identification result is determined based on the association law, wherein the association law is used to reflect the fault propagation mode and historical processing efficiency of the faulty device.

2. The method according to claim 1, characterized in that Get alarm information within the preset time period, including: The alarm information within the preset time period is obtained from the operation and maintenance guarantee center through a preset calling mechanism, wherein the operation and maintenance guarantee center stores the equipment map, the alarm map and the operation and maintenance map, and the operation and maintenance map is used to provide repair guidance for the faulty equipment.

3. The method according to claim 2, characterized in that The method further comprises: The group fault identification result is converted into a preset format, and a link to a plan document corresponding to the group fault identification result is provided, wherein the plan document link includes a fault handling measure determined according to the operation and maintenance map.

4. The method according to claim 1, characterized in that: The device map is determined by: Acquire link data of all network devices, and determine an initial device map based on the link data, wherein the link data includes device information of the network devices and link information between the network devices; Acquire the station room data and power environment data of the network device, wherein the station room data is used to represent station information and room information with clear geographic coordinates and type attributes, and the power environment data is used to represent environmental monitoring data in the station and room; The initial equipment map is optimized according to the station room data and the power environment data to obtain the equipment map.

5. The method according to claim 4, characterized in that The device map is constructed with the resource objects corresponding to the network devices as nodes and the logical relationships between the resource objects as edges, wherein the logical relationships at least include the spatial position relationship, physical connection relationship and carrying relationship between the resource objects.

6. The method according to claim 1, characterized in that The alarm graph is constructed with the alarm events corresponding to the network devices as nodes and the causal relationships between the alarm events as edges.

7. The method according to claim 1, characterized in that The fault identification model is trained in the following way: Obtain historical alarm data and historical work order data; Performing feature extraction on the historical alarm data and the historical work order data to obtain a first feature sequence corresponding to the historical alarm data and a second feature sequence corresponding to the historical work order data, wherein the first feature sequence includes time series features and spatial distribution features corresponding to the historical alarm data, and the second feature sequence is used to reflect the relationship between historical alarm events and historical processing measures; According to the first feature sequence and the second feature sequence, the initial model is trained in a target application scenario to obtain the fault identification model, wherein the target application scenario is a cluster offline fault scenario of an optical line terminal.

8. The method according to claim 1, characterized in that The association rules of the faulty device in the alarm process include: the frequent time periods, durations and co-line modes with other network events of various fault types of the faulty device, wherein the fault types at least include: cluster offline faults of optical line terminals.

9. A group obstacle identification device, characterized in that: include: An acquisition module, used to acquire alarm information within a preset time period, wherein the alarm information is used to prompt that the network device is in an abnormal state; A first determination module is used to determine a device address corresponding to the alarm information, and determine an uplink network device corresponding to the device address according to a device map, wherein the device map is used to represent an association relationship between various network devices; A second determination module is used to determine the alarm status of the uplink network device according to the alarm map, and determine the faulty device from the uplink network device according to the alarm status, wherein the alarm map is used to represent the correlation between each alarm event; The third determination module is used to determine the association law of the faulty device in the alarm process through the fault identification model, and determine the group fault identification result based on the association law, wherein the association law is used to reflect the fault propagation mode and historical processing efficiency of the faulty device.

10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory is used to store program instructions; The processor is connected to the memory and is used to execute the group obstacle identification method described in any one of claims 1 to 8.

11. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the group obstacle identification method according to any one of claims 1 to 8 by running the computer program.

12. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the group obstacle identification method described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Root cause analysis method, device and equipment and computer storage medium

    CN112152852A

  • Network fault determination method and network equipment

    CN113839804A

  • Fault reason determination method and device, equipment and storage medium

    CN117459365A

  • Telecommunication network alarm method and system

    CN119011422A

  • Optical network fault positioning method and device and related equipment

    CN119070899A