System fault locating method, system and electronic device
By acquiring and analyzing various types of data in the network management system, and utilizing mapping data tables and preset analysis models, the problem of insufficient root cause analysis of alarms in existing technologies has been solved, enabling rapid fault location and efficient handling, and improving operation and maintenance efficiency and system stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER
- Filing Date
- 2024-12-31
- Publication Date
- 2026-05-29
AI Technical Summary
Existing network management systems lack effective methods for alarm root cause analysis, resulting in slow response to faults, which affects operational efficiency and customer experience.
By acquiring network system topology data, network data, alarm data, device performance data, and log data, and using mapping data tables and preset analysis models, alarm information is transformed, organized, and analyzed to accurately locate the root cause of alarms and generate solutions, and faults are handled automatically or manually.
It improves the accuracy of fault location and operational efficiency, reduces operational costs, and ensures that the system can quickly return to normal.
Smart Images

Figure CN119814536B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of deep learning technology, and more specifically, to a system fault location method, system, and electronic device. Background Technology
[0002] The widespread application of 5G and other network services has made network operation and maintenance scenarios more complex and diverse, with maintenance personnel frequently dealing with large amounts of real-time and historical data. Existing network management systems lack the technical means for alarm root cause analysis and application processes, making it difficult to provide sufficient support for maintenance personnel. Many problems cannot be responded to and resolved quickly, leading to the spread and escalation of faults, ultimately affecting customer experience.
[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0004] The purpose of this disclosure is to provide a system fault location method, system, and electronic equipment to improve the operation and maintenance efficiency of complex optical network systems.
[0005] According to a first aspect of the present disclosure, a system fault location method is provided, comprising: acquiring preset original information of a target network system, the preset original information including at least topology data, network data, alarm data, device performance data, and log data of the target network system; converting and organizing the alarm data according to a first mapping data table to form an alarm information set; determining candidate root cause alarm devices corresponding to alarm information in the alarm information set according to a second mapping data table; determining an alarm root cause, a root cause alarm device, and a solution measure according to the alarm information set, the candidate root cause alarm devices, and the preset original information; forming a control signal for the root cause alarm device according to the solution measure; sending the control signal to the root cause alarm device, or outputting the solution measure.
[0006] According to a second aspect of the present disclosure, a system fault location system is provided, comprising: a network data collection module, an alarm data processing module, a correlation mining module, and a preset analysis model, wherein: the network data collection module is used to acquire preset raw information of a target network system and send the preset raw information to the alarm data processing module and the preset analysis model, the preset raw information including at least topology data, network data, alarm data, device performance data, and log data of the target network system; the alarm data processing module is used to extract the alarm data from the preset raw information and send the alarm data to the correlation mining module; the correlation mining module is used to acquire preset raw information of a target network system and send the preset raw information to the alarm data processing module and the preset analysis model; the network data collection module is used to acquire preset raw information of a target network system and send the preset raw information to the alarm data processing module and the preset analysis model, the network data collection module is used to acquire preset raw information, and send the preset raw information to the alarm data processing module and the preset analysis model, the network data collection module is used to acquire preset raw information, and send the preset raw information to the alarm data processing module and the preset analysis model, the network data collection module is used to acquire preset raw information, and send the preset raw information to the alarm data processing module and the alarm data processing module; the correlation mining ... The joint data mining module is used to transform and organize the alarm data according to the first mapping data table to form an alarm information set, determine the candidate root cause alarm devices corresponding to the alarm information in the alarm information set according to the second mapping data table, and send the organized alarm information set and the candidate root cause alarm devices to the preset analysis model. The preset analysis model is used to determine the alarm root cause, root cause alarm device and solution measures according to the alarm information set, the candidate root cause alarm devices and the preset original information, generate control signals for the root cause alarm devices according to the solution measures, send the control signals to the root cause alarm devices, or output the solution measures.
[0007] According to a third aspect of this disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to perform the method as described in any of the preceding methods based on instructions stored in the memory.
[0008] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided having a program stored thereon that, when executed by a processor, implements the system fault location method as described in any of the preceding claims.
[0009] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of the method as described in any of the preceding claims.
[0010] This embodiment of the disclosure automatically collects various preset raw information from the target model, transforms and organizes the alarm data in the preset raw information, and can effectively identify alarm data from a large number of devices in a complex system, improving the efficiency of subsequent analysis. By processing the preset raw information, the organized alarm information set, and the candidate root cause alarm devices using a trained preset analysis model, multiple information can be effectively integrated to accurately locate the alarm cause and the root cause alarm device, and automatically generate corresponding solutions. When the solution can be automatically executed by the root cause alarm device, a control signal is sent to the root cause alarm device to automatically eliminate or reduce the system alarm. When the solution cannot be automatically executed by the root cause alarm device, the solution is automatically output for maintenance personnel to handle, which can greatly improve maintenance efficiency and the accuracy of system fault location, enabling the system to return to normal as soon as possible and reducing maintenance costs.
[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0013] Figure 1 This is a flowchart of a system fault location method in an exemplary embodiment of this disclosure.
[0014] Figure 2 This is a sub-flowchart of step S2 in an exemplary embodiment of this disclosure.
[0015] Figure 3 This is a schematic diagram of a system fault location system in an exemplary embodiment of this disclosure.
[0016] Figure 4 This is a schematic diagram of the system fault location system 300 in an exemplary embodiment of this disclosure.
[0017] Figure 5 This is a schematic diagram of the system fault location system 300 in an exemplary embodiment of this disclosure.
[0018] Figure 6 This is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0019] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0020] Furthermore, the accompanying drawings are merely illustrative of this disclosure, and the same reference numerals in the drawings denote the same or similar parts, thus repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0021] The exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0022] Figure 1 This is a flowchart of a system fault location method in an exemplary embodiment of this disclosure.
[0023] refer to Figure 1 The system fault location method 100 may include:
[0024] Step S1: Obtain preset original information of the target network system. The preset original information includes at least the topology data, network data, alarm data, device performance data, and log data of the target network system.
[0025] Step S2: The alarm data is converted and organized according to the first mapping data table to form an alarm information set, and the candidate root cause alarm devices corresponding to the alarm information in the alarm information set are determined according to the second mapping data table.
[0026] Step S3: Determine the alarm root cause, root cause alarm device, and solution based on the alarm information set, the candidate root cause alarm device, and the preset original information.
[0027] Step S4: Generate a control signal for the root cause alarm device according to the solution, and send the control signal to the root cause alarm device, or output the solution.
[0028] This embodiment of the disclosure automatically collects various preset raw information from the target model, transforms and organizes the alarm data in the preset raw information, and can effectively identify alarm data from a large number of devices in a complex system, improving the efficiency of subsequent analysis. By processing the preset raw information, the organized alarm information set, and the candidate root cause alarm devices using a trained preset analysis model, multiple information can be effectively integrated to accurately locate the alarm cause and the root cause alarm device, and automatically generate corresponding solutions. When the solution can be automatically executed by the root cause alarm device, a control signal is sent to the root cause alarm device to automatically eliminate or reduce the system alarm. When the solution cannot be automatically executed by the root cause alarm device, the solution is automatically output for maintenance personnel to handle, which can greatly improve maintenance efficiency and the accuracy of system fault location, enabling the system to return to normal as soon as possible and reducing maintenance costs.
[0029] The method disclosed in this embodiment can be used to locate faults in a target network system. The target network system may include multiple network devices, at least one of which may be an optical network device. An optical network device refers to a collection of devices that realize functions such as generating, modulating, amplifying, transmitting, receiving, and converting optical signals to electrical signals based on the principles of optical communication technology. It may include optical transmission equipment, optical transport network equipment, optical access equipment, optical switching equipment, etc.
[0030] The following is a detailed explanation of each step in the system fault location method 100.
[0031] In step S1, the preset original information of the target network system is obtained. The preset original information includes at least the topology data, network data, alarm data, device performance data, and log data of the target network system.
[0032] The target network system typically encompasses various types and levels of devices, which continuously generate various types of data during their operation. In step S1, data can be collected from various data sources within the target network system. These data sources are widely distributed across different types of network devices, servers, and various software application systems. The raw data collected at a single sampling point can include at least topology data, network data, alarm data, device performance data, and log data.
[0033] Topology data contains network and device connectivity information, with granularity ranging from network element level, board level, port level, or time slot level. It describes the connection links, hierarchical architecture, and data transmission paths between various device nodes, including backbone links between core switches and edge routers, as well as the interconnection architecture between optical network devices and server clusters.
[0034] The network data records detailed configuration information and key identifiers for each network element, such as its model, configuration parameters, and region, which are used to accurately locate problematic devices later.
[0035] Alarm data provides real-time feedback on various emergencies occurring on the device, such as port failures, overheating, and bandwidth congestion. Its timeliness and accuracy directly affect the efficiency of troubleshooting. If the system is operating normally, alarm data may not be available at the sampling time.
[0036] The equipment performance data is continuously monitored, including key indicators such as optical power, CPU utilization, memory usage, bit error rate, and throughput.
[0037] The log data fully records every key operation, every abnormal event, and the corresponding timestamp of the device from startup, operation to maintenance, providing a basis for tracing the fault occurrence process.
[0038] Table 1 shows the types of data included in exemplary alarm data and device performance data.
[0039] Table 1
[0040]
[0041] The aforementioned types of preset raw information are diverse and numerous. When the number of devices in the target network system is large, they can provide a complete description of the target network system's state at the sampling time point when these preset raw information are collected.
[0042] After obtaining the preset raw information corresponding to a sampling time point, the obtained raw information can be purified based on preset data cleaning rules to remove potentially erroneous, duplicate, and incomplete data fragments, ensuring the data quality for subsequent analysis. For example, for alarm data, false alarms caused by temporary sensor malfunctions are removed; for log data, redundant records generated by routine system maintenance operations are filtered out. Through precise data cleaning, the data entering the subsequent analysis process becomes authentic and reliable.
[0043] In step S2, the alarm data is transformed and organized according to the first mapping data table to form an alarm information set, and the candidate root cause alarm devices corresponding to the alarm information in the alarm information set are determined according to the second mapping data table.
[0044] Because the preset raw information is complex and diverse, alarm data sources are widely distributed across different types of network devices, servers, and various software application systems, with diverse data formats, including but not limited to text strings, binary code, and vendor-specific custom formats. For example, a traditional network switch might output port fault alarms in simple text format, recorded as "Port03Down"; while optical network devices use structured XML format to report optical module anomalies. Furthermore, each vendor, product line, and product model has corresponding alarm handling suggestions. Therefore, before conducting comprehensive analysis, the alarm data must first be converted and organized. Simultaneously, if no alarm data is found in the currently collected preset raw information, the execution of the current fault location method can be suspended until the next sampling time point. In some embodiments, method 100 can be started only when alarm data is detected.
[0045] Figure 2 This is a sub-flowchart of step S2 in an exemplary embodiment of this disclosure.
[0046] refer to Figure 2 In an exemplary embodiment, step S2 may include:
[0047] Step S21: Determine the alarm category and alarm device corresponding to each alarm data according to the first mapping data table. The first mapping data table includes the mapping relationship between device model, alarm code, and alarm category.
[0048] Step S22: Based on the alarm categories and alarm devices corresponding to multiple alarm data, determine the list of alarm devices corresponding to each alarm category, and / or the list of alarm categories corresponding to each alarm device;
[0049] Step S23: Form an alarm information set based on the list of alarm devices corresponding to each alarm category and / or the list of alarm categories corresponding to each alarm device.
[0050] The first mapping data table is primarily used for translating alarm data. Alarm information from different manufacturers and models varies greatly in format and expression, some using specialized codes and others lengthy English descriptions, posing a significant obstacle to unified analysis. Based on industry standards, operational experience, and accumulated past fault cases, the first mapping data table integrates industry-standard alarm specifications from mainstream manufacturers and conversion rules derived from extensive operational practices. It can standardize and convert diverse alarm data, organizing ambiguous information into a clear and semantically precise set of alarm messages. For example, it unifies different alarm expressions regarding link interruption from various devices into the format "Link Interruption [Specific Link Identifier], Fault Level [High / Medium / Low]". For text-based alarms, it uses natural language processing techniques such as text cleaning, word segmentation, and part-of-speech tagging to remove redundant information and extract key descriptive terms. For instance, it converts "Port03Down" into "Device Port Fault [Port Number: 03]", making its semantics clearer and more intuitive. For structured data, it reorganizes the data according to predetermined field mapping relationships by writing appropriate parsing scripts. For example, converting the XML format optical module alarm into "Optical Module Alarm [Module Number: OpticalModule01, Fault Type: PowerAbnormal, Time: 2024-12-30 10:23:45]" ensures that alarms from different sources can be unified into a standard structured expression.
[0051] After format conversion, the alarm information is categorized and aggregated. Alarm information is grouped based on type, impact scope, severity, and other dimensions. For example, all alarms related to network connection interruptions are grouped into one category, covering alarms from different devices related to connection problems, such as router port outages, link failures, and network card failures. Alarms caused by device resource overload, such as high CPU utilization and memory overflow, are grouped into resource alarm groups. During this process, each alarm category can be assigned a priority, typically determined by the degree of impact on business continuity and the scope of the affected area. For instance, core link interruption alarms causing widespread business paralysis are given the highest priority, while alarms about abnormal fan speeds on individual devices are given a lower priority.
[0052] The next step is the integration of related alarm information. Considering that many faults do not occur in isolation, a single root cause may trigger alarms in multiple devices or links. In some cases, factors such as the time sequence of alarm occurrences and device topology connections correspond to fixed related alarms. For example, if an optical power decrease alarm occurs in succession within a short period of time, followed by an increased bit error rate alarm on a downstream switch port connected to it, these two alarms are highly correlated and correspond to a fixed alarm scenario: "Optical network link failure risk [Source device: Optical network device [ID: XX], Affected device: Downstream switch [Port number: YY], Related alarms: Optical power decrease, increased bit error rate]".
[0053] This process can be achieved by querying the second mapping data table. The second mapping data table is primarily used to identify potential root cause alarm devices corresponding to the alarm information based on the pre-defined alarm and device fault association logic and the translated alarm data. The alarm and device fault association logic can be provided by the equipment manufacturer or by maintenance personnel based on experience. It can identify corresponding potential root cause alarm devices based on the device connection topology and fault propagation model. For example, when an alarm is found on a certain link based on the translated alarm data, potential fault source devices such as the optical network transceiver and connected switch ports corresponding to that alarm type on that link can be quickly located, narrowing down the troubleshooting scope.
[0054] Finally, the alarm information, after being processed through multiple steps such as conversion, classification, and integration, is gathered in an orderly manner to form a complete alarm information set, as well as the candidate root cause alarm devices corresponding to this alarm information set.
[0055] It should be noted that the candidate root cause alarm devices are not the final analysis results, but only the location results obtained based on a fixed mapping relationship. The specific cause of the fault still needs further analysis.
[0056] By introducing mapping queries, alarm data can be simplified, and candidate root cause alarm devices can be provided based on fixed matching relationships, thereby reducing the computational load of subsequent analysis and improving fault location efficiency.
[0057] In step S3, the root cause of the alarm, the root cause alarm device, and the solution are determined based on the alarm information set, the candidate root cause alarm device, and the preset original information.
[0058] In an exemplary embodiment, an alarm information set, candidate root cause alarm devices, and preset raw information can be input into a preset analysis model. The preset analysis model outputs the alarm root cause, the root cause alarm device, and the corresponding solution for the alarm root cause and / or the solution for the root cause alarm device. To maintain the similarity between the training data and the input data actually processed by the model, the preset raw data can also be used to input the alarm information set and candidate root cause alarm devices.
[0059] Preset analysis models include, for example, neural network models and generative large-scale models. Neural network models construct multi-layered neuron structures to process input data layer by layer, ultimately outputting accurate prediction results. With their powerful self-learning and non-linear mapping capabilities, they can deeply uncover hidden patterns and rules when faced with massive and complex data.
[0060] The preset analysis model can be trained and updated with parameters using historical data, and then automatically obtain analysis results based on input data. The training data (historical data) of the preset analysis model includes preset raw information corresponding to multiple sampling time points, the root cause of the alarm corresponding to the preset raw information, the root cause alarm device, the solution, and the feedback information of the solution. The feedback information includes manual debugging feedback information and / or feedback parameters automatically obtained from the root cause alarm device or the target network system.
[0061] The pre-defined analysis model can use machine learning algorithms to perform deep learning on pre-defined raw information, analysis results (root cause of alarms / root cause alarm device / solution), and feedback information to uncover the root cause of alarms hidden beneath the surface. For example, multiple seemingly isolated device port alarms may belong to a pattern, namely, optical power attenuation caused by the aging of optical modules in upstream optical network equipment, which in turn triggers a chain reaction. Therefore, the root cause alarm device corresponding to the alarm information of these device ports can be set as that optical network equipment, and the root cause of the alarm can be determined as the aging of optical modules in that optical network equipment.
[0062] If the optical power of an optical network device connected to a certain optical link is continuously below a threshold, the affected service range corresponding to the optical network device is located, and the fault may be concentrated in the link and the upstream and downstream devices related to the services it carries. The root cause alarm devices of these upstream and downstream devices are set to the optical network device, and the alarm root cause is set to the optical network device's continuous optical power being below the threshold.
[0063] Furthermore, the pre-defined analysis model can extract information fragments closely related to the time of the fault, the type of fault, and the identifier of the faulty device from complex alarm texts and log records using intelligent text analysis algorithms. For example, when a large-scale service interruption occurs in the system, the model can extract a "high temperature alarm of optical module" triggered by a certain optical network device at a specific time and a series of related "automatic link switching" records from massive logs. By following this clue, the model can pinpoint the root cause of the fault as a transmission anomaly caused by the overheating of the optical module, identify the root cause alarm device as the optical network device, and determine the root cause of the alarm as the overheating of the optical module.
[0064] In some embodiments, knowledge from a knowledge base can be pre-input into a preset analysis model during training. This knowledge base can store a large number of past fault cases, fault handling experience, and corresponding solutions. Thus, the preset analysis model can compare the current fault's characteristic information, such as topology location, alarm combinations, and performance degradation trends, with each entry in the knowledge base to match the closest historical case and draw upon its mature handling procedures. For example, if the current fault exhibits characteristics similar to a previous "optical network equipment board failure causing partial service lag," the root cause alarm device can be identified as the optical network equipment, and the root cause of the alarm can be determined as a board failure.
[0065] Therefore, the root cause of the alarm can be located, along with at least one root cause alarm device involved; or, at least one root cause alarm device can be located, along with the corresponding root cause alarm device. It should be noted that the device that caused the alarm may not itself have an alarm; therefore, the root cause alarm device may not necessarily exhibit an alarm phenomenon, nor may it necessarily have a root cause. However, an alarm root cause generally involves at least one root cause alarm device that led to the generation of the alarm data.
[0066] Once the root cause alarm device is identified, the preset analysis module can automatically determine the solution based on the built-in knowledge base and the corresponding troubleshooting measures for the device model of the root cause alarm device. This includes hardware detection, hardware replacement, software reset, software parameter adjustment instructions, emergency switching plans, etc., which can accelerate the fault repair process and improve the system availability recovery efficiency.
[0067] In an exemplary embodiment, step S3 may also receive manual input information, which is used to set candidate fault influencing factors, including special events, special times, and special equipment; and determine the alarm root cause, root cause alarm equipment, and solution measures based on the candidate fault influencing factors, alarm information set, candidate root cause alarm equipment, and preset original information.
[0068] For example, major sales events like Singles' Day (November 11th) and the Spring Festival travel rush can put significant pressure on system operation and maintenance, leading to a series of obvious fault characteristics. Therefore, this embodiment allows input of candidate fault influencing factors into a preset analysis model, enabling the model to refer to these factors and provide a final analysis conclusion. To maintain consistency between use and training, these candidate fault influencing factors can be used during the training phase.
[0069] In step S4, a control signal for the root cause alarm device is generated according to the solution, and the control signal is sent to the root cause alarm device, or the solution is output.
[0070] When the solution is ready for automated execution, control signals for the root cause alarm device can be quickly generated based on instructions and delivered to the root cause alarm device in the form of electrical signals or network commands. For example, for device alarms caused by software parameter errors, the control signal carries the corrected parameter values, remotely triggering the device's software module to automatically update, instantly completing the repair action, and the system alarm is eliminated or significantly reduced, restoring smooth service. In some complex scenarios, such as when replacing hardware modules of large optical network equipment, the solution cannot be automatically executed by the device. In this case, detailed solutions can be automatically output in the form of a report, including fault description, root cause analysis, required spare parts list, operating steps, and precautions. This allows maintenance personnel to carry out on-site repairs efficiently and accurately based on the report, ensuring the system returns to normal as soon as possible, minimizing maintenance costs, and comprehensively improving the system's reliability and stability.
[0071] Corresponding to the above method embodiments, this disclosure also provides a system fault location system, which can be used to execute the above method embodiments.
[0072] Figure 3 This is a schematic diagram of a system fault location system in an exemplary embodiment of this disclosure.
[0073] refer to Figure 3 The system fault location system 300 may include: a network data collection module 1, an alarm data processing module 2, a correlation mining module 3, and a preset analysis model 4, wherein:
[0074] The network data collection module 1 is used to acquire the preset raw information of the target network system and send the preset raw information to the alarm data processing module 2 and the preset analysis model 4. The preset raw information includes at least the topology data, network data, alarm data, device performance data and log data of the target network system.
[0075] Alarm data processing module 2 is used to extract alarm data from preset raw information and send the alarm data to association mining module 3;
[0076] The correlation mining module 3 is used to transform and organize alarm data according to the first mapping data table to form an alarm information set, determine the candidate root cause alarm devices corresponding to the alarm information in the alarm information set according to the second mapping data table, and send the organized alarm information set and candidate root cause alarm devices to the preset analysis model 4.
[0077] The preset analysis model 4 is used to determine the root cause of the alarm, the root cause alarm device, and the solution based on the alarm information set, the candidate root cause alarm device, and the preset original information. Based on the solution, it generates a control signal for the root cause alarm device and sends the control signal to the root cause alarm device, or outputs the solution.
[0078] Among them, the alarm data processing module 2 is used to filter alarm data from all the collected preset raw information, and the correlation mining module is used to process the alarm data from the alarm data processing module 2 as in step S2 of method 100, and output the alarm information set and the candidate root cause alarm devices.
[0079] The output data of the correlation mining module 3 (including but not limited to alarm information sets and candidate root cause alarm devices) is shown in Table 2.
[0080] Table 2:
[0081]
[0082] In an exemplary embodiment, the network data collection module 1 and the association mining module 3 can integrate the input data from the preset analysis module 4 and form it in the following format:
[0083] Alarm data: XXXX.
[0084] Performance data: XXXX.
[0085] Topology information: XXXX.
[0086] Configuration information: XXXX.
[0087] Log information: XXXX.
[0088] RootAlarmConfidence:XXXX.
[0089] RootAlarmState:XXXX.
[0090] RootAlarmLevel:XXXX.
[0091] RootAlarmNamel:XXXX.
[0092] DerivedAlarms:XXXX.
[0093] Number of Alarms: XXXX.
[0094] NumberofAffectedConnections:XXXX.
[0095] The output data of the preset analysis model 4 (including but not limited to root cause alarm devices and alarm root causes) can be shown in Table 3.
[0096] Table 3:
[0097]
[0098] The solutions output by the preset analysis model 4 include, but are not limited to, those shown in Table 4.
[0099] Table 4:
[0100]
[0101] The specific steps for the network data collection module 1, alarm data processing module 2, correlation mining module 3, and preset analysis model 4 mentioned above are described in the above method embodiment and will not be repeated here.
[0102] Figure 4 This is a schematic diagram of the system fault location system 300 in an exemplary embodiment of this disclosure.
[0103] refer to Figure 4 In some embodiments, the system fault location system 300 further includes an information storage module 5. The information storage module 5 stores a first mapping data table, a second mapping data table, and manually input information. The manually input information is then input into a preset analysis model 4. This manually input information is used to set candidate fault influencing factors, including special events, special times, and special equipment. The information storage module 5 can be implemented as a knowledge base, pre-storing more types of preset knowledge (defined relationships). This disclosure does not impose any special limitations on this.
[0104] Figure 5 This is a schematic diagram of the system fault location system 300 in an exemplary embodiment of this disclosure.
[0105] refer to Figure 5 In some embodiments, the system fault location system 300 further includes a training data processing module 6, which is used to receive historical data and form training data for the target analysis model based on the historical data. The historical data includes at least preset original information corresponding to multiple sampling time points, alarm root causes corresponding to the preset original information, root cause alarm devices, solutions, and feedback information of the solutions. The feedback information includes manual debugging feedback information and / or feedback parameters automatically obtained from root cause alarm devices or target network systems.
[0106] Combination Figure 5 This section provides an interactive diagram illustrating the various modules within System 300.
[0107] In step S1: Network data collection module 1 obtains preset raw information from the target network system through data acquisition methods.
[0108] In step S2: The network data collection module 1 synchronously transmits topology information, network element information, alarm information, performance data and logs to the alarm data processing module 2 and the preset analysis model 4.
[0109] In step S3: the alarm data processing module 2 inputs the extracted alarm data into the correlation mining module 3.
[0110] In step S4: The association mining module 3 outputs a set of alarm information and candidate root cause alarm devices based on the alarm data.
[0111] In step S5: The preset analysis model 4 performs root cause analysis of network faults based on the received information and provides root cause analysis results, maintenance operations or repair suggestions.
[0112] In step S6: In the execution unit, if the network resources can meet the requirements of automatic operation and maintenance, the control signal is automatically output to the target network system and the direct input to the root cause alarm device is set. If the requirements cannot be met, the relevant analysis results and targeted operation and maintenance suggestions are output for reference during the manual intervention stage.
[0113] Step S7: The target network system performs relevant operations to optimize and adjust the root cause alarm devices.
[0114] Step S8: The training data processing module 6 collects the input and output data of the preset analysis model 4, as well as the feedback information of the target network system on the solution, and merges and splices the three types of data to continuously accumulate the training dataset of the large model for optimizing the analysis capability of the preset analysis model 4.
[0115] Feedback information includes manual debugging feedback and / or feedback parameters automatically obtained from root cause alarm devices or target network systems.
[0116] The method and system proposed in this disclosure have several significant characteristics and advantages. Firstly, in terms of overall architecture, the application process of alarm root cause analysis in optical networks is clearly defined. The input and output data parameters and formats required for the system process are meticulously outlined and defined, covering all aspects from system data acquisition and processing, training data acquisition and processing, alarm correlation analysis, root cause analysis to decision-making and resource allocation. This provides a comprehensive and systematic solution for optical network management. Crucially, this solution emphasizes network resource management and optimization during optical network operation and maintenance. It generates accurate prediction results through data processing, enabling scientific adjustment decisions and optimizing resource allocation and adjustment cycles. This significantly improves the management efficiency of optical networks and is widely applicable to fault root cause localization for different optical network bearer technologies. Furthermore, this method is not limited to specific prediction algorithms. During system operation, it continuously optimizes the model based on accumulated data, enhancing the model's analytical capabilities in practical applications while reducing the need for external intervention.
[0117] From a practical perspective, the embodiments disclosed herein demonstrate superior performance. On one hand, they propose a complete model and process for optical network fault root cause localization, comprehensively covering key aspects such as data collection and processing, alarm correlation analysis, alarm root cause analysis, and model optimization. Furthermore, they clearly define data input, output, and model prompt word formats, highly aligning with the usability principles of network operation. On the other hand, they propose a systematic management method for optical network operation and maintenance. This method has broad applicability, seamlessly integrating with fault alarm root cause localization and related decision-making processes in any optical transport network. It is easily integrated with existing network management controller systems, exhibiting stronger adaptability and usability compared to general root cause localization methods. Moreover, it not only generates accurate analysis results but also refines specific network adjustment technical solutions into two types: automatic maintenance optimization and manual intervention, closely linking prediction results with actual operations and effectively meeting various specific needs in network operation.
[0118] By combining collected actual decision-making data and network system parameters, carefully crafted and optimized training data is used for model reinforcement training, continuously enhancing the model's adaptability and prediction accuracy. Furthermore, the pre-defined analysis model possesses strong flexibility; depending on the specific task, only minor adjustments to the training data are needed to adapt to different optical network system management architectures and achieve the target functionality.
[0119] In summary, the embodiments disclosed herein provide a comprehensive, efficient, and practical solution for optical network alarm / fault root cause localization and operation and maintenance management, which is expected to drive the level of optical network resource management to a new level.
[0120] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0121] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided.
[0122] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: entirely in hardware, entirely in software (including firmware, microcode, etc.), or in a combination of hardware and software, collectively referred to herein as “circuit,” “module,” or “system.”
[0123] The following reference Figure 6To describe an electronic device 600 according to this embodiment of the present invention. Figure 6 The electronic device 600 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0124] like Figure 6 As shown, the electronic device 600 is manifested in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, and a bus 630 connecting different system components (including storage unit 620 and processing unit 610).
[0125] The storage unit stores program code that can be executed by the processing unit 610, causing the processing unit 610 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of the present invention. For example, the processing unit 610 can perform the method shown in the embodiments of this disclosure.
[0126] Storage unit 620 may include readable media in the form of volatile storage units, such as random access memory (RAM) 6201 and / or cache memory 6202, and may further include read-only memory (ROM) 6203.
[0127] Storage unit 620 may also include a program / utility 6204 having a set (at least one) program module 6205, such program module 6205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.
[0128] Bus 630 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0129] Electronic device 600 can also communicate with one or more external devices 700 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 600, and / or with any device that enables electronic device 600 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 650. Furthermore, electronic device 600 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 660. As shown, network adapter 660 communicates with other modules of electronic device 600 via bus 630. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0130] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0131] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of the invention may also be implemented as a program product comprising program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of the invention described in the "Exemplary Methods" section of this specification.
[0132] The program product for implementing the above-described method according to embodiments of the present invention may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0133] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0134] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0135] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0136] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0137] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0138] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and concept of this disclosure are indicated by the claims.
Claims
1. A system fault location method, characterized in that, include: Obtain preset raw information of the target network system, wherein the preset raw information includes at least the topology data, network data, alarm data, device performance data, and log data of the target network system; The alarm data is transformed and organized according to the first mapping data table to form an alarm information set. The candidate root cause alarm devices corresponding to the alarm information in the alarm information set are determined according to the second mapping data table. The first mapping data table is used to translate alarm data to unify alarm data from different sources into a standard structured expression. The second mapping data table is used to determine the candidate root cause alarm devices corresponding to the alarm information based on the preset alarm and device fault association logic and the translated alarm data. The alarm root cause, root cause alarm device, and solution are determined based on the alarm information set, the candidate root cause alarm device, and the preset original information. The solution is used to generate a control signal for the root cause alarm device, and the control signal is sent to the root cause alarm device, or the solution is output.
2. The system fault location method as described in claim 1, characterized in that, The alarm data is transformed and organized according to the first mapping data table to form an alarm information set, including: The alarm category and alarm device corresponding to each alarm data are determined according to the first mapping data table. The first mapping data table includes the mapping relationship between device model, alarm code, and alarm category. Based on the alarm categories and alarm devices corresponding to multiple alarm data, determine a list of alarm devices corresponding to each alarm category, and / or a list of alarm categories corresponding to each alarm device; The alarm information set is formed based on the list of alarm devices corresponding to each alarm category, and / or the list of alarm categories corresponding to each alarm device.
3. The system fault location method as described in claim 1, characterized in that, The step of determining the alarm root cause, root cause alarm device, and solution based on the alarm information set, the candidate root cause alarm devices, and the preset original information includes: The alarm information set, the candidate root cause alarm devices, and the preset original information are input into a preset analysis model. The preset analysis model outputs the alarm root cause, the root cause alarm device, and the corresponding solution for the alarm root cause and / or the corresponding solution for the root cause alarm device.
4. The system fault location method as described in claim 3, characterized in that, The training data of the preset analysis model includes preset raw information corresponding to multiple sampling time points, alarm root causes corresponding to the preset raw information, root cause alarm devices, solutions, and feedback information of the solutions. The feedback information includes manual debugging feedback information and / or feedback parameters automatically obtained from the root cause alarm devices or the target network system.
5. The system fault location method as described in claim 1, characterized in that, The step of determining the alarm root cause, root cause alarm device, and solution based on the alarm information set, the candidate root cause alarm devices, and the preset original information includes: Receive manual input information, which is used to set candidate fault influencing factors, including special events, special times, and special equipment; The root cause of the alarm, the root cause alarm device, and the solution are determined based on the candidate fault influencing factors, the alarm information set, the candidate root cause alarm device, and the preset original information.
6. A system fault location system, characterized in that, It includes a network data collection module, an alarm data processing module, a correlation mining module, and a preset analysis model, among which: The network data collection module is used to acquire preset raw information of the target network system and send the preset raw information to the alarm data processing module and the preset analysis model. The preset raw information includes at least the topology data, network data, alarm data, device performance data and log data of the target network system. The alarm data processing module is used to extract the alarm data from the preset original information and send the alarm data to the association mining module; The association mining module is used to transform and organize the alarm data according to the first mapping data table to form an alarm information set, determine the candidate root cause alarm devices corresponding to the alarm information in the alarm information set according to the second mapping data table, and send the organized alarm information set and the candidate root cause alarm devices to the preset analysis model. The first mapping data table is used to translate alarm data to unify alarm data from different sources into a standard structured expression. The second mapping data table is used to determine the candidate root cause alarm devices corresponding to the alarm information according to the translated alarm data based on the preset alarm and device fault association logic. The preset analysis model is used to determine the root cause of the alarm, the root cause alarm device, and the solution based on the alarm information set, the candidate root cause alarm device, and the preset original information. It generates a control signal for the root cause alarm device based on the solution and sends the control signal to the root cause alarm device, or outputs the solution.
7. The system fault location system as described in claim 6, characterized in that, It also includes an information storage module, which is used to store the first mapping data table, the second mapping data table and the manual input information, and input the manual input information into the preset analysis model. The manual input information is used to set the candidate fault influencing factors, which include special events, special times and special equipment.
8. The system fault location system as described in claim 6, characterized in that, It also includes a training data processing module, which is used to receive historical data and form training data for the preset analysis model based on the historical data. The historical data includes at least preset original information corresponding to multiple sampling time points, alarm root causes corresponding to the preset original information, root cause alarm devices, solutions, and feedback information of the solutions. The feedback information includes manual debugging feedback information and / or feedback parameters automatically obtained from the root cause alarm devices or the target network system.
9. An electronic device, characterized in that, include: Memory; as well as A processor coupled to the memory, the processor being configured to perform the method as described in any one of claims 1-5 based on instructions stored in the memory.
10. A computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the method as claimed in any one of claims 1-5.