Communication network fault analysis system and method
By designing a communication network fault analysis system, using data acquisition, preprocessing, fault analysis and evaluation and fault diagnosis modules, the problem that network fault location depends on personal abilities is solved, and fast and accurate fault analysis and positioning is achieved, improving network management efficiency.
Patent Information
- Application Number
- CN202510510143.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology relies on personal professional capabilities and experience in network failure location, and cannot quickly and accurately identify the scope of business impact and risks, and cannot quickly and intuitively judge the risk points of the entire network.
A communication network fault analysis system is designed, including a data acquisition module, a preprocessing module, a fault analysis and evaluation module, and a fault diagnosis and early warning module. The system collects network data in real time, performs data preprocessing, uses rules engines and statistical model libraries to perform fault analysis and early warning, and combines machine learning models to identify traffic patterns and fault location.
It realizes rapid fault analysis and auxiliary positioning, improves the accuracy of fault detection, enhances the flexibility and adaptability of the system, and improves the efficiency of network management and maintenance.
Smart Images

Figure CN120034424A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of communication networks, and in particular relates to a communication network fault analysis system and method. Background Art
[0002] A communication network is a structure composed of a series of connected devices and systems, used to achieve the transmission and exchange of information. It can include different types of transmission media such as wired, wireless, and optical fiber, as well as various network devices such as routers, switches, and servers. Through the collaborative work of these components, the communication network can transmit data, voice, video and other forms of communication content between different locations. The communication network fault analysis system is a system used to monitor, diagnose and solve faults in the communication network.
[0003] Problems with existing technologies: At present, network fault location is highly dependent on personal professional ability and experience, and many situations cannot be quickly located; the business scope and risks affected by network cutover, upgrade and other operations cannot be quickly and accurately identified; the risk points formed by the business distribution in the entire network cannot be quickly and intuitively judged. Summary of the invention
[0004] The object of the present invention is to provide a communication network fault analysis system and method, which can realize rapid fault analysis and auxiliary positioning.
[0005] The technical solution adopted by the present invention is as follows: A communication network fault analysis system, comprising: Data acquisition module: The data acquisition module obtains the analysis and evaluation results of the fault analysis and evaluation module, collects various data in the communication network, and exports the data; The data collection module also includes a network resource and network performance module, which is used to obtain data information and usage status of network resources; The data acquisition module connects multiple distributed network elements and collects network data in real time through the distributed network elements; Preprocessing module: connected to the data acquisition module, used to normalize the original data, extract abnormal data and feature vectors, generate a standardized time series data set, and obtain the preprocessing analysis results through the changes of abnormal data and feature vectors in the standardized time series data set; Fault analysis and evaluation module: including a rule engine, which runs standardized rules based on standardized time series data sets and compares them with the generated abnormal data to generate corresponding fault response data. It is also configured with a network protocol abnormal rule set, which is applied to different types of protocol abnormal data, and generates corresponding abnormal rule data for protocols of different network elements, and organizes them into abnormal data sets of the same category. The fault analysis and evaluation module is provided with a statistical model library, which records and analyzes the generated fault response data to generate a statistical model; The statistical model library is configured with benchmark thresholds, corresponding historical baseline thresholds, corresponding protocol baseline thresholds, and machine learning model clusters; The machine learning model cluster uses LSTM neural network to identify traffic patterns and generates corresponding neural network traffic models for the corresponding baseline threshold, historical baseline threshold, and protocol baseline threshold. Fault diagnosis and early warning module: The fault diagnosis and early warning module has a built-in fault knowledge graph, which includes the mapping relationship between equipment models and corresponding fault modes, a historical solution library, and repair action priority rules.
[0006] Failure analysis evaluation modules include: Network risk analysis module: Based on the business type modeling and fault analysis and evaluation results, the risk values of various corresponding business types are calculated, and different weight values are defined and assigned. The link is used as the basic analysis point, and the links, single disks, and network elements in the network are analyzed to obtain the business types affected by link failures. The risk point topology nodes are marked, and the total risk value of all analyzed network elements, single disks, or links in the entire network is calculated according to the business type weight, and the interrupted business volume is sorted; Among them, the risk contribution value of the fault-affected service type is expressed as: ; in, is the number of the fault point, A number indicating the type of business. express The total risk value of the failure points, Indicates The weight of the business type reflects the importance and priority of different types of business. Indicates The fault point is The impact factor of the business type; Fault simulation module: Based on the risk point topology marked by the network risk analysis module, the topology selection includes multiple point faults of ports, single disks, network elements and links, analyzes the damaged services, outputs simulation results, and the topology map shows the links affected by the fault, the affected services and the total service volume. The services affected by the fault point are presented according to the simulation results; Fault auxiliary location module: select the damaged service or service path, filter the service according to the current network element, classify the service according to landing and direct connection, analyze the suspected fault point according to the service and path, locate the suspected faulty network element on the topology map, change the color of the network element and link, the network element and link in the topology map interact, and search and locate the network element according to the network element name, logical address or IP address; Based on the standardized time series data set of the preprocessing module, the status of each network element and link is monitored, and the normal behavior is predicted using statistical models or machine learning models. The behaviors that deviate from the expected are marked as potential failure points, and the Z-score method is used to identify outliers: ; when If the value is greater than the set Z-score absolute value limit, it is considered an outlier. represents the actual value of the evaluation data point, represents the average value of the entire data set, Represents the standard deviation of the data set, indicating the degree of dispersion of the data distribution; Same-route detection module: Analyzes the primary and backup routes of LSP services in the network, screens LSPs with the same node, board or logical link in the primary and backup routes, outputs LSP service same-route details, guides service route optimization, avoids service failures caused by same-route, analyzes the interruption of optical cable and optical fiber information configured in the network management "optical cable same-route setting", and outputs the affected service scope and service detailed information.
[0007] Data acquisition module Data acquisition includes: The traffic mirroring unit is set at the core switching point and uses the sFlow protocol to collect all traffic; At least one SNMP agent unit running on the managed device, periodically obtaining real-time operating parameters of the managed device; API gateway unit to obtain software-defined network flow instrumentation status.
[0008] The data acquisition module includes an in-software data acquisition unit and an out-software data acquisition unit. The preprocessing module groups and synthesizes the in-software data and out-software data in the corresponding time period at the same time point, and groups and synthesizes the in-software data and out-software data of the same business type in the corresponding time period at the same time point.
[0009] The method for obtaining the fault point specifically includes the following steps: Determine the fault type based on the fault analysis data obtained by the fault diagnosis and early warning module; Determine the network element, disk or link where the fault is located based on the fault type; Determine the installation location according to the corresponding network element, single disk or link number; By analyzing the corresponding faults, the fault type is obtained, and the fault diagnosis and early warning modules are used to repair them first according to the fault type, and then manual or mechanical equipment is used to repair them.
[0010] A communication network fault analysis method comprises the following steps: Obtain various data indicators under the network networking status; Based on the changing relationship between various data based on the baseline threshold, historical baseline threshold and protocol baseline threshold; Obtain corresponding fault warnings based on the change relationship; According to the fault warning, the corresponding warning information is input into the fault diagnosis and warning module; Repair the faults according to the fault mode mapping relationship, historical solutions and repair action priorities; Map the above repair rules to the fault diagnosis and warning module to further improve the fault knowledge graph; Among them, various data are based on the change relationship between the benchmark threshold, the historical baseline threshold and the protocol baseline threshold. In the analysis stage, different types of thresholds are compared respectively. The comparison methods include: The baseline threshold weight value is greater than the historical baseline threshold and the protocol baseline threshold; When the fault response data is between the benchmark threshold and the historical baseline threshold or the protocol baseline threshold, an intermediate warning data set is generated; When the historical baseline threshold or the protocol baseline threshold is exceeded, a low-level warning data set is generated; When the baseline threshold is not exceeded, an advanced warning data set is generated; The corresponding faults are input into the corresponding warning data set according to the rules; To achieve automatic processing and early warning of corresponding fault information.
[0011] The various data indicators include network structure indicators, network performance indicators, network resource indicators, and network security indicators.
[0012] According to another aspect of an embodiment of the present invention, an electronic device is provided. The electronic device includes a memory and a processor; the memory is used to store programs; and the processor executes the program to implement any one of the aforementioned methods.
[0013] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, any of the aforementioned methods is implemented.
[0014] According to yet another aspect of an embodiment of the present invention, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, any one of the aforementioned methods is implemented.
[0015] The technical effects achieved by the present invention are: The present invention realizes multi-level and multi-dimensional fault analysis and evaluation through the fault analysis and evaluation module, and can also improve the accuracy of fault detection, enhance the flexibility and adaptability of the system, so that it can operate effectively in a complex and changeable network environment. By combining the rule engine and the statistical model library, the fault analysis and evaluation module can provide comprehensive and accurate fault warning and positioning services, thereby greatly improving the efficiency of network management and maintenance.
[0016] The present invention selects damaged services or service paths through a fault auxiliary positioning module, filters services according to the current network element, classifies services according to landing and direct connection, analyzes suspected fault points according to services and paths, locates suspected faulty network elements on a topology map, interacts with network elements and links in the topology map according to changes in network element and link colors, and searches and locates network elements according to network element names, logical addresses or IP addresses.
[0017] The present invention calculates the risk value of each corresponding business type based on business type modeling and fault analysis and evaluation results, defines and assigns different weight values, takes the link as the basic analysis point, analyzes the links, single disks, and network elements in the network, obtains the business type affected by the link failure, marks the risk point topology node, calculates the total risk value of all analyzed network elements, single disks or links in the entire network according to the business type weight, and sorts the interrupted business volume.
[0018] The present invention is based on the risk point topology marked by the network risk analysis module, the topology selection includes multiple point failures of ports, single disks, network elements and links, analyzes the damaged services, outputs simulation results, and the topology map displays the links affected by the failure, displays the affected services and the total business volume, and presents the services affected by the failure points according to the simulation results.
[0019] The present invention, based on the same-route detection module, analyzes the primary and backup routes of LSP services in the network, screens LSPs of the same node, single board or logical link of the primary and backup routes, outputs LSP service same-route details, guides service route optimization, avoids service failures caused by the same route, analyzes the interruption of the optical cable and optical fiber information configured in the "optical cable same-route setting" of the network management, and outputs the affected service scope and service detailed information. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a structural schematic diagram of a fault analysis and evaluation module in a communication network fault analysis system of the present invention; Figure 2 It is a flow chart of the communication network fault analysis method in the present invention; Figure 3 It is a schematic diagram of the structure of different types of thresholds in the present invention; Figure 4 It is a flowchart of different types of threshold comparison methods in the present invention. DETAILED DESCRIPTION
[0021] In order to make the purpose and advantages of the present invention more clearly understood, the present invention is specifically described below in conjunction with embodiments. It should be understood that the following text is only used to describe one or several specific embodiments of the present invention, and does not strictly limit the scope of protection of the specific claims of the present invention.
[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0023] like Figure 1-3 As shown, a communication network fault analysis system includes a data acquisition module, a preprocessing module, a fault analysis and evaluation module, and a fault diagnosis and early warning module.
[0024] Furthermore, the data collection module obtains the analysis and evaluation results of the fault by the fault analysis and evaluation module, as well as the collection of various data in the communication network, and exports the data; The data collection module also includes a network resource and network performance module, which is used to obtain data information and usage status of network resources; The data acquisition module connects multiple distributed network elements and collects network data in real time through the distributed network elements. The network data includes traffic data, equipment status logs and performance indicator data. The performance indicator data includes operation data of equipment or network elements, such as temperature, resource utilization, data transmission stability and other hardware information during operation. Although other hardware information will not affect the normal operation of network elements or equipment during operation, it will exist for a long time and will affect the equipment or network elements to a certain extent. After a certain period of time, it is combined with other hardware information and input into the data acquisition module to achieve an overall reference to the data, including the temperature change cycle of the equipment under the same task, which is used to analyze the changes in the cleaning and heat dissipation performance of the equipment, and to analyze whether the equipment is aging.
[0025] Furthermore, the data acquisition module includes: Traffic mirroring units are installed at core switching points and use the sFlow protocol to collect all traffic. Traffic mirroring units replicate traffic passing through core switching points for analysis, monitoring, or security testing. sFlow protocol is used for full traffic collection and a sampling-based method is used to collect network traffic information. The sFlow protocol is a technology used to monitor device performance and supports data packet sampling in high-speed network environments. It can provide detailed information on traffic patterns, utilization, and abnormal behavior. This technology is particularly important for large-scale networks because it can provide the necessary monitoring capabilities without affecting network performance. At least one SNMP agent unit running on the managed device periodically obtains the real-time operating parameters of the managed device. The SNMP agent unit runs on the managed device and is a software module. It periodically obtains the real-time operating parameters of these managed devices and reports the operating parameters to the SNMP management system. The operating parameters include but are not limited to CPU usage, memory usage, interface status, traffic size, etc., and can achieve remote monitoring and data management through the SNMP management system; The managed device includes a central device for managing multiple network elements or any computer management device in a communication network, and may also include an independently set network element; The API gateway unit obtains the software-defined network flow meter status and is responsible for exchanging with the software-defined network (SDN) controller or other network management systems to obtain the current status of the network flow meter. The software-defined network flow meter status includes but is not limited to the currently active flow table items, the data transmission rate of each flow, the packet loss situation and other detailed information. As a middleware, the API gateway simplifies the communication process between different systems and realizes the unification of data information obtained in fault analysis in the network, which facilitates normalized data and allows different applications or services to query the network status, which is convenient for the later upgrading of management software or management system. In addition, it can also realize functions such as automated operation and maintenance and dynamic resource allocation.
[0026] Preprocessing module: connected to the data acquisition module, used to normalize the original data, extract abnormal data and feature vectors, generate a standardized time series data set, and obtain the preprocessing analysis results through the changes of abnormal data and feature vectors in the standardized time series data set; The original data is the various data collected by the data acquisition module. The flow data normalization process uses the following formula: ; in, is the normalized vertical, is the original value, and are the minimum and maximum values of this feature respectively.
[0027] Abnormal data points are identified using the Z-score method: ; in, It represents the deviation between a data point and the mean, and is used to measure the degree to which the data point deviates from the mean. To evaluate the actual value of a data point, is the mean value of the entire data set, It is the standard deviation of the data set, indicating the degree of dispersion of the data distribution. The larger the standard deviation, the greater the difference between the data, and the smaller the standard deviation, the more concentrated the data.
[0028] In the preprocessing module, Z-score is used to monitor network performance indicators, such as latency, packet loss, etc., or device operating parameters, such as CPU usage, memory usage, etc., to see if there are any abnormalities. Specifically, for each data point in the selected data set, the Z-score is calculated by comparing the actual value with the mean and standard deviation of the data set. If the absolute value of the Z-score of a data point exceeds the set threshold, it can be regarded as a potential outlier for further monitoring and analysis, which helps to quickly locate data points that exhibit abnormal behavior, thereby achieving in-depth data analysis and problem diagnosis.
[0029] For example, when monitoring network traffic, if the traffic at a certain point in time is significantly higher or lower than the normal level, it indicates an abnormal event, such as a DDoS attack or equipment failure. Identifying these anomalies through the Z-score method can help network administrators take timely measures to prevent potential problems from spreading.
[0030] In the above system module, the data acquisition module also includes an in-software data acquisition unit and an out-software data acquisition unit. The preprocessing module groups and synthesizes the in-software data and the out-software data in the corresponding time period at the same time point, and groups and synthesizes the in-software data and the out-software data of the same business type in the corresponding time period at the same time point.
[0031] Furthermore, the data acquired by the in-software data acquisition unit directly from within a running application or service includes a software program running in a PC, a network element, or a server.
[0032] For example, application performance monitoring tools collect data on metrics such as response time, error rates, and transaction success rates.
[0033] Furthermore, the external data acquisition unit of the software collects system external or environmental data, including but not limited to status information of network devices, traffic statistics data, external environmental temperature of network devices or devices connected by third parties, etc.
[0034] For example, as pointed out above, the periodic operating temperature changes, performance changes, etc. of the network device can be obtained by using a data acquisition unit in the software to determine whether there is a serious dust accumulation problem during its operation, so as to avoid affecting the heat dissipation effect of the network device and thus affecting the performance of the network device.
[0035] The preprocessing module groups and synthesizes data at the same time point to ensure that data from inside or outside the software at the same time point can be correctly associated to form a complete time series view.
[0036] For example, a timestamp T is set. During the time period T, the data collection unit inside the software collects indicators about application performance, such as response time and data throughput. At the same time, the data collection unit outside the software records the network status, such as bandwidth utilization or network device temperature. The two data sources are merged to form a data set containing all relevant information.
[0037] Furthermore, different business types are grouped and merged, and a specific business type is analyzed. If there are multiple business types, the pre-processing module needs to first identify the data corresponding to the specific business type, and further integrate the corresponding data on this basis.
[0038] For example, in a data transmission task, the time required for data transmission is set as a time node T. The data acquisition unit in the software records the data output end and the data input end, as well as the data response speed between the network elements and central servers passed between the two, the data fluctuations during the data transmission process, and the data transmission success rate.
[0039] At the same time node T, the data collection unit outside the software records the load, packet loss rate, etc. of each device in the network environment, and compares the above data with the historical equivalent data to predict the problems that may be hidden in the relevant data changes.
[0040] Furthermore, the preprocessing module will merge the above data according to the timestamp T to generate one or more records. Any record contains information on application performance and network status. If it is synthesized by business type, it is necessary to further distinguish business-related data such as data type and packet transmission protocol, and may conduct a deeper analysis together with data from other similar businesses. According to the changes in the corresponding data, it is helpful to quickly locate the root cause of the problem, so as to optimize the configuration, improve service quality, and facilitate further analysis and evaluation of the fault analysis and evaluation module.
[0041] Fault analysis and evaluation module: including a rule engine, which runs standardized rules based on standardized time series data sets and compares them with the generated abnormal data to generate corresponding fault response data. It is also configured with a network protocol abnormal rule set, which is applied to different types of protocol abnormal data, and generates corresponding abnormal rule data for protocols of different network elements, and organizes them into abnormal data sets of the same category. The fault analysis and evaluation module is provided with a statistical model library, which records and analyzes the generated fault response data to generate a statistical model; The statistical model library is configured with benchmark thresholds, corresponding historical baseline thresholds, corresponding protocol baseline thresholds, and machine learning model clusters; The machine learning model cluster uses the LSTM neural network to identify traffic patterns and generates corresponding neural network traffic models for the corresponding baseline thresholds, historical baseline thresholds, and protocol baseline thresholds.
[0042] In addition to directly analyzing and presenting apparent faults, the fault analysis and evaluation module can also analyze potential faults through in-depth data analysis and generate corresponding measures.
[0043] The rule engine runs a series of predefined standardized rules based on the standardized time series data set, including but not limited to: The network protocol anomaly rule set is used to monitor different types of network protocol anomalies, such as errors in the TCP / IP protocol stack, HTTP request timeouts, etc. Generate corresponding exception rule data for protocols of different network elements, and configure specific protocol anomaly detection rules for any network element, switch, server, etc. to accurately capture problems of specific devices.
[0044] For example, in a network environment, a large amount of network traffic data is collected through sFlow or NetFlow, and the rule engine sets the following rules: If the HTTP response time of a server exceeds 500ms, a warning is triggered; If the packet loss rate of a link exceeds 1%, it is marked as abnormal and recorded.
[0045] Based on the data corresponding to the warning triggered by the above rules, a fault response data set is statistically generated to reflect the parameter change status of the corresponding device data during the warning process of the system, which helps to identify the difference from normal data.
[0046] In addition, the baseline threshold is set as a standard value as a reference point, and any data exceeding this range is considered abnormal; Historical baseline thresholds are calculated based on past data, reflecting the stability of the system's long-term operation and used to achieve benchmark comparisons with existing acquired data; Protocol baseline threshold sets a value within a rule for different network protocols, such as the time limit for establishing a TCP connection.
[0047] For example, the LSTM (Long Short-Term Memory) model is used for traffic pattern recognition. LSTM is a deep learning algorithm suitable for processing sequence data, especially for time series prediction and anomaly detection tasks.
[0048] Take the following scenario as an example: Use LSTM to train a model that learns normal traffic patterns based on past network traffic data (including normal and abnormal situations).
[0049] When real-time traffic data is input into this model, the traffic conditions in the future are predicted and compared with the actual traffic. If the actual traffic deviates significantly from the predicted value (for example, it exceeds the benchmark threshold or historical baseline threshold), it is considered an anomaly.
[0050] Furthermore, the machine learning model cluster contains multiple models, among which the LSTM neural network is used to perform traffic pattern recognition. The model not only relies on the baseline threshold, historical baseline threshold and protocol baseline threshold, but also adjusts its parameters according to the specific situation to adapt to the new data pattern.
[0051] For example, a data center network monitoring case: Data centers experience peak and off-peak hours every day, and the traffic patterns during these two periods are very different.
[0052] Using the LSTM model cluster, two models can be trained based on historical data from different time periods: one for peak hours and the other for off-peak hours.
[0053] In actual application, the system will automatically select the appropriate model for traffic prediction and anomaly detection based on the current time. For example, the peak period model is used between 9 am and 6 pm, and the off-peak period model is used at other times.
[0054] According to the fault analysis and evaluation module, multi-level and multi-dimensional fault analysis and evaluation can be realized, and the accuracy of fault detection can be improved, and the flexibility and adaptability of the system can be enhanced, so that it can operate effectively in a complex and changeable network environment. By combining the rule engine and the statistical model library, the fault analysis and evaluation module can provide comprehensive and accurate fault warning and positioning services, thereby greatly improving the efficiency of network management and maintenance.
[0055] Fault diagnosis and early warning module: The fault diagnosis and early warning module has a built-in fault knowledge graph, which records common fault forms and operating characteristics for various models of routers, switches, servers, etc. For example, for a certain manufacturer's router, overheating or network traffic thresholds are prone to occur under high load conditions. In actual applications, it can quickly locate and warn of corresponding faults, and can quickly provide an effective solution. The knowledge graph includes the mapping relationship between device models and corresponding fault modes, a historical solution library, and repair action priority rules.
[0056] As an optional embodiment, refer to the attached Figure 1 As shown, the fault analysis and evaluation module includes a network risk analysis module, a fault simulation and drill module, a fault auxiliary location module and a same-route detection module.
[0057] Among them, the network risk analysis module calculates the risk value of each corresponding business type based on business type modeling and fault analysis and evaluation results, defines and assigns different weight values, takes the link as the basic analysis point, analyzes the links, single disks, and network elements in the network, obtains the business type affected by the link failure, marks the risk point topology node, calculates the total risk value of all analyzed network elements, single disks or links according to the business type weight, and sorts the interrupted business volume; Among them, the risk contribution value of the fault-affected service type is expressed as: ; in, is the number of the fault point, A number indicating the type of business. express The total risk value of the failure points, Indicates The weight of the business type reflects the importance and priority of different types of business. Indicates The fault point is The influencing factors of the business type, the business type weight value in the above parameters can be adjusted accordingly based on the salesperson's experience.
[0058] The risk contribution value of the above-mentioned fault-affected business type can also be applied to calculate the risk value of a single fault point; To evaluate and analyze the amount of service interruption caused by each fault point, use: ; in, Indicates The traffic level of this type of business reflects the activity volume of this type of business, where Indicates Link pairj The influencing factors of this business.
[0059] The algorithm flow of the network risk analysis module includes the following steps: S101. Determine the weight of each business type based on its importance and enterprise needs. , construct the impact factor matrix , make corresponding adjustments and settings based on salespersons and historical data; S102, perform fault analysis on links, single disks, and network elements in the network, identify all potential fault points, and mark them in the topology diagram; S103. For each fault point, calculate its risk value using the risk contribution value expression of the service type affected by the fault; S104: Based on the interruption traffic volume Arrange all fault points in descending order, output the sorted fault point list and its related information in the application, so as to give priority to major fault points.
[0060] The fault simulation module is based on the risk point topology marked by the network risk analysis module. The topology selection includes multiple point faults of ports, single disks, network elements and links. The damaged services are analyzed and the simulation results are output. The topology diagram shows the links affected by the fault, the affected services and the total business volume. The services affected by the fault point are presented according to the simulation results.
[0061] By obtaining the complete topology diagram of the network, including information on ports, single disks, network elements and links, all potential risk points are marked in the topology diagram based on the results of the network risk analysis module, and traffic data of each service type is collected and prepared, including path information for each service.
[0062] The fault simulation algorithm includes the following steps: S201, randomly selecting or selecting multiple risk points from the marked risk points as the fault points; S202, for each selected fault point, simulate its failure situation and calculate the affected business volume; S203, displaying the simulation results in a visual form.
[0063] The fault auxiliary location module selects the damaged service or service path, filters the service according to the current network element, classifies the service according to landing and direct connection, analyzes the suspected fault point according to the service and path, locates the suspected faulty network element on the topology map, and searches and locates the network element according to the network element name, logical address or IP address according to the change of network element and link color. Based on the standardized time series data set of the preprocessing module, the status of each network element and link is monitored, and the normal behavior is predicted using statistical models or machine learning models. The behaviors that deviate from the expected are marked as potential failure points, and the Z-score method is used to identify outliers: ; when If the value is greater than the set Z-score absolute value limit, it is considered an outlier. represents the actual value of the evaluation data point, represents the average value of the entire data set, Represents the standard deviation of a data set, indicating the degree of dispersion of the data distribution.
[0064] Same-route detection module: Analyzes the primary and backup routes of LSP services in the network, screens LSPs with the same node, board or logical link in the primary and backup routes, outputs LSP service same-route details, guides service route optimization, avoids service failures caused by same-route, analyzes the interruption of optical cable and optical fiber information configured in the network management "optical cable same-route setting", and outputs the affected service scope and service detailed information.
[0065] Based on the above, network topology data is obtained, including but not limited to nodes, links and their connection relationships, LSP configuration data is obtained, including the primary and backup routing path information of each LSP, and optical cable co-routing setting data is obtained.
[0066] Evaluate the scope and details of the impact on services when physical connections are interrupted. The method includes the following steps: S301. Extract the primary and backup routing path information of each LSP from the LSP configuration data.
[0067] S302: For each LSP, compare whether the primary and backup routes have common nodes, boards or logical links; S303, record and output all LSPs with the same routing and their common elements (nodes, boards or logical links); S304. Based on the data in the "optical cable co-routing setting", the scope of services that may be affected if a specific optical cable or optical fiber fails is evaluated, and a mapping is constructed to associate the optical cable / optical fiber at the physical layer with the LSP.
[0068] According to the above steps, the co-routing problems existing in LSP services can be identified, and the specific impact of optical cable or optical fiber failure on the service can be evaluated. This not only helps to guide service route optimization and avoid the risk of service interruption due to co-routing, but also enables advance planning of emergency response measures to ensure network reliability and stability.
[0069] Furthermore, the method for obtaining the fault point specifically includes the following steps: S401, determining the fault type according to the fault analysis data obtained by the fault diagnosis and early warning module; S402, determining the network element, single disk or link where the fault is located according to the fault type; S403, determining the installation location according to the corresponding network element, single disk or link number; S404, obtaining the fault type by analyzing the corresponding fault, and using the fault diagnosis and early warning module to repair it first according to the fault type, and then using manual or mechanical equipment to repair it.
[0070] It should be noted that the above modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0071] According to an embodiment of the present invention, a method embodiment of a communication network fault analysis method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0072] Reference Figure 2-4 , a communication network fault analysis method, comprising the following steps: S1. Obtain various data indicators under the network networking status; S2. Based on the change relationship of each data based on the baseline threshold, historical baseline threshold and protocol baseline threshold; S3. Obtain corresponding fault warning according to the change relationship; S4. According to the fault warning, the corresponding warning information is input into the fault diagnosis and warning module; S5. Repair the fault according to the fault mode mapping relationship, historical solutions and repair action priority; S6. Map the above repair rules to the fault diagnosis and warning module to further improve the fault knowledge graph; Among them, various data are based on the change relationship between the benchmark threshold, the historical baseline threshold and the protocol baseline threshold. In the analysis stage, different types of thresholds are compared respectively. The comparison methods include: S11, the baseline threshold weight value is greater than the historical baseline threshold and the protocol baseline threshold; S12, when the fault response data is between the benchmark threshold and the historical baseline threshold or the protocol baseline threshold, an intermediate warning data set is generated; S13. When the historical baseline threshold or the protocol baseline threshold is exceeded, a low-level warning data set is generated; S14, when the baseline threshold is not exceeded, generating an advanced warning data set; S15, inputting the corresponding faults into the corresponding warning data set according to the rules; S16, to realize automatic processing and early warning of corresponding fault information.
[0073] As an optional embodiment, the various data indicators include network structure indicators, network performance indicators, network resource indicators, and network security indicators.
[0074] According to another aspect of an embodiment of the present invention, an electronic device is provided. The electronic device includes a memory and a processor; the memory is used to store programs; and the processor executes the program to implement any one of the aforementioned methods.
[0075] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, any of the aforementioned methods is implemented.
[0076] According to yet another aspect of an embodiment of the present invention, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, any one of the aforementioned methods is implemented.
[0077] The above is only a preferred embodiment of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principles of the present invention, and these improvements and modifications should also be considered as the protection scope of the present invention. The structures, devices and operating methods not specifically described and explained in the present invention shall be implemented according to the conventional means in the art unless otherwise specified and limited.
Claims
1. A communication network fault analysis system, characterized in that: include: Data acquisition module: The data acquisition module obtains the analysis and evaluation results of the fault by the fault analysis and evaluation module, collects various data in the communication network, and exports the data; The data acquisition module also includes a network resource and network performance module, which is used to obtain data information and usage status of network resources; The data acquisition module is connected to a plurality of distributed network elements and collects network data in real time through the distributed network elements; Preprocessing module: connected to the data acquisition module, used to normalize the original data, extract abnormal data and extract feature vectors, generate a standardized time series data set, and obtain the preprocessing analysis results through the changes of abnormal data and feature vectors in the standardized time series data set; Fault analysis and evaluation module: including a rule engine, which runs standardized rules based on a standardized time series data set, compares the generated abnormal data to generate corresponding fault response data, and is configured with a network protocol abnormal rule set, which is applied to different types of protocol abnormal data, and generates corresponding abnormal rule data for protocols of different network elements, and organizes them into abnormal data sets of the same category; The fault analysis and evaluation module is provided with a statistical model library, which records and analyzes the generated fault response data to generate a statistical model; The statistical model library is configured with a benchmark threshold, a corresponding historical baseline threshold, a corresponding protocol baseline threshold, and a machine learning model cluster; The machine learning model cluster uses an LSTM neural network to perform traffic pattern recognition, and generates a corresponding neural network traffic model for corresponding benchmark thresholds, historical baseline thresholds, and protocol baseline thresholds; Fault diagnosis and early warning module: The fault diagnosis and early warning module has a built-in fault knowledge graph, which includes a mapping relationship between equipment models and corresponding fault modes, a historical solution library, and repair action priority rules.
2. A communication network fault analysis system according to claim 1, characterized in that: The fault analysis and evaluation module comprises: Network risk analysis module: Based on the business type modeling and fault analysis and evaluation results, the risk values of various corresponding business types are calculated, and different weight values are defined and assigned. The link is used as the basic analysis point, and the links, single disks, and network elements in the network are analyzed to obtain the business types affected by link failures. The risk point topology nodes are marked, and the total risk value of all analyzed network elements, single disks, or links in the entire network is calculated according to the business type weight, and the interrupted business volume is sorted; Among them, the risk contribution value of the fault-affected service type is expressed as: ; in, is the number of the fault point, A number indicating the type of business. express The total risk value of the failure points, Indicates The weight of the business type reflects the importance and priority of different types of business. Indicates The fault point is The impact factor of the business type; Fault simulation module: Based on the risk point topology marked by the network risk analysis module, the topology selection includes multiple point faults of ports, single disks, network elements and links, analyzes the damaged services, outputs simulation results, and the topology map shows the links affected by the fault, the affected services and the total service volume. The services affected by the fault point are presented according to the simulation results; Fault auxiliary location module: select the damaged service or service path, filter the service according to the current network element, classify the service according to landing and direct connection, analyze the suspected fault point according to the service and path, locate the suspected faulty network element on the topology map, change the color of the network element and link, the network element and link in the topology map interact, and search and locate the network element according to the network element name, logical address or IP address; Based on the standardized time series data set of the preprocessing module, the status of each network element and link is monitored, and the normal behavior is predicted using statistical models or machine learning models. The behaviors that deviate from the expected are marked as potential failure points, and the Z-score method is used to identify outliers: ; when If the value is greater than the set Z-score absolute value limit, it is considered an outlier. represents the actual value of the evaluation data point, represents the average value of the entire data set, Represents the standard deviation of the data set, indicating the degree of dispersion of the data distribution; Same-route detection module: Analyzes the primary and backup routes of LSP services in the network, screens LSPs with the same node, board or logical link in the primary and backup routes, outputs LSP service same-route details, guides service route optimization, avoids service failures caused by same-route, analyzes the interruption of the optical cable and optical fiber information configured in the "optical cable same-route setting" of the network management, and outputs the affected service scope and service detailed information.
3. A communication network fault analysis system according to claim 1, characterized in that: The data acquisition module data acquisition includes: Traffic mirroring unit, installed at the core switching point, uses sFlow protocol to collect full traffic; At least one SNMP agent unit running on the managed device, periodically obtaining real-time operating parameters of the managed device; API gateway unit to obtain software-defined network flow instrumentation status.
4. A communication network fault analysis system according to claim 1, characterized in that: The data acquisition module includes an in-software data acquisition unit and an out-software data acquisition unit. The preprocessing module groups and synthesizes the in-software data and out-software data in the corresponding time period at the same time point, and groups and synthesizes the in-software data and out-software data of the same business type in the corresponding time period at the same time point.
5. A communication network fault analysis system according to claim 2, characterized in that: The method for obtaining the fault point specifically comprises the following steps: Determine the fault type based on the fault analysis data obtained by the fault diagnosis and early warning module; Determine the network element, disk or link where the fault is located based on the fault type; Determine the installation location according to the corresponding network element, single disk or link number; By analyzing the corresponding faults, the fault type is obtained, and the fault diagnosis and early warning modules are used to repair them first according to the fault type, and then manual or mechanical equipment is used to repair them.
6. A communication network fault analysis method, characterized in that: Using the system as claimed in any one of claims 1 to 5 comprises the following steps: Obtain various data indicators under the network networking status; Based on the changing relationship between various data based on the baseline threshold, historical baseline threshold and protocol baseline threshold; Obtain corresponding fault warnings based on the change relationship; According to the fault warning, the corresponding warning information is input into the fault diagnosis and warning module; Repair the faults according to the fault mode mapping relationship, historical solutions and repair action priorities; Map the above repair rules to the fault diagnosis and warning module to further improve the fault knowledge graph; Among them, various data are based on the change relationship between the benchmark threshold, the historical baseline threshold and the protocol baseline threshold. In the analysis stage, different types of thresholds are compared respectively. The comparison methods include: The baseline threshold weight value is greater than the historical baseline threshold and the protocol baseline threshold; When the fault response data is between the benchmark threshold and the historical baseline threshold or the protocol baseline threshold, an intermediate warning data set is generated; When the historical baseline threshold or the protocol baseline threshold is exceeded, a low-level warning data set is generated; When the baseline threshold is not exceeded, an advanced warning data set is generated; The corresponding faults are input into the corresponding warning data set according to the rules; To achieve automatic processing and early warning of corresponding fault information.
7. A communication network fault analysis method according to claim 6, characterized in that: The various data indicators include network structure indicators, network performance indicators, network resource indicators, and network security indicators.
8. An electronic device, characterized in that: The electronic device comprises a memory and a processor; The memory is used to store programs; The processor executes the program to implement the method according to claim 6 or 7.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to claim 6 or 7 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to claim 6 or 7 is implemented.
Citation Information
Patent Citations
Positioning and analyzing method and system for fault of transmission network based on big data analysis
CN108111361A
Method and equipment for analyzing same routing hidden dangers in transmission network
CN110858777A
Label switching path active-standby same-route optimization method and label switching path active-standby same-route optimization system
CN112291143A
Network traffic simulation and prediction method and device
CN112887148A
Network fault risk analysis method and device
CN114172784A
Cited By
Fault positioning method and system for communication network
CN121151203A
Vehicle cloud cooperative vehicle-mounted network fault processing method and system
CN121691032A