Network equipment fault rapid positioning method and system based on historical data

Through the network equipment fault location method based on historical data, the SNMP protocol and decision tree model are used to summarize three types of application scenarios, data cleaning and feature screening, and priority is calculated, which solves the problems of insufficient manual experience dependence and global analysis in the traditional method, and achieves fast and accurate fault location and processing.

CN120263628AActive Publication Date: 2025-07-04EXANDS INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510379327.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-04
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

Existing network equipment fault location methods rely on manual experience, making it difficult to quickly and accurately detect complex faults, and lack global analysis capabilities, and cannot effectively process multidimensional network data, resulting in inefficient troubleshooting.

Method used

Based on historical data, the operation data of network equipment is collected through the SNMP protocol, summarized into three types of application scenarios, cleaned and preprocessed, and used the decision tree machine learning model to filter and match features, calculate the priority of application scenarios, select the most matching positioning strategy, and record the priority order, expand the newly identified subscenarios as the main scenario.

Benefits of technology

It realizes fast and accurate fault location of network equipment, reduces the workload of manual inspection, adapts to changes in the network environment, and improves fault handling efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263628A_ABST
    Figure CN120263628A_ABST
Patent Text Reader

Abstract

The invention discloses a network equipment fault rapid positioning method and system based on historical data, and relates to the technical field of network equipment fault detection, and the method comprises the following steps: collecting historical operation data of network equipment; the method comprises the following steps: summarizing network faults into three types of application scenes, and positioning according to different logics; cleaning and preprocessing the collected network equipment data, performing recursive feature elimination and screening based on a preprocessing result, matching real-time features with a historical database, and calculating priorities of application scenes; the matching degree of the current fault feature and the three types of application scenes is evaluated, the most matched scene is selected as a fault positioning strategy, and an execution priority sequence is recorded to a historical scene library; for a newly-recognized fault scene, the newly-recognized fault scene serves as a sub-scene to be expanded, the occurrence frequency of the newly-recognized fault scene is counted according to the occurrence frequency, whether the newly-recognized fault scene needs to be upgraded into a main scene or not is judged, and rapid network equipment fault positioning is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network device fault detection, and specifically to a method and system for quickly locating network device faults based on historical data. Background Art

[0002] In today's digital age, the network has been deeply integrated into all aspects of society. Whether it is enterprise operation, daily life, industrial production, or scientific research activities, they all highly rely on a stable and efficient network environment. As the core component of the network, the stable operation of network devices is crucial. Once a network device fails, it may lead to service interruption and data loss, bringing huge economic losses and inconveniences to enterprises and users.

[0003] However, the existing means for locating network device faults have many limitations. Firstly, traditional fault location mainly relies on the experience of operation and maintenance personnel to check the data of network devices one by one, and it is difficult to quickly discover complex fault problems. With the continuous development of network technology, the network structure and business models are becoming increasingly complex, and network data includes information in multiple dimensions, such as traffic changes in different regions, performance parameters of various devices, etc. Traditional methods cannot mine potential fault risks from the associations of these multi-dimensional data. Secondly, network data processing lacks the ability of global review. Traditional methods often only focus on the faults of a single device or a local network, without comprehensively analyzing the overall operation status of the network, including the change trends of network traffic in different regions and the mutual influence of various device performance indicators. In addition, traditional fault location methods mostly rely on manual experience judgment. Facing the complex and changeable network environment and the increasing network scale, the efficiency of manual troubleshooting is low, and it is difficult to quickly and accurately determine the fault point. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for quickly locating network device faults based on historical data to solve the problems raised in the prior art.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A method for quickly locating network device faults based on historical data, the method comprising the following steps:

[0006] Collect historical operation data of network devices;

[0007] Classify network faults into three application scenarios and locate them according to different logics;

[0008] Clean and preprocess the collected network device data, perform recursive feature elimination and screening based on the preprocessing results, match real-time features with the historical database, and calculate the priority of the application scenario;

[0009] Evaluate the matching degree of the current fault feature with three types of application scenarios, select the most matching scenario as the fault location strategy, and record the execution priority order in the historical scenario library;

[0010] For the newly identified fault scenario, expand it as a sub-scenario, and determine whether it needs to be upgraded to the main scenario according to the occurrence times. Among them, the sub-scenario refers to the fault scenario newly identified during the fault location process and not yet summarized into the existing main scenario. The sub-scenario is generated based on the fact that the current fault feature does not completely match the existing fault scenario pattern but has similarity or relevance; the main scenario refers to the set of representative and typical fault scenarios formed after classifying and summarizing network faults. The set of fault scenarios is summarized based on historical fault data and actual experience. Each main scenario corresponds to specific fault features, location logics, and processing strategies.

[0011] Collect the historical operation data of network devices through the SNMP protocol, including device logs, traffic information, device performance configurations, and fault records. The specific steps are as follows:

[0012] Collect the historical operation data of network devices through the SNMP protocol, including device logs, traffic information, network device performance configurations, and fault records; among them, the traffic information includes network traffic involving multiple regions of offline and online platforms, as well as network traffic demand data;

[0013] The network traffic demand data represents the data on the demand for data transmission traffic and the traffic usage behavior characteristics in the network within a specific time, expressed as: [F1, F2,..., Fz]; where z is a positive integer representing the number of traffic demand types, and F1, F2,..., Fz respectively represent the 1st, 2nd,..., zth traffic demand types; the traffic demand data is obtained through network traffic monitoring devices and network service providers; the historical data at different times of the traffic demand has the same time interval Δf; the data is arranged in chronological order [f1, f2,..., fw] to form the network traffic demand time series data; where w represents the number of moments, and f1, f2,..., fw respectively represent the 1st, 2nd,..., wth traffic demand data collection moments;

[0014] The network device performance configuration data represents the operating performance status presented by the network device, the dynamic process of configuration parameter adjustment, and the device health status data, expressed as [P1, P2, …, Pu]; where u is a positive integer representing the type of network device performance configuration data, and P1, P2, …, Pu respectively represent the 1st, 2nd, …, u-th types of network device performance configuration data; the network device performance configuration data is obtained from the device's own monitoring system and the device manufacturer; the data at different moments of the network device performance configuration history has the same time interval Δp; the data is arranged in chronological order [p1, p2, …, pv] to form the network device performance configuration time series data; where v represents the number of moments, and p1, p2, …, pv respectively represent the 1st, 2nd, …, v-th network device performance configuration data collection moments.

[0015] Network failures are classified into three types of application scenarios and located according to different logics. The specific steps include:

[0016] In the first type of application scenario: when the traffic in the i-th area drops below the designed traffic threshold and there is no switch error reported in the device log, in the first type of application scenario, historical link interruption cases (optical cables being dug up) are preferentially matched, and the link detection tool (Traceroute) is directly triggered, where i is used to identify a specific area, and the designed traffic threshold is determined by historical data;

[0017] In the second type of application scenario: when there are abnormal conditions in the traffic of multiple areas and the device log shows that the switch CPU / memory is overloaded or the port is down, in the second type of application scenario, the status of the core switch (port status, load threshold) is preferentially checked, and the recovery strategies for historical similar faults (restarting or switching to a standby switch) are compared;

[0018] In the third type of application scenario: when local user connections fail and authentication errors or ARP conflicts frequently appear in the log, in the third type of application scenario, the configuration snapshot of the edge device (store switch) is preferentially checked, and compared with the historical correct configuration version;

[0019] The first type of application scenario, the second type of application scenario, and the third type of application scenario together constitute the initial application scenario set for network device fault location.

[0020] The network device-related data collected is cleaned and preprocessed to remove noise data. For traffic data, when the traffic value at a certain moment is missing, linear interpolation is used to fill it according to the traffic data at adjacent moments before and after, and the data format is standardized, and all data is stored according to a unified database table structure;

[0021] Features are extracted from the preprocessed data, denoted as [T1, T2, …, Tn];

[0022] Select a decision tree machine learning model for fault prediction, use the preprocessed and feature-extracted data as input, and the historical fault records as labels to train the initial model;

[0023] Measure the contribution of each feature to the model performance by calculating the information gain of the feature;

[0024] Remove the feature with the smallest contribution from the current feature set, retrain the model, and evaluate the model performance. When the model performance fluctuates within 2%, retain the feature set after removing the feature; otherwise, restore the feature, where q represents the threshold parameter for measuring the degree of change in model performance;

[0025] Repeat the above process until the influence of all features in the feature set on the model performance reaches the set influence threshold, where the influence threshold is determined by cross-validation and the accuracy metric of model performance evaluation;

[0026] Match the real-time features with the historical database, sort them by priority, analyze the priority order of the three types of application scenarios, construct an application scenario priority calculation formula based on the influence range and repair time involved in the application scenario, and sort the three types of application scenarios according to the priority size according to the application scenario priority calculation formula. The application scenario priority calculation formula is as follows:

[0027] Asp = α * Si + β * Rt;

[0028] Where Asp represents the application scenario priority, Si represents the influence range involved in the application scenario, Rt represents the repair time involved in the application scenario, and α, β represent weight coefficients.

[0029] Calculate the matching degree between the current fault feature and the three types of application scenarios. The specific steps include:

[0030] Calculate the matching degree between the current fault feature and the three types of application scenarios. When the real-time data is compared with the three types of application scenarios one by one in the specified priority order, it is recorded as m1, m2, and m3. Where m1 represents the matching degree of the real-time data compared with the first type of application scenario, m2 represents the matching degree of the real-time data compared with the second type of application scenario, and m3 represents the matching degree of the real-time data compared with the third type of application scenario;

[0031] Sort the calculated matching degrees m1, m2, and m3 between the current fault feature and the three types of application scenarios, and select the one with the highest matching degree as the localization strategy to be adopted for the current fault feature;

[0032] Determine whether the selected positioning strategy with the highest matching degree successfully solves the current fault. If the positioning measure with the highest matching degree fails to solve the current fault, then select other scenarios in descending order of priority as the positioning measures to be taken for the current scenario;

[0033] Among them, for the positioning measure with the highest matching degree to solve the current fault, the specific steps are as follows:

[0034] Sort the sub-scenario types in the application scenario involved in the positioning measure with the highest matching degree according to the number of times the sub-scenario appears in the application scenario. Analyze them one by one in the order of the number of times the current sub-scenario is sorted under the current application scenario until the current fault is solved by the sub-scenario in the application scenario, and record the sub-scenario in the application scenario as the solution measure for the current fault;

[0035] Add the priority order of the current fault feature to the historical scenario library;

[0036] Classify the current fault feature under the three types of application scenarios, that is, expand the newly added scenario as a sub-scenario, and record the number of times the current sub-scenario appears as 1. When the current sub-scenario appears subsequently, increase the number of times the current sub-scenario appears by 1, and finally count the number of times the current sub-scenario appears within a period of time;

[0037] Set the number of appearances U1 required for the sub-scenario when the sub-scenario is adjusted to the main scenario. When the number of appearances of the current sub-scenario exceeds the set number of appearances U1 required for the sub-scenario when the sub-scenario is adjusted to the main scenario, adjust the current sub-scenario to the main scenario. Among them, the calculation formula for the number of appearances required for the sub-scenario when the sub-scenario is adjusted to the main scenario is as follows: U1 = γ * TNF;

[0038] Among them, TNF represents the total number of faults, and γ represents the weight coefficient.

[0039] A rapid fault location system for network devices based on historical data. The system includes a data acquisition module, a data processing module, a fault scenario analysis module, a fault location decision module, and a scenario management module. The data acquisition module is used to collect the historical operation data of network devices; the data processing module is used to clean and preprocess the collected data related to network devices, and extract features based on the preprocessed data; the fault scenario analysis module is used to classify network faults into three types of application scenarios, locate them according to different logics, match real-time features with the historical database, and calculate the priority of the application scenarios; the fault location decision module is used to evaluate the matching degree between the current fault features and the three types of application scenarios, select the most matching scenario as the fault location strategy, and record the executed priority order in the historical scenario library; the scenario management module is used to expand a newly identified fault scenario as a sub-scenario, and judge whether it needs to be upgraded to a main scenario according to the occurrence times.

[0040] The data acquisition module includes a data collection unit and a time series construction unit. The data collection unit uses the SNMP protocol to collect data from multiple channels such as network devices, network traffic monitoring devices, network service providers, device self-monitoring systems, and device manufacturers. The collected content includes device logs, traffic information, device performance configuration data, and fault records. The time series construction unit is used to construct network traffic demand time series data and network device performance configuration time series data for the collected traffic demand data and device performance configuration data respectively according to their respective time intervals.

[0041] The data processing module includes a preprocessing unit, a feature extraction unit, and a feature screening unit. The preprocessing unit is used to remove the noise data in the collected data, fill in the missing values of the traffic data using the linear interpolation method, standardize all data formats, and uniformly store them in the database table structure. The feature extraction unit is used to extract features from the preprocessed data, denoted as [T1, T2,..., Tn]. The feature screening unit is used to select a decision tree machine learning model, use the preprocessed and feature-extracted data as input, and the historical fault records as labels to train the initial model. By calculating the information gain of the features to measure their contribution to the model performance, recursively remove the features with the smallest contribution and retrain and evaluate the model. When the model performance fluctuates within the set threshold, retain the feature set after removal, otherwise restore the feature until the impact of all features on the model performance reaches the set threshold, and screen out the features that have the most influence on fault prediction.

[0042] The fault scenario analysis module includes a scenario classification and location unit and a priority calculation unit. The scenario classification and location unit is used to classify network faults into three types of application scenarios and locate them according to different logics. The priority calculation unit is used to construct an application scenario priority calculation formula based on the influence range and repair time involved in the application scenario, calculate and analyze the priority order of the three types of application scenarios through this formula, and determine the order of fault handling. The fault location decision module includes a matching evaluation unit and a policy execution verification unit. The matching evaluation unit is used to calculate the matching degrees of the current fault features with the three types of application scenarios, denoted as m1, m2, and m3 respectively, sort these matching degrees, and select the scenario with the highest matching degree as the initial fault location strategy. The policy execution verification unit is used to execute the location strategy with the highest matching degree, judge whether the current fault is successfully resolved. If not, select the location measures of other scenarios in descending order of priority and continue to try. For the case where the fault is successfully resolved, record the relevant sub-scenarios as the solution measures of the fault.

[0043] The scenario management module includes a historical scenario recording unit and a scenario extension and adjustment unit. The historical scenario recording unit is used to add the priority order executed by the current fault features to the historical scenario library and record the information during the fault handling process. The scenario extension and adjustment unit is used to extend the newly identified fault scenario as a sub-scenario under the three types of application scenarios, count the number of times the sub-scenario appears, and when the number of times the sub-scenario appears exceeds the set number of times, adjust it to the main scenario.

[0044] Compared with the prior art, the beneficial effects of the present invention are:

[0045] 1. With the continuous emergence of new fault scenarios, the present invention extends them as sub-scenarios to the historical scenario library, and statistically calculates their occurrence frequencies according to the number of occurrences. When the occurrence frequency of a certain sub-scenario reaches a certain level, it is upgraded to the main scenario, enabling it to continuously adapt to new network fault modes;

[0046] 2. By means of recursive feature elimination and screening, features are extracted from multi-dimensional network data, and these features reflect the operating status of network devices, helping to discover potential fault risks;

[0047] 3. By collecting the historical operation data of network devices, establishing a database, and calling the database information during the occurrence of a fault, combining with the real-time monitored network information, the fault troubleshooting range is narrowed. Compared with the traditional manual troubleshooting method, there is no need to check a large number of network devices and data one by one. Brief Description of the Drawings

[0048] Figure 1 It is a schematic flowchart of a method for quickly locating network device faults based on historical data of the present invention;

[0049] Figure 2 This is a schematic structural diagram of a network device fault rapid location system based on historical data according to the present invention. Specific embodiments

[0050] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0051] In the embodiment: as Figure 1 - Figure 2 shown, the present invention provides a technical solution, a method for rapid location of network device faults based on historical data, and the method includes the following steps:

[0052] Collect historical operation data of network devices;

[0053] Classify network faults into three types of application scenarios and perform location according to different logics;

[0054] Clean and preprocess the collected network device data, perform recursive feature elimination and screening based on the preprocessing results, match real-time features with the historical database, and calculate the priority of application scenarios;

[0055] Evaluate the matching degree between the current fault features and the three types of application scenarios, select the most matching scenario as the fault location strategy, and record the execution priority order in the historical scenario library;

[0056] For newly identified fault scenarios, expand them as sub-scenarios, and judge whether they need to be upgraded to main scenarios according to the occurrence times. Among them, the sub-scenario refers to a fault scenario newly identified during the fault location process and not yet classified into existing main scenarios. The sub-scenario is generated based on the fact that the current fault features do not completely match the existing fault scenario patterns, but have similarity or relevance; the main scenario refers to a set of representative and typical fault scenarios formed after classifying and summarizing network faults. The set of fault scenarios is summarized based on historical fault data and actual experience. Each main scenario corresponds to specific fault features, location logics, and processing strategies.

[0057] Collect historical operation data of network devices through the SNMP protocol, including device logs, traffic information, device performance configurations, and fault records. The specific steps include:

[0058] Collect historical operation data of network devices through the SNMP protocol, including device logs, traffic information, network device performance configurations, and fault records; among them, the traffic information includes network traffic involving multiple regions of offline and online platforms, as well as network traffic demand data;

[0059] The network traffic demand data represents the data on the demand for data transmission traffic in the network and the traffic usage behavior characteristics within a specific time, expressed as: [F1, F2, …, Fz]; where z is a positive integer representing the number of traffic demand types, and F1, F2, …, Fz respectively represent the 1st, 2nd, …, zth traffic demand types; the traffic demand data is obtained through network traffic monitoring devices and network service providers; the data at different historical moments of traffic demand has the same time interval Δf; the data is arranged in chronological order [f1, f2, …, fw] to form network traffic demand time series data; where w represents the number of moments, and f1, f2, …, fw respectively represent the 1st, 2nd, …, wth traffic demand data collection moments;

[0060] The network device performance configuration data represents the operating performance state presented by the network device, the dynamic process of configuration parameter adjustment, and the device health status data, expressed as [P1, P2, …, Pu]; where u is a positive integer representing the type of network device performance configuration data, and P1, P2, …, Pu respectively represent the 1st, 2nd, …, u th network device performance configuration data; the network device performance configuration data is obtained from the device's own monitoring system and device manufacturers; the data at different historical moments of the network device performance configuration has the same time interval Δp; the data is arranged in chronological order [p1, p2, …, pv] to form network device performance configuration time series data; where v represents the number of moments, and p1, p2, …, pv respectively represent the 1st, 2nd, …, v th network device performance configuration data collection moments.

[0061] Specifically, in a certain clothing store that has multiple offline stores (taking 3 stores as an example, marked as Store A, Store B, and Store C respectively) and conducts business relying on an online platform, its network architecture covers multiple regional networks, and is supported by various network devices to support business operations. During the actual operation process, a large amount of historical operation data of network devices is generated;

[0062] Traffic information: With the help of network traffic monitoring devices and network service providers, network traffic demand data for the past month has been obtained. During the period from 10:00 to 11:00 am on weekdays at Store A, the network traffic demand data F1, F2, F3 (representing the data traffic of the sales system, the data traffic of surveillance videos, and the data traffic of the office system respectively), at 10:00 on a certain weekday, the traffic values are 5Mbps, 2Mbps, and 1Mbps, and the data time interval Δf at different historical moments is 1 minute. A time series data is formed in the order of f1, f2,..., f60 according to time;

[0063] Network device performance configuration data has been collected from the device's own monitoring system and the device manufacturer. Taking the core switch of Store A as an example, its network device performance configuration data P1, P2, P3 (representing CPU usage rate, memory usage rate, and port bandwidth utilization rate respectively), at 10:10 on a certain day, the data is 60%, 70%, 50%, and the data time interval Δp is 10 minutes. A time series data is formed in the order of p1, p2,..., p144 according to time

[0064] Device logs and fault records are collected through the SNMP protocol. The switch of Store B records "abnormal connection of port 5" at a certain moment, as well as various types of fault information that occurred in the past, including fault time, fault description, etc.

[0065] Network faults are classified into three types of application scenarios and located according to different logics. The specific steps include:

[0066] In the first type of application scenario: When the traffic in the i-th area drops below the designed traffic threshold and there is no switch error reported in the device log, in the first type of application scenario, historical link interruption cases (optical cables being dug up) are preferentially matched, and the link detection tool (Traceroute) is directly triggered. Here, i is used to identify a specific area, and the designed traffic threshold is determined through historical data;

[0067] In the second type of application scenario: When there are abnormal situations in multiple areas and traffic, and the device log shows that the switch CPU / memory is overloaded or the port is down, in the second type of application scenario, the status of the core switch (port status, load threshold) is preferentially checked, and the recovery strategies for historical similar faults (restarting or switching to a standby switch) are compared;

[0068] In the third type of application scenario: When local user connections fail and authentication errors or ARP conflicts frequently appear in the log, in the third type of application scenario, the configuration snapshot of the edge device (store switch) is preferentially checked, and compared with the historical correct configuration version;

[0069] The first type of application scenario, the second type of application scenario, and the third type of application scenario together constitute the initial application scenario set for network device fault location.

[0070] Clean and preprocess the data related to network devices collected, remove the noise data. For the traffic data, when the traffic value at a certain moment is missing, according to the traffic data at the adjacent moments before and after, use the linear interpolation method to fill it, and standardize the data format, and store all the data according to the unified database table structure;

[0071] Extract features from the preprocessed data, denoted as [T1, T2,..., Tn];

[0072] Select a decision tree machine learning model for fault prediction, use the data after preprocessing and feature extraction as input, and the historical fault records as labels to train the initial model;

[0073] Measure the contribution degree of each feature to the model performance by calculating the information gain of the feature;

[0074] Remove the feature with the smallest contribution degree from the current feature set, retrain the model, and evaluate the model performance. When the model performance fluctuates within 2%, retain the feature set after removing the feature; otherwise, restore the feature, where q represents the threshold parameter for measuring the degree of change in model performance;

[0075] Repeat the above process until the influence of all features in the feature set on the model performance reaches the set influence threshold, where the influence threshold is determined by cross-validation and the accuracy index of model performance evaluation;

[0076] Match the real-time features with the historical database, sort them by priority, analyze the priority order of the three types of application scenarios, construct an application scenario priority calculation formula based on the influence range and repair time involved in the application scenario, and according to the application scenario priority calculation formula, sort the three types of application scenarios according to the priority size, where the application scenario priority calculation formula is as follows:

[0077] Asp = α * Si + β * Rt;

[0078] Among them, Asp represents the application scenario priority, Si represents the influence range involved in the application scenario, Rt represents the repair time involved in the application scenario, and α, β represent weight coefficients.

[0079] Specifically, set a reasonable threshold range for the collected traffic data. The normal range of the traffic data of the sales system is 2 Mbps - 10 Mbps. When the traffic value at a certain moment is 20 Mbps, it is judged as noise data and removed;

[0080] Missing value filling: When the monitoring video data traffic value of store C is missing at a certain moment, linear interpolation is used to fill it according to the traffic data at the adjacent moments before and after. Given that the traffic at the previous moment is 1.8 Mbps and the traffic at the next moment is 2.2 Mbps, the missing value is filled as (1.8 + 2.2) ÷ 2 = 2 Mbps;

[0081] Unify all data formats, unify the time format to "YYYY-MM-DD HH:MM:SS", unify the traffic data unit to Mbps, standardize the device performance data in percentage form, and store it according to the unified database table structure;

[0082] Extract features T1, T2, T3 from the preprocessed data, such as traffic change amplitude, error log type, and device load. Taking the data traffic of the sales system of store A as an example, calculate the traffic change amplitude within a certain period. The traffic at the previous moment is 5 Mbps and the current moment is 6 Mbps, so the traffic change amplitude is (6 - 5) ÷ 5 × 100% = 20%; count the error log types, such as "authentication failure" and "port connection timeout"; the device load is calculated based on the CPU usage rate, memory usage rate, etc. to calculate the comprehensive load index;

[0083] Select the decision tree machine learning model, use the preprocessed and feature-extracted data as input, and the historical fault records as labels to train the initial model. Measure the contribution degree by calculating the information gain of the features. Given the initial feature set T1, T2, T3, it is found that the contribution degree of T3 is the smallest. After removing T3, retrain the model and evaluate the model performance. If the model accuracy fluctuates from 80% to 79% (the fluctuation is within the set 2% threshold q), then retain the feature set after removing T3, and repeat this process until the impact of all features on the model performance reaches the set impact threshold.

[0084] Calculate the matching degree between the current fault feature and the three types of application scenarios. The specific steps include:

[0085] Calculate the matching degree between the current fault feature and the three types of application scenarios. When the real-time data is compared with the three types of application scenarios one by one in the specified priority order, they are denoted as m1, m2, and m3. Among them, m1 represents the matching degree of the real-time data compared with the first type of application scenario, m2 represents the matching degree of the real-time data compared with the second type of application scenario, and m3 represents the matching degree of the real-time data compared with the third type of application scenario;

[0086] Sort the calculated matching degrees m1, m2, and m3 between the current fault feature and the three types of application scenarios, and select the one with the highest matching degree as the positioning strategy to be adopted for the current fault feature;

[0087] Determine whether the selected positioning strategy with the highest matching degree successfully solves the current fault. If the positioning measure with the highest matching degree fails to solve the current fault, select other scenarios in descending order of priority as the positioning measures to be taken for the current scenario;

[0088] Among them, for the positioning measure with the highest matching degree to solve the current fault, the specific steps are as follows:

[0089] Sort the sub-scenario types in the application scenario involved in the positioning measure with the highest matching degree according to the number of times the sub-scenario appears in the application scenario. Analyze them one by one in the order of the number of times the current sub-scenario is sorted under the current application scenario until the current fault is solved by the sub-scenario in the application scenario, and record the sub-scenario in the application scenario as the solution measure for the current fault;

[0090] Add the priority order executed by the current fault feature to the historical scenario library;

[0091] Classify the current fault feature into the three types of application scenarios, that is, expand the newly added scenario as a sub-scenario, and record the number of times the current sub-scenario appears as 1. When the current sub-scenario appears later, add 1 to the number of times the current sub-scenario appears, and finally count the number of times the current sub-scenario appears within a certain period of time;

[0092] Set the number of appearances U1 required for the sub-scenario when the sub-scenario is adjusted to the main scenario. When the number of appearances of the current sub-scenario exceeds the set number of appearances U1 required for the sub-scenario when the sub-scenario is adjusted to the main scenario, adjust the current sub-scenario to the main scenario. Among them, the calculation formula for the number of appearances required for the sub-scenario when the sub-scenario is adjusted to the main scenario is as follows: U1 = γ * TNF;

[0093] Among them, TNF represents the total number of faults, and γ represents the weight coefficient.

[0094] Specifically, at a certain moment, a network fault occurs in store B, some users' connections fail, and at the same time, authentication errors frequently appear in the device logs;

[0095] Scenario matching: Determine that this fault conforms to the third type of application scenario. First, check the configuration snapshot of the edge device (switch of store B), compare it with the historical correct configuration version, and find that a certain VLAN configuration of the switch is incorrect, resulting in some users being unable to authenticate correctly;

[0096] Priority calculation: Based on the scope of influence and repair time involved in the application scenario, the priority calculation formula for the application scenario Asp = α * Si + β * Rt is constructed. The scope of influence Si is determined to be 8 (full score 10) according to the number of affected users, the number of business functions, etc. The repair time Rt is expected to be 2 (the repair time for simple configuration errors is relatively short). The weight coefficients are α = 0.6 and β = 0.4. Then the priority of this scenario Asp = 0.6 × 8 + 0.4 × 2 = 5.6. After comparing with the priorities of other possible fault scenarios, the processing order is determined;

[0097] Fault handling: According to the comparison result of the configuration snapshot, correct the VLAN configuration of the switch, the network returns to normal, the user connection is successful, and the authentication error disappears;

[0098] Recording and updating: Add the priority order executed by the current fault feature to the historical scenario library. This fault scenario is extended as a sub-scenario of the third type of application scenario. The initial occurrence count is recorded as 1. When a similar fault scenario appears again later, the occurrence count of the sub-scenario is incremented by 1. Set the number of occurrences U1 required for the sub-scenario to be adjusted to the main scenario as U1 = γ × TNF (the total number of faults TNF = 100, the weight coefficient γ = 0.1, then U1 = 10). When the occurrence count of this sub-scenario exceeds 10 times, upgrade it to the main scenario and optimize the fault location strategy.

[0099] A network device fault rapid location system based on historical data, the system includes a data acquisition module, a data processing module, a fault scenario analysis module, a fault location decision module, and a scenario management module. The data acquisition module is used to collect the historical operation data of network devices; the data processing module is used to clean and preprocess the collected data related to network devices, and extract features based on the preprocessed data; the fault scenario analysis module is used to classify network faults into three types of application scenarios, locate them according to different logics, match real-time features with the historical database, and calculate the priority of the application scenario; the fault location decision module is used to evaluate the matching degree of the current fault feature with the three types of application scenarios, select the most matching scenario as the fault location strategy, and record the executed priority order in the historical scenario library; the scenario management module is used to expand the newly identified fault scenario as a sub-scenario, and judge whether it needs to be upgraded to the main scenario according to the occurrence count.

[0100] The data acquisition module includes a data collection unit and a time series construction unit. The data collection unit uses the SNMP protocol to collect data from multiple channels, such as network devices, network traffic monitoring devices, network service providers, the device's own monitoring system, and device manufacturers. The collected content includes device logs, traffic information, device performance configuration data, and fault records. The time series construction unit is used to construct network traffic demand time series data and network device performance configuration time series data for the collected traffic demand data and device performance configuration data respectively according to their respective time intervals.

[0101] The data processing module includes a preprocessing unit, a feature extraction unit, and a feature screening unit. The preprocessing unit is used to remove noise data from the collected data, fill in missing values in the traffic data using the linear interpolation method, standardize all data formats, and uniformly store them in the database table structure. The feature extraction unit is used to extract features from the preprocessed data, denoted as [T1, T2,..., Tn]. The feature screening unit is used to select a decision tree machine learning model, use the preprocessed and feature-extracted data as input, and the historical fault records as labels to train the initial model. By calculating the information gain of the features to measure their contribution to the model performance, recursively remove the features with the smallest contribution and retrain and evaluate the model. When the model performance fluctuates within the set threshold, retain the feature set after removal; otherwise, restore the feature until the impact of all features on the model performance reaches the set threshold, and screen out the features that have the most influence on fault prediction.

[0102] The fault scenario analysis module includes a scenario classification and location unit and a priority calculation unit. The scenario classification and location unit is used to classify network faults into three types of application scenarios and locate them according to different logics. The priority calculation unit is used to construct an application scenario priority calculation formula based on the impact range and repair time involved in the application scenario, calculate and analyze the priority order of the three types of application scenarios through this formula, and determine the order of fault handling. The fault location decision module includes a matching evaluation unit and a strategy execution verification unit. The matching evaluation unit is used to calculate the matching degree of the current fault features with the three types of application scenarios, denoted as m1, m2, and m3 respectively, sort these matching degrees, and select the scenario with the highest matching degree as the initial fault location strategy. The strategy execution verification unit is used to execute the location strategy with the highest matching degree and determine whether the current fault is successfully resolved. If not, select the location measures of other scenarios in descending order of priority and continue to try. For the case where the fault is successfully resolved, record the relevant sub-scenarios as the fault solution measures.

[0103] The scene management module includes a historical scene recording unit and a scene extension and adjustment unit. The historical scene recording unit is used to add the priority order executed by the current fault feature to the historical scene library and record the information during the fault handling process. The scene extension and adjustment unit is used to extend the newly identified fault scene as a sub-scene to three types of application scenarios, count the number of times the sub-scene appears, and when the number of times the sub-scene appears exceeds the set number of times, adjust it to the main scene.

[0104] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. A method for quickly locating network device faults based on historical data, characterized in that: The method includes the following steps: Collect historical operation data of network devices; Classify network faults into three types of application scenarios and locate them according to different logics; Clean and preprocess the collected network device data, perform recursive feature elimination and screening based on the preprocessing results, match real-time features with the historical database, and calculate the priorities of application scenarios; Evaluate the matching degree between the current fault features and the three types of application scenarios, select the most matching scenario as the fault location strategy, and record the executed priority order in the historical scenario library. Among them, the three types of application scenarios belong to the category of main scenarios, and the three types of application scenarios are the initial application scenario set; For newly identified fault scenarios, expand them as sub-scenarios, and judge whether to upgrade them to main scenarios according to the occurrence times. Among them, the main scenario represents a set of representative and typical fault scenarios formed after classifying and summarizing network faults; the sub-scenario represents a fault scenario newly identified during the fault location process that has not been summarized into the existing main scenarios.

2. The method for quickly locating network device faults based on historical data according to claim 1, characterized in that: Collect historical operation data of network devices through the SNMP protocol, including device logs, traffic information, device performance configurations, and fault records. The specific steps include: Collect historical operation data of network devices through the SNMP protocol, including device logs, traffic information, network device performance configurations, and fault records; among them, the traffic information includes network traffic involving multiple regions of offline and online platforms, as well as network traffic demand data; The network traffic demand data represents the data of the demand for data transmission traffic in the network and the traffic usage behavior characteristics within a specific time, expressed as: [F1, F2, …, Fz]; where z is a positive integer representing the number of traffic demand types, and F1, F2, …, Fz respectively represent the 1st, 2nd, …, zth traffic demand types; the traffic demand data is obtained through network traffic monitoring devices and network service providers; the historical data at different moments of traffic demand has the same time interval Δf; the data is arranged in chronological order [f1, f2, …, fw] to form a network traffic demand time series data; where w represents the number of moments, and f1, f2, …, fw respectively represent the 1st, 2nd, …, wth traffic demand data collection moments; The network device performance configuration data represents the operating performance state presented by the network device, the dynamic process of configuration parameter adjustment, and the device health status data, expressed as [P1, P2, …, Pu]; where u is a positive integer representing the type of network device performance configuration data, and P1, P2, …, Pu respectively represent the 1st, 2nd, …, u types of network device performance configuration data; the network device performance configuration data is obtained from the device's own monitoring system and device manufacturers; the historical data at different moments of network device performance configuration has the same time interval Δp; the data is arranged in chronological order [p1, p2, …, pv] to form a network device performance configuration time series data; where v represents the number of moments, and p1, p2, …, pv respectively represent the 1st, 2nd, …, vth network device performance configuration data collection moments.

3. A method for quickly locating network device faults based on historical data according to claim 2, characterized in that: Network failures are classified into three application scenarios and located according to different logics. The specific steps are as follows: In the first application scenario: when the traffic in the i-th area drops to the designed traffic threshold and there is no switch error reported in the device log, historical link interruption cases are preferentially matched in the first application scenario, and the link detection tool is directly triggered, where i is used to identify a specific area, and the designed traffic threshold is determined by historical data; In the second application scenario: when there are abnormal situations in the traffic of multiple areas and the device log shows that the switch CPU / memory is overloaded or the port is down, the status of the core switch is preferentially checked in the second application scenario, and the recovery strategies for historical similar faults are compared; In the third application scenario: when local user connections fail and authentication errors or ARP conflicts frequently appear in the log, the configuration snapshot of the edge device is preferentially checked in the third application scenario, and compared with the historical correct configuration version; The first application scenario, the second application scenario, and the third application scenario together constitute the initial application scenario set for network device fault location.

4. The method for quickly locating network device faults based on historical data according to claim 3, characterized in that: The network device-related data collected is cleaned and preprocessed to remove noise data. For traffic data, when the traffic value at a certain moment is missing, linear interpolation is used to fill it according to the traffic data at the adjacent moments before and after, and the data format is standardized, and all data is stored according to a unified database table structure; Extract features from the preprocessed data, denoted as [T1, T2,..., Tn]; Select a decision tree machine learning model for fault prediction. Use the preprocessed and feature-extracted data as input and historical fault records as labels to train the initial model; Measure the contribution of each feature to the model performance by calculating the information gain of the feature; Remove the feature with the smallest contribution from the current feature set, retrain the model, and evaluate the model performance. When the model performance fluctuates within q, retain the feature set after removing the feature; Otherwise, restore the feature, where q represents the threshold parameter for measuring the degree of change in model performance; Repeat the above process until the influence of all features in the feature set on the model performance reaches the set influence threshold, where the influence threshold is determined by cross-validation and the accuracy index of model performance evaluation; Match the real-time features with the historical database, sort them by priority, analyze the priority order of the three application scenarios, construct an application scenario priority calculation formula based on the influence range and repair time involved in the application scenario, and sort the three application scenarios according to the priority size according to the application scenario priority calculation formula.

5. A method for quickly locating network device faults based on historical data according to claim 4, characterized in that: Calculate the matching degree between the current fault feature and the three application scenarios. The specific steps are as follows: Calculate the matching degree between the current fault feature and the three types of application scenarios. When the real-time data is compared with the three types of application scenarios one by one in the specified priority order, it is denoted as m1, m2, and m3. Among them, m1 represents the matching degree of the real-time data compared with the first type of application scenario, m2 represents the matching degree of the real-time data compared with the second type of application scenario, and m3 represents the matching degree of the real-time data compared with the third type of application scenario; Sort the calculated matching degrees m1, m2, and m3 between the current fault feature and the three types of application scenarios, and select the one with the highest matching degree as the positioning strategy to be adopted for the current fault feature; Judge whether the positioning strategy adopted with the highest matching degree can successfully solve the current fault. If the positioning measure with the highest matching degree fails to solve the current fault, then select other scenarios in descending order of priority as the positioning measure required for the current scenario; Among them, for the positioning measure with the highest matching degree to solve the current fault, the specific steps are as follows: Sort the sub-scenario types under the application scenario involved in the positioning measure with the highest matching degree according to the number of occurrences of the sub-scenario under the application scenario. Analyze them one by one in the order of the number of occurrences of the current sub-scenario under the current application scenario until the current fault is solved by the sub-scenario under the application scenario, and then stop the analysis. Denote the sub-scenario under the application scenario as the solution measure for the current fault; Add the priority order executed by the current fault feature to the historical scenario library; Classify the current fault feature under the three types of application scenarios, that is, expand the newly added scenario as a sub-scenario, and record the number of occurrences of the current sub-scenario as 1. When the current sub-scenario appears later, add 1 to the number of occurrences of the current sub-scenario, and finally count the number of occurrences of the current sub-scenario within a certain period of time; Set the number of occurrences U1 required for the sub-scenario to be adjusted to the main scenario. When the number of occurrences of the current sub-scenario exceeds the set number of occurrences U1 required for the sub-scenario to be adjusted to the main scenario, adjust the current sub-scenario to the main scenario.

6. A rapid fault location system for network devices based on historical data, characterized in that: The system includes a data acquisition module, a data processing module, a fault scenario analysis module, a fault location decision module, and a scenario management module. The data acquisition module collects the historical operation data of network devices; the data processing module is used to clean and preprocess the collected data related to network devices, and perform feature extraction based on the preprocessed data; The fault scenario analysis module is used to classify network faults into three types of application scenarios, locate them according to different logics, match real-time features with the historical database, and calculate the priority of the application scenarios; The fault location decision module is used to evaluate the matching degree between the current fault feature and the three types of application scenarios, select the most matching scenario as the fault location strategy, and record the executed priority order in the historical scenario library; The scenario management module is used to expand the newly identified fault scenario as a sub-scenario, and judge whether it needs to be upgraded to the main scenario according to the number of occurrences.

7. The rapid fault location system for network devices based on historical data according to claim 6, characterized in that: The data acquisition module includes a data collection unit and a time series construction unit. The data collection unit uses the SNMP protocol to collect data from multiple channels, including network devices, network traffic monitoring devices, network service providers, the device's own monitoring system, and device manufacturers. The data collected includes device logs, traffic information, device performance configuration data, and fault records. The time series construction unit is used to construct network traffic demand time series data and network device performance configuration time series data for the collected traffic demand data and device performance configuration data respectively according to their respective time intervals.

8. The quick fault location system of network equipment based on historical data according to claim 7, characterized in that: The data processing module includes a preprocessing unit, a feature extraction unit, and a feature screening unit. The preprocessing unit is used to remove noise data from the collected data, fill in missing values in traffic data using linear interpolation, standardize all data formats, and uniformly store them in the database table structure. The feature extraction unit is used to extract features from the preprocessed data, denoted as [T1, T2,..., Tn]. The feature screening unit is used to select a decision tree machine learning model, use the preprocessed and feature-extracted data as input, and historical fault records as labels to train the initial model. By calculating the information gain of the features to measure their contribution to the model performance, recursively remove the features with the smallest contribution and retrain and evaluate the model. When the model performance fluctuates within the set threshold, retain the feature set after removal; otherwise, restore the feature until the impact of all features on the model performance reaches the set threshold, and screen out the features that have the most influence on fault prediction.

9. The fast fault location system for network devices based on historical data according to claim 8, wherein: The fault scenario analysis module includes a scenario classification and location unit and a priority calculation unit. The scenario classification and location unit classifies network faults into three types of application scenarios and locates them according to different logics. The priority calculation unit constructs an application scenario priority calculation formula based on the impact range and repair time involved in the application scenario, analyzes the priority order of the three types of application scenarios, and determines the order of fault handling. The fault location decision module includes a matching evaluation unit and a policy execution verification unit. The matching evaluation unit is used to calculate the matching degrees of the current fault features with the three types of application scenarios, denoted as m1, m2, and m3 respectively, sort these matching degrees, and select the scenario with the highest matching degree as the initial fault location policy. The policy execution verification unit is used to execute the location policy with the highest matching degree and determine whether the current fault is successfully resolved. If not, select other location measures for other scenarios in descending order of priority and continue to try. For the case where the fault is successfully resolved, record the relevant sub-scenarios as the fault resolution measures.

10. A rapid fault location system for network devices based on historical data according to claim 9, characterized in that: The scenario management module includes a historical scenario recording unit and a scenario extension and adjustment unit. The historical scenario recording unit is used to add the priority order executed by the current fault features to the historical scenario library and record the information during the fault handling process. The scenario extension and adjustment unit is used to extend the newly identified fault scenarios as sub-scenarios under the three types of application scenarios, count the number of times the sub-scenarios appear, and when the number of times the sub-scenarios appear exceeds the set number of times, adjust them to the main scenarios.

Citation Information

Patent Citations

  • Fault processing method and device and storage medium

    CN117640338A

  • Scene-adaptive fault first-aid repair scheduling method and system

    CN118735208A

  • Communication network operation and maintenance fault positioning and tracking method and system

    CN119420639A

  • Fault positioning method and system

    CN119473858A

  • Fault monitoring in a communications network

    GB2597920A

Cited By

  • Network equipment fault intelligent detection method and system based on data transmission

    CN120934995A