Intelligent network fault detection method and system
By building a labeled topological subgraph database and calculating the chain similarity index, the intelligent network fault detection method can effectively predict the chain reaction of network faults, improve the accuracy and efficiency of fault handling, and solve the problem of low accuracy in the prediction of chain faults in the prior art.
Patent Information
- Application Number
- CN202510351350.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-05-27
AI Technical Summary
The prior art is good for single fault detection in network fault detection, but lacks effective prediction of chain reactions caused by failure of a certain node, and does not fully utilize historical data for multi-dimensional data evaluation and analysis, resulting in low prediction accuracy.
An intelligent network fault detection method is proposed. By collecting historical fault cases, a labeled topological subgraph database is constructed, the fault node is identified when a fault occurs, and its neighbor node is determined. Based on similar chain failures, similar historical fault cases are matched in the database, the time zone similarity index, performance similarity index and environmental similarity index are calculated, and the degree of similarity between the current fault and historical similar chain failure is comprehensively evaluated, so as to predict possible hidden danger chain failures.
It effectively improves the diagnosis ability of multi-node chain faults, reduces dependence on manual experience, improves fault processing efficiency and accuracy, and reduces the probability of misjudgment.
Smart Images

Figure CN120050160A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network detection, and in particular to an intelligent network fault detection method and system. Background Art
[0002] With the rapid development of network technology, the scale of networks continues to expand and the network structure becomes increasingly complex. In this context, network failures become more frequent, with various types of failures, making it more difficult to locate and repair them.
[0003] However, the network fault detection method in the prior art still has the following deficiencies in practical application: The existing network fault detection methods are good at detecting single faults, but lack effective prediction of the chain reaction caused by a node failure, such as the spread of neighbor node failures; In addition, when predicting cascading failures, historical data was not fully utilized to evaluate and analyze multi-dimensional data, resulting in the accuracy of the prediction not being guaranteed.
[0004] Therefore, an intelligent network fault detection method and system are introduced. Summary of the invention
[0005] The purpose of the present invention is to solve the problems mentioned in the background technology and to provide an intelligent network fault detection method and system.
[0006] The purpose of the present invention can be achieved by the following technical solution: An intelligent network fault detection method, comprising: Fault trigger diagnosis: Collect historical fault cases and build a labeled topological subgraph database. When a fault occurs at a certain node, identify the location of the faulty node in the network topology map, and then determine the nodes directly connected to the faulty node as neighbor nodes. Based on the neighbor nodes corresponding to the faulty node, determine similar chain faults that are adapted to the current fault in the topological subgraph database, and analyze the chain similarity index between each similar chain fault and the current fault. , determine the hidden danger chain failure caused by the current faulty node; Diagnosis result processing: Based on the chain similarity index between each similar chain fault and the current fault , execute the corresponding steps to determine the hidden danger chain failure caused by the current faulty node.
[0007] As a preferred implementation of the present invention, similar chain failures adapted to the current failure are determined in the topology subgraph database based on the neighbor nodes corresponding to the failed node, specifically: Identify the initial code of the current faulty node, set an initial code corresponding to each node, represented by Ci, where i represents the node number, i=1,2,...,k, k is the total number of nodes; count the number of neighboring nodes of the faulty node, represented by r, to form the label corresponding to the current faulty node, that is, Ci-r; The label corresponding to the current fault node is input into the topology subgraph database and matched with the labels of each historical fault case, and the successfully matched historical fault case is used as a similar chain fault adapted to the current fault.
[0008] As a preferred embodiment of the present invention, the chain similarity index between each similar chain failure and the current failure is analyzed. , specifically: Identify the time point of the fault occurrence of the current fault node and match it with each set of preset time windows, each set of time windows corresponds to a time coefficient; determine the time coefficient corresponding to the current fault node, calculate the difference between the time coefficient of the current fault node and the time coefficient of each similar chain fault, and take the absolute value as the time zone similarity index p1 between each similar chain fault and the current fault node; Extract the CPU, memory and traffic data of the faulty node when it fails; and calculate the ratio with the preset CPU standard value, memory standard value and traffic standard value, that is, take the CPU, memory and traffic data of the faulty node when it fails as the numerator, and the CPU standard value, memory standard value and traffic standard value as the denominator; obtain the CPU evaluation value, memory evaluation value and traffic evaluation value of the faulty node; The CPU evaluation value, memory evaluation value and traffic evaluation value of the faulty node are used as the length, width and height of the three-dimensional rectangular model respectively. A three-dimensional rectangular model of the current faulty node is constructed when a fault occurs. The cosine similarity is calculated between the three-dimensional rectangular models of similar chain faults to obtain the performance similarity index p2 between each similar chain fault and the current faulty node.
[0009] As a preferred embodiment of the present invention, the chain similarity index between each similar chain failure and the current failure is analyzed. , also includes: Extract the room temperature, room humidity, and equipment power supply voltage when the faulty node fails, and mark them as u1, u2, and u3 respectively; Extract the room temperature, room humidity, and equipment power supply voltage corresponding to each similar chain failure, and record them as e1, e2, and e3; The room temperature, room humidity and equipment power supply voltage at the time of the faulty node failure are normalized and entered into the formula Perform weighted calculation to obtain the environmental similarity index p3 between each similar chain failure and the current fault node; a1, a2, and a3 are the influence weight factors of the computer room temperature, computer room humidity, and equipment power supply voltage, respectively; Extract the time zone similarity index p1, performance similarity index p2, and environment similarity index p3 between each similar chain failure and the current fault node, normalize them, and then enter them into the formula Perform weighted calculation to determine the chain similarity index of each similar chain failure ; Among them, n1, n2 and n3 are the influencing weight factors of the time zone similarity index p1, the performance similarity index p2 and the environment similarity index p3 respectively.
[0010] As a preferred implementation of the present invention, corresponding steps are performed to determine the hidden danger chain failure caused by the current faulty node, specifically: Preset chain similarity index The corresponding index reference range, if a similar chain failure chain similarity index If the number of highly similar data is greater than one, the similarity index is taken. Larger highly similar data is used as the hidden danger chain failure caused by the current fault node and the corresponding fault solution is extracted as the response plan for the current fault node; If there is no chain failure similarity index If the fault is higher than the index reference range, similar chain faults within the index reference range are identified, and the chain similarity index within the index reference range is taken. A large similar chain failure is sent to any technician, who will make a judgment and select a group of similar chain failures as the hidden chain failure caused by the current fault node, or trigger a search signal to select the technician with the highest performance index to search the current fault node; If the chain similarity index of each similar chain failure If both are lower than the index reference range, the search signaling is directly triggered to select the technician with the highest performance index to search the current faulty node.
[0011] As a preferred embodiment of the present invention, the performance index of the technicians obtained by selection is specifically: Extract the initial code Ci of the current fault node, obtain the historical processing times of each technician processing Ci, and select the top three technicians with the highest historical processing times as pre-selected personnel; Obtain the working time, historical processing records and assessment results of each pre-selected personnel; extract the processing time used by each pre-selected personnel for each Ci fault from the historical processing records, and calculate the average processing time of each pre-selected personnel after calculating the average value; Preset the intervals of working hours and average processing time for each group, corresponding to the working hours and average processing time, respectively. Each working hours interval corresponds to an experience evaluation value, and each average processing time interval corresponds to an efficiency evaluation value; convert the working hours and average processing time of each pre-selected person into experience evaluation values and efficiency evaluation values respectively; Based on regular internal technical assessments, obtain the internal assessment results of each pre-selected personnel closest to the time point of the current fault occurrence, and record them as the technical assessment value; The technical assessment value, experience evaluation value and efficiency evaluation value of each pre-selected candidate are multiplied by the corresponding preset weight coefficient respectively, and then the sum is obtained to obtain the performance index of each pre-selected candidate.
[0012] As a preferred implementation of the present invention, the technician determines whether a cascading failure occurs after searching the current faulty node, and if a cascading failure occurs, the technician updates the topology subgraph database after searching and maintaining.
[0013] An intelligent network fault detection system, comprising: Topology map construction module: each device in the network is regarded as a node, and the connection links between devices are regarded as edges to construct a network topology map; Fault diagnosis module: collect historical fault cases and build a labeled topological subgraph database. When a fault occurs at a certain node, identify the location of the faulty node in the network topology map, and determine the nodes directly connected to the faulty node as neighbor nodes. Based on the neighbor nodes corresponding to the faulty node, determine similar chain faults that are adapted to the current fault in the topological subgraph database, and analyze the chain similarity index between each similar chain fault and the current fault. ; Fault processing module: based on the chain similarity index between each similar chain fault and the current fault , execute the corresponding steps to determine the hidden danger chain failure caused by the current faulty node Compared with the prior art, the present invention has the following beneficial effects: 1. When a node fails, the present invention matches similar historical failure cases in the database according to the label of the failed node to determine similar chain failures. By calculating the time zone similarity index, performance similarity index and environment similarity index, and comprehensively obtaining the chain similarity index, the similarity between the current failure and the historical similar chain failures is evaluated from multiple dimensions, thereby predicting the potential chain failures that may be caused by the current failure, effectively improving the diagnostic capability of multi-node chain failures, and solving the problem of lack of effective prediction of chain reactions caused by a certain node failure and low prediction accuracy in the prior art; 2. The present invention calculates the chain similarity index by using multi-dimensional data, and automatically makes fault handling decisions based on the comparison results between the index and the preset reference range. If the chain similarity index is higher than the reference range, the corresponding fault solution is directly extracted; if it is within the reference range, the technician determines or selects the technician with the highest performance index to handle it; if both are lower than the reference range, the technician with the highest performance index is also selected to handle it. This method reduces the reliance on manual experience, improves the efficiency of fault handling, and reduces the probability of misjudgment; 3. The present invention extracts the initial code of the current fault node, obtains the historical processing times, working hours, historical processing records, and assessment performance information of each technician, calculates the average processing time, and converts the working hours and average processing time into experience evaluation values and efficiency evaluation values. Combined with the technical assessment value, the performance index of the technician is obtained through weighted calculation. The most suitable technician is selected to handle the fault based on the performance index, thereby realizing the rational use of technician resources and further improving the efficiency and accuracy of fault handling. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.
[0015] Figure 1 is a flow chart of the present invention; Figure 2 It is a principle block diagram of the present invention. DETAILED DESCRIPTION
[0016] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0017] Example 1 See also Figure 1 As shown, an intelligent network fault detection method includes: Construct a network topology diagram: Take each device in the network as a node and the connection links between devices as edges to construct a network topology diagram; It should be noted that Node determination: Comprehensively sort out all devices in the network, define routers, switches, servers and other key devices with network connection functions as nodes in the network topology diagram, and assign a unique identifier to each node, such as the IP address or MAC address of the device, so as to accurately identify and distinguish them in subsequent analysis; Edge construction: Obtain the physical and logical connection information between devices through network scanning tools or device management systems. For physical connections, such as devices connected by Ethernet cables, optical fibers, etc., they are connected by edges in the topology diagram; for logical connections, such as connections established through virtual private networks (VPNs) or network tunnels, corresponding edges are also constructed. Each edge records the type of connection (such as Ethernet links, wireless links, etc.), bandwidth, latency and other attribute information; GNN topology analysis: Using the node embedding algorithm in the graph neural network (GNN), random walks are performed from each node on the pre-built network topology. The walk process is repeated continuously to generate a node sequence of the current network topology. These node sequences are used as input and trained using the word vector model in natural language processing (such as the Skip-Gram model) to obtain a low-dimensional vector representation of each node. This vector contains the location information of the node in the network topology and the connection relationship information with other nodes, thereby capturing the dependencies between devices. Using a graph convolution-based method, the node embedding vector is used as input to calculate the weighted sum of the node embedding vectors at both ends of each edge. The specific value of the weight is preset by the technician. Through the transformation of the graph convolution layer, a low-dimensional vector representation of the edge is obtained. The edge embedding vector reflects the relationship strength, connection type and other information between the two devices connected by the edge, which helps to explore potential fault propagation paths. Fault trigger diagnosis: Collect historical fault cases and build a labeled topological subgraph database. When a fault occurs at a certain node, identify the location of the faulty node in the network topology map, and then determine the nodes directly connected to the faulty node as neighbor nodes. Based on the neighbor nodes corresponding to the faulty node, determine similar chain faults that are adapted to the current fault in the topological subgraph database, and analyze the chain similarity index between each similar chain fault and the current fault. ; It should be noted that all the fault data that has occurred in the past should be collected and sorted, including the node where the fault occurred, the fault type, the time of the fault, the changes in equipment performance indicators, and the chain failures that eventually resulted. These data should be cleaned and pre-processed to remove duplicate, erroneous or incomplete data to ensure data quality; Specifically: S1: Identify the initial code of the current faulty node, set an initial code corresponding to each node, represented by Ci, where i represents the node number, i=1,2,...,k, k is the total number of nodes; count the number of neighboring nodes of the faulty node, represented by r, to form the label corresponding to the current faulty node, that is, Ci-r; The label corresponding to the current fault node is input into the topology subgraph database to match with the labels of each historical fault case, and the successfully matched historical fault case is used as a similar chain fault adapted to the current fault; that is, it is matched with the fault-causing node label of each historical fault case; It should be noted that cleaning and preprocessing the historical fault data to remove duplicate, erroneous or incomplete data ensures the accuracy and reliability of the data in the topological subgraph database. High-quality data is the basis for subsequent fault diagnosis and analysis, and can avoid misdiagnosis or misjudgment caused by data problems, thereby improving the performance and credibility of the entire fault diagnosis system. By setting initial codes for different nodes to form labels in a unified format, fault data is structured and standardized. This standardized label makes data easier to manage, store, and retrieve, and facilitates efficient matching operations in the topology subgraph database. S2: Identify the time point of the fault occurrence of the current fault node and match it with each set of preset time windows. Each set of time windows corresponds to a time coefficient. The time coefficient range is set to 0.5-1.5. Determine the time coefficient corresponding to the current fault node, calculate the difference between the time coefficient of the current fault node and the time coefficient of each similar chain fault, and take the absolute value as the time zone similarity index p1 between each similar chain fault and the current fault node; S3: Extract the CPU, memory and flow data of the faulty node when the fault occurs; calculate the ratio with the preset CPU standard value, memory standard value and flow standard value; set based on historical data; that is, take the CPU, memory and flow data of the faulty node when the fault occurs as the numerator, and the CPU standard value, memory standard value and flow standard value as the denominator; obtain the CPU evaluation value, memory evaluation value and flow evaluation value of the faulty node; The CPU evaluation value, memory evaluation value and traffic evaluation value of the faulty node are used as the length, width and height of the three-dimensional rectangular model respectively, and a three-dimensional rectangular model of the current faulty node is constructed when the fault occurs. The cosine similarity is calculated between the three-dimensional rectangular models of similar chain faults to obtain the performance similarity index p2 between each similar chain fault and the current faulty node. S4: extract the room temperature, room humidity and equipment power supply voltage when the faulty node fails, and mark them as u1, u2 and u3 respectively; Extract the room temperature, room humidity, and equipment power supply voltage corresponding to each similar chain failure, and record them as e1, e2, and e3; The room temperature, room humidity and equipment power supply voltage at the time of the faulty node failure are normalized and entered into the formula Perform weighted calculation to obtain the environmental similarity index p3 between each similar chain failure and the current fault node; a1, a2, and a3 are the influence weight factors of the computer room temperature, computer room humidity, and equipment power supply voltage, respectively; S5: Extract the time zone similarity index p1, performance similarity index p2, and environment similarity index p3 between each similar chain failure and the current fault node, perform normalization, and then enter the formula Perform weighted calculation to determine the chain similarity index of each similar chain failure ; Where n1, n2 and n3 are the influence weight factors of time zone similarity index p1, performance similarity index p2 and environment similarity index p3 respectively; It should be noted that by comprehensively considering the time zone similarity index, performance similarity index and environment similarity index, the similarity between the current fault and similar historical chain faults can be comprehensively evaluated from multiple dimensions, avoiding the inaccurate evaluation caused by analyzing only from a single perspective, making the judgment of fault similarity more accurate.
[0018] By using the preset time window and time coefficient, the time coefficient corresponding to the current fault node can be quickly determined and compared with historical fault cases. This can quickly screen out chain faults that are similar to the current fault in the time dimension, narrow the scope of fault diagnosis, and improve diagnostic efficiency. Considering the performance data of CPU, memory, traffic, etc. when the fault occurs, as well as the environmental data such as room temperature, humidity, and equipment power supply voltage, it can more comprehensively reflect the various conditions when the fault occurs. These data combined with historical data can help accurately identify the feature matching degree between the current fault and similar historical faults, so as to more accurately determine the cause of the fault and possible chain reactions, and improve the accuracy of fault diagnosis; Assume that an enterprise network consists of 10 nodes (including routers, switches, servers and other devices), and the nodes are numbered from 1 to 10, that is, k = 10; Build a topology subgraph database: After a period of data collection and collation, we obtained various fault data that occurred in the history of the network. After cleaning and preprocessing these data, we generated corresponding labels for each historical fault case, for example: Historical failure case 1: Node 4 (C4) fails, and the number of neighboring nodes is 3 (r=3), and its label is C4-3; Historical failure case 2: Node 8 (C8) fails, and the number of neighboring nodes is 5 (r=5), and its label is C8-5; Similarly, such labels are generated for all historical fault cases, and a topological subgraph database is constructed; Current fault occurrence and handling: One day, node 6 (C6) fails; Step S1: First, identify the initial code of the faulty node as C6; Count the number of neighbor nodes of node 6 and find that there are 4 neighbor nodes, that is, r=4; In this way, the label corresponding to the current faulty node is C6-4; The label C6-4 is input into the topology subgraph database and matched with the labels of each historical fault case. Assuming that two historical fault cases are found in the database, with labels C6-4 and C6-3, the two historical fault cases are determined to be similar chain faults that are adapted to the current fault. Step S2: Identify that the fault time of the current faulty node 6 is 3:00 pm (15:00); The preset time windows and corresponding time coefficients are as follows: time window 8:00-12:00, time coefficient 0.5; Time window 12:00-16:00, time coefficient 1.0; Time window 16:00-20:00, time coefficient 1.2; Time window 20:00-8:00 (next day), time coefficient 0.8; Since the fault occurred at 15:00, which is within the time window of 12:00-16:00, the time coefficient corresponding to the current faulty node is 1.0; For the similar cascading fault C6-4 found, its historical fault occurred at 14:00, and the corresponding time coefficient is also 1.0; for the similar cascading fault C6-3, its historical fault occurred at 17:00, and the corresponding time coefficient is 1.2; Calculate the time zone similarity index p1: For similar chain fault C6-4, p1=|1.0-1.0|=0; For similar cascading fault C6-3, p1=|1.0-1.2|=0.2; Step S3: Extract the CPU usage of faulty node 6 when it fails, which is 80%, the memory usage is 70%, and the traffic is 500Mbps; The default CPU standard value is 60%, the memory standard value is 60%, and the traffic standard value is 400Mbps; Calculate the evaluation value of the faulty node: CPU evaluation value = 80% / 60%≈1.33; Memory evaluation value = 70% / 60%≈1.17; Traffic assessment value = 500Mbps / 400Mbps = 1.25; The CPU evaluation value 1.33, the memory evaluation value 1.17, and the traffic evaluation value 1.25 are used as the length, width, and height of the three-dimensional rectangular model, respectively, to construct a three-dimensional rectangular model when the current fault node fails; For the similar cascading failure C6-4, its CPU evaluation value at that time was 1.2, memory evaluation value was 1.1, and traffic evaluation value was 1.3; for the similar cascading failure C6-3, its CPU evaluation value at that time was 1.4, memory evaluation value was 1.0, and traffic evaluation value was 1.1; The performance similarity index p2 is calculated using cosine similarity: the vector of the current fault is A=(1.33, 1.17, 1.25), the vector of the similar chain fault C6-4 is B=(1.2, 1.1, 1.3), and the vector of the similar chain fault C6-3 is C=(1.4, 1.0, 1.1); For similar cascading failure C6-4, it is calculated to be approximately 0.92; For similar cascading failure C6-3, it is calculated to be approximately 0.85; Step S4: extract the room temperature u1=28, room humidity u2=50, and equipment power supply voltage u3=220V when the faulty node 6 fails; For similar chain fault C6-4, when it occurs, the room temperature e1=27, the room humidity e2=52, and the equipment power supply voltage e3=225V; for similar chain fault C6-3, when it occurs, the room temperature e1=29, the room humidity e2=48, and the equipment power supply voltage e3=215V; Normalize the data (assuming linear normalization to the interval [0,1]): For the current fault: the normalized room temperature u1=0.8 (assuming the room temperature range is 20°C to 30°C); The normalized humidity in the computer room is u2=0.5 (assuming the humidity in the computer room ranges from 40% to 60%). The normalized device power supply voltage u3=0.5 (assuming the device power supply voltage range is 200V to 240V); For similar cascading fault C6-4: the normalized room temperature e1=0.7; The normalized humidity in the computer room is e2=0.6; The normalized device supply voltage is e3=0.625; For similar cascading fault C6-3: the normalized room temperature e1=0.9; The normalized room humidity after normalization is e2=0.4; The normalized device supply voltage is e3=0.375; Assume a1=0.4, a2=0.3, a3=0.3, and calculate the environmental similarity index p3: For similar chain failure C6-4: p3 is calculated to be approximately 0.0975; For similar cascading failures C6-3: p3 is calculated to be approximately 0.0875; Diagnosis result processing: Based on the chain similarity index between each similar chain fault and the current fault , execute the corresponding steps to determine the hidden danger chain failure caused by the current faulty node; Specifically: M1: Preset chain similarity index The corresponding index reference range, if a similar chain failure chain similarity index If the number of highly similar data is greater than one, the similarity index is taken. Larger highly similar data is used as the hidden danger chain failure caused by the current fault node and the corresponding fault solution is extracted as the response plan for the current fault node; M2: Chain similarity index if there is no similar chain failure If the fault is higher than the index reference range, similar chain faults within the index reference range are identified, and the chain similarity index within the index reference range is taken. A large similar chain failure is sent to any technician, who will make a judgment and select a group of similar chain failures as the hidden chain failure caused by the current fault node, or trigger a search signal to select the technician with the highest performance index to search the current fault node; M3: If the chain similarity index of each similar chain failure If both are lower than the index reference range, the search signaling is directly triggered to select the technician with the highest performance index to search the current fault node; M4: After searching the current faulty node, the technician determines whether a cascading failure has occurred. If a cascading failure has occurred, the technician searches and maintains the node and updates the topology subgraph database. Extract the initial code Ci of the current fault node, obtain the historical processing times of each technician processing Ci, and select the top three technicians with the highest historical processing times as pre-selected personnel; Obtain the working time, historical processing records and assessment results of each pre-selected personnel; extract the processing time used by each pre-selected personnel for each Ci fault from the historical processing records, and calculate the average processing time of each pre-selected personnel after calculating the average value; The intervals of working hours and average processing time corresponding to the preset working hours and average processing time of each group are respectively corresponding to an experience evaluation value, and the interval of average processing time of each group is respectively corresponding to an efficiency evaluation value; the ranges of experience evaluation value and efficiency evaluation value are both set at 1-10, and the higher the working hours, the higher the experience evaluation value, and the shorter the average processing time, the higher the efficiency evaluation value; the working hours and average processing time of each pre-selected person are respectively converted into experience evaluation value and efficiency evaluation value; Based on regular internal technical assessments, the internal assessment scores of each pre-selected personnel closest to the time point of the current fault occurrence are obtained and recorded as the technical assessment value; the score range is set to 1-10, and the higher the score, the better the score; The technical assessment value, experience assessment value and efficiency assessment value of each pre-selected candidate are multiplied by the corresponding preset weight coefficients respectively, and then the sum is calculated to obtain the performance index of each pre-selected candidate; It should be noted that by presetting the chain similarity index reference range, it is possible to quickly and accurately screen out situations that are highly similar to the current fault from similar chain faults. When there is highly similar data that is higher than the index reference range, the data with a larger similarity index is taken as a hidden chain fault and a response plan is extracted, which provides clear and targeted guidance for fault handling, helps to quickly resolve the current fault and avoid further spread of the fault causing greater losses.
[0019] In the case that there is no chain similarity index higher than the reference range, similar chain faults with larger chain similarity index within the index reference range will be sent to technical personnel for judgment. This gives full play to the professional experience and judgment ability of technical personnel. They can comprehensively consider various factors according to the actual situation and select the most suitable one for the current situation from these similar faults as the hidden danger chain fault, which improves the accuracy and reliability of fault diagnosis.
[0020] The triggering of the search signaling selects the technician with the highest performance index to search the current fault node, which can ensure that the most capable and experienced technicians are assigned to the most needed fault processing, improve the utilization efficiency of technician resources, avoid the waste of manpower, and also help improve the efficiency and quality of fault processing; After searching the current faulty node, technicians will search and maintain if a chain failure occurs, and update the topology subgraph database. This enables the system to continuously learn and accumulate new fault cases and processing experience, optimize the data in the database, and improve the accuracy and efficiency of subsequent fault diagnosis. This allows the entire fault diagnosis system to be continuously improved and evolved over time, and better cope with various complex fault situations. Example 2 See also Figure 2As shown, based on an intelligent network fault detection method provided by embodiment 1 of the present application, embodiment 2 of the present application proposes an intelligent network fault detection system. Embodiment 2 is only a preferred method of embodiment 1, and the implementation of embodiment 2 will not affect the independent implementation of embodiment 1.
[0021] Specifically, the difference of an intelligent network fault detection system provided in Embodiment 2 of the present application is that it includes: Topology map construction module: each device in the network is regarded as a node, and the connection links between devices are regarded as edges to construct a network topology map; Fault diagnosis module: collect historical fault cases and build a labeled topological subgraph database. When a fault occurs at a certain node, identify the location of the faulty node in the network topology map, and determine the nodes directly connected to the faulty node as neighbor nodes. Based on the neighbor nodes corresponding to the faulty node, determine similar chain faults that are adapted to the current fault in the topological subgraph database, and analyze the chain similarity index between each similar chain fault and the current fault. ; Fault processing module: based on the chain similarity index between each similar chain fault and the current fault , execute the corresponding steps to determine the hidden danger chain failure caused by the current faulty node; The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only specific implementation methods. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. An intelligent network fault detection method, characterized in that: include: Fault trigger diagnosis: Collect historical fault cases and build a labeled topological subgraph database. When a fault occurs at a certain node, identify the location of the faulty node in the network topology map, and then determine the nodes directly connected to the faulty node as neighbor nodes. Based on the neighbor nodes corresponding to the faulty node, determine similar chain faults that are adapted to the current fault in the topological subgraph database, and analyze the chain similarity index between each similar chain fault and the current fault. , determine the hidden danger chain failure caused by the current faulty node; Diagnosis result processing: Based on the chain similarity index between each similar chain fault and the current fault , execute the corresponding steps to determine the hidden danger chain failure caused by the current faulty node.
2. The intelligent network fault detection method according to claim 1, characterized in that: Based on the neighboring nodes corresponding to the faulty node, similar chain failures that are adapted to the current failure are determined in the topology subgraph database, specifically: Identify the initial code of the current faulty node, set an initial code corresponding to each node, represented by Ci, where i represents the node number, i=1,2,...,k, k is the total number of nodes; count the number of neighboring nodes of the faulty node, represented by r, to form the label corresponding to the current faulty node, that is, Ci-r; The label corresponding to the current fault node is input into the topology subgraph database and matched with the labels of each historical fault case, and the successfully matched historical fault case is used as a similar chain fault adapted to the current fault.
3. The intelligent network fault detection method according to claim 2, characterized in that: Analyze the chain similarity index between each similar chain failure and the current failure , specifically: Identify the time point of the fault occurrence of the current fault node and match it with each set of preset time windows, each set of time windows corresponds to a time coefficient; determine the time coefficient corresponding to the current fault node, calculate the difference between the time coefficient of the current fault node and the time coefficient of each similar chain fault, and take the absolute value as the time zone similarity index p1 between each similar chain fault and the current fault node; Extract the CPU, memory and traffic data of the faulty node when it fails; and calculate the ratio with the preset CPU standard value, memory standard value and traffic standard value, that is, the CPU, memory and traffic data of the faulty node when it fails are used as the numerator, and the CPU standard value, memory standard value and traffic standard value are used as the denominator; Get the CPU evaluation value, memory evaluation value, and traffic evaluation value of the faulty node; The CPU evaluation value, memory evaluation value and traffic evaluation value of the faulty node are used as the length, width and height of the three-dimensional rectangular model respectively. A three-dimensional rectangular model of the current faulty node is constructed when a fault occurs. The cosine similarity is calculated between the three-dimensional rectangular models of similar chain faults to obtain the performance similarity index p2 between each similar chain fault and the current faulty node.
4. The intelligent network fault detection method according to claim 3, characterized in that: Analyze the chain similarity index between each similar chain failure and the current failure , also includes: Extract the room temperature, room humidity, and equipment power supply voltage when the faulty node fails, and mark them as u1, u2, and u3 respectively; Extract the room temperature, room humidity, and equipment power supply voltage corresponding to each similar chain failure, and record them as e1, e2, and e3; The room temperature, room humidity and equipment power supply voltage at the time of the faulty node failure are normalized and entered into the formula Perform weighted calculation to obtain the environmental similarity index p3 between each similar chain failure and the current fault node; a1, a2, and a3 are the influence weight factors of the computer room temperature, computer room humidity, and equipment power supply voltage, respectively; Extract the time zone similarity index p1, performance similarity index p2, and environment similarity index p3 between each similar chain failure and the current fault node, normalize them, and then enter them into the formula Perform weighted calculation to determine the chain similarity index of each similar chain failure ; Among them, n1, n2 and n3 are the influencing weight factors of the time zone similarity index p1, the performance similarity index p2 and the environment similarity index p3 respectively.
5. The intelligent network fault detection method according to claim 4, characterized in that: Execute the corresponding steps to determine the hidden danger chain failure caused by the current faulty node, specifically: Preset chain similarity index The corresponding index reference range, if a similar chain failure chain similarity index If the number of highly similar data is greater than one, the similarity index is taken. Larger highly similar data is used as the hidden danger chain failure caused by the current fault node and the corresponding fault solution is extracted as the response plan for the current fault node; If there is no chain failure similarity index If the fault is higher than the index reference range, similar chain faults within the index reference range are identified, and the chain similarity index within the index reference range is taken. A large similar chain failure is sent to any technician, who will make a judgment and select a group of similar chain failures as the hidden chain failure caused by the current fault node, or trigger a search signal to select the technician with the highest performance index to search the current fault node; If the chain similarity index of each similar chain failure If both are lower than the index reference range, the search signaling is directly triggered to select the technician with the highest performance index to search the current faulty node.
6. An intelligent network fault detection method according to claim 5, characterized in that: The performance index of the technician is specifically: Extract the initial code Ci of the current fault node, obtain the historical processing times of each technician processing Ci, and select the top three technicians with the highest historical processing times as pre-selected personnel; Obtain the working time, historical processing records and assessment results of each pre-selected personnel; extract the processing time used by each pre-selected personnel for each Ci fault from the historical processing records, and calculate the average processing time of each pre-selected personnel after calculating the average value; Preset the intervals of working hours and average processing time for each group, corresponding to the working hours and average processing time, respectively. Each working hours interval corresponds to an experience evaluation value, and each average processing time interval corresponds to an efficiency evaluation value; convert the working hours and average processing time of each pre-selected person into experience evaluation values and efficiency evaluation values respectively; Based on regular internal technical assessments, obtain the internal assessment results of each pre-selected personnel closest to the time point of the current fault occurrence, and record them as the technical assessment value; The technical assessment value, experience evaluation value and efficiency evaluation value of each pre-selected candidate are multiplied by the corresponding preset weight coefficient respectively, and then the sum is obtained to obtain the performance index of each pre-selected candidate.
7. The intelligent network fault detection method according to claim 6, characterized in that: After searching the current faulty node, the technician determines whether a cascading failure occurs. If a cascading failure occurs, the technician updates the topology subgraph database after searching and maintaining.
8. An intelligent network fault detection system, applied to an intelligent network fault detection method according to any one of claims 1 to 7, characterized in that: include: Topology map construction module: each device in the network is regarded as a node, and the connection links between devices are regarded as edges to construct a network topology map; Fault diagnosis module: collect historical fault cases and build a labeled topological subgraph database. When a fault occurs at a certain node, identify the location of the faulty node in the network topology map, and determine the nodes directly connected to the faulty node as neighbor nodes. Based on the neighbor nodes corresponding to the faulty node, determine similar chain faults that are adapted to the current fault in the topological subgraph database, and analyze the chain similarity index between each similar chain fault and the current fault. ; Fault processing module: based on the chain similarity index between each similar chain fault and the current fault , execute the corresponding steps to determine the hidden danger chain failure caused by the current faulty node.