Fault locating method and device, electronic equipment, program product and storage medium

By acquiring network quality parameters and utilizing time-series characteristics and multi-level topology analysis, network faults can be quickly and accurately located, solving the problem of high location complexity in existing technologies and improving network operation and maintenance efficiency and user satisfaction.

CN118802499BActive Publication Date: 2026-01-20CHINA MOBILE GROUP DESIGN INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410428913.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-10
Publication Date
2026-01-20
Estimated Expiration
2044-04-10

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly and accurately locate network faults, especially in large-scale networks, leading to high complexity in network operations and maintenance and impacting the quality of data services.

Method used

By acquiring network quality parameters of user equipment in the network, using time-series features and threshold analysis to identify faulty user equipment, constructing a multi-level faulty network topology, and locating network faults step by step.

Benefits of technology

It enables rapid and accurate location of network faults, reduces the complexity of fault location, and improves network operation and maintenance efficiency and user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118802499B_ABST
    Figure CN118802499B_ABST
Patent Text Reader

Abstract

The present disclosure provides a fault positioning method, device, electronic equipment, program product and storage medium. The positioning method comprises: acquiring network quality parameters of one or more user equipment in a network; determining one or more faulty user equipment in the one or more user equipment based on the network quality parameters; determining one or more groups of faulty user equipment groups based on time sequence characteristics of the network quality parameters of the one or more faulty user equipment; determining a multi-level fault network topology for one or more faulty user equipment in the faulty user equipment groups, and positioning the fault in the network level by level based on the multi-level fault network topology. The present disclosure identifies the faulty user equipment by acquiring and analyzing the network quality parameters of the user equipment, groups the faulty user equipment based on the time sequence characteristics, determines the multi-level fault network topology, and realizes the positioning of the fault in the network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of communication technology, and in particular, to a network fault locating method and device, electronic equipment and computer program product. BACKGROUND

[0002] With the development of data services, the scale of the network carrying data services is getting larger and larger, and the difficulty of network operation and maintenance is also increasing, and the proportion of data services damaged by network faults is also getting higher and higher.

[0003] In the prior art, it is difficult to discover and locate the fault for network service quality degradation. Generally, network fault identification relies on passive optical network management alarm, but this method is mainly for service interruption type faults, and cannot achieve fault detection and positioning for most devices in the network; some other methods need to use more than 100 kinds of data, and need to iterate repeatedly on the basis of modeling the target associated device of the user, and the positioning process is complex.

[0004] Therefore, how to quickly and accurately locate network faults and reduce the positioning complexity is a problem to be solved. SUMMARY

[0005] The present disclosure is proposed in view of the above problems. The present disclosure provides a network fault locating method, device, electronic equipment and computer program product.

[0006] According to one aspect of the present disclosure, a network fault locating method is provided, comprising: obtaining network quality parameters of one or more user devices in the network; determining one or more faulty user devices in the one or more user devices based on the network quality parameters; determining one or more groups of faulty user device groups based on the time sequence characteristics of the network quality parameters of the one or more faulty user devices; determining a multi-level fault network topology for one or more faulty user devices in the faulty user device groups, and locating the fault in the network step by step based on the multi-level fault network topology.

[0007] In addition, according to the network fault locating method of one aspect of the present disclosure, determining one or more faulty user devices in the one or more user devices based on the network quality parameters comprises: determining a threshold of the network quality parameters based on the network quality parameters; determining one or more faulty user devices in the one or more user devices based on the network quality parameters and the threshold.

[0008] In addition, the network fault locating method according to one aspect of the present disclosure determines the threshold of the network quality parameter based on the network quality parameter, including: drawing a box plot of the network quality parameter based on the network quality parameter; taking the quartiles and interquartile ranges of the network quality parameter in the box plot as the criterion for judging outliers to identify outliers in the network quality parameter; and determining the threshold of the network quality parameter based on the outliers.

[0009] In addition, the network fault locating method according to one aspect of the present disclosure determines one or more groups of faulty user equipment groups based on the time sequence characteristics of the network quality parameters of the one or more faulty user equipment, including: performing binary processing on the time sequence characteristics of the network quality parameters of the one or more faulty user equipment based on the threshold to obtain binary vectors corresponding to the one or more faulty user equipment; and performing cluster analysis on the binary vectors corresponding to the one or more faulty user equipment to determine the one or more groups of faulty user equipment groups.

[0010] In addition, the network fault locating method according to one aspect of the present disclosure determines a multi-level fault network topology for one or more faulty user equipment in the faulty user equipment group, and locates the fault in the network level by level based on the multi-level fault network topology, including:

[0011] determining a multi-level network topology corresponding to the one or more user equipment;

[0012] determining a multi-level fault network topology for one or more faulty user equipment in the faulty user equipment group based on the multi-level network topology;

[0013] locating the fault in the network level by level based on the multi-level network topology and the multi-level fault network topology.

[0014] In addition, the network fault locating method according to one aspect of the present disclosure locates the fault in the network level by level based on the multi-level network topology and the multi-level fault network topology, including:

[0015] comparing the multi-level network topology and the multi-level fault network topology level by level to determine the number of nodes at the same level in the multi-level network topology and the multi-level fault network topology;

[0016] locating the fault in the network level by level based on the number of nodes.

[0017] In addition, the network fault locating method according to one aspect of the present disclosure, the network quality parameter includes the average delay data of the first two handshakes of the transmission control protocol (TCP) flow.

[0018] According to another aspect of the present disclosure, there is provided a fault locating apparatus of a network, comprising: a network quality parameter obtaining module configured to obtain network quality parameters of one or more user equipments in the network; a faulty user equipment determining module configured to determine one or more faulty user equipments among the one or more user equipments based on the network quality parameters; a faulty user equipment group determining module configured to determine one or more faulty user equipment groups based on time sequence features of the network quality parameters of the one or more faulty user equipments; and a fault locating module configured to determine a multi-level fault network topology for one or more faulty user equipments in the faulty user equipment groups, and locate the fault in the network level by level based on the multi-level fault network topology.

[0019] According to a further aspect of the present disclosure, there is provided an electronic device, comprising: a memory configured to store computer readable instructions; and a processor configured to execute the computer readable instructions to cause the electronic device to perform the fault locating method of a network as described above.

[0020] According to a further aspect of the present disclosure, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the fault locating method of a network as described above.

[0021] According to a further aspect of the present disclosure, there is provided a computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the fault locating method of a network as described above.

[0022] As will be described in detail below, according to the fault locating method of a network, apparatus, electronic device and computer program product of the present disclosure, the present disclosure can quickly identify faulty user equipments by obtaining and analyzing network quality parameters of user equipments, group the faulty user equipments based on time sequence features, determine a multi-level fault network topology, and then locate the fault in the network, which can narrow down the fault range level by level, realize accurate positioning of the fault in the network, improve the accuracy and efficiency of fault locating, reduce the positioning complexity, thereby optimizing the network performance management, and improving the efficiency of network operation and user satisfaction.

[0023] It is to be understood that both the foregoing general description and the following detailed description are exemplary, and are intended to provide further explanation of the subject technology. BRIEF DESCRIPTION OF DRAWINGS

[0024] The above and other objects, features and advantages of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings in which:

[0025] Figure 1 is a schematic diagram illustrating an application scenario of a network fault locating method according to an embodiment of the present disclosure.

[0026] Figure 2 is a flowchart illustrating a network fault locating method according to an embodiment of the present disclosure.

[0027] Figure 3 is a flowchart further illustrating a network fault locating method according to an embodiment of the present disclosure.

[0028] Figure 4 is a structural schematic diagram illustrating a box plot according to an embodiment of the present disclosure.

[0029] Figure 5 is a schematic diagram illustrating IP address resource planning resolution according to an embodiment of the present disclosure.

[0030] Figure 6 is a schematic diagram illustrating a multi-level fault network topology according to an embodiment of the present disclosure.

[0031] Figure 7 is a functional block diagram illustrating a network fault locating apparatus according to an embodiment of the present disclosure.

[0032] Figure 8 is a hardware block diagram illustrating an electronic device according to an embodiment of the present disclosure.

[0033] Figure 9 is a schematic diagram illustrating a computer program product according to an embodiment of the present disclosure.

[0034] Figure 10 is a schematic diagram illustrating a computer readable storage medium according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0035] In order to make the objectives, technical solutions and advantages of the present disclosure more apparent, the following will describe example embodiments according to the present disclosure in detail with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the example embodiments described herein.

[0036] As Figure 1As shown, a schematic diagram of an application scenario of a network fault locating method according to an embodiment of the present disclosure.

[0037] The following is an application scenario description of a network fault locating method according to an embodiment of the present disclosure. The application scenario 100 mainly involves the network connection between a home user equipment terminal and a service ISP (Internet Service Provider), and realizes the transmission and exchange of network signals through a series of network devices and network components.

[0038] Specifically, the application scenario 100 can include a user equipment 101, an optical network unit (ONU) 102, a secondary optical splitter 103, a primary optical splitter 104, an optical line terminal (OLT) 105, a broadband remote access server (BRAS) 106, a core router 107, and an Internet service provider 108.

[0039] Among them, the user equipment 101 can refer to the terminal equipment of the user, such as a computer, a mobile phone, a smart home device, etc. The user initiates a network request and receives Internet services through the user equipment 101. The optical network unit (ONU) 102 is a user-side device, which is located at the network access end of the user equipment. The ONU converts optical signals into electrical signals so that the user equipment 101 can access the network. The secondary splitter 103 and the primary splitter 104 distribute optical signals from the OLT to different ONUs, ensuring that each ONU can receive signals. The primary splitter 104 is usually located closer to the OLT 105, while the secondary splitter 103 is closer to the user equipment 101. They are connected by optical fibers to ensure that optical signals can accurately reach each ONU. The optical line terminal (OLT) 105 is responsible for sending and receiving optical signals and communicating with the ONU on the user side. The OLT 105 is usually connected to a broadband remote access server (BRAS) or other core network equipment. The broadband remote access server (BRAS) 106 is located at the convergence layer of the network. The BRAS is responsible for user authentication, address allocation, traffic control and other functions to ensure that users can safely and efficiently access the Internet. The core router 107 is a key device in the network, used to forward data packets in high-speed backbone networks. It connects different network segments to ensure that data packets can be quickly and accurately transmitted in the network. The Internet service provider (ISP) 108 provides Internet access services, including network infrastructure, device maintenance, technical support, etc. In this scenario, the ISP 108 connects to the core router 107 through its core network to provide Internet access services for users. In this scenario, the user equipment 101 connects to the Internet through a series of network devices and components such as optical network units (ONUs), splitters, optical line terminals (OLTs), broadband remote access servers (BRASs), etc. to access the network of the Internet service provider (ISP), thereby accessing the Internet.

[0040] Specifically, the user equipment 101 is first connected to the ONU 102, which communicates with the OLT 105 through the secondary splitter 103 and the primary splitter 104. The OLT 105 forwards data to the BRAS 106 for user authentication and address allocation. Finally, the data packet is forwarded to the ISP 108 through the core router 107, and the user can access the Internet and perform various online activities.

[0041] Figure 1 Only the network, device or server components related to the fault locating method of the network are shown schematically, and the application scenario of the fault locating method of the network according to the embodiments of the present disclosure is not limited thereto. Those skilled in the art can understand that, Figure 1 The structure shown in the above embodiment does not constitute a limitation on the fault locating method of the network.

[0042] The application scenario of the fault locating method of the network according to the embodiments of the present disclosure will be described in detail with reference to the followingFigures 2 to 6 The detailed description is a method for locating a fault of a network according to an embodiment of the present disclosure.

[0043] As shown in Figure 2 The method for locating a fault of a network according to an embodiment of the present disclosure includes the following steps.

[0044] In step S201, the network quality parameters corresponding to one or more user devices in the network are obtained.

[0045] It can be understood that the network can refer to the infrastructure connecting user devices and other network devices, used to realize the transmission and exchange of data. This network can be a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), or the Internet. When user devices (such as computers, mobile phones, smart home devices, etc.) access the network, they will communicate with other devices in the network to obtain or send data. In this process, network quality parameters are important indicators for measuring network performance, which can include packet loss rate, delay, bandwidth utilization, jitter, signal strength, etc. According to these parameters, the health status of the network and the user's experience when using the network can be understood.

[0046] In an embodiment of the present disclosure, the network quality parameters corresponding to one or more user devices in the network can be obtained in various ways, such as through software or hardware probes on user devices, network management systems (NMS), network performance monitoring tools, etc. After obtaining these network quality parameters, they can be further analyzed to determine possible problems or bottlenecks in the network, so as to perform corresponding optimization or fault location. Further, it ensures that the network can provide high-quality services for users.

[0047] In step S202, one or more faulty user devices among the one or more user devices are determined based on the network quality parameters.

[0048] It can be understood that determining one or more faulty user devices according to network quality parameters is a key step in fault location. Among them, network quality parameters are important indicators reflecting the running status and performance of the network, and faulty user devices usually show abnormal or unexpected characteristics on network quality parameters.

[0049] Specifically, after collecting the network quality parameters of one or more user devices in the network, these network quality parameters are analyzed and compared. This usually involves comparing the actual parameter values with preset thresholds or standards, or using advanced techniques such as statistical methods, machine learning, etc. to identify abnormal patterns. If the network quality parameters of a certain user device are outside the normal range, or show a clear downward trend or fluctuation, then this device is likely to have a fault or problem. For example, if the packet loss rate of a certain user device is consistently high, or the delay time is abnormally long, then this device is likely to be a faulty user device. After determining the faulty user device, further fault localization and analysis can be carried out to find out the specific cause of the fault and take appropriate repair measures. This helps to quickly restore the normal operation of the network and improve user satisfaction and service quality.

[0050] In step S203, one or more groups of faulty user device groups are determined based on the time series characteristics of the network quality parameters of one or more faulty user devices.

[0051] It can be understood that the time series characteristics of the network quality parameters refer to the patterns and trends of the network quality parameters over time. Since the network is dynamic, network quality parameters such as delay, packet loss rate, bandwidth utilization, etc. also change over time. By analyzing the time series characteristics of these parameters, similarities or correlations between faulty user devices can be identified, and they can be grouped. For example, if the network quality parameters of a group of faulty user devices all show similar abnormal changes in the same time period, then they may be affected by the same network fault or problem. By grouping these devices, the fault can be analyzed and solved more targetedly. This grouping method helps to narrow down the scope of fault localization and improve the efficiency of fault handling. At the same time, by analyzing the time series characteristics of the network quality parameters of different groups of faulty user devices, the characteristics and rules of different fault types can be further understood, providing valuable reference for future fault prevention and handling.

[0052] In one embodiment of the present disclosure, determining the faulty user device group by using the time series characteristics of the network quality parameters of one or more faulty user devices is a scientific and effective method, which can more accurately locate and solve network faults.

[0053] In step S204, for one or more faulty user devices in the faulty user device group, a multi-level fault network topology is determined, and based on the multi-level fault network topology, the fault in the network is located step by step.

[0054] It can be understood that determining the multi-level fault network topology and locating the fault in the network step by step for one or more faulty user devices in the group of faulty user devices is a key link in the fault locating process. The multi-level fault network topology refers to the transmission path and flow direction of the faulty user device in the network constructed according to the connection relationship and layout between different levels of devices in the network. This topology not only shows the physical connection relationship between various network devices, but also reflects the transmission path and flow direction of data in the network. It includes the path from the faulty user device to the core network device, to other key nodes, and the connection mode and hierarchical relationship between these nodes. By constructing such a structure, the hierarchical relationship and mutual dependence between various components in the network can be clearly understood, so that the fault can be located more accurately.

[0055] In an embodiment of the present disclosure, by collecting and analyzing the connection information, routing information, configuration information, etc. between network devices, a network topology graph about the faulty user device can be constructed. This topology graph not only shows the physical connection relationship of various network devices, but also reflects the transmission path and flow direction of the data transmitted by the faulty user device in the network. According to the multi-level fault network topology, the fault in the network can be located step by step. For example: starting from the faulty user device, tracing back along the data transmission path, step by step analyzing the status and performance of each level device. By checking the log information, performance indicators, error codes, etc. of the device, it can be judged whether the fault occurs at this level device. If so, further analyze the fault cause and take appropriate repair measures; if not, continue to trace back to the upper level device, repeat the process of fault locating and analyzing.

[0056] In an embodiment of the present disclosure, according to the multi-level fault network topology, starting from the root node of the structure, tracing down along the data transmission path, step by step analyzing the node number of each node, when the network fails, the node number of the faulty node or its adjacent nodes may change abnormally. If the number of child nodes under a certain root node is normal, it may indicate that the root node has an abnormality, affecting the performance of its child nodes. For example, the root node may be unable to normally process or forward data due to configuration errors, performance degradation or attacks, etc. Although the number of its child nodes appears to be normal, the normal transmission of data flow has been affected. If the child nodes are different, the next level needs to be investigated, and then appropriate repair measures are taken.

[0057] The above-mentioned step-by-step locating method can accurately determine the specific location and cause of the fault, avoiding blind investigation and wasting time. At the same time, by constructing the multi-level fault network topology, the running mechanism and performance bottleneck of the network can also be better understood, providing strong support for the optimization and upgrading of the network. It is of great significance to improve the efficiency and accuracy of fault handling.

[0058] As Figure 3 illustrated, the method for locating faults of a network according to an embodiment of the present disclosure further comprises the following steps. Wherein, step S301 is the same as step S201, which will not be repeated here.

[0059] In step S302, based on the network quality parameter, a threshold of the network quality parameter is determined.

[0060] It can be understood that the network quality parameter, such as the average latency data of the first two handshakes of the Transmission Control Protocol (TCP) flow, is an important indicator for evaluating network performance. The threshold is set to distinguish between normal network state and state that may have faults or performance degradation.

[0061] In an embodiment of the present disclosure, according to the obtained network quality parameter, the threshold of the network quality parameter is determined, including collecting network quality parameter data in the past period of time, calculating its mean, standard deviation and other statistical quantities, and then selecting a suitable threshold according to the normal distribution or other applicable distribution characteristics. For example, the mean plus twice the standard deviation can be selected as the threshold, so that the situation exceeding the threshold can be considered as abnormal or not as expected. It also includes setting the threshold according to specific business needs. For example, if a certain business has strict requirements on latency, the threshold of latency can be set relatively low to ensure the smooth progress of the business. It also includes real-time monitoring of network quality parameters according to actual conditions, and adjusting the threshold according to feedback information, which can ensure that the threshold setting is always matched with the current network environment.

[0062] In an embodiment of the present disclosure, based on the network quality parameter, the threshold of the network quality parameter is determined, which further includes drawing a box plot of the network quality parameter based on the network quality parameter; taking the quartiles and interquartile ranges of the network quality parameter in the box plot as the standard for judging abnormal values to identify abnormal values in the network quality parameter; and determining the threshold of the network quality parameter based on the abnormal values.

[0063] It can be understood that the statistical analysis method of box plot is used to identify abnormal values in the network quality parameter. The box plot is a standardized way to show data distribution, which provides an overview of the data by depicting the minimum value, the first quartile (Q1), the median (Q2), the third quartile (Q3) and the maximum value of the data, as Figure 4The key is that it uses the so-called "interquartile range" (IQR = Q3-Q1) to determine the limit of outliers. In the embodiments of the present disclosure, data points higher than the threshold Threshold = Q3+3xIQR are classified as outliers, according to which the threshold of the network quality parameter can be determined as Q3+3xIQR. Among them, the coefficient 3 of IQR in the formula is obtained based on the normal distribution assumption and the 3σ empirical rule (under the normal distribution assumption, about 99.7% of the data should be within 3 standard deviations from the mean).

[0064] Specifically, the threshold of the network quality parameter can be determined according to the average latency data obtained from the first two handshakes of the transmission control protocol (TCP) flow. For example, taking a home broadband network as an example, the "TCP connection two handshake average latency" data of all user equipment in a province every hour for five days is obtained as the network quality parameter, and a box plot is applied to analyze the network quality parameter, and the Threshold value is calculated and taken as the threshold of the network quality parameter, according to which the quality of the home broadband network can be judged. Each province can calculate its own threshold according to this method.

[0065] In step S303, one or more faulty user equipment in the one or more user equipment is determined based on the network quality parameter and the threshold.

[0066] It can be understood that after the threshold is determined, it can be compared with the real-time network quality parameter to determine whether the one or more user equipment is in a normal state. If the network quality parameter of the certain user equipment or certain user equipment exceeds the set threshold, it can be reasonably inferred that these devices have faults, so as to determine one or more faulty user equipment in the one or more user equipment. This comparison and determination process helps to quickly identify the faulty user equipment in the network. Once the faulty user equipment is determined, further fault locating and repairing work can be carried out to ensure the stability and performance of the network.

[0067] In step S304, the time sequence characteristics of the network quality parameter of the one or more faulty user equipment are binarized based on the threshold, and a binary vector corresponding to the one or more faulty user equipment is obtained.

[0068] It can be understood that the timing characteristics of the network quality parameters refer to the values of the network quality parameters changing over time, which form a time series. Specifically, the network quality parameters can include bandwidth, delay, packet loss rate, jitter, etc., and the timing characteristics of these parameters can exhibit periodic changes, trend changes, random fluctuations, etc. For example, bandwidth can exhibit different change patterns at different times of the day, and delay can exhibit a significant increase when the network is congested. By analyzing these timing characteristics, the network performance can be more accurately evaluated, and potential faults in the network can be detected in a timely manner. Binary processing is a data processing method that converts continuous or discrete numerical data into data with only two states (usually 0 and 1). Binary processing of the timing characteristics of the network quality parameters of one or more faulty user equipment refers to comparing the timing characteristics of the network quality parameters of one or more faulty user equipment with a set threshold value. If the value of the network quality parameter changing over time exceeds the threshold value, the corresponding binary bit is set to 1 (indicating an abnormal or faulty state), otherwise it is set to 0 (indicating a normal state). In this way, the network quality parameter changing over time for each faulty user equipment corresponds to a binary bit. Combining the binary bits of all network quality parameter timing characteristics of the faulty user equipment can form a binary string, and the binary string corresponding to the faulty user equipment is converted into a binary vector. This binary vector succinctly represents the network quality state of the faulty user equipment, facilitating subsequent analysis and processing.

[0069] In an embodiment of the present disclosure, for example, the gateway soft probe reports network quality parameters by default every 10 minutes, and if the selected time interval is 2 hours, the network quality parameters of the user equipment correspond to a 12-bit binary string in the two hours, such as 000011110000 and 000011011000. Further converting the binary string into a binary vector can obtain a binary vector corresponding to one or more faulty user equipment.

[0070] In an embodiment of the present disclosure, the timing characteristic data of the network quality parameters of each faulty user equipment can also be converted into a binary vector of length 24 for each vector, and each bit in the vector represents the network state (0 or 1) of the faulty user equipment at a specific time point.

[0071] The above binary processing is particularly useful in network fault detection and diagnosis, as it can highlight data points that exceed the normal range (defined by the threshold value), making it easier to identify potential fault patterns or abnormal behavior.

[0072] In step S305, the binary vector corresponding to one or more faulty user equipment is subjected to cluster analysis to determine one or more groups of faulty user equipment.

[0073] It can be appreciated that clustering analysis is a process of grouping a collection of data objects into multiple classes of similar objects, with the purpose of collecting data on a similar basis for classification.

[0074] In one embodiment of the present disclosure, binary vectors corresponding to all faulty user equipment are collected. These binary vectors have been binarized based on the time-series features of network quality parameters through previous steps. According to the characteristics of the data and the requirements of the problem, a suitable clustering method is selected. The selected clustering method is applied to the binary vectors, which are grouped according to the similarity between them (e.g., based on distance metrics), so that the binary vectors within the same group have high similarity, while the vectors between different groups have low similarity. According to the clustering results, the faulty user equipment is divided into different groups. These groups represent a collection of user equipment with similar failure modes or network quality problems. The determined groups of faulty user equipment are taken as the output of the clustering analysis. These groups will serve as the basis for subsequent fault localization steps, helping to further narrow down the scope of the fault and improve the accuracy of localization.

[0075] In one embodiment of the present disclosure, the cosine similarity formula is used to calculate the similarity between the binary vectors of each two faulty user equipment.

[0076] The formula for calculating the cosine similarity is:

[0077]

[0078] where A i and B i are the values of the two vectors at the i-th position.

[0079] After calculating the cosine similarity between all the binary vectors of the faulty user equipment, the similarity scores are obtained, which are used as input for the clustering method, in order to group the faulty user equipment with highly similar patterns into a group, forming groups of faulty user equipment. This method is particularly suitable for the application scenario of the present disclosure, because the present disclosure is more concerned about the similarity of the trends and patterns of the network quality parameter data of the faulty user equipment, rather than their size or numerical difference.

[0080] Before clustering, the optimal number of clusters K can be determined. Preferably, the within-cluster sum of squares (WCSS) under different K values is calculated and plotted. As K increases, WCSS will decrease, but the decrease will become insignificant at a certain point. This point is the optimal K value. After determining the optimal K value, the faulty user equipment is classified. Each faulty user equipment is assigned to the cluster with the highest cosine similarity, forming one or more faulty user equipment groups. From this point, the user equipment groups with quality of service degradation in similar time periods are identified.

[0081] In an embodiment of the present disclosure, the binary vectors corresponding to one or more faulty user equipment are subjected to cluster analysis, and the binary vectors corresponding to two faulty user equipment can also be compared bit by bit. If more than 80% of the bits of the binary vectors are aligned, it can be considered that the two faulty user equipment have high similarity, and thus the two faulty user equipment are assigned to a faulty user equipment group, achieving the determination of one or more faulty user equipment groups. The above threshold (80%) is set according to the specific application scenario and needs, and may need to be adjusted according to the actual situation.

[0082] In step S306, a multi-level network topology structure corresponding to one or more user equipment is determined.

[0083] It can be understood that the topology information of the network in which the user equipment is located needs to be collected. This includes the connection relationship of the equipment, the configuration information of the network equipment, etc. According to the collected topology information, different levels in the network are identified. Based on the identified network levels and topology information, a multi-level network topology structure is constructed. The multi-level network topology structure can be a detailed network topology graph, containing a structured data table of network levels and connection information.

[0084] In an embodiment of the present disclosure, the multi-level network topology structure can also be constructed using management data. Considering that the link and device ownership information of the user equipment cannot be completely accurate in management, preferably, the nine-level network topology graph of a province is established using the authentication, authorization and accounting (AAA) call data and management data of the province. The nine-level network topology graph is the multi-level network topology structure of the present disclosure. The AAA call data is carried according to the actual network information and is more reliable.

[0085] Specifically, the ownership information of the user gateway is obtained based on the AAA call data, including the belonging city, district, Bras, OLT, passive optical network (PON) port. The Internet Protocol (IP) address of the network between the Bras to which the user equipment belongs can be directly obtained from the AAA data, i.e., Bras_IP. The acquisition of other information needs the following two steps:

[0086] Parsing the Internet Protocol version 6 address IPV6_ADDRESS obtains the city and county information of the user equipment.

[0087] According to the China Mobile IP address resource planning (1.9.1 version) about China Mobile IP address resource planning, the first 40 bits (binary) of IPV6 can be located to the county, and the specific parsing specification refers to Figure 5 .

[0088] The first 12 bits of IPV6 (hexadecimal) in the bill data can be used to locate the county information (Region). According to the corresponding relationship table between the county and the city in the specification, the city information (City) can be obtained.

[0089] The logical port number logicalportno is parsed to obtain the OLT and PON port information. Specifically, the example format of the Logicalportno field is: trunk 3 / 0 / 7:2125.18172.31.251.84 / 0 / 0 / 7 / 0 / 16 / CMDCB238F0FA GP. This field has a fixed combination specification, and the OLT_IP and PON information can be obtained according to the specification. In this example, the OLT_IP is "172.31.251.84", and the PON port information is "0 / 0 / 7 / 0 / 16". Since different cities may have repeated OLT_IP, to uniquely determine an OLT, the city information (City) needs to be combined, that is, "City+OLT_IP" as the unique key of the OLT, and the PON port needs to include the OLT information to be uniquely determined, so the unique key of the PON port is: "City+OLT_IP+PON".

[0090] Get the resource management data within the province, from which the corresponding secondary splitter port of the user gateway can be obtained, and the port connection information between the secondary splitter and the primary splitter can be traced upwards, and the connection information between the primary splitter and the OLT PON port can be obtained.

[0091] Combining the other user attribution information (city, county, Bras, OLT, PON port) obtained before, taking the PON port as the association factor, if the PON port information obtained from the AAA data of the same user is consistent with the PON port information obtained from the management data, the device attribution information of the user can be extended down by two layers, that is, the information of the primary splitter and the secondary splitter is added.

[0092] Based on all the home information of the user gateway, the user gateway is taken as the connection basis, a huge multi-level network topology structure (named as Tree1) can be established for the province. The multi-level network topology structure is a tree structure, the root node of the tree is the province, and the following layers are respectively the city, the county, the Bras, the OLT, the PON port, the first-level optical splitter, the second-level optical splitter, and the user gateway. It can be seen that the gateway is the leaf node of the tree, and all the gateways of the province are reflected on the leaf node.

[0093] In step S307, based on the multi-level network topology structure, the multi-level fault network topology structure is determined for one or more fault user devices in the fault user device group.

[0094] It can be understood that the multi-level fault network topology structure shows the connection relationship and path of one or more fault user devices in the fault user device group between different levels in the network. For one or more fault user devices in the fault user device group, referring to the multi-level network topology structure Tree1 obtained in the last step, the multi-level fault network topology structure corresponding to the one or more fault user devices is constructed. Specifically, according to Tree1, the specific positions of the fault user devices in the network are identified, i.e., which level of the network they belong to (such as the access layer, the aggregation layer, or the core layer). Starting from the fault user device, the connection path is traced upwards until the network device of the higher level is reached. In this process, the connection relationship between each level needs to be recorded, including the link, the device type, the port, and other information of the connection.

[0095] Based on the connection path and connection relationship traced above, a multi-level fault network topology structure reflecting the multi-level connection relationship of the fault user devices in the network is constructed. This multi-level fault network topology structure should be able to clearly show how the fault user devices are connected to the network through network devices of different levels.

[0096] In an embodiment of the present disclosure, the multi-level fault network topology structure construction process can be: starting from the leaf node (i.e., each user device) and finding the parent node upwards layer by layer, tracing the source upwards layer by layer, and all related network devices or regional attributions will be constructed on the tree layer by layer until the root node (the province) of the uppermost layer. For the user without the splitter information, the PON port can be found as the parent node, and the tree structure is the multi-level fault network topology structure. When the multi-level fault network topology structure (named as Tree2) is constructed, the number of child nodes of each non-leaf node on the tree is calculated and saved as an attribute of the node. Similar calculation and saving of the number of child nodes also need to be done for Tree1.

[0097] In step S308, based on the multi-level network topology structure and the multi-level fault network topology structure, the fault in the network is located level by level.

[0098] It can be understood that, according to the multi-level network topology structure and the multi-level fault network topology structure, the network level and the connection relationship in which the fault user equipment is located are analyzed. Starting from the access layer, the possible fault points and the influence range in each level can be analyzed by tracing back to the higher level network components such as the convergence layer, the core layer and the like. Or, starting from the high network level, the propagation path and the influence range of the fault of the multi-level fault network topology structure corresponding to the fault equipment group are further analyzed with reference to the dependency relationship and the communication path between the components in the multi-level network topology structure. This helps to determine which components may be affected by the fault and which components may be the fault source. Once the fault position is determined, corresponding measures can be taken for repair and recovery of the normal operation of the network.

[0099] In an embodiment of the present disclosure, as shown in Figure 6 Fig. 2 shows an example of the constructed multi-level fault network topology structure, which omits the first-level splitting and the second-level splitting information, and directly hangs the user equipment under the PON port. In this example, the OLTs connected by the dashed lines on the left and right sides do not actually exist in Tree2, but exist in Tree1. In this example, it can be simply analyzed that the quality degradation problem is not caused by the Bras and the devices above, and the situation of the OLT with the OLT_IP of 10.101.251.29 and the subordinate devices thereof needs to be specifically analyzed, and the specific analysis process will be described below.

[0100] It can be understood that, based on the network quality parameters, one or more fault user equipment in one or more user equipment is determined, including comparing the multi-level network topology structure with the multi-level fault network topology structure level by level, determining the number of nodes at the same level of the multi-level network topology structure and the multi-level fault network topology structure, and locating the fault in the network based on the number of nodes.

[0101] In determining the multi-level network topology structure and the multi-level fault network topology structure (i.e. the network topology in which the fault user equipment is located). These two structures respectively represent the connection relationship in the normal state and the fault state of the entire network. These two structures can be compared level by level. The comparison contents include the number of nodes at the same level, the connection relationship, the device state and the like. It should be noted that "same level" here refers to the same level in the network structure. In the comparison process, special attention is paid to the difference in the number of nodes. If the number of nodes of a certain level in the multi-level fault network topology structure is less than that of the corresponding level in Tree1, it may mean that there is a fault or a missing device in that level. On the contrary, if the number of nodes is more than that of Tree1, it may mean that there is an abnormal connection or a newly added fault device. If the number of nodes of a certain level in the multi-level fault network topology structure is the same as that of Tree1, it is possible that the upper level of that level has a fault, and the device and link state of the upper level need to be further analyzed.

[0102] In one embodiment of the present disclosure, after the multi-level network topology and the multi-level fault network topology are constructed, the network fault can be located according to the process in Table 1 (the home broadband network quality degradation problem positioning process table) below. For the constructed Tree2, starting from serial number 1, that is, starting from the root node province, if the number of child nodes of the current provincial node in Tree2 is the same as the number of child nodes of the current provincial node in Tree1, it indicates that all cities are covered in Tree2, indicating that the fault affects all provincial users, and the reasons can be two: 1. abnormal provincial trunk or inter-provincial link, 2. abnormal content source. At this time, the problem positioning process ends. On the contrary, if only part of the cities are affected, it is necessary to separately check the abnormalities of each city, that is, to follow the branch of serial number 2, and to jump to the next level for further checking. For a single city, it is necessary to compare whether the number of child nodes of the city in Tree2 and Tree1 is equal, and if so, the problem positioning result of serial number 3 is obtained, otherwise it is necessary to jump to the next level for further checking. By analogy, as long as the problem positioning result has not been obtained at the current node, it means that further analysis needs to be done at the next level until the problem positioning process ends.

[0103] The judgment process embodied in the table is completely matched with the constructed user topology graph. According to the judgment process, the possible problem reason can be found at a certain level, and the process is terminated.

[0104] Table 1: Home broadband network quality degradation problem positioning process table

[0105]

[0106]

[0107]

[0108] The present disclosure exhibits the network quality parameters for evaluating the online quality of user equipment by using binary coding, and forms a binary vector by binary coding the time sequence characteristics of the network quality parameters. By using the K-means clustering method based on cosine similarity on the binary vector, a group of fault user equipment with degraded service quality in a similar time period is identified. According to the group of fault user equipment, a multi-level fault network topology is constructed, so as to locate the fault in the network step by step.

[0109] The above describes the network fault positioning method according to the embodiments of the present disclosure. In the following, the network fault positioning apparatus for implementing the above network fault positioning method will be further described. Figure 7 is a functional block diagram illustrating the network fault positioning apparatus according to the embodiments of the present disclosure.

[0110] As Figure 7 shown in FIG. 7, the fault locating apparatus 700 of the network according to the embodiments of the present disclosure comprises: a network quality parameter obtaining module 701, a faulty user equipment determining module 702, a faulty user equipment group determining module 703, and a fault locating module 704.

[0111] Specifically, the network quality parameter obtaining module 701 is configured to obtain network quality parameters of one or more user equipments in the network.

[0112] Specifically, the faulty user equipment determining module 702 is configured to determine one or more faulty user equipments from the one or more user equipments based on the network quality parameters.

[0113] Specifically, the faulty user equipment group determining module 703 is configured to determine one or more faulty user equipment groups based on the time sequence characteristics of the network quality parameters of the one or more faulty user equipments.

[0114] Specifically, the fault locating module 704 is configured to determine a multi-level fault network topology for one or more faulty user equipments in the faulty user equipment groups, and locate the fault in the network level by level based on the multi-level fault network topology.

[0115] Those skilled in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present disclosure.

[0116] The basic principles of the present disclosure are described above in combination with specific embodiments, but it should be noted that the advantages, advantages, effects and the like mentioned in the present disclosure are only examples and are not limiting, and these advantages, advantages, effects and the like cannot be considered as the various embodiments of the present disclosure must have. In addition, the above specific details are only for the purpose of example and for the purpose of understanding, and the above details do not limit the present disclosure to the above specific details.

[0117] The block diagrams of the devices, apparatuses, equipment, systems involved in the present disclosure are only illustrative examples and are not intended to require or imply that the connection, arrangement, configuration must be as shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, systems can be connected, arranged, configured in any manner. Words such as "include", "contain", "have", and the like are open-ended words, mean "including but not limited to", and can be used interchangeably. The words "or" and "and" used herein mean the word "and / or", and can be used interchangeably, unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as but not limited to", and can be used interchangeably.

[0118] In addition, as used herein, "or" used in the listing of items, starting with "at least one of, indicates a disjunctive list, such that, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "such as" is not meant to limit the examples that follow to a specific list of examples.

[0119] It is also noted that in the systems and methods of the present disclosure, the various components and steps can be rearranged and / or combined. These rearrangements and / or combinations are also contemplated in the present disclosure.

[0120] Various changes, modifications and alterations in the techniques described herein can be made without departing from the teachings of the attached claims. Moreover, the scope of the claims of the present disclosure is not limited to specific aspects described herein. The current existing or later developed processes, machines, manufactures, compositions of matter, means, methods, or steps that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein are intended to be within the scope of the claims of the present disclosure. Accordingly, the appended claims include within their scope such processes, machines, manufactures, compositions of matter, means, methods, or steps.

[0121] Figure 8 is a hardware block diagram illustrating an electronic device 800 according to an embodiment of the present disclosure. The electronic device according to an embodiment of the present disclosure includes at least a processor; and a memory for storing computer readable instructions. When the computer readable instructions are loaded and run by the processor, the processor performs the verification method as above.

[0122] Figure 8The electronic device 800 shown specifically includes a central processing unit (CPU) 801, a graphics processing unit (GPU) 802, and a main memory 803. These units are connected to each other through a bus 804. The central processing unit (CPU) 801 and / or the graphics processing unit (GPU) 802 can be used as the above-described processor, and the main memory 803 can be used as the above-described memory storing computer readable instructions. In addition, the electronic device 800 can further include a communication unit 805, a storage unit 806, an output unit 807, an input unit 808, and an external device 809, which are also connected to the bus 804.

[0123] Figure 9 is a schematic diagram illustrating a computer program product according to an embodiment of the present disclosure. As shown in Figure 9 the computer program product 900 according to the embodiment of the present disclosure has a computer program 901 stored thereon. The computer program is executed by a processor to implement the verification method as described above. The computer program product includes, but is not limited to, system software, application software, and games, etc. System software is the basic software of a computer, responsible for managing the hardware and application programs of the computer, including operating systems, device drivers, etc. Application software is software designed to meet specific needs, such as office software, image processing software, etc. Games are software for entertainment, providing various gaming experiences. In addition, the computer program product can also include embedded software, firmware, etc., for controlling and operating various hardware devices.

[0124] Figure 10 is a schematic diagram illustrating a computer readable storage medium according to an embodiment of the present disclosure. As shown in Figure 10 the computer readable storage medium 1000 according to the embodiment of the present disclosure has computer readable instructions 1001 stored thereon. When the computer readable instructions 1001 are run by a processor, the fault locating method of the network according to the embodiment of the present disclosure described with reference to the above figures is executed. The computer readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may, for example, include random access memory (RAM) and / or cache memory, etc. Non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, optical disk, magnetic disk, etc.

[0125] As described in detail above, according to the network fault positioning method, apparatus, electronic device and computer program product of embodiments of the present disclosure, the present disclosure can quickly identify the faulty user equipment by acquiring and analyzing the network quality parameters of the user equipment, group the faulty user equipment based on the timing characteristics, determine the multi-level fault network topology, and further realize the positioning of the network fault, which can gradually narrow down the fault range and realize accurate positioning of the fault in the network, thereby improving the accuracy and efficiency of fault positioning, reducing the positioning complexity, optimizing the network performance management, and improving the efficiency of network operation and user satisfaction.

[0126] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0127] The above description has been presented for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present disclosure to forms disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions and sub-combinations thereof.

Claims

1. A method for fault location in a network, characterized in that, include: Obtain network quality parameters for one or more user devices in the network; Based on the network quality parameters, identifying one or more faulty user equipments among the one or more user equipments includes: Based on the network quality parameters, determine the threshold of the network quality parameters; Based on the network quality parameters and the threshold, one or more faulty user equipments among the one or more user equipments are identified; Based on the temporal characteristics of the network quality parameters of the one or more faulty user equipments, one or more groups of faulty user equipments are determined, including: Based on the threshold, the temporal characteristics of the network quality parameters of the one or more faulty user equipments are binarized to obtain the binary vectors corresponding to the one or more faulty user equipments. Cluster analysis is performed on the binary vectors corresponding to the one or more faulty user equipment to determine one or more groups of faulty user equipment; For one or more faulty user equipment in the faulty user equipment group, a multi-level faulty network topology is determined, and based on the multi-level faulty network topology, the faults in the network are located level by level.

2. The network fault location method according to claim 1, characterized in that, Determining the threshold of the network quality parameters based on the network quality parameters includes: Based on the network quality parameters, draw a box plot of the network quality parameters; The interquartiles and interquartile ranges of the network quality parameters in the box plot are used as criteria for judging outliers to identify outliers in the network quality parameters. Based on the outliers, the threshold values ​​for the network quality parameters are determined.

3. The network fault location method according to claim 1, characterized in that, The step of determining a multi-level fault network topology for one or more faulty user equipment in the faulty user equipment group, and locating faults in the network level by level based on the multi-level fault network topology, includes: Determine the multi-level network topology corresponding to the one or more user equipments; Based on the multi-level network topology, a multi-level fault network topology is determined for one or more faulty user equipment in the faulty user equipment group. Based on the multi-level network topology and the multi-level fault network topology, faults in the network are located level by level.

4. The network fault location method according to claim 3, characterized in that, The step-by-step location of faults in the network based on the multi-level network topology and the multi-level fault network topology includes: The multi-level network topology is compared with the multi-level faulty network topology level by level to determine the number of nodes at the same level as the multi-level network topology and the multi-level faulty network topology. Based on the number of nodes, faults in the network are located step by step.

5. The network fault location method according to any one of claims 1-4, characterized in that, The network quality parameters include the average delay data of the first two handshakes of the Transmission Control Protocol (TCP) stream.

6. A network fault location device, characterized in that, include: The network quality parameter acquisition module is configured to acquire network quality parameters corresponding to one or more user equipments in the network. The faulty user equipment determination module is configured to determine one or more faulty user equipments among the one or more user equipments based on the network quality parameters, including: Based on the network quality parameters, determine the threshold of the network quality parameters; Based on the network quality parameters and the threshold, one or more faulty user equipments among the one or more user equipments are identified; The faulty user equipment group determination module is configured to determine one or more groups of faulty user equipment based on the temporal characteristics of the network quality parameters of the one or more faulty user equipments, including: Based on the threshold, the temporal characteristics of the network quality parameters of the one or more faulty user equipments are binarized to obtain the binary vectors corresponding to the one or more faulty user equipments. Cluster analysis is performed on the binary vectors corresponding to the one or more faulty user equipment to determine one or more groups of faulty user equipment; The fault location module is configured to determine a multi-level fault network topology for one or more faulty user equipments in the faulty user equipment group, and to locate the faults in the network level by level based on the multi-level fault network topology.

7. An electronic device, characterized in that, include: Memory, used to store computer-readable instructions; as well as A processor for executing the computer-readable instructions, causing the electronic device to perform a network fault location method as described in any one of claims 1 to 5.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the fault location method of the network according to any one of claims 1 to 5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the fault location method for the network as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Electric power communication network fault analysis and positioning method based on convolutional neural network

    CN110943857A

  • Power communication network fault detection method based on transfer learning

    CN110995475A