Method for fault location, electronic device and storage medium

Through the fault location model of the graph neural network structure, the object relationship structure diagram of the network to be tested is used to solve the problem of inaccurate fault location caused by independent network element feature information, and more efficient and accurate fault identification is achieved.

CN114221857BActive Publication Date: 2025-07-18ZTE CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010921946.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-04
Publication Date
2025-07-18
Estimated Expiration
2040-09-04

AI Technical Summary

Technical Problem

In the existing network fault location method based on machine learning, network element feature information is independent of each other, resulting in inaccurate fault location.

Method used

The fault location model of the graph neural network structure is adopted, and the object relationship structure diagram is generated by obtaining the object, feature data and association relationship in the network to be tested, and the surrounding nodes are used to encode the central node to perform fault location.

Benefits of technology

Improve the accuracy and efficiency of fault location, and enable more accurate identification of fault objects in the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114221857B_ABST
    Figure CN114221857B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the field of computers, and in particular to a method, an electronic device, and a storage medium for fault location. The method for fault location provided by the embodiments of the present application includes: obtaining at least two objects to be measured, characteristic data of the objects to be measured, and the association relationships between the objects to be measured in the network to be measured; generating an object relationship structure diagram according to the at least two objects to be measured, the characteristic data of the objects to be measured, and the association relationships; and locating a faulty object in the network to be measured according to the object relationship structure diagram and a preset fault location model, where the fault location model is a graph neural network structure. It can accurately and efficiently locate faults in the network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computers, and in particular to a fault location method, electronic device, and storage medium. Background Art

[0002] As the scale of networks becomes larger and larger, and the structure of networks becomes more and more complex, when a network failure occurs, it is very important to quickly locate the network failure. In related technologies, a method based on machine learning is usually used to locate the failure.

[0003] However, in the network fault location method based on machine learning, the characteristic information of network elements is mainly extracted for fault location. The characteristic information between each network element is independent of each other, resulting in inaccurate fault location. Summary of the invention

[0004] The main purpose of the embodiments of the present application is to provide a fault location method, electronic device and storage medium, which can accurately and efficiently locate faults in a network.

[0005] To achieve the above-mentioned objectives, an embodiment of the present application provides a method for fault location, including: obtaining at least two objects to be tested in the network to be tested, characteristic data of the objects to be tested, and the association relationship between the objects to be tested; generating an object relationship structure diagram based on at least two of the objects to be tested, the characteristic data of the objects to be tested, and the association relationships; locating the fault object in the network to be tested according to the object relationship structure diagram and a preset fault location model, wherein the fault location model is a graph neural network structure.

[0006] To achieve the above objective, an embodiment of the present application further provides an electronic device, comprising: at least one processor; and

[0007] A memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the above-mentioned fault location method.

[0008] To achieve the above objectives, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which implements the above fault locating method when executed by a processor.

[0009] The method for fault location proposed in this application obtains the objects to be measured in the network to be measured, the characteristic data of the objects to be measured, and the association relationships between the objects to be measured. According to the objects to be measured, the characteristic data of the objects to be measured, and the association relationships between the objects to be measured, an object relationship structure diagram is generated, thereby converting the network to be measured into a graph form representing the relationships between the objects to be measured. According to the preset fault location model, the fault objects in the network to be measured are located; the fault location model is obtained by training based on the graph neural network structure. The graph neural network encodes the central node through the surrounding nodes and fully utilizes the relationship between the surrounding nodes and the central node during training, making the located fault nodes more accurate. By using this graph neural network model and the object relationship structure diagram for fault location, the located fault objects are more accurate. Description of the Drawings

[0010] Figure 1 is a flowchart of the method for fault location in the first embodiment of the present invention;

[0011] Figure 2 is a flowchart of the method for fault location in the second embodiment of the present invention;

[0012] Figure 3 is a flowchart of the method for fault location in the third embodiment of the present invention;

[0013] Figure 4 is a schematic diagram of an edge in the method for fault location in the third embodiment of the present invention;

[0014] Figure 5 is a schematic diagram of an edge in the method for fault location in the third embodiment of the present invention;

[0015] Figure 6 is a schematic diagram of the object relationship structure diagram in the method for fault location in the third embodiment of the present invention;

[0016] Figure 7 is a schematic diagram of the aggregated sample nodes in the method for fault location in the third embodiment of the present invention;

[0017] Figure 8 is a block diagram of the structure of the electronic device in the fourth embodiment of the present invention. Detailed Embodiments

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will elaborate on each embodiment of this application in conjunction with the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of this application, many technical details are provided to help readers better understand this application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in this application can still be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation to the specific implementation of this application. Each embodiment can be combined and cross-referenced with each other on the premise of not being contradictory.

[0019] The first embodiment of the present invention relates to a method for fault location, and its process is as Figure 1 shown:

[0020] Step 101: Obtain at least two objects to be measured, the characteristic data of the objects to be measured, and the association relationships between the objects to be measured in the network to be measured.

[0021] Step 102: Generate an object relationship structure diagram according to at least two objects to be measured, the characteristic data of the objects to be measured, and the association relationships.

[0022] Step 103: Locate the faulty object in the network to be measured according to the object relationship structure diagram and a preset fault location model, and the fault location model is a graph neural network structure.

[0023] The fault location method proposed in this application obtains the objects to be measured in the network to be measured and the association relationships between the objects to be measured, generates an object relationship structure diagram according to the objects to be measured and the association relationships between the objects to be measured, converts the network to be measured into a graph form, and locates the faulty object in the network to be measured according to a preset fault location model; the fault location model is obtained by training based on a graph neural network, and the graph neural network encodes the central node through the surrounding nodes, making full use of the relationship between the surrounding nodes and the central node during training, so that the located faulty node is more accurate. Fault location is performed through this graph neural network model and the object relationship structure diagram, making the located faulty object more accurate.

[0024] The second embodiment of the present invention relates to a method for fault location, and this fault location method is applied to an electronic device, such as a server, etc. The second embodiment is a specific description of steps 101-103 in the first embodiment, and its process is specifically as Figure 2 shown:

[0025] Step 201: Obtain at least two objects to be measured, the characteristic data of the objects to be measured, and the association relationships between the objects to be measured in the network to be measured.

[0026] Specifically, the method for fault location in this example is mainly used to locate faults in a network. The network to be measured can be any type of network, such as a clock and time synchronization network, a Synchronous Digital Hierarchy (SDH) network, a Packet Transport Network (PTN) network, an IPRAN network, an optical transport network (OTN) network, an IP network, etc. The object to be measured is the object in the network to be measured where the fault needs to be located, such as network elements, optical fiber links, etc. in the network to be measured.

[0027] The characteristic data of the object to be measured can be obtained through collection. The characteristic data can be alarm data or performance data in the object to be measured, or can include both alarm data and performance data, etc. The association relationship can be the connection relationship between the objects to be measured, and the association relationship between each object to be measured can be detected through data transmission.

[0028] Step 202: Use the object to be measured as the corresponding node.

[0029] Specifically, the object to be measured can be used as a node. For example, if the object to be measured is a network element, then each network element can be converted into a corresponding node; if the object to be measured also includes an optical fiber link, then the optical fiber link can be converted into a corresponding node; the object to be measured can be network elements and optical fiber links, abstract the network elements in the network to be measured as the corresponding nodes, and use the optical fiber link as the corresponding node.

[0030] Step 203: Generate node information for the node corresponding to the object to be measured according to the characteristic data of the object to be measured.

[0031] In one example, according to the number of types of the characteristic data, convert the characteristic data into a characteristic vector with a dimension equal to the number of types, and use the characteristic vector as the node information of the node.

[0032] Specifically, obtain the number of types of the characteristic data. For example, if there are multiple alarms in the characteristic data of the object to be measured, obtain the number of types N in the alarm data. If the characteristic data includes performance data, obtain the number of types M of the performance data. If the characteristic data includes both alarm data and performance data, the number of types in the alarm data is N, and the number of types in the performance data is M, then the number of types of the characteristic data is N + M.

[0033] In one example, if the characteristic data includes alarm data, convert the alarm data into digital coding data; use the digital coding data as the characteristic vector, and use the characteristic vector as the node information of the node.

[0034] Specifically, since the alarm data can be character data, in order to unify the data form, the alarm data can be converted into digital encoded data, and the digital encoded data is used as feature data. For example, the alarm data can be represented in the form of 0 and 1 encoding; each type of alarm corresponds to a feature dimension. If the alarm exists, this dimension is represented as 1, otherwise it is represented as 0. If there are N different types of alarm data, an N-dimensional feature vector will be obtained, and this N-dimensional feature vector is denoted as the node information of this node.

[0035] In an example, if the feature data includes the performance data of the object to be measured, the performance data is normalized, and the normalized performance data is used as the feature vector; alternatively, according to at least two preset discrete numerical intervals, the performance data is dispersed into each discrete numerical interval; the value corresponding to the discrete numerical interval where the performance data is located is obtained as the feature vector.

[0036] Specifically, if the feature data includes performance data, the value of the performance data can be directly taken as the value in the feature vector. For the convenience of representing the feature vector, the performance data can also be normalized or discretized. Normalization processing means normalizing the representation interval of the performance data to between 0 and 1. Discretization processing means setting one or more thresholds. Through the thresholds, multiple discrete numerical intervals can be obtained. According to the thresholds, the performance data can be dispersed into each discrete numerical area, and the value corresponding to the discrete numerical interval where the performance data is located is used as the value of the feature vector. For example, if there is one threshold and two corresponding discrete numerical intervals, the performance value exceeding the threshold is assigned to discrete numerical interval 1, and the value corresponding to this discrete numerical interval 1 is 1. Otherwise, the performance value is assigned to discrete numerical interval 0, and the value corresponding to this discrete numerical interval 0 is represented as 0. If there are multiple thresholds, according to the multiple thresholds, it is divided into different discrete numerical intervals. From low to high, each discrete numerical interval corresponds to a value. If the performance data is in this discrete numerical interval, the value corresponding to this discrete numerical interval is obtained. For example, if three thresholds are set, it can be divided into four discrete numerical intervals, corresponding to the four values of 0, 1, 2, and 3 respectively.

[0037] Step 204: Generate an edge between every two nodes according to the association relationship between the objects to be measured.

[0038] Specifically, the association relationship between the objects to be measured is used as the edge between every two nodes. For example, if network element A and network element B are connected, then an edge is generated between the corresponding node A and node B.

[0039] Step 205: Form an object relationship structure diagram according to each node, the node information of each node, and each edge.

[0040] The node, the node information of the node, and each edge are combined to form the object relationship structure diagram.

[0041] Step 206: Locate the faulty object in the network under test according to the object relationship structure diagram and the preset fault location model.

[0042] Specifically, the fault location model is pre-trained based on the graph neural network structure, and the fault location model can be trained based on the node classification model of the graph neural network.

[0043] To facilitate the understanding of this example, the graph neural network is introduced below.

[0044] The mathematical description of the propagation mechanism of the graph neural network is as shown in formula (1):

[0045]

[0046] where h represents the representation vector of the node (denoted as "Embedding"), the subscript v represents the index of the current node, u represents the index of the node adjacent to the V node, the superscript k represents the adjacent node at the k-th layer, σ represents the activation function, W k and B k represent matrices, N(v) represents the set of adjacent nodes of node v, and AGG(*) represents the aggregation operation. When k = 0, where, x v is the input feature vector of node v.

[0047] The basic version of the propagation mechanism of the graph neural network is that when aggregating the information of the adjacent nodes of a node, the average method is adopted, and the neural network is used for the aggregation operation. Among them, the node information of each node can be the representation vector h of the node. The mathematical description of the propagation mechanism is as shown in formula (2):

[0048]

[0049] It can be understood that formula (2) is the basic version of the propagation mechanism of the graph neural network, and other versions of the propagation mechanism of the graph neural network can also be adopted. According to different propagation mechanisms, the graph neural network can be divided into Graph Convolutional Networks (GCN), Graph Attention Networks (GAN), Gated Graph Neural Network, Graph Isomorphism Network (GIN), and Graph SAGE network, etc.

[0050] In this example, according to a preset fault location model trained based on a graph neural network, the sample structure diagrams in the training set are trained to obtain the network parameters in the fault location model. The fault location model is verified through a test sample set, and then the network parameters are adjusted to obtain the fault location model.

[0051] Specifically, the sample structure diagrams in the training set may include nodes and node relationship information, node labels, and node feature data, etc. According to the needs of the fault location scenario, the node labels can be specifically divided into fault nodes and normal nodes, or divided into fault root nodes, fault impact nodes, and normal nodes. The sources of the sample networks to be measured in the sample structure diagrams include a laboratory simulation network environment and a network environment in actual use. The sample structure diagrams can also be formed based on data of the same sample network to be measured obtained at different times.

[0052] According to the needs of fault location, an end-to-end fault location model based on a graph neural network is established. The graph neural network can select a graph neural network model with node classification capabilities, including but not limited to the following graph neural network models, such as Graph Convolutional Networks (GCN), GraphAttention Networks (GAN), Gated Graph Neural Network, Graph Isomorphism Network (GIN), and Graph SAGE network, etc.

[0053] To perform node classification, it is also necessary to append a classifier after the Embedding of the nodes to complete the task, which is responsible for mapping the Embedding obtained after each node propagates through the graph neural network into the corresponding class output. In this example, the softmax classifier is taken as an example, and of course, other classifiers can also be selected. In addition, a corresponding loss function needs to be designed to measure the deviation degree between the model prediction value and the actual value, so as to train the established node classification model based on the graph neural network. Commonly used loss functions include cross-entropy loss function, 0-1 loss function, square loss function, absolute loss function, logarithmic loss function, and exponential loss function, etc. In this example, the cross-entropy loss function is taken as an example.

[0054] The established fault location model based on the graph neural network structure is trained using the training sample set. After the model training is completed, the test sample set is used to verify the effect of node classification of the model. When the accuracy of node classification reaches the standard for actual business use, the parameters obtained from the model training can be solidified for specific fault location applications.

[0055] In one example, an object relationship structure diagram is displayed, as well as the position of the faulty object in the object relationship structure diagram.

[0056] Specifically, the results of each object in the network to be tested can be displayed in the form of a topology diagram or a list. When presented in the form of a topology diagram for intuitive visualization, different background colors need to be used for rendering different node types. For example, for faulty nodes, a red background color can be used for rendering; for the nodes and edges that will be affected in the direction of fault propagation, a yellow or orange background color can be used for rendering; for normal nodes and edges, a green background color or no background color can be used for rendering. When presented in the form of a list, the faulty nodes and normal nodes can be distinguished by adding a column.

[0057] The third embodiment of the present invention relates to a method for fault location. The third embodiment is a further improvement of the second embodiment. The main improvement lies in: according to whether there is directionality in the association relationship, if there is directionality, the association relationship is converted into a directional edge. The process is as Figure 3 shown.

[0058] Step 301: Obtain at least two objects to be tested, the characteristic data of the objects to be tested, and the association relationships between the objects to be tested in the network to be tested.

[0059] Step 302: Use the objects to be tested as the corresponding nodes.

[0060] Step 303: Generate node information for the nodes corresponding to the objects to be tested according to the characteristic data of the objects to be tested.

[0061] In this example, steps 301 to 303 are substantially the same as steps 201 to 203 in the second embodiment, and will not be elaborated here.

[0062] Step 304: Perform the following processing for each association relationship: Determine whether the association relationship has directionality. If it has directionality, convert the association relationship into an edge representing directionality.

[0063] Specifically, determine whether the association relationship has directionality. If there is reverse directionality, abstract the association relationship into a directional edge. For example, if network element A transmits data unidirectionally to network element B, the association relationship between network element A and network element B has directionality. Then an edge as shown in Figure 4 can be formed, that is, node A points to node B. If there are both optical fiber links and network elements among the objects to be tested, and the data of node A passes through the input port 1 of the optical fiber and is transmitted to node C through the output port 2, with the optical fiber link as node B, an edge as shown in Figure 5 can be formed. Each edge has corresponding edge information, and the edge information includes the source end and the destination end. For example, the source end is node A and the destination end is the input port 1 of the optical fiber link.

[0064] If there is no directionality in the association relationship, directly convert the association relationship into the corresponding edge to form an undirected graph.

[0065] Step 305: Generate an object relationship structure diagram based on each node, the node information of each node, and each edge.

[0066] This step is substantially the same as step 205 in the second embodiment, and will not be elaborated here.

[0067] Step 306: Locate the faulty object in the network under test according to the object relationship structure diagram and the preset fault location model.

[0068] In one example, the training process of the fault location model may include: obtaining a sample structure diagram in the training set, where the sample structure diagram is generated from sample objects in the sample network, the feature data of each sample object, and the association relationships between each sample object; if the edges in the sample structure diagram have directionality, during the training process of the sample structure diagram, the following aggregation process is performed for each sample node in the sample structure diagram: aggregating the node information of the sample node in the propagation direction and the node information of the adjacent nodes of the sample node, where the adjacent nodes are other sample nodes whose distance from the sample node is within a preset distance.

[0069] Specifically, the adjacent nodes of each sample node can be other sample nodes whose distance from the sample node is within a preset distance, and the distance between the sample nodes can be represented by the layer number of the layer where the adjacent nodes are located; for example, as Figure 6 and Figure 7 shown, if an aggregation operation is performed on node A, Figure 7 the rectangles in represent aggregation, and assuming k is 2, then the adjacent nodes are sample nodes within 2 layers from the A node.

[0070] The 0th layer is the input layer, and the representation vector of the 0th layer is the initial Embedding of each sample node; the 1st layer: the Embedding of sample node B comes from the propagation of the Embeddings of its adjacent nodes A and C. The Embedding of sample node C comes from the propagation of the Embeddings of nodes A, B, E, and F. The Embedding of sample node D comes from the propagation of the Embedding of its sample node A. The 2nd layer: the Embedding of sample node A comes from the propagation of the Embeddings of its nodes B, C, and D. That is to say, for the adjacent nodes of sample node A, they refer to node B, C, and D in the 1st layer; and sample nodes A, C, B, E, and F in the 0th layer. During the aggregation operation, the node information of sample node A and the node information of its adjacent nodes are aggregated according to the calculation method of formula (1). Among them, since the preset distance is set to 2, k = 2 in formula (1) can be preset.

[0071] It is worth mentioning that if the fault propagation is directional and a directed graph is established, when using a graph neural network to aggregate the feature information of surrounding adjacent nodes, the feature information of adjacent nodes in the fault propagation direction can be aggregated and calculated, which can reduce the calculation cost and avoid the interference of adjacent but irrelevant node information. Further improving the accuracy of the fault location model.

[0072] Next, in this example, the entire fault location process is introduced with a specific network to be measured.

[0073] Scenario 1: The network to be measured is a clock synchronization network.

[0074] The clock synchronization network is mainly responsible for synchronizing the clock frequencies of each network device in the network and controlling the clock frequency deviation of each network device within the required range. Clock synchronization is directional, that is, the clock frequency is synchronized from upstream network devices to downstream network devices. When a fault occurs, there is a directionality of fault propagation, that is, a fault in an upstream network device will cause abnormalities in downstream network devices. The goal of fault location is to quickly find the faulty network device when a fault occurs in the clock synchronization network.

[0075] The relationship between the objects to be measured in the clock synchronization network fault diagnosis is converted into nodes and edges to generate an object relationship structure diagram. For the clock synchronization network, the objects to be measured include network elements, physical optical fiber links, and external clock sources, which are converted into nodes. The connection relationships between network elements, physical optical fiber links, and external clock sources are converted into edges. For example, if the object to be measured is a network element, an edge is established according to the link relationship between network elements. Since the fault propagation in the clock synchronization network has a directionality, the association relationship is converted into a directed edge. By collecting the information of network elements, physical optical fiber links, and external clock sources in the clock synchronization network, the corresponding nodes, edges, and node information of each node are obtained to form a directional object relationship structure diagram.

[0076] The clock-related alarms and performance data of network elements can be collected as the characteristic data of the objects to be measured. For physical optical fiber links, the alarms and performance data generated by the physical ports at both ends of the physical optical fiber link can be used as its characteristic data. For external clock sources, by collecting the network elements connected to the external clock source, the relevant clock-related alarms and performance data can be used as its characteristic data.

[0077] The characteristic data of each node is converted into a feature vector. For alarm data, in this example, it is converted into a feature vector by using the 0 and 1 coding method. Specifically, for each type of alarm related to the clock, it corresponds to a feature dimension. If the alarm exists, it is represented by the numerical value 1, otherwise, it is represented by the numerical value 0. For each node: if there are N different types of characteristic alarms, an N-dimensional feature vector will be obtained. For performance data, in this example, the performance data is processed by normalization, and the performance value is normalized to the range of 0 to 1 for representation. One type of performance data corresponds to a feature dimension. If there are M types of characteristic performance indicators, then an M-dimensional feature vector is obtained. The feature vectors of alarm data and performance data are concatenated to obtain an N + M-dimensional feature vector, and this N + M-dimensional feature vector is used as the node information of this node.

[0078] According to the needs of the fault location scenario, the types of nodes in the object relationship structure diagram can be specifically divided. For example, they can be divided into fault nodes and normal nodes, or divided into fault root cause nodes, fault impact nodes, and normal nodes. In this example, the types of nodes are divided into fault nodes and normal nodes.

[0079] Before training the fault location model, multiple sample structure diagrams are obtained to form a training set. The sample structure diagrams include nodes, edge information, node labels, and node information, etc. The sources of the sample structure diagrams include the laboratory simulation network environment and the real network environment in actual use.

[0080] After obtaining the training set, an end-to-end fault localization model based on a graph neural network is preset. The graph neural network model can be a graph attention network, the node classifier can be a softmax classifier, and the loss function can be a cross-entropy loss function. The input of this fault localization model is the feature vectors of each node, edge, and node in the graph, and the output is the type label of each node. It should be noted that when performing the aggregation calculation of the graph neural network, adjacent nodes in the aggregation fault propagation direction are aggregated.

[0081] Use the training set to train the fault localization model with the established node classification model based on the graph neural network. A training sample structure diagram includes each node, edge, node label, and feature vector in the graph. After the model training is completed, use the test set to verify the effect after model training. A test sample structure diagram also includes each node, edge, node label, and feature vector in the graph. Continuously update the network parameters of this fault localization model through training and verification; when the accuracy of fault localization reaches the standard for actual business use, the parameters obtained from model training can be solidified to generate a fault localization model.

[0082] The process of locating a fault object in the network under test based on the trained fault localization model includes: First, collect the information of the clock synchronization network where the fault needs to be located, including information such as the object to be tested, the relationships between the objects to be tested, and the feature data of each object to be tested. Then, preprocess the collected feature data and convert it into digital encoded data. Use the converted data encoded data as the feature vector of the node, use the object to be tested as the node, and use the relationships between the objects as the edges. According to the nodes, node information, and edges, form an object relationship structure diagram; input the obtained object relationship structure diagram into the fault localization model to obtain the node types of each node. This object relationship structure diagram includes nodes, node information, and edges. If the node type is a fault node type, obtain the location of this fault node to complete the localization of the fault location.

[0083] It can be presented in the form of a topology graph or a list. Presenting it in the form of a topology graph is intuitive. Different background colors need to be used to render different node types. For example, for fault nodes, a red background color can be used for rendering. For nodes and edges that will be affected in the fault propagation direction, a yellow or orange background color can be used for rendering. For normal nodes and edges, a green background color or no background color can be used for rendering. Presenting it in the form of a list, the fault nodes and normal nodes can be distinguished by adding a column.

[0084] The fault location model in this example utilizes the characteristic information of clock synchronization network nodes and also fully utilizes the characteristic information of adjacent nodes around the nodes. Through the graph neural network, the fault characteristic information of the object to be measured in the network can be utilized more fully, and the accuracy of fault location is higher. At the same time, the clock synchronization network has the directionality of fault propagation. When aggregating the characteristic information of adjacent nodes around based on the graph neural network, the node information of the nodes in the fault propagation direction and the node information of the adjacent nodes are aggregated and calculated for this node, which can reduce the computational cost and also reduce the interference of irrelevant node information.

[0085] Scenario 2: The network to be measured is a bearer network.

[0086] The bearer network is mainly responsible for transmitting service data and can provide services such as L2VPN or L3VPN. The data transmitted in the bearer network includes two directions: sending and receiving. Therefore, the direction of fault propagation includes two directions. When a fault occurs in a network element node or a physical optical fiber link in the bearer network, it will affect the surrounding adjacent data transmission nodes, and the fault propagation may exist in both directions. Therefore, the fault propagation in this example is non-directional. The goal of fault location is to quickly find the fault node when a fault occurs in the bearer network.

[0087] For the bearer network, the objects to be measured include network elements and physical optical fiber links, and the objects to be measured are regarded as nodes. The association relationship between network elements and physical optical fiber links is regarded as an edge. The fault propagation of the bearer network is non-directional, and an undirected edge is generated. By collecting the information of network element nodes and physical optical fiber links in the bearer network, the node information of the corresponding nodes is obtained, and according to the nodes, node information, and edges, the object relationship structure diagram of this bearer network is generated, and this object relationship structure diagram is an undirected graph.

[0088] The clock-related alarms and performance data of network elements can be collected as the characteristic data of the objects to be measured. For physical optical fiber links, the alarms and performance data generated by the physical ports at both ends of the physical optical fiber link can be used as its characteristic data. For external clock sources, by collecting the network elements connected to the external clock source, the relevant clock-related alarms and performance data can be used as its characteristic data.

[0089] Convert the feature data of each node into a feature vector. For the alarm data, in this example, it is converted into a feature vector by using 0 and 1 encoding. Specifically, for each type of alarm related to the clock, it corresponds to a feature dimension. If the alarm exists, it is represented by the numerical value 1, otherwise, it is represented by the numerical value 0. For each node: if there are N different types of feature alarms, an N-dimensional feature vector will be obtained. For the performance data, in this example, the performance data is processed by normalization, and the performance value is normalized to the range of 0 to 1 for representation. One type of performance data corresponds to a feature dimension. If there are M types of feature performance metrics, then an M-dimensional feature vector is obtained. Concatenate the feature vectors of the alarm data and the performance data respectively to obtain an N+M-dimensional feature vector, and use this N+M-dimensional feature vector as the node information of this node.

[0090] According to the needs of the fault location scenario, the types of nodes in the object relationship structure diagram can be specifically divided. For example, it can be divided into fault nodes and normal nodes, or divided into fault root cause nodes, fault impact nodes and normal nodes. In this example, the types of nodes are divided into fault nodes and normal nodes.

[0091] Before training the fault location model, obtain multiple sample structure diagrams to form a training set. The sample structure diagrams include node and edge information, node labels, node information, etc. The sources of the sample structure diagrams include the laboratory simulation network environment and the real network environment in actual network use.

[0092] After obtaining the training set, preset an end-to-end fault location model based on a graph neural network. The graph neural network model can be a graph attention network, the node classifier can be a softmax classifier, and the loss function can be a cross-entropy loss function. The input of this fault location model is the feature vectors of each node, edge and node in the graph, and the output is the type label of each node. It should be noted that when performing the aggregation calculation of the graph neural network, since the fault propagation is not directional, when performing the aggregation operation on each node, it is necessary to aggregate this node and all its adjacent surrounding nodes for the aggregation calculation.

[0093] Use the training set to train the established fault location model based on the graph neural network. A training sample structure diagram contains each node, edge, node label and feature vector in the graph. After the model training is completed, use the test set to verify the effect after the model training. A test sample structure diagram also contains each node, edge, node label and feature vector in the graph. Continuously update the network parameters of this fault location model through training and verification; when the accuracy of the fault location reaches the standard for actual business use, the parameters obtained by the model training can be solidified to generate a fault location model.

[0094] The process of locating a fault object in a network under test based on a trained fault location model includes: First, collect information on the clock synchronization network where the fault needs to be located, including information such as the objects under test, the relationships between the objects under test, and the characteristic data of each object under test. Then, preprocess the collected characteristic data, convert it into digital encoded data, use the converted data encoded data as the feature vector of the node, use the object under test as the node, use the relationships between the objects as the edges, and form an object relationship structure diagram according to the nodes, node information, and edges; input the obtained object relationship structure diagram into the fault location model to obtain the node types of each node. The object relationship structure diagram includes nodes, node information, and edges. If the node type is a fault node type, obtain the location of the fault node to complete the location of the fault position.

[0095] It can be presented in the form of a topology graph or a list. Presenting it in the form of a topology graph is intuitive, and different background colors need to be used to render different node types. For example, for fault nodes, a red background color can be used for rendering; for the nodes and edges that will be affected in the fault propagation direction, a yellow or orange background color can be used for rendering; for normal nodes and edges, a green background color or no background color can be used for rendering. Presenting it in the form of a list can distinguish between fault nodes and normal nodes by adding a column.

[0096] The fourth embodiment of the present invention relates to an electronic device, and its structural block diagram is as Figure 8 shown. The electronic device includes: at least one processor 401; and a memory 402 communicatively connected to the at least one processor 401; wherein, the memory 402 stores instructions executable by the at least one processor 401, and the instructions are executed by the at least one processor 401 to enable the at least one processor 401 to execute the above-mentioned fault location method.

[0097] Among them, the memory and the processor are connected by a bus. The bus can include any number of interconnected buses and bridges, and the bus links various circuits of one or more processors and memories together. The bus can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits together, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one component or multiple components, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor.

[0098] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. The memory can be used to store the data used by the processor during operation.

[0099] The fifth embodiment of the present invention relates to a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-mentioned fault location method.

[0100] Those skilled in the art can understand that all or part of the steps in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0101] Those of ordinary skill in the art can understand that the above embodiments are specific examples for implementing the present invention, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present invention.

Claims

1. A method for fault location, characterized in that Including: Obtain at least two objects to be tested, the characteristic data of the objects to be tested, and the association relationships between the objects to be tested in the network to be tested. The objects to be tested are the objects in the network to be tested for which faults need to be located, and the objects to be tested include network elements and / or optical fiber links in the network to be tested. The characteristic data includes alarm data and / or performance data of the objects to be tested; Generate an object relationship structure diagram according to at least two of the objects to be tested, the characteristic data of the objects to be tested, and the association relationships. Among them, the object relationship structure diagram includes nodes, node information, and edges. The nodes are the objects to be tested, the edges are the association relationships, and the node information is the characteristic data of the objects to be tested; Locate the faulty object in the network to be tested according to the object relationship structure diagram and a preset fault location model. The fault location model is a graph neural network structure, and the fault location model is trained based on a node classification model of the graph neural network.

2. The method for fault location according to claim 1, wherein The step of obtaining at least two objects to be tested, the characteristic data of the objects to be tested, and the association relationships between the objects to be tested in the network to be tested includes: Use the object to be tested as the corresponding node; Generate the node information of the node corresponding to the object to be tested according to the characteristic data of the object to be tested; Generate an edge between every two nodes according to the association relationships between the objects to be tested; Form the object relationship structure diagram according to each node, the node information of each node, and each edge.

3. The method for fault location according to claim 2, wherein The step of generating an edge between every two nodes according to the association relationships between the objects to be tested includes: Perform the following processing for each of the association relationships: Judge whether the association relationship has directionality. If it has directionality, convert the association relationship into an edge representing the directionality.

4. The method for fault location according to claim 2, wherein, The step of obtaining the characteristic data of the object to be tested and generating the node information of the node corresponding to the object to be tested includes: Obtain the number of types of the characteristic data; Convert the characteristic data into a feature vector with a dimension equal to the number of types; Use the feature vector as the node information of the node.

5. The method for fault location according to claim 4, wherein If the characteristic data includes the performance data of the object to be tested; The step of converting the characteristic data into a feature vector with a dimension equal to the number of types includes: Perform normalization processing on the performance data, and use the normalized performance data as the feature vector; or, According to at least two preset discrete numerical intervals, disperse the performance data into each of the discrete numerical intervals; obtain the numerical value corresponding to the discrete numerical interval where the performance data is located as the feature vector.

6. The method for fault location according to claim 4, wherein If the characteristic data includes alarm data; The step of converting the characteristic data into a feature vector with a dimension equal to the number of types includes: Convert the alarm data into digital encoded data; Use the digital encoded data as the feature vector.

7. The method for fault location according to claim 4, characterized in that Before locating the faulty object in the network to be tested according to the object relationship structure diagram and a preset fault location model, the method further includes: Obtain the sample structure diagram in the training set, where the sample structure diagram is generated from sample objects in the sample network and the association relationships between the sample objects. If the edges in the sample structure diagram are directional, during the training process of the sample structure diagram, the following aggregation processing is performed for each sample node: Aggregate the node information of the sample node and the node information of the adjacent nodes of the sample node in the propagation direction, where the adjacent nodes are other sample nodes whose distance from the sample node is within a preset distance.

8. The method for fault location according to any one of claims 1 to 7, characterized in that, After locating the faulty object in the network under test according to the object relationship structure diagram and the preset fault location model, it includes: Display the object relationship structure diagram and the position of the faulty object in the object relationship structure diagram.

9. An electronic device, characterized in that, It includes: At least one processor; And a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the fault location method as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the fault location method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and device for presentation of transmission network topology figures

    CN105634763A

  • Covert fault troubleshooting method and apparatus of 4G base station

    CN109618361A