Fault diagnosis method based on wide-depth fusion graph convolutional neural network under distributed autonomous network architecture

By combining digital twin networks with wide-deep fusion graph convolutional neural networks, the real-time and accuracy issues of fault diagnosis in distributed autonomous networks are solved, efficient fault location and diagnosis are achieved, and network overhead is reduced.

CN120676392APending Publication Date: 2025-09-19CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510760340.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The existing technology has the following problems in distributed autonomous networks: the fault diagnosis model is difficult to achieve both high accuracy and real-time performance, there is a lack of training methods for distributed dynamic topology structures, and intrusive fault injection affects network performance.

Method used

The digital twin network is used to non-invasively acquire fault feature data. Combined with a wide-deep fusion graph convolutional neural network, a distributed online training model is used to identify the fault propagation path using the graph neural network to locate the fault.

Benefits of technology

It improves the real-time and accuracy of fault diagnosis, reduces network storage and computing overhead, and avoids negative impact on network services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120676392A_ABST
    Figure CN120676392A_ABST
Patent Text Reader

Abstract

The invention relates to a fault diagnosis method based on a wide-depth fusion graph convolutional neural network under a distributed autonomous network architecture, and belongs to the technical field of mobile communication. The method comprises the following steps: establishing a digital twin network according to a distributed autonomous network architecture, and injecting faults into different network element types in the digital twin network to collect fault feature data; dividing the working states of the network elements into three types, namely a normal state, an abnormal state and a fault state, by considering network fault propagation characteristics; taking the width-depth fusion graph convolutional neural network as a basic fault diagnosis model of a distributed network, firstly performing offline training on the width-depth fusion graph convolutional neural network, and then performing online training on a linear part of the width-depth fusion graph convolutional neural network until convergence to obtain a distributed online fault diagnosis model; and identifying a fault propagation path in the network by using a message passing mechanism of the graph neural network, and positioning a fault network element in the network. According to the invention, the network storage and calculation overhead can be reduced, and the real-time performance and accuracy of fault diagnosis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of mobile communication technology and relates to network fault diagnosis, and in particular to a fault diagnosis method based on a wide-deep fusion graph convolutional neural network in a distributed autonomous network architecture. Background Art

[0002] Distributed autonomous networks (DANs) are a key enabler of 6G. Through distributed network technology and artificial intelligence (AI), they effectively address network challenges faced by operators, such as long response times, poor flexibility, high O&M costs, and low efficiency. Given the inherent complexity and dynamism of DANs, leveraging AI to accurately locate potential network element failures and implement intelligent fault diagnosis is crucial for building a highly reliable and efficient 6G network intelligent O&M system.

[0003] As network operations and maintenance management extends from physical devices to virtual machines, the types of potential network element failures increase, and the shared nature of network elements strengthens, resulting in a dramatic increase in the complexity of network fault location. On the one hand, the deployment of virtual machines expands the paths for network fault propagation. Faults not only propagate along service paths, but also spread based on the hierarchical relationships of network elements (for example, a physical node failure will cause the virtual machines it hosts to become unavailable). On the other hand, the highly shared nature of network elements significantly increases the likelihood of network failure. Intelligent network fault diagnosis technology leverages machine learning models to automatically diagnose network faults and improve location efficiency. Given the powerful modeling capabilities of graph neural networks for network topology and dependencies, graph neural networks can achieve more accurate fault location by aggregating feature information from adjacent nodes and learning the propagation pattern of faults.

[0004] However, the existing technology still has some shortcomings:

[0005] First, research on online fault diagnosis models based on time-series network operation data is still relatively limited. The contradiction between achieving high-precision diagnostic performance and maintaining real-time performance in existing methods is becoming increasingly prominent. Second, distributed network management architecture increases the complexity of fault location, and customized service paths further increase the fault propagation path. Existing research lacks sufficient exploration of distributed dynamic topology structures and distributed training methods for fault diagnosis models. Finally, fault injection is a necessary means to obtain sufficient fault feature data, but injecting faults into the existing network in an invasive manner will have a negative impact on network service performance. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a fault diagnosis method based on a wide-deep fusion graph convolutional neural network under a distributed autonomous network architecture. By utilizing a digital twin network to obtain fault feature data in a non-invasive manner, a distributed online fault diagnosis model is established, thereby improving the real-time and accuracy of fault diagnosis while reducing network storage and computing overhead.

[0007] In order to achieve the above object, the present invention provides the following technical solutions:

[0008] A fault diagnosis method based on a wide-deep fusion graph convolutional neural network in a distributed autonomous network architecture, the method comprising:

[0009] S1. Establish a digital twin network based on a distributed autonomous network architecture, inject faults into different network element types in the digital twin network, and collect fault feature data in a non-invasive manner;

[0010] S2. Considering the network fault propagation characteristics, the working status of network elements is divided into three types: normal state, abnormal state and fault state, providing prior information for the training of fault diagnosis model;

[0011] S3. Use the wide-deep fusion graph convolutional neural network as the basic fault diagnosis model for distributed networks and construct an adjacency matrix based on network topology information. First, train the wide-deep fusion graph convolutional neural network offline based on distributed network topology information and historical network data. Then, train the linear part of the wide-deep fusion graph convolutional neural network online based on dynamic topology perception information and real-time network operation data until convergence, thus obtaining a distributed online fault diagnosis model.

[0012] S4. Use the message passing mechanism of graph neural networks to identify potential fault propagation paths in the network and locate faulty network elements in the network.

[0013] Furthermore, in step S1, the network element types include physical nodes, physical links, virtual machines, virtual links, CPU resource modules, storage resource modules and disk resource modules; the fault characteristic data of the network element types include CPU usage, memory usage, disk read / write speed, network input / output speed, average load and average number of blocked processes.

[0014] Furthermore, in step S2, the abnormal state is defined as being caused by abnormal propagation of the faulty network element or its own problems, which does not affect the normal operation of other network elements; the fault state is defined as being caused by the faulty network element affecting other network elements and causing abnormal working states of other network elements through the fault propagation path.

[0015] Furthermore, in step S3, the wide-deep fusion graph convolutional neural network includes a linear graph filter and a nonlinear graph convolutional neural network, which can be expressed as:

[0016] Ψ(X;A,B,C)=α L B(X;A,B)+α NL Φ(X;A,C)+β

[0017] Where Ψ(X; A, B, C) represents the wide-deep fusion graph convolutional neural network; B(X; A, B) represents the linear graph filter; Φ(X; A, C) represents the nonlinear graph convolutional neural network; X represents the data feature set of all network elements; A represents the adjacency matrix; B and C represent the parameters of the linear graph filter and the nonlinear graph convolutional neural network, respectively; α L , α NL and β are the joint weights of the width part and the depth part.

[0018] Furthermore, in step S3, the wide-depth fusion graph convolutional neural network is trained offline based on the distributed network topology information and historical network data, including: i In the example, the adjacency matrix A is obtained based on the network topology. i , combined with historical data X i Offline training of the wide-deep fusion graph convolutional neural network until convergence is performed to obtain the optimal graph filter parameters and graph convolutional neural network parameters; the optimization goal in the offline training phase is:

[0019]

[0020] Where, represents the dimension of the offline training dataset; J represents the cross entropy loss function; y i represents the label of the i-th training sample; c represents the output dimension of the wide-deep fusion graph convolutional neural network; j represents the j-th dimension of the wide-deep fusion graph convolutional neural network.

[0021] Furthermore, in step S3, the linear part of the wide-deep fusion graph convolutional neural network is trained online based on the dynamic topology perception information and the real-time network operation data until convergence, including: each sub-network P i Initialize its linear graph filter parameters represents the optimal graph filter parameters;

[0022] Based on dynamic topology perception, the adjacency matrix A of the i subnet at time t is obtained i,t , and combined with the real-time operation data X of the i-th subnet at time t i,t Train the linear graph filters of the wide-deep fusion graph convolutional neural network; keep the nonlinear graph convolutional neural network unchanged during online training;

[0023] The optimization goal of the online training phase is:

[0024]

[0025] Where m represents the number of sub-networks in the distributed network; J i,t (·) represents the local loss function of the i-th subnetwork at time t; X i,tA represents the real-time data characteristics of the i-th subnet at time t; i,t represents the adjacency matrix of the i-th subnet at time t; represents the optimal graph convolutional neural network parameters obtained by offline training; Represents the subnetwork P i All neighbor subnetworks of ; and They represent the parameters of the linear graph filters in the i-th and j-th sub-networks respectively; C1 represents the consistency constraint condition of the linear graph filter parameters between adjacent networks.

[0026] Furthermore, in step S3, during the online training phase, each sub-network P i Update its local parameters online based on iterative method To aggregate the linear graph filter parameters of adjacent sub-networks, as shown in the following formula:

[0027]

[0028] Where, and Represent the linear graph filter parameters of the i-th sub-network at time t+1 and time t respectively; W t represents the local linear graph filter aggregation weight matrix at time t; γ represents the update step size; represents the model gradient.

[0029] The beneficial effects of the present invention are as follows: the present invention proposes a fault diagnosis method based on a wide-deep fusion graph convolutional neural network under a distributed autonomous network architecture. The method constructs a digital twin of a distributed network by utilizing digital twin technology, and injects faults into different network element types in the digital twin to collect fault data features. By defining the fault type, a data set of fault type-fault data features is obtained. This method can not only obtain sufficient fault data features for the training of the diagnostic model, but also realize non-invasive fault injection to avoid negative impacts on network service performance. In addition, the present invention proposes a fault diagnosis model based on a wide-deep fusion graph convolutional neural network, and trains the model in a distributed training method based on real-time network operation data, which reduces network storage and computing overhead. At the same time, the graph convolutional neural network can aggregate the feature information of adjacent nodes and learn the propagation mode of faults, thereby achieving more accurate fault location and improving the real-time and accuracy of fault diagnosis.

[0030] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0032] Figure 1 Schematic diagram of the fault management framework based on digital twin network;

[0033] Figure 2 This is a schematic diagram of the distributed online fault diagnosis model principle;

[0034] Figure 3 Flowchart of the training and fault localization method for wide-deep fusion graph convolutional neural network. DETAILED DESCRIPTION

[0035] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0036] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0037] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0038] See also Figures 1 to 3 This embodiment provides a fault diagnosis method based on a wide-deep fusion graph convolutional neural network in a distributed autonomous network architecture, including the following steps:

[0039] S1. In a distributed autonomous network architecture, a digital twin network is established using the dynamic topology, service request data, and resource information of the actual network. Faults are injected into different network element types in the digital twin network, and fault feature data is collected in a non-invasive manner to provide a data basis for training fault diagnosis models.

[0040] The digital twin network is composed of three modules: data management, virtual space, and model management. The data management module collects and stores network resource information, service request data, and topology information, providing a data foundation for the establishment of the digital twin network. The virtual space module is a digital representation of the real network, used to monitor the operation of the network in real time and collect fault characteristic data through fault injection to support the training of the fault diagnosis model. The model management module contains a basic model for fault diagnosis. The fault diagnosis model locates the specific network element that has a fault based on the network's adjacency matrix and fault characteristic matrix.

[0041] Different network element types include physical nodes, physical links, virtual machines, virtual links, CPU resource modules, storage resource modules, and disk resource modules. Fault signature data is obtained by injecting faults into different network element types. Fault signature data for different network element types includes CPU usage, memory usage, disk read / write speed, network input / output speed, average load, and average number of blocked processes.

[0042] S2. Considering the network fault propagation characteristics, the working status of the network element is divided into three types: normal state, abnormal state and fault state, providing prior information for the training of the fault diagnosis model.

[0043] Abnormal states are caused by abnormal propagation of faulty network elements or their own problems, but do not affect the normal operation of other network elements. A faulty network element can affect other network elements, causing abnormal operation of other network elements through the fault propagation path. These three states and their fault propagation processes are modeled in the digital twin network, and the corresponding data features of these three states are annotated to provide prior information and data foundation for training fault diagnosis models.

[0044] S3. Design a wide-deep fusion graph convolutional neural network as the basic fault diagnosis model for distributed networks, construct an adjacency matrix based on network topology information, and establish a distributed online learning mechanism to train the wide-deep fusion graph convolutional neural network. First, train the wide-deep fusion graph convolutional neural network offline based on distributed network topology information and historical network data, and then train the linear part of the wide-deep fusion graph convolutional neural network online based on dynamic topology perception information and real-time network operation data until convergence, to obtain a distributed online fault diagnosis model.

[0045] 1) In each sub-network, network elements (nodes, links, etc.) are used as graph nodes, and the connection relationships between network elements are used as edges to construct a network topology model. Based on the network topology model, an adjacency matrix is ​​obtained.

[0046] 2) The wide-deep fusion graph convolutional neural network is composed of a linear graph filter and a nonlinear graph convolutional neural network. Its mathematical representation is:

[0047]

[0048] Where, represents a wide and deep fusion graph convolutional neural network; is the width part, representing the linear graph filter; Φ(X; A, C) is the depth part, representing the nonlinear graph convolutional neural network; X represents the data feature set of all network elements; A represents the adjacency matrix; and Represent the parameters of the width part and the depth part respectively; α L , α NL and β are the joint weights of the width part and the depth part.

[0049] 3) Offline training of wide-deep fusion graph convolutional neural networks based on distributed network topology information and historical network data, including:

[0050] In each subnetwork P i In the example, the adjacency matrix A is obtained based on the network topology. i , combined with historical data X i Offline training of the wide-deep fusion graph convolutional neural network until convergence to obtain the optimal graph filter parameters and graph convolutional neural network parameters The optimization goal in the offline training phase is:

[0051]

[0052] in, represents the dimension of the offline training dataset; J represents the cross entropy loss function; y i represents the label of the i-th training sample; exp(·) and log(·) represent the exponential function and logarithmic function respectively; c is the output dimension of the wide-deep fusion graph convolutional neural network; j represents the j-th dimension of the wide-deep fusion graph convolutional neural network.

[0053] 4) Online training of the linear part of the wide-deep fusion graph convolutional neural network until convergence based on dynamic topology perception information and real-time network operation data, including:

[0054] The optimal parameters obtained in the offline training phase are used as the initial parameters in the online training phase. Specifically, each sub-network P iInitialize its linear graph filter parameters

[0055] Based on dynamic topology perception, the adjacency matrix A of the i subnet at time t is obtained i,t , and combined with the real-time operation data X of the i-th subnet at time t i,t Train the linear part of the wide-deep fusion graph convolutional neural network; during the online training process, to ensure convergence, the nonlinear part of the wide-deep fusion graph convolutional neural network remains unchanged.

[0056] The optimization goal of the online training phase is:

[0057]

[0058] Where m represents the number of subnets in the distributed network; J i,t (·) represents the local loss function of the i-th subnetwork at time t; X i,t is the real-time data feature of the ith subnet at time t; A i,t represents the adjacency matrix of the i-th subnet at time t; The optimal graph convolutional neural network parameters obtained by offline training; Represents the subnetwork P i All neighbor subnetworks of ; and They represent the parameters of the linear graph filters in the i-th and j-th sub-networks respectively; C1 represents the consistency constraint of the model parameters between adjacent networks. In the optimization goal of distributed model training, the consistency constraint of the model parameters between adjacent networks is mainly used to coordinate the learning process of each sub-network, ensure that the linear graph filter parameters of each sub-network remain synchronized during update, avoid deviations in the training process, improve global convergence, and enable the wide-deep fusion graph convolutional neural network model to quickly converge to a global optimal solution.

[0059] During the online training phase, each sub-network P i Update its local parameters online based on iterative method As shown in the following formula:

[0060]

[0061] in, and Represent the linear graph filter parameters of the i-th sub-network at time t+1 and time t respectively; W t is the local model aggregation weight matrix at time t; γ represents the update step size; represents the model gradient.

[0062] S4. Utilize the message passing mechanism of graph neural networks to identify potential fault propagation paths within the network and accurately locate faulty network elements. The identified faulty network elements are fed back to the physical network via the digital twin network. The physical network then restores network service performance through self-healing mechanisms such as activating backup devices or reconfiguring service paths, ensuring a continuous, high-quality user experience for all network services.

[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A fault diagnosis method based on wide-deep fusion graph convolutional neural network in a distributed autonomous network architecture, characterized by: The method includes: S1. Establish a digital twin network based on a distributed autonomous network architecture, inject faults into different network element types in the digital twin network, and collect fault feature data in a non-invasive manner; S2. Considering the network fault propagation characteristics, the working status of network elements is divided into three types: normal state, abnormal state and fault state, providing prior information for the training of fault diagnosis model; S3. Use the wide-deep fusion graph convolutional neural network as the basic fault diagnosis model for distributed networks and construct an adjacency matrix based on network topology information. First, train the wide-deep fusion graph convolutional neural network offline based on distributed network topology information and historical network data. Then, train the linear part of the wide-deep fusion graph convolutional neural network online based on dynamic topology perception information and real-time network operation data until convergence, thus obtaining a distributed online fault diagnosis model. S4. Use the message passing mechanism of graph neural networks to identify potential fault propagation paths in the network and locate faulty network elements in the network.

2. The method according to claim 1, characterized in that In step S1, the network element types include physical nodes, physical links, virtual machines, virtual links, CPU resource modules, storage resource modules, and disk resource modules; Fault characteristic data of network element types includes CPU usage, memory usage, disk read / write speed, network input / output speed, average load, and average number of blocked processes.

3. The method according to claim 1, characterized in that In step S2, the abnormal state is defined as being caused by abnormal propagation of the faulty network element or its own problems, which does not affect the normal operation of other network elements; the fault state is defined as being caused by the faulty network element affecting other network elements and causing abnormal working states of other network elements through the fault propagation path.

4. The method according to claim 1, wherein The wide-deep fusion graph convolutional neural network includes a linear graph filter and a nonlinear graph convolutional neural network, which can be expressed as: Where, represents a wide and deep fusion graph convolutional neural network; represents a linear graph filter; Φ(X; A, C) represents a nonlinear graph convolutional neural network; X represents the data feature set of all network elements; A represents the adjacency matrix; and C represent the parameters of the linear graph filter and the nonlinear graph convolutional neural network respectively; α L , α NL and β are the joint weights of the linear graph filter and the nonlinear graph convolutional neural network.

5. The method according to claim 4, characterized in that Offline training of wide-depth fusion graph convolutional neural network based on distributed network topology information and historical network data includes: i In the example, the adjacency matrix A is obtained based on the network topology. i , combined with historical data X i Offline training of the wide-deep fusion graph convolutional neural network until convergence is performed to obtain the optimal graph filter parameters and graph convolutional neural network parameters; the optimization goal in the offline training phase is: Where, represents the dimension of the offline training dataset; J represents the cross entropy loss function; y i represents the label of the i-th training sample; c represents the output dimension of the wide-deep fusion graph convolutional neural network; j represents the j-th dimension of the wide-deep fusion graph convolutional neural network.

6. The method according to claim 5, characterized in that Based on dynamic topology perception information and real-time network operation data, the linear part of the wide-deep fusion graph convolutional neural network is trained online until convergence, including: each sub-network P i Initialize its linear graph filter parameters represents the optimal graph filter parameters; Based on dynamic topology perception, the adjacency matrix A of the i subnet at time t is obtained i,t , and combined with the real-time operation data X of the i-th subnet at time t i,t Train the linear graph filters of the wide-deep fusion graph convolutional neural network; keep the nonlinear graph convolutional neural network unchanged during online training; The optimization goal of the online training phase is: Where m represents the number of sub-networks in the distributed network; J i,t (·) represents the local loss function of the i-th subnetwork at time t; X i,t A represents the real-time data characteristics of the i-th subnet at time t; i,t represents the adjacency matrix of the i-th subnet at time t; C * represents the optimal graph convolutional neural network parameters obtained by offline training; Represents the subnetwork P i All neighbor subnetworks of ; and They represent the parameters of the linear graph filters in the i-th and j-th sub-networks respectively; C1 represents the consistency constraint condition of the linear graph filter parameters between adjacent networks.

7. The method according to claim 6, characterized in that During the online training phase, each sub-network P i Online update of local linear graph filter parameters based on iterative method To aggregate the linear graph filter parameters of adjacent sub-networks, as shown in the following formula: Where, and Represent the linear graph filter parameters of the i-th sub-network at time t+1 and time t respectively; W t represents the local model aggregation weight matrix at time t; γ represents the update step size; represents the model gradient.