Fault operation and maintenance data transmission method and system based on artificial intelligence

By using artificial intelligence methods in network operation and maintenance, combined with operation and maintenance knowledge topology structure and information processing algorithms, the problems of insufficient information and insufficient correlation utilization in traditional fault detection methods are solved, and more accurate and efficient fault analysis and processing are achieved.

CN120263623AActive Publication Date: 2025-07-04SICHUAN BODA ZHENGHENG INFORMATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510734203.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

Traditional network operation and maintenance fault detection methods rely on simple operation and maintenance information flow analysis, and cannot comprehensively and accurately detect the essence and potential correlation factors of the fault, resulting in inaccurate detection results and lack effective utilization of the correlation between failure modes, making it difficult to achieve accurate diagnosis and effective resolution.

Method used

Using the fault operation and maintenance data transmission method based on artificial intelligence, the knowledge topology branches are determined by obtaining the operation and maintenance information flow, and the pre-deployed operation and maintenance knowledge topology structure is used to determine the knowledge topology branches, and the features are acquired in combination with the embedding layer of the operation and maintenance information processing algorithm, aggregation features are established, and fault identification layer is used to perform fault analysis.

Benefits of technology

It improves the accuracy and efficiency of fault detection, can have a more comprehensive understanding of the causes and possible impact of faults, provides more accurate fault analysis results, and helps operation and maintenance personnel quickly locate and deal with network equipment problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263623A_ABST
    Figure CN120263623A_ABST
Patent Text Reader

Abstract

The invention provides a fault operation and maintenance data transmission method and system based on artificial intelligence. The method comprises the steps of obtaining a to-be-processed operation and maintenance information stream; if the operation and maintenance information flow comprises the first fault mode, a first knowledge topological branch is determined in an operation and maintenance knowledge topological structure deployed in advance, and the first knowledge topological branch comprises a first topological unit, topological units within a preset distance obtained based on the first topological unit and a topological chain between the topological units; based on the first knowledge topology branch, obtaining a first knowledge topology feature through a first embedding layer; based on the operation and maintenance information flow, operation and maintenance information features are obtained through a second embedding layer in an operation and maintenance information processing algorithm; establishing an aggregation feature according to the first knowledge topological feature and the operation and maintenance information feature; and obtaining a fault analysis result through a fault identification layer in an operation and maintenance information processing algorithm based on the aggregation features. The technical problem of inaccurate fault detection result caused by insufficient instantaneous information amount of the operation and maintenance information flow can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and specifically to a method and system for transmitting fault operation and maintenance data based on artificial intelligence. Background Art

[0002] As the scale of networks continues to expand and their complexity increases, the probability of network equipment failure also increases accordingly. How to efficiently and accurately detect and handle network failures has become a key issue that needs to be urgently addressed in the field of network operation and maintenance.

[0003] Traditional network operation and maintenance fault detection methods mainly rely on simple analysis of operation and maintenance information flows, usually only making isolated judgments on equipment operating parameters, log information, etc. at a single point in time. This approach is often limited by the insufficient amount of instantaneous information in the operation and maintenance information flow, and it is difficult to fully and accurately grasp the nature and overall picture of the fault. For example, when a network device fails, due to data limitations, the detection results may only reflect the surface fault phenomenon, but cannot deeply analyze the root cause of the fault and potential related factors, resulting in inaccurate fault detection results, which in turn affects the efficiency and effectiveness of fault handling. In addition, traditional methods lack effective use of the correlation between fault modes. In actual network environments, there are often complex connections between different fault modes, and the occurrence of a fault may trigger a series of related faults. However, traditional detection methods fail to fully consider these correlations, and only treat each fault mode as an independent individual for processing, which makes the fault detection and processing process lack of systematicity and comprehensiveness, and it is difficult to achieve accurate diagnosis and effective resolution of network faults. At the same time, because traditional methods cannot improve the contextual features of the operation and maintenance information flow, they have poor adaptability and flexibility when facing complex and changeable network fault scenarios. Summary of the invention

[0004] The present application provides a method and system for transmitting fault operation and maintenance data based on artificial intelligence.

[0005] According to one aspect of the present application, a method for transmitting fault operation and maintenance data based on artificial intelligence is provided, which is applied to a server. The server is communicatively connected to an operation and maintenance terminal. The method includes: obtaining the operation and maintenance information flow to be processed uploaded by the operation and maintenance terminal; if the operation and maintenance information flow includes a first fault mode, determining a first knowledge topology branch in a pre-deployed operation and maintenance knowledge topology structure, where each topology unit in the operation and maintenance knowledge topology structure represents a fault mode, each topology chain in the operation and maintenance knowledge topology structure represents the relevance of fault modes, the first fault mode belongs to the fault mode represented by the first topology unit in the operation and maintenance knowledge topology structure, and the first knowledge topology branch includes the first topology unit, topology units within a preset distance obtained based on the first topology unit, and the topology chains between the topology units; based on the first knowledge topology branch, obtaining a first knowledge topology feature through a first embedding layer included in the operation and maintenance information processing algorithm; based on the operation and maintenance information flow, obtaining an operation and maintenance information feature by using a second embedding layer included in the operation and maintenance information processing algorithm; establishing an aggregated feature according to the first knowledge topology feature and the operation and maintenance information feature; based on the aggregated feature, obtaining a fault analysis result of the operation and maintenance information flow by using a fault identification layer included in the operation and maintenance information processing algorithm, and transmitting the fault analysis result to the operation and maintenance terminal.

[0006] According to another aspect of the present application, a fault operation and maintenance data transmission system is provided, including a server and an operation and maintenance terminal that communicate with each other. The server includes: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.

[0007] This application obtains the operation and maintenance information flow to be processed. If it is detected that the operation and maintenance information flow includes a first failure mode, a first knowledge topology branch is determined in the pre-deployed operation and maintenance knowledge topology structure. Each topology unit in the operation and maintenance knowledge topology structure represents a failure mode, and each topology chain in the operation and maintenance knowledge topology structure represents the relevance of failure modes. Among them, the first failure mode belongs to the failure mode represented by the first topology unit in the operation and maintenance knowledge topology structure. The first knowledge topology branch includes the first topology unit, the topology units within a preset distance obtained based on the first topology unit, and the topology chains between the topology units. Based on this, based on the first knowledge topology branch, the first knowledge topology feature is obtained through the first embedding layer included in the operation and maintenance information processing algorithm. And, based on the operation and maintenance information flow, the operation and maintenance information feature is obtained through the second embedding layer included in the operation and maintenance information processing algorithm. Then, based on the first knowledge topology feature and the operation and maintenance information feature, an aggregated feature is established, and based on the aggregated feature, the failure analysis result of the operation and maintenance information flow is obtained through the failure identification layer included in the operation and maintenance information processing algorithm. This application combines the failure modes and the relevance of failure modes in the operation and maintenance knowledge topology structure to establish a knowledge topology branch, and improves the context features of the operation and maintenance information flow through the knowledge topology branch, which can overcome the current situation that the failure detection result is inaccurate due to insufficient instantaneous information volume in the operation and maintenance information flow. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 FIG. shows a schematic diagram of a fault operation and maintenance data transmission system according to an embodiment of the present application.

[0009] Figure 2 FIG. shows a flowchart of a fault operation and maintenance data transmission method based on artificial intelligence according to an embodiment of the present application.

[0010] Figure 3 FIG. shows a schematic diagram of the composition of a server according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0011] Figure 1 FIG. shows a schematic diagram of a fault operation and maintenance data transmission system 100 provided according to an embodiment of the present application. The fault operation and maintenance data transmission system 100 includes one or more operation and maintenance terminals 101, a server 120, and one or more communication networks 110 that couple the one or more operation and maintenance terminals 101 to the server 120. The operation and maintenance terminal 101 can be configured to execute one or more application programs.

[0012] In Figure 1In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or a combination thereof that can be executed by one or more processors. The user of operation and maintenance terminal 101 can sequentially utilize one or more applications to interact with server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which may be different from the fault operation and maintenance data transmission system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.

[0013] The user can use operation and maintenance terminal 101 to receive the operation and maintenance information stream sent by operation and maintenance terminal 101. Operation and maintenance terminal 101 can be a general-purpose computer (such as a personal computer and a laptop computer), a workstation computer, a thin client, etc. It can also be an operation and maintenance toolbox designed according to actual needs. These computer devices can run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android.

[0014] The fault operation and maintenance data transmission system 100 may also include one or more databases 130. In some embodiments, these databases can be used to store data and other information. Database 130 can reside in various locations.

[0015] Please refer to Figure 2 , which is a flowchart of the artificial intelligence-based fault operation and maintenance data transmission method provided by this application. The method includes: Step S100: Obtain the operation and maintenance information stream to be processed uploaded by the operation and maintenance terminal.

[0016] The operation and maintenance terminal is an operation and maintenance toolbox for network operation and maintenance. In actual applications, when operation and maintenance personnel perform operation and maintenance operations on network devices, the operation and maintenance terminal will record relevant operation information, device status information, etc. These information are aggregated to form an operation and maintenance information flow and uploaded to the server. The operation and maintenance information flow to be processed refers to an information collection that has not been processed and analyzed by the server and contains various types of data during the operation and maintenance process. These data may involve device operation parameters, fault prompt information, operation logs, etc. For example, in an enterprise's network environment, operation and maintenance personnel use the operation and maintenance terminal to perform daily inspections and configuration adjustments on routers. The operation and maintenance terminal will record operation parameters such as the CPU usage rate, memory occupancy, and port connection status of the router, as well as configuration modification operations performed by operation and maintenance personnel, such as modifying the access control list and adjusting the routing policy. The recorded data is uploaded to the server in the form of an operation and maintenance information flow, waiting for further processing by the server.

[0017] The server can use network communication protocols such as HTTP, HTTPS, etc. to obtain the operation and maintenance information flow. The operation and maintenance terminal encapsulates the operation and maintenance information flow in the format of the corresponding protocol and then sends it to the server through the network. The server listens on a specific port, receives the data packets from the operation and maintenance terminal, and parses the content therein to extract the operation and maintenance information flow. For example, the operation and maintenance terminal can encode the operation and maintenance information flow in JSON format and send it to the server through an HTTP POST request. After receiving the request, the server decodes the JSON data to obtain the operation and maintenance information therein.

[0018] Alternatively, the operation and maintenance terminal sends the operation and maintenance information flow to a message queue, and the server subscribes to these messages from the message queue. The message queue can play a role in buffering and asynchronous processing to ensure the reliable transmission of the operation and maintenance information flow. For example, using a message queue system such as RabbitMQ or Kafka, the operation and maintenance terminal sends the operation and maintenance information flow to a specified queue, and the server obtains the messages from this queue for processing.

[0019] Step S200: If the operation and maintenance information flow includes a first fault mode, determine a first knowledge topology branch in the pre-deployed operation and maintenance knowledge topology structure. Each topology unit in the operation and maintenance knowledge topology structure represents a fault mode, each topology chain in the operation and maintenance knowledge topology structure represents the association of fault modes, the first fault mode belongs to the fault mode represented by the first topology unit in the operation and maintenance knowledge topology structure, and the first knowledge topology branch includes the first topology unit, topology units within a preset distance obtained based on the first topology unit, and the topology chains between the topology units.

[0020] The operation and maintenance knowledge topology structure is a topology structure that describes failure modes and their correlations, such as a graph structure, where each topology unit represents a failure mode and each topology chain represents the correlation between failure modes. The first failure mode is the failure mode contained in the operation and maintenance information flow and existing in the operation and maintenance knowledge topology structure, and the first knowledge topology branch is a sub-structure composed of the topology unit corresponding to the first failure mode, the topology units within the preset distance from it, and the topology chains between them. When determining the first knowledge topology branch, the server finds the topology unit corresponding to the first failure mode in the operation and maintenance knowledge topology structure, that is, the first topology unit. Then, according to the operation and maintenance knowledge topology structure, it obtains a group of topology units that are relevant to the first topology unit and within the preset distance. The preset distance is a pre-set parameter, and the distance can be the number of hops between two topology units, which is used to control the scope of the knowledge topology branch. For example, the failure modes that may be related to "excessive CPU usage of the router" include "excessive memory occupation of the router" and "abnormal port traffic of the router". The distance between the topology units corresponding to these failure modes and the topology unit corresponding to "excessive CPU usage of the router" is within the preset distance, and they form a group of topology units. The server can determine the first knowledge topology branch by using a graph search algorithm, such as breadth-first search (BFS) or depth-first search (DFS). Taking breadth-first search as an example, the server starts from the first topology unit, expands the search layer by layer, and adds the topology units within the preset distance to the group of topology units. During the search process, it traverses along the topology chain to determine the correlation between topology units.

[0021] In some cases, the operation and maintenance information flow may also contain the first failure mode correlation information. For example, in addition to indicating that the CPU usage of the router is too high, the operation and maintenance information flow also shows that there is a specific correlation between this failure and the abnormal port traffic of the router, that is, the first failure mode correlation. At this time, according to the operation and maintenance knowledge topology structure, the topology units in the group of topology units that have the first failure mode correlation with the first topology unit are further screened out to form the first sub-group of topology units. The server then establishes the first knowledge topology branch based on the first topology unit, the first sub-group of topology units, and the topology chains between them. By combining the failure modes and the failure mode correlations in the operation and maintenance knowledge topology structure, the server can more comprehensively understand the context information of the failures in the operation and maintenance information flow. This helps to overcome the problem that the operation and maintenance information flow may lead to inaccurate failure detection results due to insufficient instantaneous information volume, and improves the accuracy of operation and maintenance detection.

[0022] Step S300: Based on the first knowledge topology branch, obtain the first knowledge topology feature through the first embedding layer included in the operation and maintenance information processing algorithm.

[0023] The first knowledge topology branch is a sub-structure in the operation and maintenance knowledge topology structure, which is composed of topology units corresponding to the first failure mode, topology units within a preset distance, and the topology chains between them. It contains context information related to the first failure mode. The operation and maintenance information processing algorithm is a machine learning algorithm used by the server to process operation and maintenance information flows and analyze faults. Among them, the first embedding layer is a component of this algorithm, which is used to convert the information in the first knowledge topology branch into feature vectors that can be processed by a computer, that is, the first knowledge topology features.

[0024] The process by which the server obtains the first knowledge topology features using the first embedding layer is a process of numerically representing the discrete information in the first knowledge topology branch. This process can be implemented using graph embedding techniques, which can map the nodes (topology units) and edges (topology chains) in a graph structure (such as the first knowledge topology branch) into a low-dimensional vector space. For example, the server can use the DeepWalk algorithm to generate a node sequence on the first knowledge topology branch through random walks, and then input these sequences into a model similar to word embedding to learn the vector representation of each topology unit. Specifically, the server starts from a certain topology unit in the first knowledge topology branch and performs random walks. Each time it walks, it selects an adjacent topology unit to form a node sequence. This process is repeated multiple times to obtain multiple node sequences. Then, the server uses the Skip-Gram model to train these node sequences so that adjacent topology units are closer in the vector space. Finally, each topology unit is represented as a vector, and these vectors constitute the first knowledge topology features. Or, the Node2Vec algorithm can be used, which combines the strategies of depth-first search and breadth-first search to walk on the first knowledge topology branch and capture different types of relationships between topology units. When the server uses the Node2Vec algorithm, according to the preset parameters, it performs random walks on the first knowledge topology branch to generate node sequences. Different from DeepWalk, Node2Vec adjusts the probability of walking according to the local and global information of the nodes during the walking process, so as to learn a richer representation of topology units.

[0025] By converting the first knowledge topology branch into feature vectors, the server can use machine learning or deep learning models to further process and analyze these features, so as to more accurately identify failure modes and predict the development trend of faults. For example, the server can input the first knowledge topology features into a neural network model and combine other features of the operation and maintenance information flows to classify and diagnose faults, improving the accuracy and efficiency of fault analysis.

[0026] Step S400: Based on the operation and maintenance information flow, use the second embedding layer included in the operation and maintenance information processing algorithm to obtain operation and maintenance information features.

[0027] In step S400, the server obtains operation and maintenance information features based on the operation and maintenance information flow by using a second embedding layer included in the operation and maintenance information processing algorithm. The operation and maintenance information flow is a set of information uploaded by the operation and maintenance terminal that contains various types of data during the network operation and maintenance process. These data cover device operation parameters, fault prompt information, operation logs, etc. The operation and maintenance information processing algorithm is an algorithm used by the server to process and analyze the operation and maintenance information flow to achieve fault detection and diagnosis. The second embedding layer is a key component of this algorithm, and its role is to convert the original data in the operation and maintenance information flow into feature vectors that are easy for the computer to process and analyze, that is, operation and maintenance information features.

[0028] To obtain operation and maintenance information features, one method is to use word embedding technology. When the data in the operation and maintenance information flow contains text information, such as the text description in the operation log, the text information can be segmented, and then the Word2Vec algorithm can be used to map each word to a low-dimensional vector space. The Word2Vec algorithm trains a neural network model so that words with similar contexts in the text are close in the vector space. For example, in the operation log of a router, expressions such as "configure port" and "set port parameters" may often appear together. After training with the Word2Vec algorithm, their corresponding vectors will be relatively close in the vector space. After the server segments the text in the operation and maintenance information flow, it converts each word into a corresponding vector, and then combines these vectors into a feature vector representing the entire text through a certain aggregation method (such as taking the average).

[0029] For numerical data in the operation and maintenance information flow, such as the CPU usage rate and memory occupancy rate of a router, the server can use methods of normalization and feature scaling for processing. Scale these numerical data to a specific range, such as the interval [0, 1], to avoid the influence of some overly large numerical values on subsequent analysis. The server can use the min-max normalization formula: , where x is the original numerical value, and are the minimum and maximum values of this numerical value in the dataset respectively, is the normalized numerical value. The server arranges the normalized numerical data in a certain order to form a feature vector, which is used as part of the operation and maintenance information features.

[0030] The server can also use convolutional neural networks (CNNs) or recurrent neural networks (RNNs) in deep learning to process the operation and maintenance information flow. Taking CNN as an example, the server can organize the data in the operation and maintenance information flow into a two-dimensional or three-dimensional matrix, extract local features in the data through the convolutional layer, then reduce the dimensions of the features through the pooling layer, and finally combine the extracted features into a feature vector through the fully connected layer, that is, the operation and maintenance information feature.

[0031] By obtaining the operation and maintenance information feature, the server can transform the complex operation and maintenance information flow into a feature representation that is convenient for processing and analysis, providing a more effective data basis for subsequent fault analysis and diagnosis, and helping to improve the accuracy and efficiency of fault detection.

[0032] Step S500: Establish an aggregated feature based on the first knowledge topology feature and the operation and maintenance information feature.

[0033] As mentioned above, the first knowledge topology feature is obtained by the server based on the first knowledge topology branch through the first embedding layer of the operation and maintenance information processing algorithm, reflecting the context information related to the first fault mode; the operation and maintenance information feature is obtained by the server based on the operation and maintenance information flow through the second embedding layer of the operation and maintenance information processing algorithm, reflecting the specific data information in the operation and maintenance process. The aggregated feature is a new feature obtained by fusing these two features. It combines the context information of the fault mode and the operation and maintenance data information, helping to conduct more comprehensive and accurate fault analysis.

[0034] One way for the server to establish the aggregated feature is to achieve feature fusion through matrix operations. Multiply the operation and maintenance information feature by the first mapping parameter to obtain the first feature, multiply the first knowledge topology feature by the second mapping parameter and transpose the product to obtain the second feature, then multiply the first feature by the second feature to obtain the third feature, and finally add the adjustment feature to the third feature to obtain the aggregated feature. Here, the first mapping parameter and the second mapping parameter are the results of decomposing a multi-dimensional data structure. The multi-dimensional data structure is, for example, a tensor with a structure of e*g*f. The structure of the first mapping parameter is e*1*f, the structure of the second mapping parameter is 1*g*f, and the dimension of the aggregated feature is f. For example, if the dimension e of the operation and maintenance information feature is 10, the dimension g of the first knowledge topology feature is 20, and f is set to 5, the server performs matrix operations according to the above rules to obtain the aggregated feature.

[0035] Another way is to perform eigenvalue alignment processing. Each eigenvalue in the first knowledge topology feature is aligned with the eigenvalue in the operation and maintenance information feature, such as inner product calculation or summation, to obtain an aggregated feature, and the dimensions of the first knowledge topology feature, the operation and maintenance information feature, and the aggregated feature are equal. Suppose the dimensions of both the first knowledge topology feature and the operation and maintenance information feature are 15. The server performs inner product or addition operations on the corresponding position eigenvalues to obtain an aggregated feature with a dimension of 15 as well.

[0036] If the operation and maintenance information flow further includes a second failure mode, determine the second knowledge topology branch and obtain the second knowledge topology feature. At this time, when establishing the aggregated feature, the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature will be comprehensively considered. The server can first perform element-wise multiplication of the first knowledge topology feature and the second knowledge topology feature with the weight feature respectively, and then perform operations such as addition, averaging, or dot product on the results to obtain the target feature. Then, combined with the operation and maintenance information feature, the aggregated feature is obtained through matrix operations or feature splicing. For example, after performing element-wise addition to obtain the target feature, multiply the operation and maintenance information feature by the first mapping parameter, multiply the target feature by the second mapping parameter and transpose it, and add the adjustment feature after multiplying the two to obtain the aggregated feature.

[0037] Step S600: Based on the aggregated feature, use the fault identification layer included in the operation and maintenance information processing algorithm to obtain the fault analysis result of the operation and maintenance information flow, and transmit the fault analysis result to the operation and maintenance terminal.

[0038] The aggregated feature is a feature obtained by the server after fusing the first knowledge topology feature and the operation and maintenance information feature, or in the presence of a second failure mode, fusing the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature. It synthesizes the context information of the failure mode and the operation and maintenance data information, providing a more comprehensive and accurate basis for fault analysis. The fault identification layer in the operation and maintenance information processing algorithm is a module specifically used to analyze and process the aggregated feature to identify faults and give fault analysis results.

[0039] In the aforementioned step 200, the identified failure mode is only the manifestation form of the fault, which is an abstraction and generalization of the fault manifestation form, referring to the specific manifestation form of a specific problem that occurs in a device or system. For example, hardware failure, software crash, network connection interruption, etc. are all different types of failure modes. Identifying the failure mode is the first step in solving problems, which helps technicians quickly locate the possible problem range. After parsing through step 600, the obtained fault analysis result is a comprehensive and accurate fault type. In other words, the failure mode is a discrete fault type identified through pattern matching, belonging to atomic problem labels (such as: "hard disk IO overlimit", "memory leak"). And the fault analysis result is a composite diagnostic conclusion generated based on topological relevance and context features.

[0040] In an enterprise network environment, operation and maintenance personnel use operation and maintenance terminals to perform daily inspections and configuration adjustments on routers. The operation and maintenance terminals record the operating parameters, operation information, etc. of the routers and upload them to the server. The server determines that the first failure mode is "excessive CPU usage of the router", obtains the first knowledge topology feature and operation and maintenance information feature through a series of steps. If there is a second failure mode, it will also obtain the second knowledge topology feature, and then establish an aggregated feature. The server uses the failure identification layer to analyze the aggregated feature to obtain a failure analysis result.

[0041] The failure identification layer can use a variety of technical means to achieve failure analysis. One method is to use machine learning classification algorithms, such as support vector machines (SVM), decision trees, random forests, etc. Deep learning models can also be used in the failure identification layer, such as multi-layer perceptrons (MLP). The MLP consists of an input layer, a hidden layer, and an output layer. The server inputs the aggregated feature into the input layer of the MLP, and through the non-linear transformation and calculation of the hidden layer, finally obtains the prediction probability of each failure mode at the output layer. The server can select the failure mode with the highest probability as the failure analysis result.

[0042] After obtaining the failure analysis result, the server transmits it to the operation and maintenance terminal.

[0043] Through step 600, the server can accurately analyze the failure mode in the operation and maintenance information flow using the aggregated feature and the failure identification layer, and promptly feedback the result to the operation and maintenance terminal, helping the operation and maintenance personnel quickly understand the failure situation of the device, take corresponding maintenance measures, improve the efficiency and accuracy of network operation and maintenance, and complete remote maintenance.

[0044] In one implementation solution, after obtaining the operation and maintenance information flow to be processed uploaded by the operation and maintenance terminal, the method provided by the embodiments of the present application may further include: Step 101: Detect the failure mode of the operation and maintenance information flow to obtain a failure mode detection result; Step 102: If the failure mode detection result indicates that the operation and maintenance information flow includes a failure mode, compare the failure mode included in the operation and maintenance information flow with the failure modes in the operation and maintenance knowledge topology structure to obtain a failure mode comparison result; Step 103: If the failure mode comparison result indicates that the failure mode included in the operation and maintenance information flow exists in the failure modes in the operation and maintenance knowledge topology structure, use the existing failure mode as the first failure mode.

[0045] In step 101, the server performs a fault mode detection on the operation and maintenance information flow to obtain a fault mode detection result. The operation and maintenance information flow is an information set uploaded by the operation and maintenance terminal and contains various data during the network operation and maintenance process. These data cover the operating parameters of the device, operation logs, fault prompts, and other contents. The fault mode detection refers to the server using specific methods and technologies to identify possible fault modes from the operation and maintenance information flow.

[0046] In an enterprise network environment, operation and maintenance personnel use the operation and maintenance terminal to perform daily inspections and configuration adjustments on the router. The operation and maintenance terminal records the operating parameters of the router, such as CPU usage rate, memory occupancy, port connection status, etc., as well as the configuration modification operations performed by the operation and maintenance personnel, such as modifying the access control list, adjusting the routing policy, etc. The recorded data is uploaded to the server in the form of the operation and maintenance information flow. When the server performs a fault mode detection on the operation and maintenance information flow, a rule-based detection method can be adopted. The server pre-sets a series of rules. For example, when the CPU usage rate of the router continuously exceeds 80% for more than 10 minutes, it is determined as the "CPU usage rate too high" fault mode; when the packet loss rate of a certain port of the router exceeds 5%, it is determined as the "port packet loss abnormal" fault mode. The server compares the data in the operation and maintenance information flow with these rules to determine whether there is a fault mode and the specific type of the fault mode.

[0047] The server can also use machine learning algorithms for fault mode detection. In the training stage, a large amount of historical operation and maintenance information flow data with fault mode labels is used to train the machine learning model, enabling the model to learn the operation and maintenance data characteristics corresponding to different fault modes. In actual applications, the server inputs the current operation and maintenance information flow into the trained model, and the model determines whether there is a fault mode and the type of the fault mode based on the characteristics.

[0048] In step 102, when the fault mode detection result indicates that the operation and maintenance information flow includes a fault mode, the server compares the fault mode included in the operation and maintenance information flow with the fault modes in the operation and maintenance knowledge topology structure to obtain a fault mode comparison result. The operation and maintenance knowledge topology structure is a topology structure pre-deployed by the server for describing fault modes and their correlations. Each topology unit represents a fault mode, and each topology chain represents the correlation between fault modes.

[0049] In the above example of the enterprise network environment, the server detects through the failure mode detection that there are failure modes such as "too high CPU usage rate" and "abnormal port packet loss" in the operation and maintenance information flow. The server compares these failure modes with the failure modes in the operation and maintenance knowledge topology. The operation and maintenance knowledge topology is constructed by the server based on a large amount of historical operation and maintenance data and failure cases, which contains the possible failure modes of various network devices and the correlation relationships between them. Check whether there are the same failure modes as "too high CPU usage rate" and "abnormal port packet loss" in the operation and maintenance knowledge topology. If they exist, it means that the failure modes in the operation and maintenance information flow match the failure modes in the operation and maintenance knowledge topology; if not, it means they do not match.

[0050] To achieve the comparison of failure modes, a string matching method can be adopted to compare the failure mode names in the operation and maintenance information flow with the failure mode names represented by each topology unit in the operation and maintenance knowledge topology. If the names are exactly the same, it is determined to be a match. The server can also use a semantic matching method. When the failure mode names are not exactly the same, it determines whether they match by analyzing whether their semantics are similar. For example, what is recorded in the operation and maintenance information flow is "too high router CPU load", while what corresponds in the operation and maintenance knowledge topology is "too high router CPU usage rate". Through semantic analysis, it can be determined that these two failure modes match.

[0051] In step 103, when the failure mode comparison result indicates that the failure modes included in the operation and maintenance information flow exist in the failure modes of the operation and maintenance knowledge topology, the existing failure modes are used as the first failure modes. The first failure modes are the key basis for subsequent analysis and processing. The server will determine the first knowledge topology branch in the operation and maintenance knowledge topology based on the first failure modes, and then conduct a more in-depth failure analysis.

[0052] In the above example, the server finds through the failure mode comparison that both the failure modes of "too high CPU usage rate" and "abnormal port packet loss" exist in the operation and maintenance knowledge topology. The server can choose one of them as the first failure mode. It can be selected according to a preset rule. For example, the failure mode with the largest support coefficient is selected as the first failure mode. The support coefficient can represent the frequency of occurrence of this failure mode in historical data or the degree of impact on the system, etc. If the support coefficient of "too high CPU usage rate" is greater than the support coefficient of "abnormal port packet loss", then the server takes "too high CPU usage rate" as the first failure mode.

[0053] Through steps 101 to 103, the server accurately identifies the failure modes from the operation and maintenance information flow, compares them with the operation and maintenance knowledge topology, determines the first failure modes, and provides a basis for subsequent failure analysis and processing.

[0054] When the server performs fault mode detection, it can also combine multiple detection methods to improve the accuracy and reliability of detection. For example, it first uses a rule-based detection method for preliminary screening, and then uses machine learning algorithms for further verification and refinement. The server can also adjust the rules and model parameters of fault mode detection according to different device types and operation and maintenance scenarios to meet different requirements. During the fault mode comparison process, a thesaurus and a near-synonym thesaurus of fault modes can be established to improve the accuracy of semantic matching. When the fault mode name in the operation and maintenance information flow is not exactly the same as the fault mode name in the operation and maintenance knowledge topology structure, the server can determine whether there is a match by querying the thesaurus and the near-synonym thesaurus. The server can also introduce natural language processing technology to conduct a more in-depth analysis and understanding of the description of the fault mode to improve the accuracy of the comparison.

[0055] After determining the first fault mode, the server determines the first knowledge topology branch in the operation and maintenance knowledge topology structure based on the first fault mode. The first knowledge topology branch contains the context information related to the first fault mode, which helps the server to more comprehensively understand the cause of the fault and the possible impacts.

[0056] In one implementation solution, step 102 above, that is, if the fault mode detection result indicates that the operation and maintenance information flow includes a fault mode, then compare the fault mode included in the operation and maintenance information flow with the fault mode in the operation and maintenance knowledge topology structure to obtain a fault mode comparison result, including: Step 1021: If the fault mode detection result indicates that the operation and maintenance information flow includes several fault modes, obtain the support coefficient corresponding to each fault mode among the several fault modes; Step 1022: Use the fault mode corresponding to the maximum support coefficient among the several fault modes as the fault mode included in the operation and maintenance information flow; Step 1023: Compare the fault mode included in the operation and maintenance information flow with the fault mode in the operation and maintenance knowledge topology structure to obtain a fault mode comparison result.

[0057] Specifically, if the fault mode detection result indicates that the operation and maintenance information flow includes several fault modes, the server obtains the support coefficient corresponding to each fault mode among the several fault modes. The support coefficient is an index that measures the possibility or credibility of the fault mode appearing in the operation and maintenance information flow, and reflects the probability of the existence of this fault mode judged according to the existing data and model.

[0058] To obtain the support coefficient for each failure mode, one method is a machine learning-based classification algorithm such as the Naive Bayes classifier. The Naive Bayes classifier is based on Bayes' theorem and the assumption of feature conditional independence. By learning from a large amount of historical operation and maintenance data, it establishes a probability model between failure modes and features. For a new operation and maintenance information flow, the server can calculate the posterior probability of each failure mode according to this model and use the posterior probability as the support coefficient. The specific formula is as follows: Let C be the failure mode category, and F be the features in the operation and maintenance information flow. Then, according to Bayes' theorem, the posterior probability of the failure mode C is: ; Since the Naive Bayes classifier assumes conditional independence between features, that is , the above formula can be simplified to: ; where P(C) is the prior probability of the failure mode C, which can be obtained through historical data statistics; is the probability that the feature F i appears under the failure mode C and can also be learned from historical data; is the same for all failure modes and can be ignored when comparing the posterior probabilities of different failure modes. Therefore, the server can calculate the posterior probability of each failure mode according to the above formula as the support coefficient.

[0059] Another method is a deep learning-based neural network model such as a Convolutional Neural Network (CNN) or a Recurrent Neural Network (RNN). For example, for a CNN model, the server takes the operation and maintenance information flow as input, processes it through convolutional layers, pooling layers, and fully connected layers, and finally outputs the probability distribution of each failure mode through the softmax function. These probabilities are the support coefficients.

[0060] After obtaining the support coefficient for each failure mode, the server executes step 1022 and takes the failure mode corresponding to the maximum support coefficient among several failure modes as the failure mode included in the operation and maintenance information flow. This is because the maximum support coefficient indicates that this failure mode has the highest possibility of occurring in the current operation and maintenance information flow and is most likely to be the actual existing failure mode.

[0061] Finally, the server executes step 1023 to compare the fault modes included in the operation and maintenance information flow with the fault modes in the operation and maintenance knowledge topology structure to obtain the fault mode comparison result. Through the above steps 1021-1023, the server accurately determines the most likely fault mode from the operation and maintenance information flow containing multiple fault modes, and compares it with the operation and maintenance knowledge topology structure, providing a basis for subsequent fault analysis and processing. This method can improve the accuracy of fault detection, overcome the problem that the operation and maintenance information flow causes inaccurate fault detection results due to insufficient instantaneous information volume, and combine the fault modes and fault mode correlations in the operation and maintenance knowledge topology structure to provide more comprehensive information for fault analysis.

[0062] As an implementation solution of step 200, that is, if the operation and maintenance information flow includes a first fault mode, determine a first knowledge topology branch in the pre-deployed operation and maintenance knowledge topology structure, including: Step 201: If the operation and maintenance information flow includes a first fault mode, use the topology unit corresponding to the first fault mode in the operation and maintenance knowledge topology structure as the first topology unit; Step 202: According to the operation and maintenance knowledge topology structure, obtain a topology unit group that has a correlation with the first topology unit, where the distance between each topology unit in the topology unit group and the first topology unit is within a preset distance; Step 203: Establish a first knowledge topology branch based on the first topology unit, the topology unit group, and the topology chain between the topology units.

[0063] In step 201, when the operation and maintenance information flow includes a first fault mode, the server uses the topology unit corresponding to the first fault mode in the operation and maintenance knowledge topology structure as the first topology unit. For example, if the server determines through the analysis of the operation and maintenance information flow that the first fault mode is "too high router CPU usage rate", then search for the topology unit corresponding to "too high router CPU usage rate" in the operation and maintenance knowledge topology structure and designate it as the first topology unit. As mentioned above, the server can search for the corresponding topology unit by name matching, compare the name of the first fault mode with the name of the fault mode represented by each topology unit in the operation and maintenance knowledge topology structure one by one. If the names are exactly the same, determine that this topology unit is the first topology unit. For some cases where the expressions are similar but not exactly the same, the server can use semantic analysis technology to determine whether they represent the same type of fault mode. For example, "too high router CPU load" and "too high router CPU usage rate" can be determined to correspond to the same topology unit through semantic analysis.

[0064] In step 202, the server obtains a group of topological units related to the first topological unit according to the operation and maintenance knowledge topology structure, where the distance between each topological unit in the group of topological units and the first topological unit is within a preset distance. The preset distance is a pre-set parameter used to control the scope of the knowledge topology branches, and it determines the breadth of the server's search for relevant topological units in the operation and maintenance knowledge topology structure. Taking the first topological unit of "the CPU usage rate of the router is too high" as an example, the related failure modes may include "the memory occupancy of the router is too large" and "the port traffic of the router is abnormal". Starting from the first topological unit, search in the operation and maintenance knowledge topology structure to find topological units that are connected to the first topological unit through a topological chain and whose distance is within the preset distance. The server can use a graph search algorithm to implement this process. For example, starting from the first topological unit, expand the search scope layer by layer, and add the topological units within the preset distance from the first topological unit to the group of topological units in turn. During the search process, the server traverses along the topological chain and determines which topological units are related to the first topological unit according to the failure mode correlation represented by the topological chain.

[0065] In step 203, the server establishes a first knowledge topology branch according to the first topological unit, the group of topological units, and the topological chain between the topological units. The first knowledge topology branch is a sub-structure that contains context information related to the first failure mode and helps the server understand the cause and possible impact of the failure more comprehensively. The server combines the first topological unit, each topological unit in the group of topological units, and the topological chain between them to form a complete knowledge topology branch. In the above example, the first topological unit is the topological unit corresponding to "the CPU usage rate of the router is too high", and the group of topological units includes the topological units corresponding to "the memory occupancy of the router is too large", "the port traffic of the router is abnormal", etc. Integrate these topological units and the topological chain between them to construct the first knowledge topology branch. This branch not only shows the first failure mode itself but also shows other related failure modes and their relationships.

[0066] As another implementation solution of step 200, that is, if the operation and maintenance information flow includes the first failure mode, determine the first knowledge topology branch in the pre-deployed operation and maintenance knowledge topology structure, including: Step 210: If the operation and maintenance information flow includes the first failure mode, use the topological unit corresponding to the first failure mode in the operation and maintenance knowledge topology structure as the first topological unit; Step 220: According to the operation and maintenance knowledge topology structure, obtain a group of topological units related to the first topological unit, where the distance between each topological unit in the group of topological units and the first topological unit is within a preset distance; Step 230: If the operation and maintenance information flow includes the first failure mode correlation, obtain, according to the operation and maintenance knowledge topology structure, a first sub-topology unit group in the topology unit group that has the first failure mode correlation with the first topology unit. Step 240: Establish a first knowledge topology branch according to the first topology unit, the first sub-topology unit group, and the topology chain between the topology units.

[0067] In Step 210, when the operation and maintenance information flow includes the first failure mode, the server uses the topology unit corresponding to the first failure mode in the operation and maintenance knowledge topology structure as the first topology unit.

[0068] In Step 220, the server obtains, according to the operation and maintenance knowledge topology structure, a topology unit group that has a correlation with the first topology unit, where the distance between each topology unit in the topology unit group and the first topology unit is within a preset distance. The preset distance is a parameter set in advance, and its function is to control the scope of the knowledge topology branch, and it determines the breadth of the server's search for relevant topology units in the operation and maintenance knowledge topology structure.

[0069] In Step 230, when the operation and maintenance information flow includes the first failure mode correlation, the server obtains, according to the operation and maintenance knowledge topology structure, a first sub-topology unit group in the topology unit group that has the first failure mode correlation with the first topology unit. The first failure mode correlation refers to the correlation relationship between the first failure mode and other failure modes, and this correlation relationship is reflected by the topology chain in the operation and maintenance knowledge topology structure. Suppose that in addition to indicating the first failure mode of "the CPU usage rate of the router is too high" in the operation and maintenance information flow, it also shows that this failure has a specific correlation with "the router port traffic is abnormal", that is, the first failure mode correlation. Among the previously obtained topology unit group, according to the operation and maintenance knowledge topology structure, filter out the topology units that have the first failure mode correlation of "related to the abnormal router port traffic" with the first topology unit "the CPU usage rate of the router is too high" to form a first sub-topology unit group. The server can determine this correlation by checking the attributes of the topology chain. Each topology chain is set with corresponding labels or attributes to represent the correlation type it represents, and the server judges whether there is a first failure mode correlation between the topology units according to these labels or attributes. To more accurately determine the first sub-topology unit group, the server can use machine learning algorithms. The server can use a large amount of historical data with correlation labels to train the machine learning model in the training stage, so that the model learns the correlation characteristics between different failure modes. In actual applications, the server inputs the topology units in the topology unit group and the first failure mode correlation information into the trained model, and the model will judge which topology units have the first failure mode correlation with the first topology unit according to the characteristics.

[0070] When establishing the first knowledge topology branch, the server can also assign weights to topology units and topology chains. For topology units and topology chains that are more strongly associated with the first failure mode, higher weights can be assigned, so that in subsequent fault analysis, these important pieces of information will receive more attention. The server can determine the weight values based on factors such as the association frequency and impact degree between failure modes in historical data.

[0071] In step 240, the server establishes the first knowledge topology branch based on the first topology unit, the first sub-topology unit group, and the topology chains between the topology units. The first knowledge topology branch is a sub-structure that contains the context information related to the first failure mode, which helps the server to more comprehensively understand the causes and possible impacts of the fault. The server combines the first topology unit, each topology unit in the first sub-topology unit group, and the topology chains between them to construct a complete knowledge topology branch. In the above example, the first topology unit is the topology unit corresponding to "the CPU usage rate of the router is too high", and the first sub-topology unit group contains topology units that have an association relationship of "related to abnormal router port traffic" with "the CPU usage rate of the router is too high", such as the topology unit corresponding to "abnormal router port traffic". Integrating these topology units and the topology chains between them forms the first knowledge topology branch. This branch not only shows the first failure mode itself, but also shows other failure modes that have specific association relationships with it and the association situations between them. For example, through the topology chain, it can be known that "the CPU usage rate of the router is too high" may be caused by "abnormal router port traffic", which provides more targeted information for the server's subsequent fault analysis.

[0072] As another implementation solution for step 200, that is, if the operation and maintenance information flow includes the first failure mode, determining the first knowledge topology branch in the pre-deployed operation and maintenance knowledge topology structure includes: Step 2100: If the operation and maintenance information flow includes the first failure mode, then use the topology unit corresponding to the first failure mode in the operation and maintenance knowledge topology structure as the first topology unit; Step 2200: According to the operation and maintenance knowledge topology structure, obtain a topology unit group that contains a correlation with the first topology unit, where the distance between each topology unit in the topology unit group and the first topology unit is within a preset distance; Step 2300: If the operation and maintenance information flow includes the first failure mode association and the second failure mode association, then according to the operation and maintenance knowledge topology structure, obtain a second sub-topology unit group in the topology unit group that has at least one of the first failure mode association or the second failure mode association with the first topology unit; Step 2400: Establish a first knowledge topology branch based on the first topology unit, the second sub-topology unit group, and the topology chain between topology units.

[0073] In step 2100, when the operation and maintenance information flow includes the first failure mode, the server uses the topology unit corresponding to the first failure mode in the operation and maintenance knowledge topology structure as the first topology unit.

[0074] In step 2200, the server obtains a topology unit group that has an inclusion relationship with the first topology unit according to the operation and maintenance knowledge topology structure, where the distance between each topology unit in the topology unit group and the first topology unit is within a preset distance. The preset distance is a pre-set parameter, and its function is to control the scope of the knowledge topology branch and determine the breadth of the server's search for relevant topology units in the operation and maintenance knowledge topology structure.

[0075] In step 2300, when the operation and maintenance information flow includes the first failure mode correlation and the second failure mode correlation, the server obtains, according to the operation and maintenance knowledge topology structure, a second sub-topology unit group in the topology unit group that has at least one of the first failure mode correlation or the second failure mode correlation with the first topology unit. The first failure mode correlation and the second failure mode correlation refer to specific association relationships between the first failure mode and other failure modes, and these association relationships are reflected by topology chains in the operation and maintenance knowledge topology structure. Suppose that in addition to indicating the first failure mode of "excessive CPU usage of the router" in the operation and maintenance information flow, it also shows that this failure has a first failure mode correlation with "abnormal router port traffic" and a second failure mode correlation with "excessive router memory occupancy". In the previously obtained topology unit group, according to the operation and maintenance knowledge topology structure, filter out the topology units that have at least one of the first failure mode correlation of "related to abnormal router port traffic" or the second failure mode correlation of "related to excessive router memory occupancy" with the first topology unit of "excessive CPU usage of the router" to form a second sub-topology unit group. The server can determine this correlation by checking the attributes of the topology chain. Each topology chain may have corresponding tags or attributes to represent the type of association it represents, and the server determines whether there is a corresponding failure mode correlation between topology units based on these tags or attributes.

[0076] In step 2400, the server establishes the first knowledge topology branch based on the first topology unit, the second sub-topology unit group, and the topology chains between the topology units. The first knowledge topology branch is a sub-structure that contains the context information related to the first failure mode, which helps the server to more comprehensively understand the causes and possible impacts of the failure. The server combines the first topology unit, each topology unit in the second sub-topology unit group, and the topology chains between them to construct a complete knowledge topology branch. In the above example, the first topology unit is the topology unit corresponding to "the CPU usage rate of the router is too high", and the second sub-topology unit group contains topology units that have an associated relationship with "the CPU usage rate of the router is too high" such as "related to abnormal router port traffic" or "related to excessive router memory occupancy", such as the topology units corresponding to "abnormal router port traffic" and "excessive router memory occupancy". Integrating these topology units and the topology chains between them forms the first knowledge topology branch. This branch not only shows the first failure mode itself, but also shows other failure modes that have specific associated relationships with it and the association situations between them.

[0077] Through steps 2100 - 2400, the server can accurately determine the first knowledge topology branch in the operation and maintenance knowledge topology structure, which combines the first failure mode and other failure modes that have two different associated relationships with it, providing more comprehensive and detailed context information for subsequent failure analysis.

[0078] In step 2300, in order to more accurately determine the second sub-topology unit group, the server can adopt machine learning algorithms. The server can use a large amount of historical data with associated labels to train the machine learning model during the training phase, enabling the model to learn the associated characteristics between different failure modes. In actual applications, the server inputs the topology units in the topology unit group and the associated information of the first and second failure modes into the trained model, and the model will judge which topology units have the corresponding failure mode association with the first topology unit according to the characteristics.

[0079] When establishing the first knowledge topology branch, the server can also assign weights to the topology units and topology chains. For topology units and topology chains that have a stronger association with the first failure mode, higher weights can be assigned, so that these important information will receive more attention in subsequent failure analysis. The server can determine the weight values according to factors such as the association frequency and impact degree between failure modes in historical data. For example, if the association frequency between "the CPU usage rate of the router is too high" and "abnormal router port traffic" is high in historical data, and this association has a greater impact on system performance, then the topology unit corresponding to "abnormal router port traffic" and the topology chain connected to it can be assigned a higher weight.

[0080] Steps 2100 - 2400 provide a more detailed and comprehensive method for the server to determine the first knowledge topology branch in the operation and maintenance knowledge topology structure. By combining the first failure mode and two different failure mode correlations, the server can construct a knowledge topology branch that better conforms to the actual failure situation, thereby providing stronger support for failure analysis and handling, improving the efficiency and quality of network operation and maintenance, and ensuring the stable operation of the enterprise network.

[0081] In the artificial intelligence-based fault operation and maintenance data transmission method, step 200 requires determining the first knowledge topology branch in the pre-deployed operation and maintenance knowledge topology structure when the operation and maintenance information flow includes the first failure mode. There are three different implementation schemes for this step, and the following will compare and introduce these three schemes, analyze their respective advantages, similarities, and differences.

[0082] The core objectives of the above three schemes are the same, which is to determine the first knowledge topology branch from the operation and maintenance knowledge topology structure when the operation and maintenance information flow contains the first failure mode. Moreover, all are based on the operation and maintenance knowledge topology structure. First, determine the first topology unit, and then search for relevant topology units around the first topology unit to finally construct the first knowledge topology branch. When obtaining the topology unit group, the distance between the topology unit and the first topology unit is considered, and it is required that the distance is within the preset distance to ensure that the obtained topology unit has a certain correlation with the first topology unit.

[0083] The difference lies in the degree of utilization of the failure mode correlation. Scheme one does not consider the failure mode correlation and only determines the relevant topology unit group based on the distance between the topology units to construct the first knowledge topology branch; Scheme two considers one failure mode correlation, that is, the first failure mode correlation, and further screens out the first sub-topology unit group that has this correlation with the first topology unit on the basis of the topology unit group; Scheme three considers two failure mode correlations, that is, the first failure mode correlation and the second failure mode correlation, and screens out the second sub-topology unit group that has at least one correlation with the first topology unit from the topology unit group.

[0084] Scheme one does not need to consider the failure mode correlation, reducing the computational complexity and data processing volume. In the case where the failure mode correlation is not obvious or difficult to accurately obtain in the operation and maintenance knowledge topology structure, this scheme can quickly determine the first knowledge topology branch and improve the processing efficiency. For example, in some newly established operation and maintenance systems, the data of the failure mode correlation is not yet perfect, and in this case, scheme one can be used to process the operation and maintenance information flow in a timely manner.

[0085] Solution 2 takes into account the relevance of the first failure mode and can more accurately screen out the topology units related to the first topology unit, and the constructed first knowledge topology branch is more targeted. In scenarios where the relevance of failure modes is relatively clear and a high precision of failure analysis is required, Solution 2 can provide more accurate context information, which helps to improve the accuracy of failure detection and analysis. For example, in a specific network operation and maintenance scenario, if it is known that a certain failure mode has a specific relevance to some other failure modes, Solution 2 can make better use of these relevances for failure analysis.

[0086] Solution 3 takes into account the relevance of two failure modes, further refines the screening conditions for topology units, and can construct a more accurate and comprehensive first knowledge topology branch. In scenarios where the relevance of failure modes is complex and diverse and extremely high requirements are placed on failure analysis, Solution 3 can make full use of various relevance information to provide richer context features, thereby analyzing failures more accurately. For example, in the operation and maintenance of a large industrial production system, the relevance between failure modes is complex and changeable, and using Solution 3 can more effectively mine failure information and improve the accuracy of failure diagnosis.

[0087] In an implementation of step 500, that is, according to the first knowledge topology feature and the operation and maintenance information feature, an aggregated feature is established, including: Step 510: Multiply the operation and maintenance information feature by the first mapping parameter to obtain a first feature, where the dimension of the operation and maintenance information feature is e; Step 520: Multiply the first knowledge topology feature by the second mapping parameter and transpose the product to obtain a second feature, where the feature of the first knowledge topology feature is g; Step 530: Multiply the first feature by the second feature to obtain a third feature; Step 540: Add the adjustment feature to the third feature to obtain the aggregated feature.

[0088] Among them, the first mapping parameter and the second mapping parameter are the results of decomposing a multi-dimensional data structure. The structure of the multi-dimensional data structure is e*g*f, the structure of the first mapping parameter is e*1*f, the structure of the second mapping parameter is 1*g*f, and the dimension of the aggregated feature is f.

[0089] In step 510, the server multiplies the operation and maintenance information feature by the first mapping parameter to obtain the first feature, where the dimension of the operation and maintenance information feature is e. The operation and maintenance information feature is obtained by the server based on the operation and maintenance information flow using the second embedding layer included in the operation and maintenance information processing algorithm, and it reflects the specific data information in the operation and maintenance process. The first mapping parameter is a hyperparameter group with a structure of e*1*f, where e represents the dimension of the operation and maintenance information feature, and f is a preset dimension value used to control the dimension of the final aggregated feature. Assuming the operation and maintenance information feature is represented as a vector of dimension e, the server multiplies this vector by the first mapping parameter to get a new vector, that is, the first feature. For example, if the dimension e of the operation and maintenance information feature is 10 and the structure of the first mapping parameter is 10*1*5, the server multiplies the operation and maintenance information feature vector by the first mapping parameter to obtain a first feature vector of dimension 5. The server can use the basic rules of matrix multiplication to implement this operation, and for the calculation of each element, it is carried out according to the formula of matrix multiplication. Let the operation and maintenance information feature vector be , the first mapping parameter be M1, and the first feature vector be Y1, then each element y 1i of Y1 is calculated as .

[0090] In step 520, the server multiplies the first knowledge topology feature by the second mapping parameter and transposes the product to obtain the second feature, where the dimension of the first knowledge topology feature is g. The first knowledge topology feature is obtained by the server based on the first knowledge topology branch using the first embedding layer included in the operation and maintenance information processing algorithm, and it reflects the context information related to the first failure mode. The second mapping parameter is a hyperparameter matrix with a structure of 1*g*f, and g represents the dimension of the first knowledge topology feature. The server multiplies the first knowledge topology feature vector by the second mapping parameter to get an intermediate result, and then transposes this intermediate result to obtain the second feature. For example, if the dimension g of the first knowledge topology feature is 20 and the structure of the second mapping parameter is 1*20*5, the server multiplies the first knowledge topology feature vector by the second mapping parameter to get an intermediate vector, and then transposes it to finally obtain a second feature vector of dimension 5. Let the first knowledge topology feature vector be , the second mapping parameter be M2, and the intermediate result vector be Y mid , then each element y mid of Y midi is calculated as , and the second feature vector Y2 is the transpose of Y mid .

[0091] In step 530, the server multiplies the first feature by the second feature to obtain a third feature. The server performs a multiplication operation on the first feature vector obtained in step 510 and the second feature vector obtained in step 520. Here, the multiplication can be the dot product operation of vectors.

[0092] In step 540, the server adds the adjustment feature to the third feature to obtain an aggregated feature, where the dimension of the aggregated feature is f. The adjustment feature is also the bias term, and its role is to fine-tune the third feature to improve the expression ability of the aggregated feature. The server performs an element-wise addition of the third feature vector obtained in step 530 and the adjustment feature vector to obtain the final aggregated feature vector.

[0093] The first mapping parameter and the second mapping parameter here are the results of decomposing a multi-dimensional data structure, and the structure of the multi-dimensional data structure is e*g*f. By decomposing this multi-dimensional data structure into the first mapping parameter and the second mapping parameter, the server can perform flexible transformations and fusions between different dimensions, thereby effectively combining the operation and maintenance information features and the first knowledge topology features. In practical applications, the values of the first mapping parameter and the second mapping parameter are learned through an optimization algorithm during the training process of the operation and maintenance information processing algorithm, and the values of these parameters are continuously adjusted according to a large amount of training data to make the final aggregated feature better used for fault analysis.

[0094] The aggregated feature constructed through steps 510 - 540 combines the information of the operation and maintenance information features and the first knowledge topology features, including both the specific data information in the operation and maintenance process and the context information related to the first fault mode. This helps the server to more comprehensively and accurately identify the fault mode and predict the fault development trend in subsequent fault analysis. For example, when analyzing router faults, the operation and maintenance information features may reflect the current operating parameters of the router, such as CPU usage, memory occupancy, etc., while the first knowledge topology features reflect other possible fault modes related to this fault mode and their correlation relationships. By constructing the aggregated feature, the server can fuse this information to more accurately determine the cause and possible impact range of the router fault.

[0095] In another implementation of step 500, that is, according to the first knowledge topology feature and the operation and maintenance information feature, an aggregated feature is established, including: Step 501: Perform a pairwise process on each feature value in the first knowledge topology feature and the feature value in the operation and maintenance information feature to obtain an aggregated feature; the pairwise process is an inner product solution or a sum, and the dimensions of the first knowledge topology feature, the operation and maintenance information feature, and the aggregated feature are equal.

[0096] As another implementation of step 500, the server performs a pairwise processing on each eigenvalue in the first knowledge topology feature and the eigenvalue in the operation and maintenance information feature to obtain an aggregated feature. The pairwise processing method is inner product solution or summation, and at the same time, the first knowledge topology feature, the operation and maintenance information feature, and the aggregated feature have the same dimension. In an enterprise network environment, operation and maintenance personnel use an operation and maintenance terminal to perform daily inspections and configuration adjustments on the router. The operation and maintenance terminal records relevant information and uploads it to the server. After a series of operations, the server obtains the first knowledge topology feature and the operation and maintenance information feature, and then constructs the aggregated feature according to step 501.

[0097] The first knowledge topology feature is obtained by the server based on the first knowledge topology branch through the first embedding layer included in the operation and maintenance information processing algorithm. It reflects the context information related to the first failure mode and is a numerical representation of the failure mode and its associated relationships. The operation and maintenance information feature is obtained by the server based on the operation and maintenance information flow through the second embedding layer included in the operation and maintenance information processing algorithm, which reflects the specific data information in the operation and maintenance process, such as the CPU usage rate and memory occupancy of the router.

[0098] If the pairwise processing method of inner product solution is adopted, assuming that the first knowledge topology feature is represented as a vector , and the operation and maintenance information feature is represented as a vector , where n is the dimension of the feature. Then each element a i of the aggregated feature vector A is calculated as . If the pairwise processing method of summation is adopted, each element a i of the aggregated feature vector A is calculated as .

[0099] The server can implement this pairwise processing through a loop traversal operation. The server sequentially extracts the eigenvalues at the corresponding positions in the first knowledge topology feature and the operation and maintenance information feature, and calculates the eigenvalues at the corresponding positions in the aggregated feature according to the rules of inner product solution or summation. This processing method is simple and efficient, and can quickly fuse the first knowledge topology feature and the operation and maintenance information feature to obtain the aggregated feature.

[0100] The aggregated feature constructed through step 501 combines the information of the first knowledge topology feature and the operation and maintenance information feature, including both the context information of the failure mode and the specific data information in the operation and maintenance process. This helps the server to more comprehensively and accurately identify the failure mode and predict the failure development trend in subsequent failure analysis.

[0101] In one implementation, the method provided by the embodiments of the present invention may further include: Step 200A: If the operation and maintenance information flow further includes a second failure mode, determine a second knowledge topology branch in the operation and maintenance knowledge topology structure, where the second failure mode belongs to the failure modes characterized by the second topology unit in the operation and maintenance knowledge topology structure, and the second knowledge topology branch includes the second topology unit, the topology units within a preset distance obtained based on the second topology unit, and the topology chains between the topology units; Step 200B: Based on the second knowledge topology branch, use the first embedding layer included in the operation and maintenance information processing algorithm to obtain the second knowledge topology feature; At this time, in step 500, according to the first knowledge topology feature and the operation and maintenance information feature, the aggregated feature can be established, which may include: Step 500A: Obtain the aggregated feature according to the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature.

[0102] In step 200A, when the operation and maintenance information flow further includes a second failure mode, the server determines a second knowledge topology branch in the operation and maintenance knowledge topology structure. The operation and maintenance knowledge topology structure is constructed by the server based on a large amount of historical operation and maintenance data and failure cases. Each topology unit represents a failure mode, and the topology chain represents the correlation between failure modes. For example, after analyzing the operation and maintenance information flow, the server determines that the first failure mode is "too high CPU usage of the router", and the second failure mode is "too large memory occupancy of the router". Find the topology unit corresponding to "too large memory occupancy of the router" in the operation and maintenance knowledge topology structure, that is, the second topology unit. Then, the server obtains the topology units related to the second topology unit and within the preset distance according to the preset distance. These topology units and the topology chains between them together constitute the second knowledge topology branch.

[0103] In step 200B, the server uses the first embedding layer included in the operation and maintenance information processing algorithm to obtain the second knowledge topology feature based on the second knowledge topology branch. The role of the first embedding layer is to convert the topology structure information into feature vectors that can be processed by a computer. The server inputs the topology unit and topology chain information in the second knowledge topology branch into the first embedding layer. Through the mapping and transformation of the embedding layer, the second knowledge topology feature is obtained. For example, the first embedding layer may use a deep learning model, such as a graph neural network (GNN), to encode the node and edge information in the second knowledge topology branch to obtain the vector representation of each topology unit. The combination of these vectors is the second knowledge topology feature. Suppose there are m topology units in the second knowledge topology branch, and each topology unit obtains a d-dimensional vector after passing through the first embedding layer. Then the second knowledge topology feature can be represented as an m×d matrix.

[0104] Based on this, in step 500A, the server obtains an aggregated feature according to the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature. The first knowledge topology feature is obtained by the server based on the first knowledge topology branch through the first embedding layer, reflecting the context information related to the first failure mode; the operation and maintenance information feature is obtained by the server based on the operation and maintenance information flow using the second embedding layer, reflecting the specific data information in the operation and maintenance process. The server needs to fuse these three features to obtain a more comprehensive and accurate aggregated feature.

[0105] One way to implement step 500A is to perform a weighted combination of the first knowledge topology feature and the second knowledge topology feature, and then fuse it with the operation and maintenance information feature. The server can assign weights to the first knowledge topology feature and the second knowledge topology feature respectively, multiply them pairwise with the weights and then add them up to obtain a target feature. Let the first knowledge topology feature be K1, the second knowledge topology feature be K2, the weight feature be W, and the target feature be T, then (where represents pairwise multiplication). Then, the server multiplies the operation and maintenance information feature O by the first mapping parameter M1, multiplies the target feature T by the second mapping parameter M2 and transposes it, and adds the adjustment feature B after multiplying the two, to obtain the aggregated feature A.

[0106] Another implementation method is to average the weighted combination of the first knowledge topology feature and the second knowledge topology feature, and then concatenate it with the operation and maintenance information feature at the head and tail. The server first multiplies the first knowledge topology feature and the second knowledge topology feature pairwise with the weight feature to obtain a first temporary feature and a second temporary feature, and then averages them pairwise to obtain the target feature. Let the first temporary feature be , the second temporary feature be , and the target feature be T'=(T1 + T2) / 2. Then, the target feature and the operation and maintenance information feature are concatenated at the head and tail to obtain the aggregated feature. In addition, another feasible implementation method is to perform a dot product on the weighted combination of the first knowledge topology feature and the second knowledge topology feature, and then add it pairwise with the operation and maintenance information feature. The server first obtains the first temporary feature and the second temporary feature, and then performs a dot product on them to obtain the target feature . Finally, the target feature is added pairwise with the operation and maintenance information feature to obtain the aggregated feature A' = T'' + O.

[0107] Through steps 200A - 200B and step 500A, the server can make full use of the multiple fault mode information contained in the operation and maintenance information flow, construct a more comprehensive knowledge topology branch, obtain the corresponding knowledge topology features, and effectively integrate these features with the operation and maintenance information features to obtain aggregated features. Such aggregated features synthesize the context information of multiple fault modes and operation and maintenance data information, which helps the server to more accurately identify fault modes and predict the development trend of faults in subsequent fault analysis, provide more targeted fault handling suggestions for operation and maintenance personnel, thereby improving the efficiency and quality of network operation and maintenance and ensuring the stable operation of the enterprise network.

[0108] In one implementation solution, after step 100 of obtaining the operation and maintenance information flow to be processed uploaded by the operation and maintenance terminal, the method further includes: Step 1001: Perform fault mode detection on the operation and maintenance information flow to obtain a fault mode detection result; Step 1002: If the fault mode detection result indicates that the operation and maintenance information flow includes two fault modes, then compare the two fault modes included in the operation and maintenance information flow with the fault modes in the operation and maintenance knowledge topology structure respectively to obtain a fault mode comparison result for each fault mode; Step 1003: If a fault mode comparison result indicates that one of the fault modes included in the operation and maintenance information flow matches the fault mode in the operation and maintenance knowledge topology structure, then regard the matching fault mode as the first fault mode; Step 1004: If the other fault mode comparison result indicates that the other fault mode included in the operation and maintenance information flow matches the fault mode in the operation and maintenance knowledge topology structure, then regard the matching other fault mode as the second fault mode.

[0109] In step 1001, the server performs fault mode detection on the operation and maintenance information flow to obtain a fault mode detection result. The detection method can refer to the foregoing content and will not be elaborated here.

[0110] In step 1002, when the fault mode detection result indicates that the operation and maintenance information flow includes two fault modes, the server compares the two fault modes included in the operation and maintenance information flow with the fault modes in the operation and maintenance knowledge topology structure respectively to obtain a fault mode comparison result for each fault mode.

[0111] In step 1003, when a fault mode comparison result indicates that one of the fault modes included in the operation and maintenance information flow matches the fault mode in the operation and maintenance knowledge topology structure, the server regards the matching fault mode as the first fault mode.

[0112] In step 1004, when the comparison result of another failure mode indicates that another failure mode included in the operation and maintenance information flow matches the failure mode in the operation and maintenance knowledge topology, the server uses the matched failure mode as the second failure mode. When there are two failure modes in the operation and maintenance information flow that match the operation and maintenance knowledge topology, the server determines them as the first failure mode and the second failure mode respectively, so as to comprehensively consider these two failure modes and their association relationships for subsequent failure analysis.

[0113] Through steps 1001 - 1004, the server can accurately identify possible failure modes from the operation and maintenance information flow, compare them with the operation and maintenance knowledge topology, and determine the first failure mode and the second failure mode. This provides an important basis for subsequent failure analysis and processing. The server can respectively determine the first knowledge topology branch and the second knowledge topology branch in the operation and maintenance knowledge topology based on the first failure mode and the second failure mode, and combine the information in these two branches to more comprehensively understand the causes and possible impacts of the failure.

[0114] After determining the first failure mode and the second failure mode, the first knowledge topology branch and the second knowledge topology branch are respectively determined in the operation and maintenance knowledge topology based on these two failure modes. The first knowledge topology branch and the second knowledge topology branch contain context information related to the first failure mode and the second failure mode, which helps the server more comprehensively understand the causes and possible impacts of the failure.

[0115] As an implementation solution of step 500A, that is, according to the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature, an aggregated feature is obtained, including: Step 500A1: Multiply each eigenvalue in the first knowledge topology feature by the eigenvalue of the weight feature to obtain a first temporary feature; Step 500A2: Multiply each eigenvalue in the second knowledge topology feature by the eigenvalue of the weight feature to obtain a second temporary feature; Step 500A3: Add each eigenvalue in the first temporary feature to the eigenvalue in the second temporary feature in a pairwise manner to obtain a target feature; Step 500A4: Multiply the operation and maintenance information feature by a first mapping parameter to obtain a first feature, where the dimension of the operation and maintenance information feature is e; Step 500A5: Multiply the target feature by a second mapping parameter, and transpose the multiplication result to obtain a second feature, where the dimension of the target feature is g; Step 500A6: Multiply the first feature by the second feature to obtain a third feature; Step 500A7: Sum the third feature and the adjustment feature to obtain an aggregated feature.

[0116] Among them, the above first mapping parameter and second mapping parameter are the results of decomposing a multi-dimensional data structure. The structure of the multi-dimensional data structure is e*g*f, the structure of the first mapping parameter is e*1*f, the structure of the second mapping parameter is 1*g*f, and the dimension of the aggregated feature is f.

[0117] In step 500A1, each eigenvalue in the first knowledge topology feature is multiplied element-wise with the eigenvalue of the weight feature to obtain a first temporary feature. The first knowledge topology feature is obtained by the server based on the first knowledge topology branch through the first embedding layer included in the operation and maintenance information processing algorithm, and it reflects the context information related to the first failure mode. The weight feature is a pre-set vector used to weight the first knowledge topology feature and the second knowledge topology feature to highlight the importance of different features. The server can obtain the first temporary feature by looping through the corresponding elements of the first knowledge topology feature and the weight feature and performing multiplication operations in sequence.

[0118] In step 500A2, each eigenvalue in the second knowledge topology feature is multiplied element-wise with the eigenvalue of the weight feature to obtain a second temporary feature. The second knowledge topology feature is obtained by the server based on the second knowledge topology branch through the same first embedding layer, and it reflects the context information related to the second failure mode.

[0119] In step 500A3, each eigenvalue in the first temporary feature is added element-wise with the eigenvalue of the second temporary feature to obtain a target feature. The target feature synthesizes the weighted information of the first knowledge topology feature and the second knowledge topology feature. The server can obtain the target feature by looping through the corresponding elements of the first temporary feature and the second temporary feature and performing addition operations in sequence.

[0120] In step 500A4, the operation and maintenance information feature is multiplied by the first mapping parameter to obtain a first feature, where the dimension of the operation and maintenance information feature is e. The operation and maintenance information feature is obtained by the server based on the operation and maintenance information flow through the second embedding layer included in the operation and maintenance information processing algorithm, and it reflects the specific data information in the operation and maintenance process. The first mapping parameter is a hyperparameter matrix with a structure of e×1×f, where e represents the dimension of the operation and maintenance information feature, and f is a pre-set dimension value used to control the dimension of the final aggregated feature.

[0121] In step 500A5, the target feature is multiplied by the second mapping parameter, and the multiplication result is transposed to obtain a second feature, where the dimension of the target feature is g. The second mapping parameter is a hyperparameter matrix with a structure of 1×g×f, and g represents the dimension of the target feature.

[0122] In step 500A6, the first feature is multiplied by the second feature to obtain the third feature. The server multiplies the first feature vector obtained in step 500A4 and the second feature vector obtained in step 500A5, where the multiplication may be a dot product operation of the vectors. Through this multiplication operation, the server preliminarily integrates the operation and maintenance information feature and the integrated knowledge topology feature to obtain a new feature vector, namely the third feature.

[0123] In step 500A7, the third feature is summed with the adjustment feature to obtain an aggregate feature, where the dimension of the aggregate feature is f. The adjustment feature is also a bias term, which is used to fine-tune the third feature to improve the expression ability of the aggregate feature. In this way, the server completes the construction process from the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature to the aggregate feature.

[0124] The first mapping parameter and the second mapping parameter here are the results of decomposing the multidimensional data structure, and the structure of the multidimensional data structure is e×g×f. By decomposing this multidimensional data structure into the first mapping parameter and the second mapping parameter, the server can flexibly transform and merge between different dimensions, thereby effectively combining the operation and maintenance information features, the first knowledge topology features and the second knowledge topology features. In practical applications, the values ​​of the first mapping parameter, the second mapping parameter, the weight feature and the adjustment feature are usually learned through an optimization algorithm during the training process of the operation and maintenance information processing algorithm, and the values ​​of these parameters are continuously adjusted based on a large amount of training data so that the final aggregated features can be better used for fault analysis.

[0125] The aggregated features constructed through steps 500A1-500A7 integrate the information of the operation and maintenance information features, the first knowledge topology features and the second knowledge topology features, and include both specific data information in the operation and maintenance process and context information related to the first failure mode and the second failure mode.

[0126] As another implementation scheme of step 500A, that is, obtaining an aggregate feature according to the first knowledge topology feature, the second knowledge topology feature and the operation and maintenance information feature, includes: Step 500A10: perform positional multiplication on each feature value in the first knowledge topology feature and the feature value of the weight feature to obtain a first temporary feature; Step 500A20: perform positional multiplication on each eigenvalue in the second knowledge topology feature and the eigenvalue in the weight feature to obtain a second temporary feature; Step 500A30: averaging each feature value in the first temporary feature with the feature value in the second temporary feature to obtain a target feature; Step 500A40: Concatenate the target feature and the operation and maintenance information feature at the beginning and end to obtain an aggregated feature.

[0127] In step 500A10, the server multiplies each eigenvalue in the first knowledge topology feature with the eigenvalue in the weight feature element by element to obtain a first temporary feature. The first knowledge topology feature is obtained by the server based on the first knowledge topology branch through the first embedding layer in the operation and maintenance information processing algorithm. It reflects the context information related to the first failure mode, such as information about the degree of association with other failure modes related to this failure mode. The weight feature is a pre-set vector used to weight the first knowledge topology feature and the second knowledge topology feature to highlight the importance of different features. The server can obtain the first temporary feature by traversing the corresponding elements of the first knowledge topology feature and the weight feature in a loop and performing multiplication operations in sequence.

[0128] In step 500A20, the server multiplies each eigenvalue in the second knowledge topology feature with the eigenvalue in the weight feature element by element to obtain a second temporary feature. The second knowledge topology feature is obtained by the server based on the second knowledge topology branch through the first embedding layer as well. It reflects the context information related to the second failure mode.

[0129] In step 500A30, the server takes the element-wise average of each eigenvalue in the first temporary feature and each eigenvalue in the second temporary feature to obtain the target feature. The target feature is the result of weighted and averaged processing of the first knowledge topology feature and the second knowledge topology feature. It synthesizes the context information related to the two failure modes.

[0130] In step 500A40, the server concatenates the target feature and the operation and maintenance information feature at the beginning and end to obtain an aggregated feature. The operation and maintenance information feature is obtained by the server based on the operation and maintenance information flow through the second embedding layer in the operation and maintenance information processing algorithm, which reflects the specific data information in the operation and maintenance process, such as the CPU usage rate and memory occupancy of the router.

[0131] Through steps 500A10 - 500A40, the server effectively fuses the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature to obtain an aggregated feature. This aggregated feature contains both the context information related to the first failure mode and the second failure mode, as well as the specific data information in the operation and maintenance process. In subsequent fault analysis, the server can use this aggregated feature to more comprehensively and accurately identify fault modes and predict the development trend of faults, providing more targeted fault handling suggestions for operation and maintenance personnel, thereby improving the efficiency and quality of network operation and maintenance and ensuring the stable operation of the enterprise network.

[0132] As another implementation of step 500A, that is, according to the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature, an aggregated feature is obtained, including: Step 500A100: Multiply each eigenvalue in the first knowledge topology feature by the eigenvalue in the weight feature in a pairwise manner to obtain a first temporary feature; Step 500A200: Multiply each eigenvalue in the second knowledge topology feature by the eigenvalue in the weight feature in a pairwise manner to obtain a second temporary feature; Step 500A300: Perform a pairwise dot product of each eigenvalue in the first temporary feature with the eigenvalue in the second temporary feature to obtain a target feature; Step 500A400: Add each eigenvalue in the target feature to the eigenvalue in the operation and maintenance information feature in a pairwise manner to obtain an aggregated feature; the dimensions of the target feature, the operation and maintenance information feature, and the aggregated feature are equal.

[0133] In step 500A100, the server multiplies each eigenvalue in the first knowledge topology feature by the eigenvalue in the weight feature in a pairwise manner to obtain a first temporary feature. The first knowledge topology feature is obtained by the server based on the first knowledge topology branch through the first embedding layer of the operation and maintenance information processing algorithm, and it reflects the context information related to the first failure mode. For example, in a network environment, if the first failure mode is "too high router CPU usage rate", the first knowledge topology feature may include associated information such as "abnormal port traffic" related to the associated failure mode. The weight feature is a pre-set vector, and its role is to weight different knowledge topology features to highlight the importance of certain features. The server can obtain the first temporary feature by simply performing a loop traversal operation to multiply the elements at the corresponding positions of the first knowledge topology feature and the weight feature in turn.

[0134] In step 500A200, the server multiplies each eigenvalue in the second knowledge topology feature by the eigenvalue in the weight feature in a pairwise manner to obtain a second temporary feature. The second knowledge topology feature is obtained by the server based on the second knowledge topology branch through the same first embedding layer, and it reflects the context information related to the second failure mode.

[0135] In step 500A300, the server performs a pairwise dot product of each eigenvalue in the first temporary feature with the eigenvalue in the second temporary feature to obtain a target feature. The target feature is the result after integrating the weighted information of the first knowledge topology feature and the second knowledge topology feature.

[0136] In step 500A400, the server adds each feature value in the target feature and the feature values in the operation and maintenance information feature in a pairwise manner to obtain an aggregated feature, and the dimensions of the target feature, the operation and maintenance information feature, and the aggregated feature are equal. The operation and maintenance information feature is obtained by the server based on the operation and maintenance information flow using the second embedding layer of the operation and maintenance information processing algorithm, which reflects the specific data information in the operation and maintenance process, such as the real-time CPU usage rate and memory occupancy percentage of the router.

[0137] Through steps 500A100 - 500A400, the server effectively fuses the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature to obtain an aggregated feature. This aggregated feature combines the context information of the two fault modes and the specific data information in the operation and maintenance process, which helps the server to more comprehensively and accurately identify fault modes and predict the development trend of faults in subsequent fault analysis.

[0138] In the method for transmitting operation and maintenance data of fault based on artificial intelligence, when the operation and maintenance information flow includes the first fault mode and the second fault mode, step 500A involves obtaining an aggregated feature according to the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature, and it has three different implementation schemes. The following will compare these three schemes and analyze their advantages, similarities, and differences.

[0139] The common goal of the above three schemes of step 500A is to construct an aggregated feature according to the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature for subsequent fault analysis. And, first, the first knowledge topology feature and the second knowledge topology feature are respectively processed pairwise with the weight feature to obtain corresponding temporary features, then the target feature is further processed based on these temporary features, and finally the aggregated feature is obtained by combining the operation and maintenance information feature.

[0140] The above three schemes each have their unique advantages and applicable scenarios. In actual operation and maintenance data processing of faults, appropriate schemes can be determined according to factors such as specific data characteristics, computing resources, and fault analysis requirements to construct aggregated features for more efficient and accurate fault analysis.

[0141] The embodiment of the present application also provides the training process of the above-mentioned operation and maintenance information processing algorithm, which specifically includes the following steps: Step T1: Obtain a set of training operation and maintenance information flows, where each training operation and maintenance information flow in the set of training operation and maintenance information flows includes a fault mode label, and the fault mode label included in each training operation and maintenance information flow belongs to the fault mode represented by the target topology unit in the operation and maintenance knowledge topology structure; Step T2: For each training operation and maintenance information flow, determine the training knowledge topology branch corresponding to the training operation and maintenance information flow in the operation and maintenance knowledge topology structure, where the training knowledge topology branch includes target topology units, topology units within a preset distance obtained based on the target topology units, and topology chains between the topology units; Step T3: For each training operation and maintenance information flow, based on the training knowledge topology branch, use the first embedding layer included in the operation and maintenance information processing algorithm to obtain training knowledge topology features; Step T4: For each training operation and maintenance information flow, based on the training operation and maintenance information flow, use the second embedding layer included in the operation and maintenance information processing algorithm to obtain training operation and maintenance information features; Step T5: For each training operation and maintenance information flow, generate training aggregation features according to the training knowledge topology features and the training operation and maintenance information features; Step T6: For each training operation and maintenance information flow, based on the training aggregation features, use the fault identification layer included in the operation and maintenance information processing algorithm to obtain the fault type support degree distribution corresponding to the training operation and maintenance information flow; Step T7: Based on the fault type support degree distribution and the fault mode label corresponding to each training operation and maintenance information flow, train the operation and maintenance information processing algorithm based on the loss function.

[0142] In Step T1, the server obtains a set of training operation and maintenance information flows. Each training operation and maintenance information flow in the set of training operation and maintenance information flows includes a fault mode label, and the fault mode label included in each training operation and maintenance information flow belongs to the fault mode characterized by the target topology unit in the operation and maintenance knowledge topology structure. The set of training operation and maintenance information flows is the data set used by the server to train the operation and maintenance information processing algorithm. These information flows are collected from historical operation and maintenance data. Each information flow records the relevant information during an operation and maintenance process, such as the operating parameters of the device, operation records, etc. The fault mode label is an identifier for the fault mode corresponding to the operation and maintenance information flow, which clarifies the fault type that occurred during this operation and maintenance process. The operation and maintenance knowledge topology structure is pre-constructed by the server and is a topology structure used to describe the fault mode and its relevance. The target topology unit is the topology unit corresponding to the fault mode label.

[0143] In step T2, for each training operation and maintenance information flow, the server determines the training knowledge topology branch corresponding to the training operation and maintenance information flow in the operation and maintenance knowledge topology structure, where the training knowledge topology branch includes the target topology unit, the topology units within a preset distance obtained based on the target topology unit, and the topology chains between the topology units. The preset distance is a pre-set parameter, which is the number of topology chains in the graph structure and is used to control the scope of the training knowledge topology branch. Taking the target topology unit of "the CPU usage rate of the router is too high" as an example, the related failure modes may include "the memory occupancy of the router is too large" and "the port traffic of the router is abnormal". Centering on the target topology unit, search for the topology units within the preset distance from it in the operation and maintenance knowledge topology structure, and combine these topology units and the topology chains between them to form the training knowledge topology branch.

[0144] In step T3, for each training operation and maintenance information flow, based on the training knowledge topology branch, the server uses the first embedding layer included in the operation and maintenance information processing algorithm to obtain the training knowledge topology features. The role of the first embedding layer is to convert the topology structure information in the training knowledge topology branch into a feature vector that can be processed by a computer. The server inputs the topology unit and topology chain information in the training knowledge topology branch into the first embedding layer, and through the mapping and transformation of the embedding layer, obtains the training knowledge topology features. For example, the first embedding layer may use a deep learning model, such as a graph neural network (GNN), to encode the node and edge information in the training knowledge topology branch to obtain the vector representation of each topology unit, and the combination of these vectors is the training knowledge topology features.

[0145] In step T4, for each training operation and maintenance information flow, based on the training operation and maintenance information flow, the server uses the second embedding layer included in the operation and maintenance information processing algorithm to obtain the training operation and maintenance information features. The second embedding layer is used to convert the original data in the training operation and maintenance information flow into a feature vector. The training operation and maintenance information flow contains various types of data, such as numerical device operation parameters and text-based operation records. The second embedding layer processes and transforms this data to extract valuable feature information. For numerical data, the server can use methods such as normalization and feature scaling to process the data and scale it to a specific range. For text data, the server can use word embedding techniques, such as Word2Vec, to map the words in the text to a low-dimensional vector space. Through the processing of the second embedding layer, the server converts the training operation and maintenance information flow into a feature vector, that is, the training operation and maintenance information features.

[0146] In step T5, for each training operation and maintenance information flow, the server generates training aggregation features based on the training knowledge topology features and the training operation and maintenance information features. The server can use various methods to generate training aggregation features. For example, the training knowledge topology features and the training operation and maintenance information features can be concatenated, weighted and summed, etc. One method is to multiply the training knowledge topology features by a weight matrix, multiply the training operation and maintenance information features by a weight matrix, and then add the two results to obtain the training aggregation features. By generating the training aggregation features, the server fuses the context information of the fault mode and the operation and maintenance data information, providing more comprehensive information for subsequent fault identification.

[0147] In step T6, for each training operation and maintenance information flow, based on the training aggregation features, the fault type support degree distribution corresponding to the training operation and maintenance information flow is obtained by using the fault identification layer included in the operation and maintenance information processing algorithm. The fault identification layer is the core part of the operation and maintenance information processing algorithm, and its role is to judge the fault type corresponding to the training operation and maintenance information flow according to the training aggregation features. The fault type support degree distribution represents the probability that the training operation and maintenance information flow belongs to each fault type. The server can use machine learning classification algorithms, such as support vector machine (SVM), neural network, etc., to implement the fault identification layer. Taking the neural network as an example, the server inputs the training aggregation features into the neural network. After multiple calculations and transformations of the neural network, the probability values of each fault type are finally output, and these probability values constitute the fault type support degree distribution. For example, for the three fault types of "router CPU usage rate too high", "router memory occupancy too large", and "router port traffic anomaly", the fault type support degree distribution may be [0.7, 0.2, 0.1], indicating that the probability that the training operation and maintenance information flow belongs to "router CPU usage rate too high" is 0.7, the probability that it belongs to "router memory occupancy too large" is 0.2, and the probability that it belongs to "router port traffic anomaly" is 0.1.

[0148] In step T7, based on the fault type support degree distribution corresponding to each training operation and maintenance information flow and the fault mode label, the operation and maintenance information processing algorithm is trained based on the loss function. The loss function is used to measure the difference between the fault type support degree distribution and the fault mode label. The goal of the server is to minimize the value of the loss function by adjusting the parameters in the operation and maintenance information processing algorithm. The loss function is, for example, the cross-entropy loss function, the mean square error loss function, etc. The server uses optimization algorithms, such as stochastic gradient descent (SGD), Adam, etc., to update the parameters in the operation and maintenance information processing algorithm according to the gradient of the loss function, and continuously iterates the training until the value of the loss function converges to a smaller value, thus completing the training of the operation and maintenance information processing algorithm.

[0149] Please refer to Figure 3, is a structural block diagram of the server 120 of the present application. The server 120 includes a computing unit 1001, which can execute various appropriate actions and processes according to a computer program stored in the ROM 1002 (i.e., read-only memory) or a computer program loaded from the storage unit 1008 into the RAM 1003 (i.e., random access memory). In the RAM 1003, various programs and data required for the operation of the server 120 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004. The input unit 1006 can be any type of device capable of inputting information to the server 120, such as receiving input digital or character information and generating key signal inputs related to the user settings and / or function controls of the server. The output unit 1007 can be any type of device capable of presenting information, such as a display or a speaker.

[0150] The computing unit 1001 executes the various methods and processes described above. For example, in some embodiments, the method 200 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the server 120 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the methods described above can be executed.

Claims

1. A fault operation and maintenance data transmission method based on artificial intelligence, characterized in that, Applied to a server, the server is communicatively connected to an operation and maintenance terminal, and the method includes: Obtain the operation and maintenance information flow to be processed uploaded by the operation and maintenance terminal; If the operation and maintenance information flow includes a first failure mode, determine a first knowledge topology branch in a pre-deployed operation and maintenance knowledge topology structure, where each topology unit in the operation and maintenance knowledge topology structure represents a failure mode, each topology chain in the operation and maintenance knowledge topology structure represents the relevance of failure modes, the first failure mode belongs to the failure mode represented by the first topology unit in the operation and maintenance knowledge topology structure, and the first knowledge topology branch includes the first topology unit, topology units within a preset distance obtained based on the first topology unit, and the topology chains between the topology units; Based on the first knowledge topology branch, obtain a first knowledge topology feature through a first embedding layer included in an operation and maintenance information processing algorithm; Based on the operation and maintenance information flow, obtain an operation and maintenance information feature by using a second embedding layer included in the operation and maintenance information processing algorithm; According to the first knowledge topology feature and the operation and maintenance information feature, establish an aggregated feature; Based on the aggregated feature, obtain a fault analysis result of the operation and maintenance information flow by using a fault identification layer included in the operation and maintenance information processing algorithm, and transmit the fault analysis result to the operation and maintenance terminal.

2. The method according to claim 1, wherein After obtaining the operation and maintenance information flow to be processed uploaded by the operation and maintenance terminal, the method further includes: Perform a failure mode detection on the operation and maintenance information flow to obtain a failure mode detection result; If the failure mode detection result indicates that the operation and maintenance information flow includes a failure mode, compare the failure mode included in the operation and maintenance information flow with the failure modes in the operation and maintenance knowledge topology structure to obtain a failure mode comparison result; If the failure mode comparison result indicates that the failure mode included in the operation and maintenance information flow exists in the failure modes in the operation and maintenance knowledge topology structure, use the existing failure mode as the first failure mode.

3. The method according to claim 2, wherein The step of if the failure mode detection result indicates that the operation and maintenance information flow includes a failure mode, then compare the failure mode included in the operation and maintenance information flow with the failure modes in the operation and maintenance knowledge topology structure to obtain a failure mode comparison result, includes: If the failure mode detection result indicates that the operation and maintenance information flow includes several failure modes, obtain the support coefficient corresponding to each failure mode among the several failure modes; Use the failure mode corresponding to the maximum support coefficient among the several failure modes as the failure mode included in the operation and maintenance information flow; Compare the failure mode included in the operation and maintenance information flow with the failure modes in the operation and maintenance knowledge topology structure to obtain the failure mode comparison result.

4. The method according to claim 1, characterized in that, The step of if the operation and maintenance information flow includes a first failure mode, then determine a first knowledge topology branch in a pre-deployed operation and maintenance knowledge topology structure, includes: If the operation and maintenance information flow includes the first failure mode, use the topology unit corresponding to the first failure mode in the operation and maintenance knowledge topology structure as the first topology unit; According to the operation and maintenance knowledge topology structure, obtain a topology unit group that has a correlation with the first topology unit, where the distance between each topology unit in the topology unit group and the first topology unit is within the preset distance; Establish the first knowledge topology branch according to the first topology unit, the topology unit group, and the topology chain between the topology units; Alternatively, if the operation and maintenance information flow includes a first failure mode, determining a first knowledge topology branch in a pre-deployed operation and maintenance knowledge topology structure includes: If the operation and maintenance information flow includes the first failure mode, use the topology unit corresponding to the first failure mode in the operation and maintenance knowledge topology structure as the first topology unit; According to the operation and maintenance knowledge topology structure, obtain a topology unit group that has a correlation with the first topology unit, where the distance between each topology unit in the topology unit group and the first topology unit is within the preset distance; If the operation and maintenance information flow includes a first failure mode correlation, according to the operation and maintenance knowledge topology structure, obtain a first sub-topology unit group in the topology unit group that has the first failure mode correlation with the first topology unit; Establish the first knowledge topology branch according to the first topology unit, the first sub-topology unit group, and the topology chain between the topology units; Alternatively, if the operation and maintenance information flow includes a first failure mode, determining a first knowledge topology branch in a pre-deployed operation and maintenance knowledge topology structure includes: If the operation and maintenance information flow includes the first failure mode, use the topology unit corresponding to the first failure mode in the operation and maintenance knowledge topology structure as the first topology unit; According to the operation and maintenance knowledge topology structure, obtain a topology unit group that has a correlation with the first topology unit, where the distance between each topology unit in the topology unit group and the first topology unit is within the preset distance; If the operation and maintenance information flow includes a first failure mode correlation and a second failure mode correlation, according to the operation and maintenance knowledge topology structure, obtain a second sub-topology unit group in the topology unit group that has at least one of the first failure mode correlation or the second failure mode correlation with the first topology unit; Establish the first knowledge topology branch according to the first topology unit, the second sub-topology unit group, and the topology chain between the topology units.

5. The method according to claim 1, wherein The establishing an aggregation feature according to the first knowledge topology feature and the operation and maintenance information feature includes: Multiply the operation and maintenance information feature by a first mapping parameter to obtain a first feature, where the dimension of the operation and maintenance information feature is e; Multiply the first knowledge topology feature by a second mapping parameter and transpose the product to obtain a second feature, where the feature of the first knowledge topology feature is g; Multiply the first feature by the second feature to obtain a third feature; Add an adjustment feature to the third feature to obtain the aggregation feature; Among them, the first mapping parameter and the second mapping parameter are the results of decomposing a multi-dimensional data structure. The structure of the multi-dimensional data structure is e*g*f, the structure of the first mapping parameter is e*1*f, the structure of the second mapping parameter is 1*g*f, and the dimension of the aggregated feature is f; Alternatively, establishing the aggregated feature according to the first knowledge topology feature and the operation and maintenance information feature includes: Performing a bitwise operation on each eigenvalue in the first knowledge topology feature and the eigenvalue in the operation and maintenance information feature to obtain the aggregated feature; the bitwise operation is inner product solving or summation, and the dimensions of the first knowledge topology feature, the operation and maintenance information feature, and the aggregated feature are equal.

6. The method according to claim 1, characterized in that The method further includes: If the operation and maintenance information flow further includes a second failure mode, determine a second knowledge topology branch in the operation and maintenance knowledge topology structure, where the second failure mode belongs to the failure mode characterized by the second topology unit in the operation and maintenance knowledge topology structure, and the second knowledge topology branch includes the second topology unit, the topology units within the preset distance based on the second topology unit, and the topology chain between the topology units; Based on the second knowledge topology branch, use the first embedding layer included in the operation and maintenance information processing algorithm to obtain a second knowledge topology feature; Establishing the aggregated feature according to the first knowledge topology feature and the operation and maintenance information feature includes: Obtaining the aggregated feature according to the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature.

7. The method according to claim 6, characterized in that, After obtaining the operation and maintenance information flow to be processed uploaded by the operation and maintenance terminal, the method further includes: Performing a failure mode detection on the operation and maintenance information flow to obtain a failure mode detection result; If the failure mode detection result indicates that the operation and maintenance information flow includes two failure modes, respectively compare the two failure modes included in the operation and maintenance information flow with the failure modes in the operation and maintenance knowledge topology structure to obtain a failure mode comparison result for each failure mode; If a failure mode comparison result indicates that a failure mode included in the operation and maintenance information flow matches a failure mode in the operation and maintenance knowledge topology structure, use the matching failure mode as the first failure mode; If another failure mode comparison result indicates that another failure mode included in the operation and maintenance information flow matches a failure mode in the operation and maintenance knowledge topology structure, use the matching another failure mode as the second failure mode.

8. The method according to claim 6, wherein Obtaining the aggregated feature according to the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature includes: Performing a bitwise multiplication on each eigenvalue in the first knowledge topology feature and the eigenvalue of the weight feature to obtain a first temporary feature; Performing a bitwise multiplication on each eigenvalue in the second knowledge topology feature and the eigenvalue of the weight feature to obtain a second temporary feature; Performing a bitwise addition on each eigenvalue in the first temporary feature and the eigenvalue in the second temporary feature to obtain a target feature; Multiply the operation and maintenance information feature by a first mapping parameter to obtain a first feature, where the dimension of the operation and maintenance information feature is e; Multiply the target feature by a second mapping parameter, transpose the multiplication result to obtain a second feature, where the dimension of the target feature is g; Multiply the first feature by the second feature to obtain a third feature; Sum the third feature and an adjustment feature to obtain the aggregated feature; The first mapping parameter and the second mapping parameter are the results of decomposing a multi-dimensional data structure, the structure of the multi-dimensional data structure is e*g*f, the structure of the first mapping parameter is e*1*f, the structure of the second mapping parameter is 1*g*f, and the dimension of the aggregated feature is f; Alternatively, obtaining the aggregated feature according to the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature includes: Perform pairwise multiplication of each eigenvalue in the first knowledge topology feature and the eigenvalue of the weight feature to obtain a first temporary feature; Perform pairwise multiplication of each eigenvalue in the second knowledge topology feature and the eigenvalue in the weight feature to obtain a second temporary feature; Perform pairwise averaging of each eigenvalue in the first temporary feature and the eigenvalue in the second temporary feature to obtain a target feature; Concatenate the target feature and the operation and maintenance information feature at the head and tail to obtain the aggregated feature; Alternatively, obtaining the aggregated feature according to the first knowledge topology feature, the second knowledge topology feature, and the operation and maintenance information feature includes: Perform pairwise multiplication of each eigenvalue in the first knowledge topology feature and the eigenvalue in the weight feature to obtain a first temporary feature; Perform pairwise multiplication of each eigenvalue in the second knowledge topology feature and the eigenvalue in the weight feature to obtain a second temporary feature; Perform pairwise multiplication of each eigenvalue in the first temporary feature and the eigenvalue in the second temporary feature to obtain a target feature; Perform pairwise addition of each eigenvalue in the target feature and the eigenvalue in the operation and maintenance information feature to obtain the aggregated feature; the dimensions of the target feature, the operation and maintenance information feature, and the aggregated feature are equal.

9. The method according to claim 1, wherein The method further includes: Obtain a training operation and maintenance information flow set, where each training operation and maintenance information flow in the training operation and maintenance information flow set includes a fault mode label, and the fault mode label included in each training operation and maintenance information flow belongs to the fault mode represented by the target topology unit in the operation and maintenance knowledge topology structure; For each training operation and maintenance information flow, determine a training knowledge topology branch corresponding to the training operation and maintenance information flow in the operation and maintenance knowledge topology structure, where the training knowledge topology branch includes the target topology unit, the topology units within the preset distance based on the target topology unit, and the topology chain between the topology units; For each training operation and maintenance information flow, based on the training knowledge topology branch, use the first embedding layer included in the operation and maintenance information processing algorithm to obtain a training knowledge topology feature; For each of the training and operation and maintenance information flows, based on the training and operation and maintenance information flow, use the second embedding layer included in the operation and maintenance information processing algorithm to obtain training and operation and maintenance information features; For each of the training and operation and maintenance information flows, generate training aggregation features according to the training knowledge topology features and the training and operation and maintenance information features; For each of the training and operation and maintenance information flows, based on the training aggregation features, use the fault identification layer included in the operation and maintenance information processing algorithm to obtain the fault type support degree distribution corresponding to the training and operation and maintenance information flow; Based on the fault type support degree distribution and the fault mode label corresponding to each training and operation and maintenance information flow, train the operation and maintenance information processing algorithm based on the loss function.

10. A fault operation and maintenance data transmission system, characterized in that, It includes a server and an operation and maintenance terminal that communicate with each other. Among them, the server includes: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Cable television network equipment failure detection method and device

    CN105516710A

  • Bayesian network-based fault detection method

    CN109270461A

  • Oil and gas field production fault determination and knowledge graph establishment method and device

    CN111475654A

  • High-end complex equipment reliability analysis method and system

    CN115829200A

  • Fault alarm processing method and device, electronic equipment and storage medium

    CN116545835A