Fault analysis method and device based on artificial intelligence, equipment and storage medium

By employing an AI-based fault analysis method, utilizing keyword extraction and knowledge graph retrieval, and combining alarm types to filter out target fault sources, the problem of low fault location accuracy and long processing cycles in wireless communication networks is solved, achieving rapid and accurate fault root cause analysis and efficiency improvement.

CN121099359APending Publication Date: 2025-12-09BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511157430.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing technologies for handling wireless communication network faults suffer from low fault location accuracy, long processing cycles, difficulty in quickly filtering invalid information and performing root cause analysis, and are particularly inefficient in scenarios with complex network topologies and massive alarm data.

Method used

An AI-based fault analysis method is adopted, which extracts keywords and features from alarm information, uses knowledge graphs for retrieval, and combines alarm types to filter out target fault sources. This includes feature extraction models and graph attention networks for feature matching, and dynamically adjusts weight coefficients to improve analysis accuracy.

Benefits of technology

It enables rapid and accurate root cause analysis of faults, improves the accuracy of fault location and processing efficiency, and breaks through the processing bottleneck of complex network topology and massive alarm data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121099359A_ABST
    Figure CN121099359A_ABST
Patent Text Reader

Abstract

The invention provides a fault analysis method and device based on artificial intelligence, equipment and a storage medium, and relates to the technical field of computers, in particular to the technical fields of artificial intelligence, knowledge maps, large models and the like. According to the specific implementation scheme, at least one keyword is extracted from alarm information; and inputting the alarm information into the feature extraction model to obtain a target feature of the alarm information, based on the keywords and the target features, performing retrieval in a knowledge graph to obtain a plurality of candidate fault sources; and screening out at least one target fault source from the plurality of candidate fault sources based on the alarm type of the alarm information.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, and particularly relates to the technical field of artificial intelligence, knowledge graph, large model, etc. BACKGROUND

[0002] With the rapid development of information technology, the scale of wireless communication network continues to expand, the network topology structure becomes increasingly complex, and the coupling degree of various network elements and business applications is also deepening.

[0003] Therefore, the stable operation of the network is directly related to user experience and business continuity, which puts forward more stringent requirements for the prevention and disposal of network faults. SUMMARY

[0004] The present disclosure provides an artificial intelligence-based fault analysis method, device, equipment and storage medium.

[0005] According to an aspect of the present disclosure, an artificial intelligence-based fault analysis method is provided, comprising:

[0006] extracting at least one keyword from the alarm information; and inputting the alarm information into a feature extraction model to obtain target features of the alarm information;

[0007] based on the keyword and the target feature, searching in a knowledge graph to obtain a plurality of candidate fault sources;

[0008] based on the alarm type of the alarm information, screening at least one target fault source from the plurality of candidate fault sources.

[0009] According to another aspect of the present disclosure, an artificial intelligence-based fault analysis device is provided, comprising:

[0010] an extraction module configured to extract at least one keyword from the alarm information; and input the alarm information into a feature extraction model to obtain target features of the alarm information;

[0011] a search module configured to search in a knowledge graph based on the keyword and the target feature to obtain a plurality of candidate fault sources;

[0012] a screening module configured to screen at least one target fault source from the plurality of candidate fault sources based on the alarm type of the alarm information.

[0013] According to another aspect of the present disclosure, an electronic device is provided, comprising:

[0014] at least one processor; and

[0015] a memory in communication with the at least one processor; wherein

[0016] The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any of the embodiments of the present disclosure.

[0017] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to enable the computer to perform the method according to any of the embodiments of the present disclosure.

[0018] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to any of the embodiments of the present disclosure.

[0019] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0020] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:

[0021] Figure 1 is a flowchart of an artificial intelligence-based fault analysis method according to an embodiment of the present disclosure;

[0022] Figure 2 is a flowchart of screening out a target fault source according to an embodiment of the present disclosure;

[0023] Figure 3 is a flowchart of obtaining a weight coefficient of to-be-processed information according to an embodiment of the present disclosure;

[0024] Figure 4 is a flowchart of optimization of the weight coefficient of to-be-processed information according to an embodiment of the present disclosure;

[0025] Figure 5 is a flowchart of an artificial intelligence-based fault analysis method according to an embodiment of the present disclosure;

[0026] Figure 6 is a structural diagram of an artificial intelligence-based fault analysis device according to an embodiment of the present disclosure;

[0027] Figure 7 is a block diagram of an electronic device for implementing an artificial intelligence-based fault analysis method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, which are cited by way of example only. The present disclosure is therefore not limited to the embodiments described herein, but encompasses all embodiments within the scope of the present disclosure. As such, various changes and modifications can be suggested to one skilled in the art, and it is intended that the present disclosure encompass such changes and modifications as fall within the scope of the appended claims. Also, in the following description, well-known functions or constructions are not described in detail because they can obscure the understanding of the present disclosure.

[0029] The terms "first", "second", and the like in the present disclosure are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, inclusion of a series of steps or units. The method, system, product or device is not necessarily limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0030] It should be noted that, unless it is explicitly stated that there is a sequence of execution between different operations or there is a sequence of execution between different operations in the technical implementation, the execution sequence between multiple operations can not be in sequence, and multiple operations can be executed simultaneously.

[0031] In the related art of wireless communication network fault processing, it is currently mostly dependent on artificial experience or simple rule engine. This brings a series of obvious disadvantages: on the one hand, the fault positioning accuracy is low, and it is often difficult to accurately lock the problem core; on the other hand, the processing period is long, and the best opportunity for fault repair is easily missed. Moreover, with the increasing complexity of network structure and the explosive growth of alarm information, these problems are more prominent in the scene of complex network topology and massive alarm data.

[0032] The processing method in the related art not only cannot quickly filter out invalid information, but also cannot realize rapid analysis of the fault root cause, which greatly hinders the overall efficiency of fault processing.

[0033] Therefore, in the embodiments of the present disclosure, an artificial intelligence-based fault analysis method is provided, which realizes rapid fault root cause analysis by a hybrid graph embedding retrieval method, and improves the efficiency and accuracy of fault analysis.

[0034] As shown in FIG. 1, a flowchart of an artificial intelligence-based fault analysis method provided by the embodiments of the present disclosure is shown, which includes the following contents: Figure 1

[0035] S101, extracting at least one keyword from the alarm information; and inputting the alarm information into a feature extraction model to obtain target features of the alarm information. ​

[0036] The alarm information refers to prompt information generated by a device, system or network when an abnormality occurs during operation.

[0037] At least one keyword is extracted from the alarm information, that is, the core content of the alarm information is obtained from the alarm information. For example, in the alarm information generated by the wireless communication network, "base station decommissioning" and "optical module abnormality" can be used as keywords in the alarm information.

[0038] In implementation, the alarm information can be preprocessed to obtain the keywords. The preprocessing method can be a rule-based method, for example, by matching the relevant words in the alarm information as initial keywords through a pre-set fault keyword library. To further highlight the important keywords that can better represent the alarm information, the initial keywords obtained can be weighted by the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm, so as to filter out the keywords with more semantic expression as the final keywords.

[0039] In the embodiments of the present disclosure, the feature extraction model can be a natural language processing model for extracting high-level semantic features from the text of the alarm information. The feature extraction model can also be a multi-modal model. When the alarm information contains both text content and image or video content, the target feature of the alarm information can be extracted by the multi-modal model. For example, the image modal extraction module in the multi-modal model can process the images and / or videos in the alarm information to obtain visual features; the text modal extraction module in the multi-modal model can process the text information in the alarm information to obtain text features. Then, the fusion module in the multi-modal model can fuse the visual features and the text features to obtain the target features. The visual features extracted by the image modal extraction module can be mapped to the feature space of the text features to realize cross-modal feature alignment, and on this basis, the fusion module can use a large language model to fuse the features of the two modalities to extract the core key features. By inputting the alarm information into such a model, the alarm information can be converted into a low-dimensional vector form of the target features.

[0040] In S102, based on the keywords and the target features, a search is performed in the knowledge graph to obtain a plurality of candidate fault sources.

[0041] In the embodiments of the present disclosure, the knowledge graph includes a plurality of nodes and the association relationship between the nodes. The nodes in the knowledge graph include fault sources, device identifiers, etc., and the edges represent the association relationship between the nodes. The knowledge graph stores a large amount of fault-related knowledge, including the connection relationship between the fault sources, the fault phenomenon and the causal relationship, etc.

[0042] In addition, the nodes of the knowledge graph are also associated with corresponding network topology structures as supplementary knowledge to enrich the knowledge graph. In implementation, when searching in the knowledge graph based on the keywords, the corresponding nodes can be found in the knowledge graph by calculating the keyword matching degree. The matching degree between the keywords in the alarm information and the keywords in the related text of each node in the knowledge graph can be compared. Specifically, the matching degree can be measured and searched in at least one dimension such as the intersection size of the common characters and the semantic similarity. For example, the keywords such as “base station decommissioning” and “optical module overload” in the alarm information are compared with the keywords such as “base station failure” and “optical module failure” contained in a node in the knowledge graph in terms of semantic similarity. If the matching degree reaches a preset threshold, the node can be located as a related node, and the corresponding fault source is taken as a candidate fault source.

[0043] When searching in the knowledge graph based on the target feature, in some embodiments, the node features of each node can be extracted and then compared with the target feature, so as to search for a node similar to the target feature.

[0044] In other embodiments, in order to improve the accuracy of the search, the neighbor nodes within a preset neighborhood range of a target node in the knowledge graph can be obtained. The preset neighborhood range is defined by the distance between nodes, that is, the number of nodes contained in the shortest path from the neighbor node to the target node. Then, the graph attention network (GAT) is used to extract the features of the target node and the neighbor nodes in the preset neighborhood range based on the attention mechanism, to generate an embedding vector of the node as the feature of the node. Finally, the cosine similarity, Euclidean distance and other methods can be used to measure the similarity between the target feature and the features of the nodes in the knowledge graph, and the fault source corresponding to the node with higher similarity is selected as the candidate fault source.

[0045] S103, based on the alarm type of the alarm information, at least one target fault source is selected from the plurality of candidate fault sources.

[0046] The alarm type can be classified according to the nature, severity and the like of the fault, such as application fault alarm, network fault alarm and the like. By combining the alarm type to screen the candidate fault source, the fault troubleshooting range can be narrowed, and the screening result is more suitable for the characteristics of the current alarm classification, so as to further improve the accuracy of fault positioning.

[0047] In the embodiments of the present disclosure, keywords are extracted from the alarm information, target features are obtained by using a feature extraction model, candidate fault sources are retrieved in a knowledge graph by combining the two, and the target fault source is filtered from the candidate fault sources according to the alarm type. In this way, the drawbacks of relying only on artificial experience or a simple rule engine can be effectively solved, and the deep semantic of the alarm information and the correlation of the fault sources in the knowledge graph can be fully mined. Further, combined with accurate filtering according to the alarm type, rapid and accurate analysis of the fault root cause is realized, the fault positioning accuracy is effectively improved, the processing period is shortened, the processing bottleneck in the complex network topology and massive alarm data scenarios is broken through, and the overall efficiency of fault processing is improved.

[0048] In the embodiments of the present disclosure, based on the alarm type of the alarm information, at least one target fault source is filtered from a plurality of candidate fault sources, comprising:

[0049] In S201, for each candidate fault source, information under the alarm type associated with the candidate fault source in the knowledge graph is obtained, and a plurality of to-be-processed information is obtained.

[0050] In the knowledge graph, each candidate fault source is a node, and its associated information is often multi-dimensional, which may include information associated with the node itself (such as attribute information) and information associated with other nodes directly connected to it (such as attribute information).

[0051] The information associated with the alarm type of the candidate fault source is obtained, that is, from all the associated information of the candidate fault source node, the information having a clear correlation with the current alarm type is filtered out and used as the to-be-processed information.

[0052] In implementation, the association standard of the alarm type and the information can be pre-defined in combination with domain knowledge, and the relevant information is directly matched and extracted by a rule engine as the to-be-processed information. For example, the attribute information of the node itself, the node directly connected to the node, and the attribute information of the directly connected node can be selected.

[0053] In some embodiments, the attribute information of the candidate fault source node and the attribute information of the node directly connected to it can also be analyzed by using a related pre-trained large model to filter out information associated with the current alarm type as to-be-processed information. In this way, the to-be-processed information that may affect the fault can be mined by the large model, and the accuracy of information mining is improved.

[0054] In S202, feature extraction is performed on the plurality of to-be-processed information respectively, and a plurality of to-be-processed features is obtained.

[0055] In implementation, the plurality of to-be-processed information can be respectively subjected to feature extraction based on the feature extraction model described in the foregoing, and a plurality of to-be-processed features is obtained.

[0056] S203, determine the weight coefficient of each of the plurality of to-be-processed information under the alarm type.

[0057] Since the same to-be-processed information has different importance under different alarm types for judging the candidate fault source, the same to-be-processed information can be assigned a corresponding adjustment coefficient for different alarm types. Thus, the weight of the same to-be-processed information under different alarm types is differentiated, which is adjusted based on the difference of the alarm types, so that the weight distribution is more reasonable, thereby improving the accuracy of fault positioning.

[0058] In implementation, the weight distribution model is trained by collecting training samples to predict the weight coefficient of each to-be-processed information under the corresponding alarm type. For example, the alarm type and to-be-processed information can be selected from historical data, and whether the to-be-processed information affects the fault root cause analysis of the alarm information is artificially labeled as a classification standard, and the weight distribution model is trained to enable the weight distribution model to extract effective classification features from the to-be-processed information based on the alarm type. The classification features are used to identify whether the to-be-processed information has an impact on the fault root cause analysis of the alarm information under the alarm type. In the weight distribution model, the classification features are used to calculate the weight coefficient, the weight coefficient is used for binary classification, and the loss is calculated with the corresponding labeled classification result. Thus, the weight distribution model can be trained to reasonably assign isolated weight coefficients to to-be-processed information under different alarm types. By training the weight distribution model, the weight distribution model can continuously learn the correlation between to-be-processed information and fault judgment results under different alarm types, so that the weight distribution model can grasp the influence degree of the same to-be-processed information on the fault judgment result under different alarm types. The trained weight distribution model is used to analyze the weight coefficient of the to-be-processed information under different alarm types.

[0059] In addition to the weight distribution model that can be used to assign the weight coefficient of the to-be-processed information, in some embodiments, the weight coefficient of the to-be-processed information can also be dynamically determined based on the associated information of the to-be-processed information in the current knowledge graph. In implementation, a weight association table of different to-be-processed information in the knowledge graph under each alarm type can be established. In implementation, for each to-be-processed information, if the known weight of the to-be-processed information under the alarm type is found in the weight association table, the weight coefficient can be used for subsequent processes. If the known weight of the to-be-processed information under the alarm type is not found in the weight association table, that is, the weight coefficient of the to-be-processed information under the alarm type is not found, the weight coefficient of the to-be-processed information can be determined based on the adjustment coefficient of the to-be-processed information and the basic coefficient of the to-be-processed information; the weight coefficient has a positive correlation with the adjustment coefficient and the basic coefficient; the adjustment coefficient is obtained based on the coefficient table corresponding to the to-be-processed information under the alarm type. The weight coefficient determined based on the adjustment coefficient of the to-be-processed information and the basic coefficient of the to-be-processed information can be updated to the weight association table for subsequent iterative optimization and searching.

[0060] In the embodiments of the present disclosure, in the case that the to-be-processed information has no known weight under a specific alarm type, by combining the basic coefficient thereof and the adjustment coefficient obtained based on the coefficient table of the alarm type, the weight coefficient of the information under the current alarm type can still be reasonably determined in the case of lacking direct reference weight, which not only retains the influence of the basic characteristics of the information itself, but also adapts to the specific alarm classification requirements, ensures the integrity and applicability of the weight system, and provides a reliable basis for subsequent fault source analysis.

[0061] The basic coefficient is a coefficient possessed by the to-be-processed information itself, which represents the coefficient of the to-be-processed information under a general alarm scenario. The higher the value is, the greater the influence of the to-be-processed information on the fault source judgment in the general alarm type is. In implementation, the basic coefficient of the to-be-processed information can be determined based on the following steps:

[0062] Step A1, determining the highest association degree between the to-be-processed information and at least one keyword as a first association degree, and determining a second association degree between the to-be-processed information and the target feature;

[0063] In implementation, the matching degree of the to-be-processed information and each keyword can be compared by semantic similarity, and the highest value is taken as the first association degree. The second association degree is obtained by calculating the similarity between the feature vector of the to-be-processed information and the feature vector of the target feature.

[0064] Step A2, in the case that the first association degree is higher than the second association degree, selecting a first preset weight associated with the keyword as the basic coefficient;

[0065] In implementation, a first association table between reference words and base coefficients can be established in advance, and the first preset weights of different reference words are determined. For a keyword not directly included in the first association table, the first preset weight corresponding to a reference word with a semantic similarity higher than a set threshold can be selected as the base coefficient of the keyword by calculating the semantic similarity between the keyword and the reference words in the first association table.

[0066] In step A3, when the second association degree is higher than the first association degree, the second preset weight of the target feature association is selected as the base coefficient.

[0067] In implementation, a second association table between multiple reference features and base coefficients can be established in advance to record the second preset weights corresponding to different reference features. For a target feature not directly included in the second association table, the second preset weight corresponding to a reference feature with a feature similarity higher than a set threshold can be selected as the base coefficient of the target feature by calculating the feature similarity between the target feature and the reference features in the second association table.

[0068] In the embodiments of the present disclosure, by comparing the first association degree of the to-be-processed information with the keyword and the second association degree of the to-be-processed information with the target feature, one of the first preset weight and the second preset weight is selected as the base coefficient according to the association degree, which can make the base coefficient reflect the importance of the to-be-processed information in analyzing the fault source under the general alarm type or the general alarm scene, thereby quickly assigning the weight coefficient to the to-be-processed information and improving the rationality and efficiency of the weight coefficient assignment.

[0069] In implementation, the first association table and the second association table can be determined according to empirical values. They can also be obtained by analyzing historical data, which is not limited in the embodiments of the present disclosure.

[0070] The adjustment coefficient is a parameter for dynamically correcting the base coefficient according to the current alarm type, and represents the relative importance of the to-be-processed information under different alarm types. The adjustment coefficient is obtained from a coefficient table corresponding to the to-be-processed information under the alarm type. The coefficient table is constructed in advance and records the correction ratios of different to-be-processed information under different alarm types.

[0071] In the embodiments of the present disclosure, the weight coefficient of the to-be-processed information is determined based on the adjustment coefficient of the to-be-processed information and the base coefficient of the to-be-processed information, which can include the following contents as shown in FIG. 6. Figure 3

[0072] S301, the normalization reference value is determined based on the adjustment coefficient of the to-be-processed information and the base coefficient of the to-be-processed information; the normalization reference value has a positive correlation with the adjustment coefficient and the base coefficient.

[0073] ​S302, determine the sum value of the normalized reference values of the plurality of to-be-processed information.

[0074] S303, divide the normalized reference value of the to-be-processed information by the sum value to obtain the weight coefficient of the to-be-processed information.

[0075] In the embodiments of the present disclosure, the weight coefficient of the to-be-processed information can be expressed by formula (1):

[0076]

[0077] In formula (1), W i (S) represents the weight coefficient of the i-th to-be-processed information under the alarm type S; base i represents the base coefficient of the i-th to-be-processed information; k i (S) represents the adjustment coefficient of the i-th to-be-processed information under the alarm type S; base i ·k i (S) represents the normalized reference value obtained based on the adjustment coefficient of the to-be-processed information and the base coefficient of the to-be-processed information; represents the sum value of the normalized reference values of the plurality of to-be-processed information; n represents the total number of to-be-processed information, that is, there are n to-be-processed information participating in the weight calculation.

[0078] In the embodiments of the present disclosure, by combining the base coefficient of the to-be-processed information and the adjustment coefficient for the specific alarm type, and after normalization processing, the weight proportion of each information in the corresponding alarm type is reasonably allocated, the basic importance of the information itself is retained, and the influence weight is dynamically adjusted according to the specific alarm type, so that the weight result is more consistent with the contribution degree of different information to the specific alarm in the actual application, the accuracy and pertinence of the weight decision are improved, and thus the accuracy of the fault root cause analysis is improved.

[0079] In the embodiments of the present disclosure, based on the calculation method of formula (1), the normalization processing of the weight coefficient can be directly realized. In actual implementation, to further more directly obtain the weight coefficient of the to-be-processed information, formula (1) can be optimized, as shown in the following formula (2): Figure 4 The optimization process includes the following contents:

[0080] S401, based on the adjustment coefficient of the to-be-processed information and the base coefficient of the to-be-processed information, determine the normalized reference value; the normalized reference value has a positive correlation with the adjustment coefficient and the base coefficient.

[0081] S402, determine the sum value of the normalized reference values of the plurality of to-be-processed information.

[0082] S403, map the sum value to a target coordinate system, so that the value of the sum value in the target coordinate system is a target value.

[0083] The target coordinate system is a reference scale for regulating a numerical range or meeting specific constraint conditions, and essentially converts the original calculation result to a numerical space meeting the preset requirements through mapping rules such as scaling, linear conversion, etc.

[0084] The target value is a preset standard value that needs to be met by the sum of the normalized reference values of the plurality of to-be-processed information, for example, the sum of the normalized reference values is 1.

[0085] Therefore, by mapping the sum value to the target coordinate system, based on the mapping rule, the sum value in the target coordinate system can be made to be the target value.

[0086] S404, map the normalized reference value to the target coordinate system to obtain the mapping value of the normalized reference value of the to-be-processed information as the weight coefficient of the to-be-processed information.

[0087] Similarly, by mapping the normalized reference value to the target coordinate system, the mapping value of the normalized reference value of the to-be-processed information can be obtained. Since the value of the sum value in the target coordinate system is already the target value, the mapping value of the normalized reference value of the to-be-processed information can be directly used as the weight coefficient of the to-be-processed information.

[0088] In the embodiments of the present disclosure, on the basis of retaining the comprehensive influence of the base coefficient and the adjustment coefficient on the weight, by increasing the mapping of the sum of the normalized reference values to the target coordinate system, it can be ensured that the sum of the weight coefficients of the final to-be-processed information meets the preset standard, further improving the convenience and reliability of the weight-based decision.

[0089] S204, based on the weight coefficients of the plurality of to-be-processed information respectively, weighted sum of the plurality of to-be-processed features is performed to obtain a fault source score of the candidate fault source.

[0090] In implementation, each to-be-processed feature can be normalized first, as shown in formula (2):

[0091]

[0092] In formula (2), F i represents the to-be-processed feature corresponding to the i-th to-be-processed information; max(F i ) represents the maximum value in each to-be-processed feature; F’ i represents the normalized feature value of the to-be-processed feature corresponding to the i-th to-be-processed information, and the range is compressed to [0, 1].

[0093] The fault source score of the candidate fault source can be expressed by formula (3):

[0094]

[0095] In formula (3), Score represents the fault source score of the candidate fault source; W i (S) represents the weight coefficient of the i-th to-be-processed information under the alarm type S; F i represents the normalized feature value of the to-be-processed feature corresponding to the i-th to-be-processed information.

[0096] S205, based on the fault source scores of the plurality of candidate fault sources respectively, screening out a target fault source.

[0097] That is, the target fault source is screened out from the plurality of candidate fault sources according to the fault source scores. In implementation, the scores of all candidate fault sources can be sorted, and the candidate fault source with the highest score or exceeding a preset threshold is selected as the target fault source.

[0098] In the embodiments of the present disclosure, the related information of the candidate fault source and the current alarm type is associated through the knowledge graph, and after feature extraction, the weight coefficients of each information under the specific alarm type are combined for weighted calculation to obtain the fault source score of the candidate fault source. Finally, the target fault source is screened out according to the score. In this way, the information directly related to the current alarm type can be focused on, the importance difference of different information is reflected through weight distribution, and the quantitative score is used as the screening basis, so as to accurately locate the fault source most likely to cause the alarm and improve the pertinence and accuracy of fault source identification.

[0099] In the embodiments of the present disclosure, in order to continuously improve the accuracy of the fault analysis method based on artificial intelligence in actual application, the weight coefficients of the to-be-processed information can be dynamically adjusted through feedback learning, which can be implemented as follows:

[0100] Step B1, for each to-be-processed information in the plurality of to-be-processed information, obtaining a diagnosis result of a target fault source to which the to-be-processed information belongs; the diagnosis result is used to indicate whether the target fault source is finally determined as a final fault source;

[0101] Step B2, based on the diagnosis result, optimizing the weight coefficient of the to-be-processed information;

[0102] Among them, the diagnosis result indicates that, in the case that the target fault source is determined as the final fault source, the weight coefficient is improved based on a preset optimization mode; in the case that the target fault source is not determined as the final fault source, the weight coefficient is reduced based on the preset optimization mode.

[0103] Among them, the preset optimization mode includes a learning rate; the optimized weight coefficient has a positive correlation with the learning rate and a target difference; the target difference is the difference between the value corresponding to the diagnosis condition and the fault source score of the target fault source.

[0104] In implementation, the target fault source obtained by the previous fault analysis method based on artificial intelligence can be judged whether it is the real cause of the fault after actual investigation and manual confirmation. In the case that the target fault source is confirmed as the final fault source, the weight of the related to-be-processed information needs to be strengthened; in the case that the target fault source is not confirmed as the final fault source, the weight of the related to-be-processed information needs to be weakened.

[0105] In the embodiments of the present disclosure, the process of iteratively optimizing the weight coefficient can be expressed by formula (4):

[0106] W i (S) 新 = W i (S) + ∩ · (R-Score) · F i (4)

[0107] In formula (4), W i (S) 新 represents the optimized weight coefficient, that is, the new weight coefficient of the to-be-processed information obtained by optimization according to whether the target fault source to which the i th to-be-processed information belongs under the alarm type S is finally determined as the final fault source; W i (S) represents the original weight coefficient of the i th to-be-processed information under the alarm type S; ∩ represents the learning rate, which can be 0.1-0.4; R represents whether the diagnosis is successful, for example, R=1 represents that the diagnosis is successful, that is, it is the final fault source, and R=0 represents that the diagnosis fails, that is, it is not the final fault source; (R-Score) represents the target difference value; F i represents the normalized feature value of the to-be-processed feature corresponding to the i th to-be-processed information.

[0108] In the embodiments of the present disclosure, by taking the actual diagnosis result as the feedback basis, combining the preset optimization mode containing the learning rate, dynamically adjusting the weight coefficient of the to-be-processed information according to the target difference value of the diagnosis condition value and the fault source score, the weight coefficient can be continuously fitted to the real law of the actual diagnosis scene, the influence weight of each to-be-processed information in the artificial intelligence fault analysis method is continuously optimized, thereby gradually improving the accuracy of fault analysis and enhancing the adaptability and reliability of the method in actual application.

[0109] In summary, the overall process of the fault analysis method based on artificial intelligence provided in the embodiments of the present disclosure is shown in Figure 5 , which includes:

[0110] S501, constructing a knowledge graph.

[0111] During implementation, structured data (such as equipment configuration and alarm logs) and unstructured data (such as maintenance manuals) can be collected through asset management systems, alarm platforms, and work order systems. A multimodal large model is used to automatically mine potential patterns and construct potential connections in the knowledge graph. The knowledge graph is then dynamically updated in real time based on the network graph topology.

[0112] S502, acquire alarm data, extract at least one keyword from the alarm information; and input the alarm information into a feature extraction model to obtain the target features of the alarm information.

[0113] S503, based on keywords and target features, searches are performed in the knowledge graph to obtain multiple candidate fault sources.

[0114] S504, based on the alarm type of the alarm information, selects at least one target fault source from multiple candidate fault sources.

[0115] S505, after identifying the target fault source, summarizes established experience in fault localization based on large model experts (i.e., pre-trained large models). During implementation, prompt words can be constructed, and root cause answers can be output based on the large model experts. For example, the output root cause answer may include "alarm name," "fault cause (level 1)," "fault cause (level 2)," "occurrence probability," and "repair plan / process."

[0116] After experimental measurement and statistical analysis, the results are shown in Table 1:

[0117]

[0118] Table 1

[0119] In summary, the AI-based fault analysis method provided in this embodiment can fully explore the deep semantics of alarm information and the correlation of fault sources in the knowledge graph. Combined with precise screening of alarm types, it can achieve rapid and accurate analysis of the root cause of the fault, effectively improve the accuracy of fault location, shorten the processing cycle, and thus break through the processing bottleneck in complex network topology and massive alarm data scenarios, thereby improving the overall efficiency of fault handling.

[0120] Based on the same technical concept, this disclosure also provides an artificial intelligence-based fault analysis device 600, such as... Figure 6 As shown, it includes:

[0121] Extraction module 601 is used to extract at least one keyword from alarm information; and to input alarm information into feature extraction model to obtain target features of alarm information.

[0122] The retrieval module 602 is used to perform retrieval in the knowledge graph based on keywords and target features to obtain multiple candidate fault sources;

[0123] The screening module 603 is configured to screen at least one target fault source from the plurality of candidate fault sources based on the alarm type of the alarm information.

[0124] In some embodiments, the screening module comprises:

[0125] The obtaining unit is configured to obtain, for each candidate fault source, information under an alarm type associated with the candidate fault source in the knowledge graph, to obtain a plurality of pieces of to-be-processed information.

[0126] The extraction unit is configured to perform feature extraction on the plurality of pieces of to-be-processed information respectively, to obtain a plurality of to-be-processed features.

[0127] The weight coefficient determination unit is configured to determine a respective weight coefficient of each piece of to-be-processed information under the alarm type.

[0128] The scoring unit is configured to perform weighted summation on the plurality of to-be-processed features based on the respective weight coefficients of the plurality of pieces of to-be-processed information, to obtain a fault source score of the candidate fault source.

[0129] The screening unit is configured to screen the target fault source based on the respective fault source scores of the plurality of candidate fault sources.

[0130] In some embodiments, the weight coefficient determination unit is specifically configured to:

[0131] In a case where no known weight of the to-be-processed information under the alarm type is found for each piece of to-be-processed information, the weight coefficient of the to-be-processed information is determined based on an adjustment coefficient of the to-be-processed information and a basic coefficient of the to-be-processed information; the weight coefficient has a positive correlation with the adjustment coefficient and the basic coefficient; and the adjustment coefficient is obtained based on a coefficient table corresponding to the to-be-processed information under the alarm type.

[0132] In some embodiments, the weight coefficient determination unit is specifically configured to:

[0133] determine a normalization reference value based on the adjustment coefficient of the to-be-processed information and the basic coefficient of the to-be-processed information; the normalization reference value has a positive correlation with the adjustment coefficient and the basic coefficient.

[0134] determine a sum value of the respective normalization reference values of the plurality of pieces of to-be-processed information.

[0135] divide the normalization reference value of the to-be-processed information by the sum value, to obtain the weight coefficient of the to-be-processed information.

[0136] In some embodiments, the weight coefficient determination unit is specifically configured to:

[0137] Determine a normalization reference value based on an adjustment coefficient of the to-be-processed information and a basic coefficient of the to-be-processed information; the normalization reference value has a positive correlation with the adjustment coefficient and the basic coefficient;

[0138] Determine a sum value of the normalization reference values of the plurality of to-be-processed information;

[0139] Map the sum value to the target coordinate system, so that the value of the sum value in the target coordinate system is a target value;

[0140] Map the normalization reference value to the target coordinate system, and obtain a mapping value of the normalization reference value of the to-be-processed information as a weight coefficient of the to-be-processed information.

[0141] In some embodiments, the method further comprises a basic coefficient determination module configured to determine a basic coefficient of the to-be-processed information based on the following method:

[0142] Determine a highest correlation degree between the to-be-processed information and at least one keyword as a first correlation degree, and determine a second correlation degree between the to-be-processed information and a target feature;

[0143] If the first correlation degree is higher than the second correlation degree, select a first preset weight associated with the keyword as the basic coefficient;

[0144] If the second correlation degree is higher than the first correlation degree, select a second preset weight associated with the target feature as the basic coefficient.

[0145] In some embodiments, the method further comprises an optimization module configured to:

[0146] For each to-be-processed information in the plurality of to-be-processed information, obtain a diagnosis result of a target fault source to which the to-be-processed information belongs; the diagnosis result is used to indicate whether the target fault source is finally determined as a final fault source;

[0147] Optimize the weight coefficient of the to-be-processed information based on the diagnosis result;

[0148] If the diagnosis result indicates that the target fault source is determined as the final fault source, increase the weight coefficient based on a preset optimization manner; if the diagnosis result indicates that the target fault source is not determined as the final fault source, decrease the weight coefficient based on the preset optimization manner.

[0149] The specific functions and examples of each module and sub-module of the apparatus of the embodiments of the present disclosure are described in the above method embodiments, and will not be described here.

[0150] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0151] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0152] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.

[0153] As shown, Figure 7 the device 700 includes a computing unit 701 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded into a random access memory (RAM) 703 from a storage unit 708. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0154] Various components in the device 700 are connected to the I / O interface 705, including an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; the storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0155] The computing unit 701 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 701 performs various methods and processes described above, such as the artificial intelligence-based fault analysis method. For example, in some embodiments, the artificial intelligence-based fault analysis method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the artificial intelligence-based fault analysis method described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the artificial intelligence-based fault analysis method by any other appropriate means, such as by means of firmware.

[0156] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0157] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine or server, or entirely on a remote machine or server.

[0158] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0159] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0160] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0161] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0162] It should be understood that the various forms of flow shown above can be re-ordered, steps added or removed, etc. For example, the steps recited in the present disclosure can be performed in parallel, in series, in a different order, etc., so long as the desired results of the technology disclosed herein are achieved, and are not limited herein.

[0163] The specific embodiments discussed above do not constrain the scope of the present disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and alternatives can be made to the specific embodiments without departing from the principles of the present disclosure. Any such modifications, alternatives, equivalents, and / or alternatives are intended to fall within the scope of the present disclosure.

Claims

1. An artificial intelligence-based fault analysis method, comprising: extracting at least one keyword from alarm information; and, inputting the alarm information into a feature extraction model to obtain target features of the alarm information; based on the keyword and the target features, searching in a knowledge graph to obtain a plurality of candidate fault sources; based on the alarm type of the alarm information, screening at least one target fault source from the plurality of candidate fault sources.

2. The method of claim 1, wherein, The method of screening at least one target fault source from the plurality of candidate fault sources based on the alarm type of the alarm information comprises: for each candidate fault source, obtaining information associated with the candidate fault source in the knowledge graph under the alarm type to obtain a plurality of to-be-processed information; performing feature extraction on the plurality of to-be-processed information to obtain a plurality of to-be-processed features; determining the respective weight coefficients of the plurality of to-be-processed information under the alarm type; based on the respective weight coefficients of the plurality of to-be-processed information, performing weighted summation on the plurality of to-be-processed features to obtain a fault source score of the candidate fault source; based on the respective fault source scores of the plurality of candidate fault sources, screening the target fault source.

3. The method of claim 2, wherein the method of determining the respective weight coefficients of the plurality of to-be-processed information under the alarm type comprises: for each to-be-processed information, if no known weight of the to-be-processed information under the alarm type is found, determining a weight coefficient of the to-be-processed information based on an adjustment coefficient of the to-be-processed information and a base coefficient of the to-be-processed information; the weight coefficient has a positive correlation with the adjustment coefficient and the base coefficient; and the adjustment coefficient is obtained based on a coefficient table corresponding to the to-be-processed information under the alarm type.

4. The method of claim 3, wherein, The method of determining the weight coefficient of the to-be-processed information based on the adjustment coefficient of the to-be-processed information and the base coefficient of the to-be-processed information comprises: determining a normalized reference value based on the adjustment coefficient of the to-be-processed information and the base coefficient of the to-be-processed information; the normalized reference value has a positive correlation with the adjustment coefficient and the base coefficient; determining a sum value of the respective normalized reference values of the plurality of to-be-processed information; dividing the normalized reference value of the to-be-processed information by the sum value to obtain the weight coefficient of the to-be-processed information.

5. The method of claim 3, wherein, The method of determining the weight coefficient of the to-be-processed information based on the adjustment coefficient of the to-be-processed information and the base coefficient of the to-be-processed information comprises: determining a normalized reference value based on the adjustment coefficient of the to-be-processed information and the base coefficient of the to-be-processed information; the normalized reference value has a positive correlation with the adjustment coefficient and the base coefficient; determining a sum value of the respective normalized reference values of the plurality of to-be-processed information; mapping the sum value to a target coordinate system so that the value of the sum value in the target coordinate system is a target value; mapping the normalized reference value to the target coordinate system to obtain a mapping value of the normalized reference value of the to-be-processed information as the weight coefficient of the to-be-processed information.

6. The method of claim 3, wherein, The base coefficient of the to-be-processed information is determined based on the following method: determine a highest correlation degree between the to-be-processed information and the at least one keyword as a first correlation degree, and determine a second correlation degree between the to-be-processed information and the target feature; when the first correlation degree is higher than the second correlation degree, select a first preset weight associated with the keyword as the basic coefficient; when the second correlation degree is higher than the first correlation degree, select a second preset weight associated with the target feature as the basic coefficient.

7. The method of any one of claims 3-6, further comprising: for each to-be-processed information in the plurality of to-be-processed information, obtaining a diagnosis result of a target fault source to which the to-be-processed information belongs; the diagnosis result is used to indicate whether the target fault source is finally determined as a final fault source; based on the diagnosis result, optimizing a weight coefficient of the to-be-processed information; wherein, when the diagnosis result indicates that the target fault source is determined as the final fault source, the weight coefficient is increased based on a preset optimization manner; and when the target fault source is not determined as the final fault source, the weight coefficient is decreased based on the preset optimization manner.

8. An artificial intelligence-based fault analysis apparatus, comprising: an extraction module configured to extract at least one keyword from alarm information; and input the alarm information into a feature extraction model to obtain a target feature of the alarm information; a retrieval module configured to retrieve in a knowledge graph based on the keyword and the target feature to obtain a plurality of candidate fault sources; a screening module configured to screen at least one target fault source from the plurality of candidate fault sources based on an alarm type of the alarm information.

9. An electronic device, comprising: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of claims 1-7.

10. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to execute the method of any one of claims 1-7.

11. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-7.

Citation Information

Cited By

  • Root cause speculation method and device based on host state abnormity, equipment and medium

    CN121547343A

  • Methods, apparatus, equipment, and media for inferring the root causes of host status anomalies.

    CN121547343B